Probability Measure Notes
Probability Measure Notes
P.J.C. Spreij
3 Integration 16
3.1 Integration of simple functions . . . . . . . . . . . . . . . . . . . . . . 16
3.2 A general definition of integral . . . . . . . . . . . . . . . . . . . . . . 19
3.3 Integrals over subsets . . . . . . . . . . . . . . . . . . . . . . . . . . . . 23
3.4 Expectation and integral . . . . . . . . . . . . . . . . . . . . . . . . . . 24
3.5 Lp -spaces of random variables . . . . . . . . . . . . . . . . . . . . . . . 26
3.6 Lp -spaces of functions . . . . . . . . . . . . . . . . . . . . . . . . . . . 27
3.7 Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 30
4 Product measures 32
4.1 Product of two measure spaces . . . . . . . . . . . . . . . . . . . . . . 32
4.2 Application in Probability theory . . . . . . . . . . . . . . . . . . . . . 35
4.3 Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 37
5 Derivative of a measure 41
5.1 Absolute continuity and singularity . . . . . . . . . . . . . . . . . . . . 41
5.2 The Radon-Nikodym theorem . . . . . . . . . . . . . . . . . . . . . . . 42
5.3 Decomposition of a distribution function . . . . . . . . . . . . . . . . . 43
5.4 Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 43
6 Conditional expectation 46
6.1 A simple, finite case . . . . . . . . . . . . . . . . . . . . . . . . . . . . 46
6.2 Conditional expectation for X ∈ L1 (Ω, F, P) . . . . . . . . . . . . . . . 47
6.3 Conditional probabilities . . . . . . . . . . . . . . . . . . . . . . . . . . 50
6.4 Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 53
1 σ-algebras and measures
In this chapter we lay down the measure theoretic foundations of probability
theory. We start with some general notions and show how these are instrumental
in a probabilistic environment.
1.1 σ-algebras
Definition 1.1 Let S be a non-empty set. A collection Σ0 ⊂ 2S is called an
algebra (on S) if
(i) S ∈ Σ0 ,
(ii) E ∈ Σ0 ⇒ E c ∈ Σ0 ,
(iii) E, F ∈ Σ0 ⇒ E ∪ F ∈ Σ0 .
Alternatively, a collection
S∞ Σ is a σ-algebra if (i) and (ii) of Definition 1.1 are
valid together with n=1 En ∈ Σ as soon as En ∈ Σ (n = 1, 2 . . .). It follows
that an algebra is closed under countably many of the usual set operations.
If Σ is a σ-algebra on S, then (S, Σ) is called a measurable space and the
elements of Σ are called measurable sets. We shall ‘measure’ them in the next
section.
If C is any collection of subsets of S, then by σ(C) we denote the smallest σ-
algebra containing C. This means that σ(C) is the intersection of all σ-algebras
that contain C (see Exercise 1.1). If Σ = σ(C), we say that C generates Σ. The
union of two σ-algebras Σ1 and Σ2 on a set S is usually not a σ-algebra. We
write Σ1 ∨ Σ2 for σ(Σ1 ∪ Σ2 ).
One of the most relevant σ-algebras of this course is B = B(R), the Borel sets
of R. Let O be the collection of all open subsets of R with respect to the usual
topology (in which all intervals (a, b) are open). Then B := σ(O). Of course,
one similarly defines the Borel sets of Rd , and in general, for a topological space
(S, O), one defines the Borel-sets as σ(O). Borel sets can in principle be rather
‘wild’, but it helps to understand them a little better, once we know that they
are generated by simple sets. The σ-algebra B(R) is also generated by all open
intervals, or all closed intervals. An alternative generating collection is given
next.
1
Proof We prove the two obvious inclusions, starting with σ(I) ⊂ B. Since
(−∞, x] = ∩n (−∞, x + n1 ) ∈ B, we have I ⊂ B and then also σ(I) ⊂ B, since
σ(I) is the smallest σ-algebra that contains I. (Below we will use this kind of
arguments repeatedly).
For the proof of the reverse inclusion we proceed in three steps. First we
observe that (−∞, x) = ∪n (−∞, x − n1 ] ∈ σ(I). Knowing this, we conclude
that (a, b) = (−∞, b) \ (−∞, a] ∈ σ(I). Let then G be an arbitrary open set.
Since G is open, for every x ∈ G there exists a rational εx > 0 such that
(x − 2εx , x + 2εx ) ⊂ G. Consider (x − εx , x + εx ) and choose a rational qx in
this interval, note that |x − qx | ≤ εx . It follows that x ∈ (qx − εx , qx + εx ) ⊂
(x − 2εx , x + 2εx ) ⊂ G. Hence G ⊂ ∪x∈G (qx − εx , qx + εx ) ⊂ G, and so
G = ∪x∈G (qx − εx , qx + εx ). But the union here is in fact a countable union,
since there are only countably many qx and εx . (Note that the arguments
above can be used for any metric space with a countable dense subset to get
that an open G is a countable union of open balls.) It follows that G ∈ σ(I),
hence O ⊂ σ(I), and therefore (recall B is the smallest σ-algebra containing O)
B ⊂ σ(I).
1.2 Measures
Let Σ0 be an algebra on a set S, and Σ be a σ-algebra on S. We consider
mappings µ0 : Σ0 → [0, ∞] and µ : Σ → [0, ∞]. Note that ∞ is allowed as a
possible value.
We call µ0 finitely additive if µ0 (∅) = 0 and if µ0 (E ∪ F ) = µ0 (E) + µ0 (F )
for every pair of disjoint sets E and F in Σ0 . Of course this addition rule then
extends to arbitrary finite unions of disjoint sets. The mapping P µ0 is called
σ-additive or countably additive, if µ0 (∅) = 0 and if µ0 (∪n En ) = n µ0 (En ) for
every sequence (En ) of disjoint sets of Σ0 whose union is also in Σ0 . σ-additivity
is defined similarly for µ, but then we don’t have to require that ∪n En ∈ Σ.
This is true by definition.
2
special case) be the counting measure: τ (E) = |E|, the cardinality of E. One
easily verifies that τ is a measure, and it is σ-finite, because N = ∪n {1, . . . , n}.
A very simple measure is the Dirac measure. Consider a measurable space
(S, Σ) and single out a specific x0 ∈ S. Define δ(E) = 1E (x0 ), for E ∈ Σ (1E is
the indicator function of the set E, 1E (x) = 1 if x ∈ E and 1E (x) = 0 if x ∈ / E).
Check that δ is a measure on Σ.
Another example is Lebesgue measure, whose existence is formulated below.
It is the most natural candidate for a measure on the Borel sets on the real line.
Theorem 1.5 There exists a unique measure λ on (R, B) with the property
that for every interval I = (a, b] with a < b it holds that λ(I) = b − a.
The proof of this theorem can be found in the literature, or in the extended
version of these notes. We take this existence result for granted. One remark
is in order. One can show that B is not the largest σ-algebra for which the
measure λ can coherently be defined. On the other hand, on the power set of
R it is impossible to define a measure that coincides with λ on the intervals.
Here are the first elementary properties of a measure.
Proposition 1.6 Let (S, Σ, µ) be a measure space. Then the following hold
true (all the sets below belong to Σ).
(i) If E ⊂ F , then µ(E) ≤ µ(F ).
(ii) µ(E ∪ F ) ≤ µ(E) + µ(F ).
Pn
(iii) µ(∪nk=1 Ek ) ≤ k=1 µ(Ek )
If µ is finite, we also have
(iv) If E ⊂ F , then µ(F \ E) = µ(F ) − µ(E).
(v) µ(E ∪ F ) = µ(E) + µ(F ) − µ(E ∩ F ).
Proof The set F can be written as the disjoint union F = E ∪ (F \ E). Hence
µ(F ) = µ(E) + µ(F \ E). Property (i) now follows and (iv) as well, provided µ
is finite. To prove (ii), we note that E ∪ F = E ∪ (F \ (E ∩ F )), a disjoint union,
and that E ∩ F ⊂ F . The result follows from (i). Moreover, (v) also follows, if
we apply (iv). Finally, (iii) follows from (ii) by induction.
Measures have certain continuity properties.
Proposition 1.7 Let (En ) be a sequence in Σ.
(i) If the sequence is increasing, with limit E = ∪n En , then µ(En ) ↑ µ(E) as
n → ∞.
(ii) If the sequence is decreasing, with limit E = ∩n En and if µ(En ) < ∞
from a certain index on, then µ(En ) ↓ µ(E) as n → ∞.
Proof (i) Define D1 = E1 and Dn = En \ ∪n−1 k=1 Ek for n ≥ 2. Then the
Dn are disjoint,
Pn E n = ∪
P∞
n
k=1 D k for n ≥ 1 and E = ∪∞k=1 Dk . It follows that
µ(En ) = k=1 µ(Dk ) ↑ k=1 µ(Dk ) = µ(E).
To prove (ii) we assume without loss of generality that µ(E1 ) < ∞. Define
Fn = E1 \ En . Then (Fn ) is an increasing sequence with limit F = E1 \ E. So
(i) applies, yielding µ(E1 ) − µ(En ) ↑ µ(E1 ) − µ(E). The result follows.
3
Corollary 1.8 Let (S, Σ, µ) be a measure P∞space. For an arbitrary sequence
(En ) of sets in Σ, we have µ(∪∞ E
n=1 n ) ≤ n=1 µ(En ).
Remark 1.9 The finiteness condition in the second assertion of Proposition 1.7
is essential. Consider N with the counting measure τ . Let Fn = {n, n + 1, . . .},
then ∩n Fn = ∅ and so it has measure zero. But τ (Fn ) = ∞ for all n.
4
E ∪ F = (E c ∩ F c )c ∈ Σ, because we have just shown that complements remain
in Σ and because Σ is a π-system. Then Σ is also closed under finite unions.
Let E1 , E2 , . . . be a sequence in Σ. We have just showed that the sets Fn =
∪ni=1 Ei ∈ Σ. But since the Fn form an increasing sequence, also their union
is in Σ, because Σ is a d-system. But ∪n Fn = ∪n En . This proves that Σ is a
σ-algebra. Of course the other implication is trivial.
Proof Suppose that we would know that d(I) is a π-system as well. Then
Proposition 1.12 yields that d(I) is a σ-algebra, and so it contains σ(I). Since
the reverse inclusion is always true, we have equality. Therefore we will prove
that indeed d(I) is a π-system.
Step 1. Put D1 = {B ∈ d(I) : B ∩ C ∈ d(I), ∀C ∈ I}. We claim that
D1 is a d-system. Given that this holds and because, obviously, I ⊂ D1 , also
d(I) ⊂ D1 . Since D1 is defined as a subset of d(I), we conclude that these
two collections are the same. We now show that the claim holds. Evidently
S ∈ D1 . Let B1 , B2 ∈ D1 with B1 ⊂ B2 and C ∈ I. Write (B2 \ B1 ) ∩ C as
(B2 ∩ C) \ (B1 ∩ C). The last two intersections belong to d(I) by definition of
D1 and so does their difference, since d(I) is a d-system. For Bn ↑ B, Bn ∈ D1
and C ∈ I we have (Bn ∩ C) ∈ d(I) which then converges to B ∩ C ∈ d(I). So
B ∈ D1 .
Step 2. Put D2 = {C ∈ d(I) : B ∩ C ∈ d(I), ∀B ∈ d(I)}. We claim, again,
(and you check) that D2 is a d-system. The key observation is that I ⊂ D2 .
Indeed, take C ∈ I and B ∈ d(I). The latter collection is nothing else but D1 ,
according to step 1. But then B ∩ C ∈ d(I), which means that C ∈ D2 . It now
follows that d(I) ⊂ D2 , but then we must have equality, because D2 is defined
as a subset of d(I). The equality D2 = d(I) and the definition of D2 together
imply that d(I) is a π-system, as desired.
All these efforts lead to the following very useful theorem. It states that any
finite measure on Σ is characterized by its action on a rich enough π-system.
We will meet many occasions where this theorem is used.
5
Theorem 1.15 Let I be a π-system and Σ = σ(I). Let µ1 and µ2 be finite
measures on Σ with the properties that µ1 (S) = µ2 (S) and that µ1 and µ2
coincide on I. Then µ1 = µ2 (on Σ).
Proof The whole idea behind the proof is to find a good d-system that contains
I. The following set is a reasonable candidate. Put D = {E ∈ Σ : µ1 (E) =
µ2 (E)}. The inclusions I ⊂ D ⊂ Σ are obvious. If we can show that D is a
d-system, then Corollary 1.14 gives the result. The fact that D is a d-system is
straightforward to check, we present only one verification. Let E, F ∈ D such
that E ⊂ F . Then (use Proposition 1.6 (iv)) µ1 (F \ E) = µ1 (F ) − µ1 (E) =
µ2 (F ) − µ2 (E) = µ2 (F \ E) and so F \ E ∈ D.
Remark 1.16 In the above proof we have used the fact that µ1 and µ2 are
finite. If this condition is violated, then the assertion of the theorem is not
valid in general. Here is a counterexample. Take N with the counting measure
µ1 = τ and let µ2 = 2τ . A π-system that generates 2N is given by the sets
Gn = {n, n + 1, . . .} (n ∈ N).
6
probability measure P on this F with the nice property that for instance the set
{ω ∈ Ω : ω1 = ω2 = 1} (in the previous example it would have been denoted by
{11}) has probability 14 .
Having the interpretation of F as a collection of events, we now introduce two
special events. Consider a sequence of events E1 , E2 , . . . and define
∞ [
\ ∞
lim sup En := En
m=1 n=m
[∞ \ ∞
lim inf En := En .
m=1 n=m
Note that the sets Fm = ∩n≥m En form an increasing sequence and the sets
Dm = ∪n≥m En form a decreasing sequence. Clearly, F is closed under taking
limsup and liminf. The terminology is explained by (i) of Exercise 1.4. In prob-
abilistic terms, lim sup En is described as the event that the En occur infinitely
often, abbreviated by En i.o. Likewise, lim inf En is the event that the En occur
eventually. The former interpretation follows by observing that ω ∈ lim sup En
iff for all m, there exists n ≥ m such that ω ∈ En . In other words, a particular
outcome ω belongs to lim sup En iff it belongs to some (infinite) subsequence of
(En ). S∞ T ∞
The terminology to call m=1 n=m En the lim inf of the sequence is justified
in Exercise 1.4. In this exercise, indicator functions of events are used, of which
we here recall the definition. If E is an event, then the function 1E is defined
by 1E (ω) = 1 if ω ∈ E and 1E (ω) = 0 if ω ∈ / E.
1.6 Exercises
1.1 Prove the following statements.
(a) The intersection of an arbitrary family of d-systems is again a d-system.
(b) The intersection of an arbitrary family of σ-algebras is again a σ-algebra.
(c) If C1 and C2 are collections of subsets of Ω with C1 ⊂ C2 , then d(C1 ) ⊂ d(C2 ).
1.2 Prove Corollary 1.8.
1.3 Prove the claim that D2 in the proof of Lemma 1.13 forms a d-system.
1.4 Consider a measure space (S, Σ, µ). Let (En ) be a sequence in Σ.
(a) Show that 1lim inf En = lim inf 1En .
(b) Show that µ(lim inf En ) ≤ lim inf µ(En ). (Use Proposition 1.7.)
(c) Show also that µ(lim sup En ) ≥ lim sup µ(En ), provided that µ is finite.
1.5 Let (S, Σ, µ) be a measure space. Call a subset N of S a (µ, Σ)-null set
if there exists a set N 0 ∈ Σ with N ⊂ N 0 and µ(N 0 ) = 0. Denote by N the
collection of all (µ, Σ)-null sets. Let Σ∗ be the collection of subsets E of S for
which there exist F, G ∈ Σ such that F ⊂ E ⊂ G and µ(G \ F ) = 0. For E ∈ Σ∗
and F, G as above we define µ∗ (E) = µ(F ).
7
(a) Show that Σ∗ is a σ-algebra and that Σ∗ = Σ ∨ N (= σ(N ∪ Σ)).
(b) Show that µ∗ restricted to Σ coincides with µ and that µ∗ (E) doesn’t
depend on the specific choice of F in its definition.
(c) Show that the collection of (µ∗ , Σ∗ )-null sets is N .
1.6 Let G and H be two σ-algebras on Ω. Let C = {G ∩ H : G ∈ G, H ∈ H}.
Show that C is a π-system and that σ(C) = σ(G ∪ H).
Ω
1.7
P Let Ω be a countable set.P Let F = 2 and let p : Ω → [0, 1] satisfy
ω∈Ω p(ω) = 1. Put P(A) = ω∈A p(ω) for A ∈ F. Show that P is a probability
measure.
1.8 Let Ω be a countable set. Let A be the collection of A ⊂ Ω such that A
or its complement has finite cardinality. Show that A is an algebra. What is
d(A)?
1.12 Consider R together with the Borel sets and the Lebesgue measure λ.
(a) Show that all finite and countable sets have measure zero. What is λ(Q)?
P
(b) Is λ(R) = x∈R λ({x}) (whatever the sum could mean)?
8
2 Measurable functions and random variables
In this chapter we define random variables as measurable functions on a proba-
bility space and derive some properties.
9
Remark 2.4 There are many variations on the assertions of Proposition 2.3
possible. For instance in (ii) we could also use {h < c}, or {h > c}. Further-
more, (ii) is true for h : S → [−∞, ∞] as well. We proved (iv) by a simple
composition argument, which also applies to a more general situation. Let
(Si , Σi ) be measurable spaces (i = 1, 2, 3), h : S1 → S2 is Σ1 /Σ2 -measurable
and f : S2 → S3 is Σ2 /Σ3 -measurable. Then f ◦ h is Σ1 /Σ3 -measurable.
Theorem 2.6 Let H be a vector space of bounded functions, with the following
properties.
(i) 1 ∈ H.
(ii) If (fn ) is a nonnegative sequence in H such that fn+1 ≥ fn for all n, and
f := lim fn is bounded, then f ∈ H.
If, in addition, H contains the indicator functions of sets in a π-system I, then
H contains all bounded σ(I)-measurable functions.
10
Proof Put D = {F ⊂ S : 1F ∈ H}. One easily verifies that D is a d-system,
and that it contains I. Hence, by Corollary 1.14, we have Σ := σ(I) ⊂ D. We
will use this fact later in the proof.
Let f be a bounded, σ(I)-measurable function. Without loss of generality,
we may assume that f ≥ 0 (add a constant otherwise), and f < K for some real
constant K. Introduce the functions fn defined by fn = 2−n b2n f c. In explicit
terms, the fn are given by
n
X−1
K2
fn (s) = i2−n 1{i2−n ≤f <(i+1)2−n } (s).
i=0
11
of X, we introduce its distribution function, usually denoted by F (or FX ,
or F X ). By definition it is the function F : R → [0, 1], given by F (x) =
µ((−∞, x]) = P(X ≤ x).
Proposition 2.8 The distribution function of a random variable is right con-
tinuous, non-decreasing and satisfies limx→∞ F (x) = 1 and limx→−∞ F (x) = 0.
The set of points where F is discontinuous is at most countable.
This proposition thus states, in a different wording, that for a random variable
X, its distribution, the collection of all probabilities P(X ∈ B) with B ∈ B, is
determined by the distribution function FX .
We call any function on R that has the properties of Proposition 2.8 a distri-
bution function. Note that any distribution function is Borel measurable (sets
{F ≥ c} are intervals and thus in B). Below, in Theorem 2.10, we justify this
terminology. We will see that for any distribution function F , it is possible
to construct a random variable on some (Ω, F, P), whose distribution function
equals F . This theorem is founded on the existence of the Lebesgue measure λ
on the Borel sets B[0, 1] of [0, 1], see Theorem 1.5.
We now give a probabilistic translation of this theorem. Consider (Ω, F, P) =
([0, 1], B[0, 1], λ). Let U : Ω → [0, 1] be the identity map. The distribution of U
on [0, 1] is trivially the Lebesgue measure again, in particular the distribution
function F U of U satisfies F U (x) = x for x ∈ [0, 1] and so P(a < U ≤ b) =
F U (b) − F U (a) = b − a for a, b ∈ [0, 1] with a ≤ b. Hence, to the distribu-
tion function F U corresponds a probability measure on ([0, 1], B[0, 1]) and there
exists a random variable U on this space, such that U has F U as its distri-
bution function. The random variable U is said to have the standard uniform
distribution.
The proof of Theorem 2.10 (Skorokhod’s representation of a random variable
with a given distribution function) below is easy in the case that F is continuous
and strictly increasing (Exercise 2.6), given the just presented fact that a random
variable with a uniform distribution exists. The proof that we give below for
the general case just follows a more careful line of arguments, but is in spirit
quite similar.
12
Theorem 2.10 Let F be a distribution function on R. Then there exists a
probability space and a random variable X : Ω → R such that F is the distri-
bution function of X.
Proof Let (Ω, F, P) = ((0, 1), B(0, 1), λ). We define X − (ω) = inf{z ∈ R :
F (z) ≥ ω}. Then X − (ω) is finite for all ω and X − is Borel measurable function,
so a random variable, as this follows from the relation to be proven below, valid
for all c ∈ R and ω ∈ (0, 1),
2.3 Independence
Recall the definition of independent events. Two events E, F ∈ F are called
independent if the product rule P(E ∩ F ) = P(E)P(F ) holds. In the present
section we generalize this notion of independence to independence of a sequence
of events and to independence of a sequence of σ-algebras. It is even convenient
and elegant to start with the latter.
The above definition also applies to finite sequences. For instance, a finite se-
quence of σ-algebras F1 , . . . , Fn is called independent if the infinite sequence
F1 , F2 , . . . is independent in the sense of part (ii) of the above definition, where
Fm = {∅, Ω} for m > n. It follows that two σ-algebras F1 and F2 are indepen-
dent, if P(E1 ∩ E2 ) = P(E1 )P(E2 ) for all E1 ∈ F1 and E2 ∈ F2 . It also follows
that two events E and F are independent iff σ(E) and σ(F ) are independent
σ-algebras, this is Exercise 2.12.
13
To check independence of two σ-algebras, Theorem 1.15 is again helpful. It
tells you that for two σ-algebras to be independent, it is sufficient to check the
product rule for generating π-systems.
Proposition 2.12 Let I and J be π-systems and suppose that for all I ∈ I
and J ∈ J the product rule P(I ∩ J) = P(I)P(J) holds. Then the σ-algebras
σ(I) and σ(J ) are independent.
Proof Put G = σ(I) and H = σ(J ). We define for each I ∈ I the finite
measures µI and νI on H by µI (H) = P(H ∩I) and νI (H) = P(H)P(I) (H ∈ H).
Notice that µI and νI coincide on J by assumption and that µI (Ω) = P(I) =
νI (Ω). Theorem 1.15 yields that µI (H) = νI (H) for all H ∈ H.
Now we consider for each H ∈ H the finite measures µH and ν H on G
defined by µH (G) = P(G ∩ H) and ν H (G) = P(G)P(H). By the previous step,
we see that µH and ν H coincide on I. Invoking Theorem 1.15 again, we obtain
P(G ∩ H) = P(G)P(H) for all G ∈ G and H ∈ H.
2.4 Exercises
2.1 If h1 and h2 are Σ-measurable functions on (S, Σ, µ), then h1 h2 is Σ-
measurable too. Show this.
2.2 Let X be a random variable. Show that Π(X) := {{X ≤ x} : x ∈ R} is
a π-system and that it generates σ(X). Formulate a similar statement for a
two-dimensional random vector X.
2.3 Let {Yγ : γ ∈ C} be an arbitrary collection of random variables and {Xn :
n ∈ N} be a countable collection of random variables, all defined on the same
probability space.
(a) Show that σ{Yγ : γ ∈ C} = σ{Yγ−1 (B) : γ ∈ C, B ∈ B}.
S∞
(b) Let Xn = σ{X1 , . . . , Xn } (n ∈ N) and A = n=1 Xn . Show that A is an
algebra and that σ(A) = σ{Xn : n ∈ N}.
2.4 Prove Proposition 2.8.
14
2.5 Let F be a σ-algebra on Ω with the property that for all F ∈ F it holds
that P(F ) ∈ {0, 1}. Let X : Ω → R be F-measurable. Show that for some c ∈ R
one has P(X = c) = 1. (Hint: P(X ≤ x) ∈ {0, 1} for all x.)
2.6 Let F be a strictly increasing and continuous distribution function. Let U
be a random variable defined on some (Ω, F, P) having a uniform distribution
on [0, 1] and put X = F −1 (U ). Show that X is F-measurable and that it has
distribution function F .
2.7 Let F be a distribution function and put X + (ω) = inf{x ∈ R : F (x) > ω}.
Show that (next to X − ) also X + has distribution function F and that P(X + =
X − ) = 1 (Hint: P(X − ≤ q < X + ) = 0 for all q ∈ Q). Show also that X + is a
right continuous function and Borel-measurable.
2.8 Consider a probability space (Ω, F, P). Let I1 , I2 , I3 be π-systems on Ω
with the properties Ω ∈ Ik and Ik ⊂ F, for all k. Assume that for all Ik ∈ Ik
(k = 1, 2, 3)
15
3 Integration
In elementary courses on Probability Theory, there is usually a distinction be-
tween random variables X having a discrete distribution, on N say, and those
having a P density. In the former case we have for the expectation ER X the ex-
pression k k P(X = k), whereas in the latter case one has E X = xf (x) dx.
This distinction is annoying and not satisfactory from a mathematical point of
view. Moreover, there exist random variables whose distributions are neither
discrete, nor do they admit a density. Here is an example. Suppose Y and Z,
defined on the same (Ω, F, P), are independent random variables. Assume that
P(Y = 0) = P(Y = 1) = 12 and that Z has a standard normal distribution. Let
X = Y Z and F the distribution function of X. Easy computations (do them!)
yield F (x) = 21 (1[0,∞) (x) + Φ(x)). We see that F has a jump at x = 0 and is
differentiable on R \ {0}, a distribution function of mixed type. How to compute
E X in this case?
In this section we will see that expectations are special cases of the unifying
concept of Lebesgue integral, a sophisticated way of addition. Lebesgue integrals
have many advantages. It turns out that Riemann integrable functions (on
a compact interval) are always Lebesgue integrable w.r.t. Lebesgue measure
and that the two integrals are the same. Also sums are examples of Lebesgue
integral. Furthermore, the theory of Lebesgue integrals allows for very powerful
limit theorems. Below we work with a measure space (S, Σ, µ).
16
Definition 3.2 Let f ∈ S+ . The (Lebesgue) integral of f with respect to the
measure µ is defined as
Z X n
f dµ := ai µ(Ai ), (3.2)
i=1
It should be clear that this definition of integral is, at first sight, troublesome.
As the representation of a simple function is not unique, one might wonder if
the just defined integral takes on different values for different representations.
This would be very bad, and fortunately it is not the case.
Proposition 3.3 Let f be a nonnegative simple function. Then the value of
the integral µ(f ) is independent of the chosen representation.
Proof Step 1. Let f be given by (3.1) and define φ : S → {0, 1}n by φ(s) =
(1A1 (s), . . . , 1An (s)). Let {0, 1}n = {u1 , . . . , um } where m = 2n and put Uk =
φ−1 (uk ). Then the collection {U1 , . . . , Um } is a measurable partition of S (the
sets Uk are measurable). We will also need the sets Si = {k : Uk ⊂ Ai } and
Tk = {i : Uk ⊂ Ai }. Note that these sets are dual in the sense that k ∈ Si iff
i ∈ Tk .
Below we will use the fact Ai = ∪k∈Si Uk , when we rewrite (3.1). We obtain
by interchanging the summation order
X X X
f= ai 1Ai = ai ( 1Uk )
i i k∈Si
X X
= ( ai )1Uk . (3.3)
k i∈Tk
17
where the Bj form a measurable partition of S as well. We obtain a third
measurable partition of S by taking the collection of all intersections Ai ∩ Bj .
Notice that if s ∈ Ai ∩ Bj , then f (s) = ai = bj and so we have the implication
Ai ∩Bj 6= ∅ ⇒ ai = bj . We compute the integral of f according to thePdefinition.
Of course, this yields (3.2) by using the representation (3.1) of f , but j bj µ(Bj )
if we use (3.4). Rewrite
X X X X
bj µ(Bj ) = bj µ(∪i (Ai ∩ Bj )) = bj µ(Ai ∩ Bj )
j j j i
XX XX
= bj µ(Ai ∩ Bj ) = ai µ(Ai ∩ Bj )
i j i j
X X X
= ai µ(Ai ∩ Bj ) = ai µ(∪j (Ai ∩ Bj ))
i j i
X
= ai µ(Ai ),
i
which shows that the two formulas for the integral are the same.
Step 3: Take now two arbitrary representations of f of the form (3.1)
and (3.4). According to step 1, we can replace each of them with a repre-
sentation in terms of a measurable partition, without changing the value of the
integral. According to step 2, each of the representations in terms of the par-
titions also gives the same value of the integral. This proves the proposition.
Corollary 3.4 Let f ∈ S+ and suppose Pn that f assumes the different values
0 ≤ a1 , . . . , an < ∞. Then µ(f ) = i=1 ai µ({f = ai }). If f is indentically zero,
then µ(f ) = 0.
Pn
Proof We have the representation f = i=1 ai 1{f =ai } . The expression for
µ(f ) follows from Definition 3.2, which is unambiguous by Proposition 3.3. The
result for the zero function follows by the representation f = 0 × 1S and the
convention ab = 0 for a = 0 and b ∈ [0, ∞].
18
different representation would yield the same answer. A generalization occurs
when the set of values {fi : i ∈ N} is finite. Then f is still a simple function if
the fi are nonnegative, as in Corollary 3.4. But note that now it may happen
that τ (f ) = ∞ (think of all fi equal to one).
Example 3.6 Let (S, Σ, µ) = ([0, 1], B([0, 1]), λ) and f the indicator of the
rational numbers in [0, 1], f = 1Q∩[0,1] . We know that λ(Q ∩ [0, 1]) = 0 and it
follows that λ(f ) = 0. This f is a nice example of a function that is not Riemann
integrable, whereas its Lebesgue integral trivially exists and has a very sensible
value.
Assertion (ii) follows by a double application of (i), whereas (iii) can also be
proved by using partitions and intersections Ai ∩ Bj .
19
Definition 3.8 Let f be a nonnegative measurable function. The integral of
f is defined as µ(f ) := sup{µ(h) : h ≤ f, h ∈ S+ }, where µ(h) is as in Defini-
tion 3.2.
Notice that for functions f ∈ S+ , Definition 3.8 yields for µ(f ) the same as
Definition 3.2 in the previous section. Thus there is no ambiguity in notation
by using the same symbol µ. We immediately have some extensions of results
in the previous section.
h1N c ≤ f 1N c ≤ g1N c ≤ g.
Proof Because µ(f ) = 0, it holds that µ(h) = 0 for all nonnegative simple
functions with h ≤ f . Take hn = n1 1{f ≥1/n} , then hn ∈ S+ and hn ≤ f .
The equality µ(hn ) = 0 implies µ({f ≥ 1/n}) = 0. The result follows from
{f > 0} = ∪n {f ≥ 1/n} and Corollary 1.8.
We now present the first important limit theorem, the Monotone Convergence
Theorem.
Theorem 3.12 Let (fn ) be a sequence in Σ+ , such that fn+1 ≥ fn a.e. for
each n. Let f = lim sup fn . Then µ(fn ) ↑ µ(f ) ≤ ∞.
Proof We first consider the case where fn+1 (s) ≥ fn (s) for all s ∈ S, so (fn ) is
increasing everywhere. Then f (s) = lim fn (s) for all s ∈ S, possibly with value
20
infinity. It follows from Proposition 3.9, that µ(fn ) is an increasing sequence,
bounded by µ(f ). Hence we have ` := lim µ(fn ) ≤ µ(f ).
We show that we actually have an equality. Take h ∈ S+ with h ≤ f ,
c ∈ (0, 1) and put En = {fn ≥ ch}. The sequence (En ) is obviously increasing
and we show that its limit is S. Let s ∈ S and suppose that f (s) = 0. Then
also h(s) = 0 and s ∈ En for every n. If f (s) > 0, then eventually fn (s) ≥
cf (s) ≥ ch(s), and so s ∈ En . This shows that ∪n En = S. Consider the chain
of inequalities
21
to take limits on both side of this inequality. On the right hand side we get
lim inf µ(fn ). The sequence (gn ) is increasing, with limit g = lim inf fn , and by
Theorem 3.12, µ(gn ) ↑ µ(lim inf fn ) on the left hand side. This proves the first
assertion. The second assertion follows by considering f¯n = h − fn ≥ 0. Check
where it is used that µ(h) < ∞.
Remark 3.16 Let (En ) be a sequence of sets in Σ, and let fn = 1En and h = 1.
The statements of Exercise 1.4 follow from Lemma 3.15.
Definition 3.17 Let f ∈ Σ and assume that µ(f + ) < ∞ or µ(f − ) < ∞. Then
we define µ(f ) := µ(f + ) − µ(f − ). If both µ(f + ) < ∞ and µ(f − ) < ∞, we
say that f is integrable. The collection of all integrable functions is denoted by
L1 (S, Σ, µ). Note that f ∈ L1 (S, Σ, µ) implies that |f | < ∞ µ-a.e.
Example 3.19 Let (S, Σ, µ) = ([0, 1], B([0, 1]), λ), where λ is Lebesgue mea-
sure. Assume that f ∈ C[0, 1]. Exercise 3.5 yields that f ∈ L1 ([0, 1], B([0, 1]), λ)
R1
and that λ(f ) is equal to the Riemann integral 0 f (x) dx. This implication fails
to hold if we replace [0, 1] with an unbounded interval, see Exercise 3.6.
On the other hand, one can even show that every function that is Riemann
integrable over [0, 1], not only a continuous function, is Lebesgue integrable too.
More knowledge is required for a precise statement and its proof.
The next theorem is known as the Dominated Convergence Theorem, also called
Lebesgue’s Convergence Theorem.
Theorem 3.20 Let (fn ) ⊂ Σ and f ∈ Σ. Assume that fn (s) → f (s) for all s
outside a set of measure zero. Assume also that there exists a function g ∈ Σ+
such that supn |fn | ≤ g a.e. and that µ(g) < ∞. Then µ(|fn − f |) → 0 (often
L1
denoted fn → f ), and hence µ(fn ) → µ(f ).
22
Proof The second assertion easily follows from the first one, which we prove
now for the case that fn → f everywhere. One has the inequality |f | ≤ g,
whence |fn − f | ≤ 2g. The second assertion of Fatou’s lemma immediately
yields lim sup µ(|fn − f |) ≤ 0, which is what we wanted. The almost everywhere
version is left as Exercise 3.4.
One verifies (Exercise 3.8) that ν is a measure on (S, Σ). We want to compute
ν(h) for h ∈ Σ+ . For measurable indicator functions we have by definition
that the integral ν(1E ) equals ν(E), which is equal to µ(1E f ) by (3.7). More
generally we have
23
R R
For the measure ν above, Proposition 3.22 states that h dν = hf dµ, valid
dν
for all h ∈ L1 (S, Σ, ν). The notation f = dµ is often used and looks like
a derivative. We will return to thisR in Chapter
R 5, where we discuss Radon-
Nikodym derivatives. The equality h dν = hf dµ now takes the appealing
form Z Z
dν
h dν = h dµ.
dµ
where the equality is valid under conditions as for instance in Example 3.19.
R
Remark 3.24 If f is continuous, see Example 3.19, then x 7→ F (x) = [0,x] f dλ
defines a differentiable function on (0, 1), with F 0 (x) = f (x). This follows
from theR theory of Riemann integrals. We adopt the conventional notation
x
F (x) = 0 f (u) du. This case can be generalized as follows. If f ∈ L1 (R, B, λ),
Rx
then (using a similar notational convention) x 7→ F (x) = −∞ f (u) du is well
defined for all x ∈ R. Moreover, F is at (Lebesgue) almost all points x of
R differentiable with derivative F 0 (x) = f (x). The proof of this result, the
fundamental theorem of calculus for the Lebesgue integral, is not given here.
provided that the integral is well defined, which is certainly the case if X is
nonnegative or if P(|X|) < ∞. Other often used notations for the integral in
(3.8) are PX and E X. The latter is the favorite one among probabilists and
one speaks of the Expectation of X. Note also that E X is always defined when
X ≥ 0 almost surely. The latter concept meaning almost everywhere w.r.t. the
probability measure P. We abbreviate almost surely by a.s.
Example 3.25 Let (Ω,PF, P) = (N, 2N , P), where P is defined by P({n}) = pn ,
where all pn ≥ 0 and pn = 1. Let (xn ) be a sequence of nonnegative real
numbers and define the random variable X, a function on Ω = N, by X(n) = xn .
In a spirit
P∞similar to what we have seen in Examples 3.5 and 3.10, we get
EX = n=1 xn pn . Let us switch to a different approach. Let ξ1 , ξ2 , . . . be
24
the different elements of the set {x1 , x2 , . . .} and put Ei = {j : xj = ξi },
i ∈ N. Notice
P that {X = ξi } = Ei and that
P the Ei form a partition
P of N with
P(Ei ) = j∈Ei pj . It follows that E X = i ξi P(Ei ), or E X = i ξi P(X = ξi ),
the familiar expression for the expectation.
If h : R → R is Borel measurable, then Y := h ◦ X (we also write Y = h(X))
is a random variable as well. There are two recipes to compute E Y . One is of
course the direct application of the definition of expectation to Y . But we also
have
Proposition 3.26 Let X be a random variable, and h : R → R Borel mea-
surable. Let PX be the distribution of X. Then h ◦ X ∈ L1 (Ω, F, P) iff
h ∈ L1 (R, B, PX ), in which case
Z
E h(X) = h dPX . (3.9)
R
It
R follows from this proposition that one can also compute E Y as E Y =
Y X
R
R
y P ( dy), and of course E X = R
x P ( dx).
Example 3.27 Suppose there exists f ≥ 0, Borel-measurable such that for all
B ∈ B one has PX (B) = λ(1B f ), in which case it is said that X has a density
f . Then, provided that the expectation is well defined, Example 3.23 yields
Z
E h(X) = h(x)f (x) dx,
R
25
The next two propositions have proven to be very useful in proofs of results in
Probability Theory.
Verify this property graphically. The d(x) are also called subgradients of g. The
following proposition (Jensen’s inequality) is now easy to prove, and you check
where we use in the proof the fact that P is a probability measure.
Proposition 3.30 Let g : G → R be convex and X a random variable with
P(X ∈ G) = 1. Assume that E |X| < ∞ and E |g(X)| < ∞. Then
26
The notation || · || suggests that we deal with a norm. In a sense, this is correct,
but we will not explain this until the end of this section. It is however obvious
that Lp := Lp (Ω, F, P) is a vector space, since |X + Y |p ≤ (|X| + |Y |)p ≤
2p (|X|p + |Y |p ).
In the special case p = 2, we have for X, Y ∈ L2 , that |XY | = 12 ((|X|+|Y |)2 −
X −Y 2 ) has finite expectation and is thus in L1 . Of course we have |E (XY )| ≤
2
Proof The standard machine easily gives E (1A Y ) = P(A) · E Y for A an event
independent of Y . Assume that X ∈ S+ . Since X is integrable we can assume
that it is finite, and thus bounded by a constantP c. Since then |XY | ≤ c|Y |,
n
we
Pn obtain E |XY | < ∞. If we represent X as i=1 ai 1Ai , then E (XY ) =
i=1 ai P(Ai )E Y readily follows and thus E (XY ) = E X · E Y . The proof may
be finished by letting the standard machine operate on X.
Proof It follows from the trivial inequality |u| ≤ 1+|u|a , valid for u ∈ R and a ≥
1, that |X|p ≤ 1+|X|r , by taking a = r/p, and hence X ∈ Lp (Ω, F, P). Observe
that x → |x|a is convex. We apply Jensen’s inequality to get (E |X|p )a ≤
E (|X|pa ), from which the result follows.
27
Definition 3.35 Let 1 ≤ p < ∞ and f a measurable function on (S, Σ, µ). If
µ(|f |p ) < ∞, we write f ∈ Lp (S, Σ, µ) and ||f ||p = (µ(|f |p ))1/p .
with the convention inf ∅ = ∞. It is clear that |f | ≤ ||f ||∞ a.e. We write
f ∈ L∞ (S, Σ, µ) if ||f ||∞ < ∞.
Here is the first of two fundamental inequalities, known as Hölder’s inequality.
Put h(s) = g(s)/f (s)p−1 if f (s) > 0 and h(s) = 0 otherwise. Jensen’s inequality
gives (P(h))q ≤ P(hq ). We compute
µ(f g)
P(h) = ,
µ(f p )
and
µ(1{f >0} g q ) µ(g q )
P(hq ) = ≤ .
µ(f p ) µ(f p )
Insertion of these expressions into the above version of Jensen’s inequality yields
(µ(f g))q µ(g q )
≤ ,
(µ(f p ))q µ(f p )
whence (µ(f g))q ≤ µ(g q )µ(f p )q−1 . Take q-th roots on both sides and the result
follows.
28
Theorem 3.38 Let f, g ∈ Lp (S, Σ, µ) and p ∈ [1, ∞]. Then ||f + g||p ≤ ||f ||p +
||g||p .
Proof The case p = ∞ is almost trivial, so we assume p ∈ [1, ∞). To exclude
another triviality, we suppose ||f + g||p > 0. Note the following elementary
relations.
|f + g|p = |f + g|p−1 |f + g| ≤ |f + g|p−1 |f | + |f + g|p−1 |g|.
Now we take integrals and apply Hölder’s inequality to obtain
µ(|f + g|p ) ≤ µ(|f + g|p−1 |f |) + µ(|f + g|p−1 |g|)
≤ (||f ||p + ||g||p )(µ(|f + g|(p−1)q )1/q
= (||f ||p + ||g||p )(µ(|f + g|p )1/q ,
because (p − 1)q = p. After dividing by (µ(|f + g|p )1/q , we obtain the result,
because 1 − 1/q = 1/p.
Recall the definition of a norm on a (real) vector space X. One should have
||x|| = 0 iff x = 0, ||αx|| = |α| ||x|| for α ∈ R (homogeneity) and ||x + y|| ≤
||x|| + ||y|| (triangle inequality). For || · ||p homogeneity is obvious, the triangle
inequality has just been proved under the name Minkowski’s inequality and
we also trivially have f = 0 ⇒ ||f ||p = 0. But, conversely ||f ||p = 0 only
implies f = 0 a.e. This annoying fact disturbs || · ||p being called a genuine
norm. This problem can be circumvented by identifying a function f that is
zero a.e. with the zero function. The proper mathematical way of doing this is
by defining the equivalence relation f ∼ g iff µ({f 6= g}) = 0. By considering
the equivalence classes induced by this equivalence relation one gets the quotient
space Lp (S, Σ, µ) := Lp (S, Σ, µ)/ ∼. One can show that || · ||p induces a norm
on this space in the obvious way. We don’t care too much about these details
and just call || · ||p a norm and Lp (S, Σ, µ) a normed space, thereby violating a
bit the standard mathematical language.
A desirable property of a normed space, (a version of) completeness, holds for
Lp spaces. We give this result for Lp (Ω, F, P).
Theorem 3.39 Let p ∈ [1, ∞]. The space Lp (Ω, F, P) is complete in the fol-
lowing sense. Let (Xn ) be a Cauchy-sequence in Lp : ||Xn − Xm ||p → 0 for
n, m → ∞. Then there exists a limit X ∈ Lp such that ||Xn − X||p → 0. The
limit is unique in the sense that any other limit X 0 satisfies ||X − X 0 ||p = 0.
Proof Omitted.
Remark 3.40 Notice that it follows from Theorem 3.39, and the discussion
preceding it, that Lp (Ω, F, P) is a truly complete normed space, called a Ba-
nach space. The same is true for Lp (S, Σ, µ) (p ∈ [1, ∞]), for which you need
Exercise 3.12. For theR special case p = 2 we endow L2 (S, Σ, µ) with the in-
ner product hf, gi := f g dµ, and it is then called a Hilbert space. Likewise
L2 (Ω, F, P) is a Hilbert space with inner product hX, Y i := E XY .
29
3.7 Exercises
3.1 Let (x1 , x2 , . . .) be a sequence of nonnegative real numbers, let ` : N → N
be a bijection and define the sequence (y1 , y2 , . . .) by yk = x`(k) . Let for each
n the n-vector y n be given by y n = (y1 , . . . , yn ). Consider then for each n
a sequence of numbers xn defined by xnk = xk if xk is a coordinate of y n .
n n
Otherwise
P∞ P∞xk = 0. Show that xk ↑ xk for every k as n → ∞. Show that
put
k=1 yk = k=1 xk .
30
3.14 Let (S, Σ, µ) be a measurable space, Σ0 a sub-σ-algebra of Σ and µ0 be the
restriction of µ to Σ0 . Then also (S, Σ0 , µ0 ) is a measurable space and integrals of
Σ0 -measurable functions can be defined according to the usual procedure. Show
that µ0 (f ) = µ(f ), if f ≥ 0 and Σ0 -measurable. Show also that L1 (S, Σ, µ)∩Σ0 =
L1 (S, Σ0 , µ0 ).
31
4 Product measures
So far we have consideredR measure spaces (S, Σ, µ) and we have looked at inte-
grals of the type µ(f ) = f dµ. Here f is a function of ‘one’ variable (depends
on how you count and what the underlying set S is). Suppose that we have two
measure spaces (S1 , Σ1 , µ1 ) and (S2 , Σ2 , µ2 ) and a function f : S1 × S2 → R.
Is it possible to integrate such a function of two variables w.r.t. some measure,
that has to be defined on some Σ-algebra of S1 × S2 . There is a natural way
of constructing this σ-algebra and a natural construction of a measure on this
σ-algebra. Here is a setup with some informal thoughts.
Take f : S1 × S2 → R and assume R any good notion of measurability and
integrability. Then µ(f (·, s2 )) := f (·, s2 ) dµ1 defines a function of s2 and so
we’d like to take the integral w.r.t. µ2 . We could as well have gone the other way
round (integrate first w.r.t. µ2 ), and the questions are whether these integrals
are well defined and whether both approaches yield the same result.
Here is a simple special case, where the latter question has a negative an-
swer. We have seen that integration w.r.t. counting measure is nothing else but
addition. What we have outlined above is in this context just interchanging the
summation order. P So P if (an,m ) P
is a P
double array of real numbers, the above is
about whether n m an,m = m n an,m . This is obviously true if n and m
run through a finite set, but things can go wrong for indices from infinite sets.
Consider for example
1 if n = m + 1
an,m = −1 if m = n + 1
0 else.
P P
One
P P easily verifies m a1,m = −1, m an,m = P n ≥ 2 and hence we find
0, if P
n a
m n,m = −1. Similarly one shows that m n an,m = +1. In order
that interchanging of the summation order yields the P same
P result, additional
conditions have to be imposed. We will see that m n |an,m | < ∞ is a
sufficient condition. As a side remark we note that this case has everything
to do with a well known theorem by Riemann that says that a series of real
numbers is absolutely convergent iff it is unconditionally convergent.
32
Next to the projections, we now consider embeddings. For fixed s1 ∈ S1 we
define es1 : S2 → S by es1 (s2 ) = (s1 , s2 ). Similarly we define es2 (s1 ) = (s1 , s2 ).
One easily checks that the embeddings es1 are Σ2 /Σ-measurable and that the
es2 are Σ1 /Σ-measurable (Exercise 4.1). As a consequence we have the following
proposition.
Proof This follows from the fact that a composition of measurable functions is
also measurable.
Remark 4.2 The converse statement of Proposition 4.1 is in general not true,
but a counterexample is beyond the scope of these lecture notes; see the full
version for details. There are functions f : S → R that are not measurable
w.r.t. the product σ-algebra Σ, although the mappings s1 7→ f (s1 , s2 ) and s2 7→
f (s1 , s2 ) are Σ1 -, respectively Σ2 -measurable. Counterexamples are not obvious,
see below for a specific one. Fortunately, there are also conditions that are
sufficient to have measurability of f w.r.t. Σ, when measurability of the marginal
functions is given. See Exercise 4.8.
33
Proof We use the Monotone Class Theorem, Theorem 2.6, and so we have
to find a good vector space H. The obvious candidate is the collection of all
bounded Σ-measurable functions f that satisfy the assertions of the lemma.
First we notice that H is indeed a vector space, since sums of measurable
functions are measurable and by linearity of the integral. Obviously, the con-
stant functions belong to H. Then we have to show that if fn ∈ H, fn ≥ 0 and
fn ↑ f , where f is bounded, then also f ∈ H. Of course here the Monotone
Convergence Theorem comes into play. First we notice that measurability of
the Iif follows from measurability of the Iifn for all n. Theorem 3.12 yields
that the sequences Iifn (si ) are increasing and converging to Iif (si ). Another
application of this theorem yields that µ1 (I1fn ) converges to µ1 (I1f ) and that
µ2 (I2fn ) converges to µ2 (I2f ). Since µ1 (I1fn ) = µ2 (I2fn ) for all n, we conclude
that µ1 (I1f ) = µ2 (I2f ), whence f ∈ H.
Next we check that H contains the indicators of sets in R. A quick com-
putation shows that for f = 1E1 ×E2 one has I1f = 1E1 µ2 (E2 ), which is Σ1 -
measurable, I2f = 1E2 µ1 (E1 ), and µ1 (I1f ) = µ2 (I2f ) = µ1 (E1 )µ2 (E2 ). Hence
f ∈ H. By Theorem 2.6 we conclude that H coincides with the space of all
bounded Σ-measurable functions.
It follows from Lemma 4.3 that for all E ∈ Σ, the indicator function 1E sat-
isfies the assertions of the lemma. This shows that the following definition is
meaningful.
In Theorem 4.5 below (known as Fubini’s theorem) we assert that this defines
a measure and it also tells us how to compute integrals w.r.t. this measure in
terms of iterated integrals w.r.t. µ1 and µ2 .
Theorem 4.5 The mapping µ of Definition 4.4 has the following properties.
(i) It is a measure on (S, Σ). Moreover, it is the only measure on (S, Σ) with
the property that µ(E1 × E2 ) = µ1 (E1 )µ1 (E2 ). It is therefore called the
product measure of µ1 and µ2 and often written as µ1 × µ2 .
(ii) If f ∈ Σ+ , then
(iii) If f ∈ L1 (S, Σ, µ), then Equation (4.3) is still valid and µ(f ) ∈ R.
34
that it is true for nonnegative simple functions f and Monotone Convergence
yields the assertion for f ∈ Σ+ .
(iii) Of course, here we have to use the decomposition f = f + − f − . The
tricky details are left as Exercise 4.2.
Theorem 4.5 has been proved under the standing assumption that the initial
measures µ1 and µ2 are finite. The results extend to the case where both these
measures are σ-finite. The approach is as follows. Write S1 = ∪∞ i
i=1 S1 with the
S1 ∈ Σ1 and µ1 (S1 ) < ∞. Without loss of generality, we can take the S1i disjoint.
i i
Take a similar partition (S2j ) of S2 . Then S = ∪i,j Sij , where the Sij := S1i × S2j ,
form a countable disjoint union as well. Let Σij = {E ∩ Sij : E ∈ Σ}. On each
measurable space (Sij , Σij ) the above results apply and one has e.g. identity of
the involved integrals by splitting the integration over the sets Sij and adding
up the results.
We note that if one goes beyond σ-finite measures (often a good thing to
do if one wants to have counterexamples), the assertion may no longer be true.
Let S1 = S2 = [0, 1] and Σ1 = Σ2 = B[0, 1]. Take µ1 equal to Lebesgue measure
and µ2 the counting measure, the latter is not σ-finite. It is a nice exercise to
show that ∆ := {(x, y) ∈ S : x = y} ∈ Σ. Let f = 1∆ . Obviously I1f (s1 ) ≡ 1
and I2f (s2 ) ≡ 0 and the two iterated integrals in (4.3) are 1 and 0. So, more or
less everything above concerning product measures fails in this example.
We close the section with a few remarks on products with more than two factors.
The construction of a product measure space carries over, without any problem,
to products of more than two factors, as long as there are finitely many. This
results in product spaces of the form (S1 × . . . × Sn , Σ1 × . . . × Σn , µ1 × . . . × µn )
under conditions similar to those of Theorem 4.5. The product σ-algebra is
again defined as the smallest σ-algebra that makes all projections measurable.
Existence of product measures is proved in just the same way as before, using an
induction argument. Note that there will be many possibilities to extend (4.1)
and (4.2), since there are n! different integration orders. We leave the details to
the reader.
35
Section 1.1), and, continuing our discussion of the previous section, the product
σ-algebra B(R) × B(R).
Remark 4.7 Observe that the proof of B(R) × B(R) ⊂ B(R2 ) generalizes to
the situation, where one deals with two topological spaces with the Borel sets.
For the proof of the other inclusion, we used (and needed) the fact that R is
separable under the ordinary topology. In a general setting one might have the
strict inclusion of the product σ-algebra in the Borel σ-algebra on the product
space (with the product topology).
We now know that there is no difference between B(R2 ) and B(R) × B(R). This
facilitates the use of the term 2-dimensional random vector and we have the
following easy to prove corollary.
Corollary 4.8 Let X1 , X2 : Ω → R be given. The vector mapping X =
(X1 , X2 ) : Ω → R2 is a random vector iff the Xi are random variables.
Recall that we defined in Section 2.2 the distribution, or the law, of a random
variable. Next to random variables, one can also look at random vectors. These
are mappings X : Ω → Rn of which the components, written as Xi for i =
1, . . . , n, are random variables. Also random vectors have distributions, now
probability measures on the Borel sets of Rn . So, for the Borel sets B in Rn ,
B ∈ B(Rn ), we define PX (B) = P(X ∈ B) and as in the one-dimensional case,
PX is a probability measure on B(Rn ), the joint distribution of X, also called
(joint) law of X. The marginal distribution of e.g. X 1 (for the other components
something similar applies) is given by PX1 (E) = PX (E ×Rn−1 ), with E ∈ B(R).
Random vector have distribution functions as well, also called joint dis-
tribution functions. The latter are functions F : Rn → [0, 1], defined by
F (x1 , . . . , xn ) = P({X1 ≤ x1 } ∩ · · · {Xn ≤ xn }) for all (x1 , . . . , xn ) ∈ Rn .
We also use the notation FX for such a distribution function. Note that one
has (in obvious notation) FX1 (x1 ) = P({X1 ≤ x1 } ∩ Rn−1 ), which gives the
36
(marginal ) distribution function of X1 . For the other components one similarly
has the (marginal) distribution functions FXi .
Let us specialize to n = 2. Then we have the relations PX (B1 ×R) = P(X1 ∈
B1 ) = PX1 (B1 ) for the marginal distribution, or marginal law, of X1 . The joint
distribution function F = FX : R2 → [0, 1] can be written as
Notice that, for instance, FX1 (x1 ) = limx2 →∞ F (x1 , x2 ), also denoted F (x1 , ∞).
An important case happens Rif there exists a nonnegative B(R2 )-measurable
function f such that PX (E) = E f d(λ × λ), for all E ∈ B(R2 ). In that case,
f is called the (joint) density
R of X. The obvious marginal density fX1 of X1
is defined by fX1 (x1 ) = f (x1 , x2 ) λ( dx2 ). One similarly defines the marginal
density of X2 . Check these are indeed densities in the sense of Example 3.27.
Independence (of random variables) had to do with multiplication of probabili-
ties (see Definition 2.11), so it should in a natural way be connected to product
measures.
The results of the present section (Proposition 4.6, Corollary 4.8, Proposi-
tion 4.10) have obvious extensions to higher dimensional situations. We leave
the formulation to the reader.
4.3 Exercises
4.1 Show that the embeddings es1 are Σ2 /Σ-measurable and that the es2 are
Σ1 /Σ-measurable. Also prove Proposition 4.1.
37
4.2 Prove part (iii) of Fubini’s theorem (Theorem 4.5) for f ∈ L1 (S, Σ, µ) (you
already know it for f ∈ Σ+ ). Explain why s1 7→ f (s1 , s2 ) is in L1 (S1 , Σ1 , µ1 )
for all s2 outside a set N of µ2 -measure zero and that I2f is well defined on N c .
Define
Z
fX (x) = f (x, y) dy.
R
38
(a) Show that for all a, b, c ∈ R the function (x, y) 7→ bx + cf (a, y) is Borel-
measurable on R2 .
(b) Let ani = i/n, i ∈ Z, n ∈ N. Define
X ani − x x − ani−1
f n (x, y) = 1(ani−1 ,ani ] (x)( f (ani−1 , y) + n f (ani , y)).
i
ani n
− ai−1 ai − ani−1
measurable on R2 .
4.9 Show that for t > 0
Z ∞
1
sin x e−tx dx = .
0 1 + t2
Although x 7→ sinx x doesn’t belong L1 ([0, ∞), B([0, ∞)), λ), show that one can
use Fubini’s theorem to compute the improper Riemann integral
Z ∞
sin x π
dx = .
0 x 2
where F (s−) = limu↑s F (u). Hint: integrate 1(a,b]2 and split the square
into a lower and an upper triangle.
(b) The above displayed formula is not symmetric in F and G. Show that it
can be rewritten in the symmetric form
F (b)G(b) − F (a)G(a) =
Z Z
F (s−) dG(s) + G(s−) dF (s) + [F, G](b) − [F, G](a),
(a,b] (a,b]
P
where [F, G](t) = a<s≤t ∆F (s)∆G(s) (for t ≥ a), with ∆F (s) = F (s) −
F (s−). Note that this sum involves at most countably many terms and is
finite.
4.11 Let F be the distribution function of a nonnegative random variable X
and α >R 0. Show (use Exercise 4.10 for instance, or write E X α = E f (X), with
x
f (x) = 0 αy α−1 dy) that
Z ∞
α
EX = α xα−1 (1 − F (x)) dx.
0
39
4.12 Let I be an arbitrary uncountable index set. For each Q i there is a proba-
bility space (Ωi , Fi , Pi ). Define the product σ-algebra F on i∈I Ωi as for the
case that I is Qcountable. Call a set C a countable cylinder if it can be written
as a product i∈I Ci , with Ci ∈ Fi and Ci a strict subset of Ωi for at most
countably many indices i.
(a) Show that the collection of countable cylinders is a σ-algebra, that it con-
tains the measurable rectangles and that every set in F is in fact a count-
able cylinder.
Q
(b) Let F = i∈I Ci ∈ F and let IF beQ the set of indices i for which Ci is a
strict subset of Ωi . Define P(F ) := i∈IF Pi (Ci ). Show that this defines
a probability measure on F with the property that P(πi−1 [E]) = Pi (E) for
every i ∈ I and E ∈ Fi .
4.13 Let X be a random variable, defined on some (Ω, F, P). Let Y be random
variable defined on another probability space (Ω0 , F 0 , P0 ). Consider the product
space, with the product σ-algebra and the product probability measure. Rede-
fine X and Y on the product space by X(ω, ω 0 ) = X(ω) and Y (ω, ω 0 ) = Y (ω 0 ).
Show that the redefined X and Y are independent and that their marginal
distributions are the same as they were originally.
40
5 Derivative of a measure
The topics of this chapter are absolute continuity and singularity of a pair of
measures. The main result is a kind of converse of Proposition 3.22, known as
the Radon-Nikodym theorem, Theorem 5.4.
41
5.2 The Radon-Nikodym theorem
As an appetizer for the Radon-Nikodym theorem (Theorem 5.4) we consider
a special case. Let S be a finite or countable set and Σ = 2S . Let µ be a
σ-finite measure on (S, Σ) and ν another, finite, measure such that ν µ.
ν({x})
Define h(x) = µ({x}) if µ({x}) > 0 and zero otherwise. It is easy to verify that
1
h ∈ L (S, Σ, µ) and
Observe that we have obtained an expression like (5.2), but now starting from
the assumption ν µ. The principal theorem on absolute continuity (and
singularity) is the following.
Theorem 5.4 Let µ be a σ-finite measure and let ν be a finite measure. Then
there exists a unique decomposition ν = νa + νs and a nonnegative function
h ∈ L1 (S, Σ, µ) such that νa (E) = µ(1E h) for all E ∈ Σ (so νa µ) and
νs ⊥ µ. Moreover, h is unique in the sense that any other h0 with this property
is such that µ({h 6= h0 }) = 0. The function h is called the Radon-Nikodym
derivative of νa w.r.t. µ and is often written as
dνa
h= .
dµ
Proof Omitted.
Theorem 5.4 is often used for probability measures Q and P with Q P. Write
Z = dQ
dP and note that E Z = 1. It is immediate from Proposition 3.22 that for
X ∈ L1 (Ω, F, Q) one has
EQ X = E [XZ], (5.4)
42
5.3 Decomposition of a distribution function
In elementary probability one often distinguishes between distribution functions
that are of pure jump type (for discrete random variables) and those that admit
an ordinary density. These can both be recognized as examples of the following
result.
F = Fd + Fs + Fac ,
Rx
holds true, with Fac defined by Fac (x) = −∞
f (y) dy. Such a decomposition is
unique.
One verifies that Fd has the asserted properties, the set of discontinuities of Fd is
also D and that Fc := F −Fd is continuous. Up to the normalization constant 1−
Fd (∞), Fc is a distribution function if Fd (∞) < 1 and equal to zero if Fd (∞) = 1.
Hence there exists a subprobability measure µc on B such that µc ((−∞, x]) =
Fc (x), Theorem 2.10. According to the Radon-Nikodym theorem, we can split
µc = µac + µs , where µac is absolutely continuous w.r.t. Lebesgue measure
λ. Hence,R there exists a λ-a.e. unique function f in L1+ (R, B, λ) such that
µac (B) = B f dλ.
5.4 Exercises
5.1 Let X be a symmetric Bernoulli distributed random variable (P(X = 0) =
P(X = 1) = 21 ) and Y uniformly distributed on [0, θ] (for some arbitrary θ > 0).
Assume that X and Y are independent.
(a) Show that the laws Lθ (θ > 0) of XY are not absolutely continuous w.r.t.
Lebesgue measure on R.
43
(b) Find a fixed dominating σ-finite measure µ such that Lθ µ for all θ and
determine the corresponding Radon-Nikodym derivatives.
5.2 Let X1 , X2 , . . . be an iid sequence of Bernoulli random variables, defined on
some probability space (Ω, F, P) with P(X1 = 1) = 12 . Let
∞
X
X= 2−k Xk .
k=1
where the factor 3 only appears for esthetic reasons. Show that the distri-
bution function F : [0, 1] → R of Y is constant on ( 14 , 43 ), that F (1 − x) =
1 − F (x) and that it satisfies F (x) = 2F (x/4) for x < 41 .
(c) Make a sketch of F and show that F is continuous, but not absolutely
continuous w.r.t. Lebesgue measure.
R (Hence there is no Borel measurable
function f such that F (x) = [0,x] f (u) du, x ∈ [0, 1]).
5.3 Let f ∈ L1 (S, Σ, µ) be such that µ(1E f ) = 0 for all E ∈ Σ. Show that
µ({f 6= 0}) = 0. Conclude that the function h in the Radon-Nikodym theorem
has the stated uniqueness property.
5.4 Let µ and ν be σ-finite measures and φ an arbitrary measure on a measurable
space (S, Σ). Assume that φ ν and ν µ. Show that φ µ and that
dφ dφ dν
= .
dµ dν dµ
dν
5.5 Let ν and µ be σ-finite measures on (S, Σ) with ν µ and let h = dµ , the
standing assumptions in this exercise. Show that ν({h = 0}) = 0. Show that
µ({h = 0}) = 0 iff µ ν. What is dµ
dν if this happens?
5.6 Let µ and ν be σ-finite measures and φ a finite measure on (S, Σ). Assume
that φ µ and ν µ with Radon-Nikodym derivatives h and k respectively.
Let φ = φa + φs be the Lebesgue decomposition of φ w.r.t. µ. Show that (ν-a.e.)
dφa h
= 1{k>0} .
dν k
44
continuous w.r.t. some σ-finite measure (e.g. Lebesgue measure), with corre-
sponding Radon-Nikodym derivatives (in this context often called densities) f
and g respectively, so f, g : Rn → [0, ∞). Assume that g > 0. Show that for
F = σ(X) it holds that P Q and that (look at Exercise 5.6) the Radon-
Nikodym derivative here can be taken as the likelihood ratio
dP f (X(ω))
ω 7→ (ω) = .
dQ g(X(ω))
5.8 Show that the two displayed formulas in Exercise 4.10 are valid for functions
F and G that are of bounded variation over some interval (a, b]. The integrals
should be taken in the Lebesgue-Stieltjes sense.
5.9 Let a random variable X have distribution function F with the decomposi-
tion as in Proposition 5.7.
(a) Suppose that Fs = 0. Assume that E X is well defined. How would one
compute this expectation practically? See also the introductory paragraph
of Chapter 3 for an example where this occurs, and compute for that case
E X explicitly. Verify the answer by exploiting the independence of Y and
Z.
(b) As an example of the other extreme case, suppose that F is the distribution
of Y as in Exercise 5.2, so F = Fs . What is E Y here?
45
6 Conditional expectation
6.1 A simple, finite case
Let X be a random variable with values in {x1 , . . . , xn } and Y a random variable
with values in {y1 , . . . , ym }. The conditional probability
P(X = xi , Y = yj )
P(X = xi |Y = yj ) :=
P(Y = yj )
E X̂1Ej = E X1Ej ,
the expectation of X̂ over the set Ej is the same as the expectation of X over
that set. We show this by simple computation. Note first that the values of
X1Ej are zero and xi , the latter reached on the event {X = xi } ∩ Ej that has
probability P({X = xi } ∩ Ej ). Note too that X̂1Ej = x̂j 1Ej . We then get
The just described two properties of the conditional expectation will lie at the
heart of a more general concept, conditional expectation of a random variable
given a σ-algebra, see Section 6.2.
The random variable X̂ is a.s. the only σ(Y )-measurable random variable that
46
satisfies (6.1). Indeed, suppose that Z is σ(Y )-measurable and that E Z1E =
E X1E , ∀E ∈ σ(Y ). Let E = {Z > X̂}. Then (Z − X̂)1E ≥ 0 and has
expectation zero since E ∈ σ(Y ), so we have (Z − X̂)1{Z>X̂} = 0 a.s. Likewise
we get (Z − X̂)1{Z<X̂} = 0 a.s. and it then follows that Z − X̂ = 0 a.s.
47
Proposition 6.6 The following elementary properties hold.
(i) If X ≥ 0 a.s., then X̂ ≥ 0 a.s. If X ≥ Y a.s., then X̂ ≥ Ŷ a.s.
(ii) E X̂ = E X.
(iii) If a, b ∈ R and if X̂ and Ŷ are versions of E [X|G] and E [Y |G], then
aX̂ + bŶ is a version of E [aX + bY |G].
(iv) If X is G-measurable, then X is a version of E [X|G].
Proof (i) Let G = {X̂ < 0}. Then we have from (6.2) that 0 ≥ E 1G X̂ =
E 1G X ≥ 0. Hence 1G X̂ = 0 a.s.
(ii) Take G = Ω in (6.2).
(iii) Just verify that E 1G (aX̂ + bŶ ) = E 1G (aX + bY ), for all G ∈ G.
(iv) Obvious.
We have taken some care in formulating the assertions of the previous theorem
concerning versions. Bearing this in mind and being a bit less precise at the
same time, one often phrases e.g. (iii) as E [aX + bY |G] = aE [X|G] + bE [Y |G].
Some convergence properties are listed in the following theorem.
Proof (i) From the previous theorem we know that the X̂n form a.s. an increas-
ing sequence. Let X̂ := lim sup X̂n , then X̂ is G-measurable and X̂n ↑ X̂ a.s.
We verify that this X̂ is a version of E [X|G]. But this follows by application of
the Monotone Convergence Theorem to both sides of E 1G Xn = E 1G X̂n for all
G ∈ G.
(ii) and (iii) These properties follow by mimicking the proofs of the ordinary
versions of Fatou’s Lemma and the Dominated Convergence Theorem, Exer-
cises 6.4 and 6.5.
48
(ii) If Z is G-measurable such that ZX ∈ L1 (Ω, F, P), then Z X̂ is a version
of E [ZX|G]. We write ZE [X|G] = E [ZX|G].
(iii) Let X̂ be a version of E [X|G]. If H is independent of σ(X) ∨ G, then X̂
is a version of E [X|G ∨ H]. In particular, E X is a version of E [X|H] if
σ(X) and H are independent.
(iv) Let X be a G-measurable random variable and let the random variable
Y be independent of G. Assume that h ∈ B(R2 ) is such that h(X, Y ) ∈
L1 (Ω, F, P). Put ĥ(x) = E [h(x, Y )]. Then ĥ is a Borel function and ĥ(X)
is a version of E [h(X, Y )|G].
(v) If c : R → R is a convex function and E |c(X)| < ∞, then c(X̂) ≤ C,
a.s., where C is any version of E [c(X)|G]. We often write c(E [X|G]) ≤
E [c(X)|G] (Jensen’s inequality for conditional expectations).
(vi) ||X̂||p ≤ ||X||p , for every p ≥ 1.
49
Invoking the Monotone Convergence Theorem again results in E 1G ĥ(X) =
E 1G h(X, Y ). Since all ĥn (X) are G-measurable, the same holds for ĥ(X) and
we conclude that h ∈ V . The remainder of the proof is Exercise 6.11.
(v) Since c is convex, there are sequences (an ) and (bn ) in R such that
c(x) = sup{an x + bn : n ∈ N}, ∀x ∈ R. Hence for all n we have c(X) ≥
an X + bn and by the monotonicity property of conditional expectation, we
also have C ≥ an X̂ + bn a.s. If Nn is the set of probability zero, where this
inequality is violated, then also P(N ) = 0, where N = ∪∞n=1 Nn . Outside N we
have C ≥ supn (an X̂ + bn ) = c(X̂).
(vi) The statement concerning the p-norms follows upon choosing c(x) = |x|p
in (v) and taking expectations.
We conclude this section with the following loose statement, whose message
should be clear from the above results. A conditional expectation is a random
variable that has properties similar to those of ordinary expectation.
50
probability P(∪∞ ∞ ∞
n=1 Fn |G). So, if P̂(∪n=1 Fn ) is any version of P(∪n=1 Fn |G), then
outside a set N of probability zero, we have
∞
X
P̂(∪∞
n=1 Fn ) = P̂(Fn ). (6.3)
n=1
Proof We split the proof into two parts. First we show the existence of a
conditional distribution function, after which we show that it generates a regular
conditional distribution of X given G.
We will construct a conditional distribution function on the rational num-
bers. For each q ∈ Q we select a version of P(X ≤ q|G), call it G(q). Let
Erq = {G(r) < G(q)}. Assume that r > q. Then {X ≤ r} ⊃ {X ≤ q} and hence
G(r) ≥ G(q) a.s. and so P(Erq ) = 0. Hence we obtain that P(E) = 0, where
51
E = ∪r>q Erq . Note that E is the set where the random variables G(q) fail to be
increasing in the argument q. Let Fq = {inf r>q G(r) > G(q)}. Let {q1 , q2 , . . .}
be the set of rationals strictly bigger then q and let rn = inf{q1 , . . . , qn }.
Then rn ↓ q, as n → ∞. Since the indicators 1{X≤rn } are bounded, we
have G(rn ) ↓ G(q) a.s. It follows that P(Fq ) = 0, and then P(F ) = 0, where
F = ∪q∈Q Fq . Note that F is the event on which G(·) is not right-continuous.
Let then H be the set on which limq→∞ G(q) < 1 or limq→−∞ G(q) > 0. By a
similar argument, we have P(H) = 0. On the set Ω0 := (E ∪ F ∪ H)c , the ran-
dom function G(·) has the properties of a distribution function on the rationals.
Note that Ω0 ∈ G. Let F 0 be an arbitrary distribution function and define for
x∈R
It is easy to check that F̂ (·) is a distribution function for each hidden argument
ω. Moreover, F̂ (x) is G-measurable and since inf q>x 1{X≤q} = 1{X≤x} , we
obtain that F̂ (x) is a version of P(X ≤ x|G). This finishes the proof of the
construction of a conditional distribution function of X given G.
For every ω, the distribution function F̂ (·)(ω) generates a probability mea-
sure PX (·)(ω) on (R, B(R)). Let C be the class of Borel-measurable sets B for
which PX (B) is a version of P(X ∈ B|G). It follows that all intervals (−∞, x]
belong to C. Moreover, C is a d-system. By virtue of Dynkin’s Lemma 1.13,
C = B(R).
Proof Consider the collection H of all Borel functions h for which (6.4) is a
version of E [h(X)|G]. Clearly, in view of Theorem 6.10 the indicator functions
1B for B ∈ B(R) belong to H and so do linear combinations of them. If
h ≥ 0, then we can find nonnegative simple functions hn that convergence to h
in a monotone way. Monotone convergence for conditional expectations yields
h ∈ H. If h is arbitrary, we split as usual h = h+ − h− and apply the previous
step.
52
6.4 Exercises
6.1 Let (Ω, F, P) be a probability space and let A = {A1 , . . . , An } be a partition
1
of Ω, where the Ai belong to F. Let X P∈n L (Ω, F, P) and G = σ(A). Show that
any version of E [X|G] is of the form i=1 ai 1Ai and determine the ai .
6.2 Let Y be a (real) random variable or random vector on a probability space
(Ω, F, P). Assume that Z is another random variable that is σ(Y )-measurable.
Use the standard machine to show that there exists a Borel-measurable function
h on R such that Z = h(Y ). Conclude that for integrable X it holds that
E [X|Y ] = h(Y ) for some Borel-measurable function h.
6.3 Prove Proposition 6.9.
6.4 Prove the conditional version of Fatou’s lemma, Theorem 6.7(ii).
6.5 Prove the conditional Dominated Convergence theorem, Theorem 6.7(iii).
6.6 Let (X, Y ) have a bivariate normal distribution with E X = µX , E Y = µY ,
2
Var X = σX , Var Y = σY2 and Cov (X, Y ) = c. Let
c
X̂ = µx + (Y − µY ).
σY2
Show that E (X − X̂)Y = 0. Show also (use a special property of the bivariate
normal distribution) that E (X − X̂)g(Y ) = 0 if g is a Borel-measurable function
such that E g(Y )2 < ∞. Conclude that X̂ is a version of E [X|Y ].
6.7 Let X, Y ∈ L1 (Ω, F, P) and assume that E [X|Y ] = Y and E [Y |X] = X (or
rather, versions of them are a.s. equal). Show that P(X = Y ) = 1. Hint: Start
to work on E (X − Y )1{X>z,Y ≤z} + E (X − Y )1{X≤z,Y ≤z} for arbitrary z ∈ R.
6.8 Let X and Y be random variables and assume that (X, Y ) admits a density
f w.r.t. Lebesgue measure on (R2 , B(R2 )). Let fY be the marginal density of
Y . Define fˆ(x|y) by
(
f (x,y)
if fY (y) > 0
fˆ(x|y) = fY (y)
0 else.
53
(a) Show that X and Y are conditionally independent given G iff for every
bounded measurable function f : R → R it holds that E [f (X)|σ(Y ) ∨ G] =
E [f (X)|G].
(b) Show (by examples) that in general conditional independence is not implied
by independence, nor vice versa.
(c) If X and Y are given random variables, give an example of a σ-algebra G
that makes X and Y conditionally independent.
(d) Propose a definition of conditional independence of two σ-algebras H1 and
H2 given G that is such that conditional independence of X and Y given
G can be derived from it as a special case.
6.10 (Hölder’s inequality for conditional expectations) Let X ∈ Lp (Ω, F, P) and
Y ∈ Lq (Ω, F, P), where p, q ∈ [1, ∞], p1 + 1q = 1. Then
E [XY |G]
E [1G 1H ≤ E [1G 1H ]
UV
holds for all G ∈ G and deduce (6.6).
6.11 Finish the proof of Theorem 6.8 (iv), i.e. show that the assertion also holds
without the boundedness condition on h.
54
Index
algebra, 1 null set, 4
σ-algebra, 1
product, 31 probability, 6
conditional, 49
distribution, 11 regular conditional probability, 50
joint, 35 probability kernel, 50
regular conditional distribution, 50
distribution function, 12 Radon-Nikodym derivative, 41
decomposition, 42 random variable, 11
distribution:conditional, 49 random vector, 34
55