3/Qn), and q <4/n.
1.4 Conditional probability
Many statements about chance take the form ‘if B occurs, then the probability of A isp’,
where B and A are events (such as ‘it rains tomorrow’ and ‘the bus being on time’ respectively)
and p isa likelihood as before. To include this in our theory, we return briefly to the discussion
about proportions at the beginning of the previous section. An experiment is repeated N times,
and on each occasion we observe the occurrences or non-occurrences of two events A and
B. Now, suppose we only take an interest in those outcomes for which B occurs; all other
experiments are disregarded. In this smaller collection of trials the proportion of times that A
occurs is N(A 9 B)/N(B), since B occurs at each of them. However,
N(ANB) — N(ANB)/N
N(B) —N(B)/N
If we now think of these ratios as probabilities, we see that the probability that A occurs, given
that B occurs, should be reasonably defined as P(A 9 B)/P(B).
Probabilistic intuition leads to the same conclusion. Given that an event B occurs, itis the
case that A occurs if and only if A 0 B occurs. Thus the conditional probability of A given B1.4 Conditional probability 9
should be proportional to P(AN, B) which isto say that it equals aP(AN B) for some constant
@ he conditional probability of 2 given B must equal 1, and thus @P(2N B) = 1,
yielding « = 1/P(B).
We formalize these notions as follows.
(1) Definition, If P(B) > 0 then the conditional probability that A occurs given that B
occurs is defined to be
P(ANB)
PUL B)=
We denote this conditional probability by P(A | B), pronounced ‘the probability of A given
B’, or sometimes ‘the probability of A conditioned (or conditional) on B’.
(2) Example. Two fair dice are thrown. Given that the first shows 3, what is the probability
that the total exceeds 6? The answer is obviously }, since the second must show 4, 5, or
6. However, let us labour the point. Clearly 2 = {1,2,3, 4,5, 6}, the sett of all ordered
pairs (i, j) for i, j € {1,2,...,6), and we can take to be the set of all subsets of 2, with
P(A) = |A|/36 for any A CQ. Let B be the event that the first die shows 3, and A be the
event that the total exceeds 6. Then
B=({G,b):16}, ANB =({(,4), (3,5).G.6)},
and
PANB) _|ANB|_ 3
PB) iB. 6 U
P(A| B)=
(3) Example. A family has two children, What is the probability that both are boys, given
that at least one is a boy? The older and younger child may each be male or female, so there
are four possible combinations of sexes, which we assume to be equally likely. Hence we can
represent the sample space in the obvious way as
& = (GG, GB, BG, BB)
where P(GG) = P(BB) = P(GB) = P(BG) = }. From the definition of conditional
probability,
P(BB | one boy at least) = P(BB | GB U BG U BB)
_ P(BBN (GBUBGUBB))
~— P(GBUBGUBB)
P(BB) 1
™ B(GBUBGUBB) 3
A popular but incorrect answer to the question is 3. This is the correct answer to another
question: for a family with two children, what is the probability that both are boys given that
the younger is a boy? In this case,
P(BB | younger is a boy) = P(BB | GBU BB)
_ P(BBA GB UBB)) PBB) it
~ P(GBUBB) PGGBUBB) 2°
Remember that A x B = (a,b) :a € A, b € B) and that A x A = A?10 1.4 Events and their probabilities
‘The usual dangerous argument contains the assertion
P(BB | one child is a boy) = P(other child is a boy).
Why is this meaningless? [Hint: Consider the sample space.] e
The next lemma is crucially important in probability theory. A family By, Bz, ..., By of
events is called a partition of the set Q if
Bi B)=o when i#j, and (JB;
Each elementary event « € & belongs to exactly one set in a partition of &
(4) Lemma. For any events A and B such that 0 0 for alli. Then
P(A) =)“ P(A | B)PB))-
Proof. A = (AN B)U (AN B®). This is a disjoint union and so
P(A) = P(AN B) + P(AN BS)
= P(A | B)P(B) + P(A | B°)P(B®).
The second part is similar (see Problem (1.8.10)). .
(5) Example. We are given two urns, each containing a collection of coloured balls. Urn I
contains two white and three blue balls, whilst urn II contains three white and four blue balls.
A ball is drawn at random from urn J and put into urn Il, and then a ball is picked at random
from urn I and examined. What is the probability that its blue? We assume unless otherwise
specified that a ball picked randomly from any urn is equally likely to be any of those present.
‘The reader will be relieved to know that we no longer need to describe (2, F ,P) in detail;
we are confident that we could do so if necessary. Clearly, the colour of the final ball depends
on the colour of the ball picked from urn I. So let us ‘condition’ on this. Let A be the event
that the final ball is blue, and let B be the event that the first one picked was blue. Then, by
Lemma (4),
P(A) = P(A | B)P(B) + P(A | BY)P(BS).
We can easily find all these probabilities:
P(A | B) = P(A | urn II contains three white and five blue balls) = 3,
P(A | B°) = P(A | urn II contains four white and four blue balls) = b
P(B)=3, PCB) = 3.1.4 Conditional probability i
Hence
P(A) e
g34+3°5
Unprepared readers may have been surprised by the sudden appearance of urns in this book.
In the seventeenth and eighteenth centuries, lotteries often involved the drawing of slips from
urns, and voting was often a matter of putting slips or balls into urns. In France today, aller aux
umes is synonymous with voting. It was therefore not unnatural for the numerous Bernoulli
and others to model births, marriages, deaths, fluids, gases, and so on, using urns containing
balls of varied hue.
(6) Example. Only two factories manufacture zoggles. 20 per cent of the zoggles from factory
Land 5 per cent from factory Il are defective. Factory I produces twice as many zoggles as
factory IT each week. What is the probability that a zoggle, randomly chosen from a week’s
production, is satisfactory? Clearly this satisfaction depends on the factory of origin. Let A
be the event that the chosen zoggle is satisfactory, and let B be the event that it was made in
factory I. Arguing as before,
P(A) = P(A | B)P(B) + P(A | B°)P(BS)
4.2419 1251
3°3+20°3 =
If the chosen zoggle is defective, what is the probability that it came from factory I? In our
notation this is just P(B | A®). However,
PIBOA®) _ P(AS| BYP(B) _
PBI A = Fa = aay
This section is terminated with a cautionary example. It is not untraditional to perpetuate
errors of logic in calculating conditional probabilities. Lack of unambiguous definitions and
notation has led astray many probabilists, including even Boole, who was credited by Russell
with the discovery of pure mathematics and by others for some of the logical foundations of
computing. The well-known ‘prisoners’ paradox’ also illustrates some of the dangers here.
(7) Example. Prisoners’ paradox. Ina dark country, three prisoners have been incarcerated
without trial. Their warder tells them that the country’s dictator has decided arbitrarily to free
one of them and to shoot the other two, but he is not permitted to reveal to any prisoner the
fate of that prisoner. Prisoner A knows therefore that his chance of survival is }. In order
to gain information, he asks the warder to tell him in secret the name of some prisoner (but
not himself) who will be killed, and the warder names prisoner B. What now is prisoner A’s
assessment of the chance that he will survive? Could it be }: after all, he knows now that
the survivor will be either A or C, and he has no information about which? Could it be 4
after all, according to the rules, at least one of B and C has to be killed, and thus the extra
information cannot reasonably affect A’s earlier calculation of the odds? What does the reader
think about this? The resolution of the paradox lies in the situation when either response (B
or C) is possible.
An alternative formulation of this paradox has become known as the Monty Hall problem,
the controversy associated with which has been provoked by Marilyn vos Savant (and many
others) in Parade magazine in 1990; see Exercise (1.4.5). e12 1.4. Events and their probabilities
Exercises for Section 1.4
1, Prove that P(A | B) = P(B | A)P(A)/P(B) whenever P(A)P(B) # 0. Show that, if P(A | B) >
P(A), then P(B | A) > P(B).
2. Forevents Aj, A,..., An satisfying P(Ay 1. Az +++ Ay—1) > 0, prove that
P(A, 1A A+++ An) = P(A1)P(A2 | Ar)P(A3 | Ay Ag) <= P(An | AL Ad 1+ An).
3. Aman possesses five coins, two of which are double-headed, one is double-
normal. He shuts his eyes, picks a coin at random, and tosses it. What is the probability that the lower
face of the coin is a head?
He opens his eyes and sees that the coin is showing heads; what is the probability that the lower
face is a head?
He shuts his eyes again, and tosses the coin again. What is the probability that the lower face is
ahead?
He opens his eyes and sees that the coin is showing heads; what is the probability that the lower
face is a head?
He discards this coin, picks another at random, and tosses it. What is the probability that it shows
heads?
4. What do you think of the following ‘proof” by Lewis Carroll that an um cannot contain two balls
of the same colour? Suppose that the umn contains two balls, each of which is either black or white;
thus, in the obvious notation, (BB) = F(BW) = P(WB) = P(WW) = }. We adda black ball, so
that P(BBB) = P(BBW) = P(BWB) = F(BWW) = §. Next we pick a ball at random; the chance
that the ball is black is (using conditional probabilities) 1-}+3-4+3-4+4-4 = 3. However, if
there is probability 3 that a ball, chosen randomly from three, is black, then there must be two black
and one white, which is to say that originally there was one black and one white ballin the urn.
5. The Monty Hall problem: goats and cars. (a) Cruel fate has made you a contestant in a game
show; you have to choose one of three doors. One conceals a new car, two conceal old goats. You
choose, but your chosen door is not opened immediately. Instead, the presenter opens another door
to reveal a goat, and he offers you the opportunity to change your choice to the third door (unopened
and so far unchosen). Let p be the (conditional) probability that the third door conceals the car. The
value of p depends on the presenter’s protocol. Devise protocols to yield the values p = }, p = 3,
Show that, for a € (4, 3], there exists a protocol such that p = a. Are you well advised to change
‘your choice to the third door?
(b) Ina variant of this question, the presenter is permitted to open the first door chosen, and to reward
you with whatever lies behind. If he chooses to open another door, then this door invariably conceals
4a goat. Let p be the probability that the unopened door conceals the car, conditional on the presenter
having chosen to open a second door. Devise protocols to yield the values p = 0, p = 1, and deduce
that, for any a € [0, 1], there exists a protocol with p = a.
6. The prosecutor’s fallacy}. Let G be the event that an accused is guilty, and 7° the event that
some testimony is true. Some lawyers have argued on the assumption that P(G | T) = P(T | G).
Show that this holds if and only if (G) = P(T).
7. Urns, There are n ums of which the rth contains r ~ 1 red balls and m —r magenta balls. You
pick an urn at random and remove two balls at random without replacement. Find the probability that:
(@) the second ball is magenta;
(b) the second ball is magenta, given that the first is magenta.
+The prosecution made this error in the famous Dreyfus case of 1894.1.5 Independence 1B
1.5 Independence
In general, the occurrence of some event B changes the probability that another event A
occurs, the original probability P(A) being replaced by P(A | B). If the probability remains
unchanged, that is to seca |B) Py ihen we call A and B ‘independent’. This is
well defined only if P(B) > 0. Definition 1.4.1) of conditional probability leads us to the
following.
(1) Definition. Events A and B are called independent if
(res B) = P(A)P(B).
€ 1} is called independent if
Ie
More generally, a family {Ay :
for all finite subsets J of J.
Remark. A common student error is to make the fallacious statement that A and B are
independent if AN B = 2.
If the family (A; : i € 7} has the property that
P(A;NAj) =P(ADP(A;) — for alli # j
then it is called pairwise independent. Pairwise-independent families are not necessarily
independent, as the following example shows.
(2) Example. Suppose 2 = (abe, acb, cab, cba, bea, bac, aaa, bbb, ccc}, and each of the
nine elementary events in @ occurs with equal probability 4. Let A, be the event that the Ath
letter is a. Itis left as an exercise to show that the family (A, A2, A3} is pairwise independent
but not independent. e
(3) Example (1.4.6) revisited. The events A and B of this example are clearly dependent
because P(A | B) = $ and P(A) = 3. e
(4) Example. Choose a card at random from a pack of 52 playing cards, each being picked
with equal probability 3. We claim that the suit of the chosen card is independent of its rank.
For example,
Pking) = 4, P(king | spade) = 74.
Alternatively,
P(spade king) = 35 = j - 7 = P(spade)P(king). e
Let C be an event with P(C) > 0. To the conditional probability measure P(- | C)
corresponds the idea of conditional independence. Two events A and B are called conditionally
independent given C if
(5) P(ANB|C)=P(A| C)P(B | C);
there is a natural extension to families of events. [However, note Exercise (1.5.5).]14 1.6 Events and their probabilities
Exercises for Section 1.5
1, Let A and B be independent events; show that A°, B are independent, and deduce that A°, B°
are independent.
2. We roll adie n times. Let Aj; be the event that the ith and jth rolls produce the same number.
Show that the events {Ajj : 1 (0, 1] be given by:
QB) Pi2(Aq Xx A2) = Pi(A1)P2(A2) for Ay € Fi, Az © Fe
It can be shown that the domain of P12 can be extended from #1 x F) to the whole of
g = o(Fi x F2). The ensuing probability space (Q1 x Q2, 9, Piz) is called the product
space of (21, Fi, P1) and (Q2, F2, P2). Products of larger numbers of spaces are constructed
similarly. The measure Pj2 is sometimes called the ‘product measure” since its defining
equation (3) assumed that two experiments are independent. There are of course many other
measures that can be applied to (21 x 22, 9).
In many simple cases this technical discussion is unnecessary. Suppose that 1 and 22
are finite, and that their o-fields contain all their subsets; this is the case in Examples (1.2.4)
and (1.4.2). Then J contains all subsets of 21 x 2216 1.7 Events and their probabilities
1.7 Worked examples
Here are some more examplesto illustrate the ideas of this chapter. The reader is now equipped
to try his or her hand at a substantial number of those problems which exercised the pioneers
in probability. These frequently involved experiments having equally likely outcomes, such
as dealing whist hands, putting balls of various colours into urns and taking them out again,
throwing dice, and so on. In many such instances, the reader will be pleasantly surprised to
find that it is not necessary to write down (Q, F , P) explicitly, but only to think of Q as being
acollection {«1, a2, ..., @v} of possibilities, each of which may occur with probability 1/N.
Thus, P(A) = |A|/N for any A C Q. The basic tools used in such problems are as follows.
(a) Combinatorics: remember that the number of permutations of n objects isn! and that
the number of ways of choosing r objects from n is (").
(b) Set theory: to obtain P(A) we can compute P(A‘) = 1 — P(A) or we can partition A
by conditioning on events B;, and then use Lemma (1.4.4).
(©) Use of independence.
(1) Example. Consider a series of hands dealt at bridge. Let A be the event that in a given
deal each player has one ace. Show that the probability that A occurs at least once in seven
deals is approximately }
Solution. ‘The number of ways of dealing 52 cards into four equal hands is 52!/(13!)4. There
are 4! ways of distributing the aces so that each hand holds one, and there are 48!/(12!)* ways
of dealing the remaining cards. Thus
4148/24 1
521/034 — 10°
Now let B; be the event that A occurs for the first time on the ith deal. Clearly B; Bj = 2,
i 4 j. Thus
P(A)
1
P(A occurs in seven deals) = P(B) U---U Bz) =) P(B)) using Definition (1.3.1).
7
Since successive deals are independent, we have
P(B;) = P(A® occurs on deal 1, A° occurs on deal 2,
; A® occurs on deal i — 1, A occurs on deal ’)
= P(A‘)! P(A) using Definition (1.5.1)
=(i- 4) b
Thus
1 1
P(A occurs in seven deals) =)“ P(Bi) ~ > (
7 7
Can you see an easier way of obtaining this answer? °
(2) Example. There are two roads from A to B and two roads from B to C. Each of the four
roads has probability p of being blocked by snow, independently of all the others. What is
the probability that there is an open road from A to C?1.7 Worked examples 7
Solution.
P(open road) = P((open road from A to B) 9 (open road from B to C))
= P(open road from A to B)P(open road from B to C)
using the independence. However, p is the same for all roads; thus, using Lemma (1.3.4),
(open road) = (1 — P(no road from A to B))”
= {1 — P((first road blocked) n (second road blocked) y
= {1 — P(first road blocked)? (second road blocked) |”
using the independence. Thus
@) (open road) = (1 — p?)?.
Further suppose that there is also a direct road from A to C, which is independently blocked
with probability p. Then, by Lemma (1.4.4) and equation (3),
P(open road) = P(open road | direct road blocked) - p
+ P(open road | direct road open) - (1 — p)
= (1 =p’): p+1-(1 =p). e
(4) Example. Symmetric random walk (or ‘Gambler’s ruin’). A man is saving up to buy
anew Jaguar at a cost of N units of money. He starts with k units where 0 < k < N, and
tries to win the remainder by the following gamble with his bank manager. He tosses a fair
coin repeatedly; if it comes up heads then the manager pays him one unit, but if it comes up
tails then he pays the manager one unit. He plays this game repeatedly until one of two events
occurs: either he runs out of money and is bankrupted or he wins enough to buy the Jaguar.
What is the probability that he is ultimately bankrupted?
Solution. This is one of many problems the solution to which proceeds by the construction
of a linear difference equation subject to certain boundary conditions. Let A denote the event
that he is eventually bankrupted, and let B be the event that the first toss of the coin shows
heads. By Lemma (1.4.4),
6) Pe(A) = P(A | BYP(B) + Px(A | BYP(BS),
where F; denotes probabilities calculated relative to the starting point k. We want to find
P(A). Consider P;(A | B). If the first toss is a head then his capital increases to k + 1 units
and the game starts afresh from a different starting point. Thus P(A | B) = Px41(A) and
similarly Py(A | B®) = Pg—1(A). So, writing py = Px(A), (5) becomes:
© Pk = 5(Peri + Pkt) if O P(A| BSAC) and P(A| BNC) > P(A| BNC)
but
(12) P(A | B) < P(A| B®).
named after Simpson (1951), was remarked by Yule in 1903. ‘The nomer
instance of Stigler's law of eponymy: “No law, theorem, or discovery is named after its originator’
applies to many eponymous statements in this book, including the law itself. As remarked by A. N, Whitehead,
“Everything of importance has been said before, by somebody who did not discover it”20 1.7 Events and their probabilities
h D3 & Ds
Figure 1.1. Two unions of rectangles illustrating Simpson's paradox.
‘We may think of A as the event that treatment is successful, B as the event that drug Tis given
toa randomly chosen individual, and C as the event that this individual is female. The above
inequalities imply that B is preferred to B® when C occurs and when C® occurs, but BS is
preferred to B overall.
Setting
a=P(ANBNC), b=P(ASNBNC),
c=P(ANBNC), d=P(ASN BNC),
e=P(ANBNC), f=P(ASNBNC),
g=PIANB'NC), h=PASN BNC),
and expanding (11)-(12), we arrive at the (equivalent) inequalities
(13) ad >be, eh> fg, (at+e\(dt+h) <(b+ fi(c+s),
subject to the conditions a, b,c,...,h > Oanda+b+e+-+-+h = 1. Inequalities (13)
are equivalent to the existence of two rectangles Ry and Ro, as in Figure 1.1, satisfying
area(D}) > area(D), area(D3) > area(D4), area(R}) < area(Rp).
Many such rectangles may be found, by inspection, as for example those witha = 3,b = +,
He=hfahsaha Similar conclusions are valid for finer
€ 1) of the sample space, though the corresponding pictures are hardér to
c= $d
partitions (C;
draw.
Simpson’s paradox has arisen many times in practical situations. There are many well-
known cases, including the admission of graduate students to the University of California at
Berkeley and a clinical trial comparing treatments for kidney stones. e
(14) Example. False positives. A rare disease affects one person in 10°. A test for the
disease shows positive with probability ;%; when applied to an ill person, and with probability
7p When applied to a healthy person. What is the probability that you have the disease given
that the test shows positive?
Solution. In the obvious notation,
PC | ill) Gill)
(I)PGl) + PG | healthy)P(healthy)
_ 7-105 _ 9 2
~ 10-3 + hl — 10-5) 99+ 105-1 ~ 1011
is rather small. Indeed it is more likely that the test was incorrect. @
PAIL) = Bey
The chance of bein;1.8 Problems 21
Exercises for Section 1.7
1. There are two roads from A to B and two roads from B to C. Each of the four roads is blocked by
snow with probability p, independently of the others. Find the probability that there is an open road
from A to B given that there is no open route from A to C.
If, in addition, there is a direct road from A to C, this road being blocked with probability p
independently of the others, find the required conditional probability.
2. Calculate the probability that a hand of 13 cards dealt from a normal shuffled pack of 52 contains
exactly two kings and one ace. What is the probability that it contains exactly one ace given that it
contains exactly two kings?
3. Asymmetric random walk takes place on the integers 0, 1, 2, ..., /V with absorbing barriers at 0
and N, starting at k. Show that the probability that the walk is never absorbed is zero.
4. The so-called ‘sure thing principle’ asserts that if you prefer x to y given C, and also prefer x to
y given C®, then you surely prefer x to y. Agreed?
5. A pack contains m cards, labelled 1,2,...,m. The cards are dealt out in a random order, one
by one. Given that the label of the kth card dealt is the largest of the first k cards dealt, what is the
probability that itis also the largest in the pack?
1.8 Problems
1. A traditional fair die is thrown twice. What is the probability that:
(a) a six tums up exactly once?
(b) both numbers are odd?
(©) the sum of the scores is 4?
(@) the sum of the scores is divisible by 3?
2. A fair coin is thrown repeatedly. What is the probability that on the nth throw:
(a) ahead appears for the first time?
(b) the numbers of heads and tails to date are equal?
(©) exactly two heads have appeared altogether to date?
(@) at least two heads have appeared to date?
3. Let Fand 9 be o-fields of subsets of 2.
(a) Use elementary set operations to show that F is closed under countable intersections; that is, if
Ay, Ap,... ate in F, then so is 1; Ay
(b) Let = FG be the collection of subsets of @ lying in both Fand g. Show that J is a.o-field.
(©) Show that FUG, the collection of subsets of @ lying in either For @, is not necessarily a o-field
4, Describe the underlying probability spaces for the following experiments:
(a) a biased coin is tossed three times;
(b) two balls are drawn without replacement from an urn which originally contained two ultramarine
and two vermilion balls;
(c) a biased coin is tossed repeatedly until a head turns up.
5. Show that the probability that exactly one of the events A and B occurs is
P(A) + F(B) — 2F(A 1 B).
6. Prove that P(A U BUC) = 1 — P(A® | BSN C®)P(BE | C)B(C*).2 1.8 Events and their probabilities
7. (a) If Ais independent of itself, show that P(A) is 0 or 1
(b) IF P(A) is 0 or 1, show that A is independent of all events B.
8& Let Fbe ao-field of subsets of , and suppose P : F > {0, 1] satisfies: (i) P() = 1, and (i) P
is additive, in that P(A U B) = P(A) + P(B) whenever AN B = @. Show that P(2) = 0.
9. Suppose (2, F. P) is a probability space and B € F satisfies P(B) > 0. Let Q: F > [0, 1] be
defined by Q(A) = P(A | B). Show that (2, F, Q) is a probability space. If C € Fand Q(C) > 0,
show that Q(4 | C) = F(A | BC); discuss.
10. Let By, Bp, ... be a partition of the sample space @, each B; having positive probability, and
show that
P(A) = P(A | B)P(B))
i
LL. Prove Boole’s inequalities:
o(U ) < SO P(A), (A 4) zl - Sup.
12, Prove that
»(Aa) = OP) = OP U A) + SO PAU A; U AR)
1 7
ig icjck
moe (CIMP(A] U Ag UU An),
13. Let Aj, Az,..-, An be events, and let Ny be the event that exactly k of the A; occur. Prove the
result sometimes referred to as Waring’s theorem:
= Yanan N Ay)
iy che P(A).
17. In Problem (1.8.16) above, show that B and C are independent whenever By and Cy are inde-
pendent for all n. Deduce that if this holds and furthermore Aj, —> A, then P(A) equals either zero or
one.
18, Show that the assumption that P is countably additive is equivalent to the assumption that P is
continuous. That is to say, show that if a function P : ¥ —~ [0, 1] satisfies P(@) = 0, P(Q) = 1, and
P(AU B) = P(A) + P(B) whenever A, B € Fand AM B = @, then P is countably additive (in the
sense of satisfying Definition (1.3.1b)) if and only if P is continuous (in the sense of Lemma (1.3.5)).
19. Anne, Betty, Chloé, and Daisy were all friends at school. Subsequently each of the (3) = 6
subpairs meet up; at each of the six meetings the pair involved quarrel with some fixed probability
p. or become firm friends with probability 1 — p. Quarrels take place independently of each other.
In future, if any of the four hears a rumour, then she tells it to her firm friends only. If Anne hears a
rumour, what is the probability that:
(a) Daisy hears it?
(b) Daisy hears it if Anne and Betty have quarrelled?
(©) Daisy hears it if Betty and Chloé have quarrelled?
(@) Daisy hears it if she has quarrelled with Anne?
20. A biased coin is tossed repeatedly. Each time there is a probability p of a head turning up. Let pa
be the probability that an even number of heads has occurred after m tosses (zero is an even number).
Show that pp = 1 and that py = p(1— pp—1)+(1—p) Py ifn > 1. Solve this difference equation
21. A biased coin is tossed repeatedly. Find the probability that there is a run of r heads in a row
before there is a run of s tails, where r and s are positive integers.
22. A bowl contains twenty cherries, exactly fifteen of which have had their stones removed. A
greedy pig eats five whole cherries, picked at random, without remarking on the presence or absence
of stones. Subsequently, a cherry is picked randomily from the remaining fifteen.
(a) What is the probability that this cherry contains a stone?
(b) Given that this cherry contains a stone, what is the probability that the pig consumed at least one
stone?
23. The ‘ménages’ problem poses the following question. Some consider it to be desirable that men
and women alternate when seated at a circular table. If n couples are seated randomly according to
this rule, show that the probability that nobody sits next to his or her partner is,
ee 4 2n (2n-k
ae a( A Jew
‘You may find it useful to show first that the number of ways of selecting k non-overlapping pairs of
adjacent seats is (""-*)2n(2n — k)~!24 1.8 Events and their probabilities
24, Anum contains b blue balls and r red balls. They are removed at random and not replaced. Show
that the probability that the first red ball drawn is the (k + 1)th ball drawn equals ("0") / (ey)
Find the probability that the last ball drawn is red,
25, An um contains a azure balls and c carmine balls, where ac 4 0. Balls are removed at random
and discarded until the first time that a ball (B, say) is removed having a different colour from its
predecessor. The ball B is now replaced and the procedure restarted. This process continues until the
Jast ball is drawn from the urn, Show that this last ball is equally likely to be azure or carmine.
26. Protocols. A pack of four cards contains one spade, one club, and the two red aces. You deal
two cards faces downwards at random in front of a truthful friend. She inspects them and tells you
that one of them is the ace of hearts. What is the chance that the other card is the ace of diamonds?
Perhaps 5?
Suppose that your friend’s protocol was:
(a) with no red ace, say “no red ace”,
(b) with the ace of hearts, say “ace of hearts”,
(c) with the ace of diamonds but not the ace of hearts, say “ace of diamonds”.
Show that the probability in question is 4.
Devise a possible protocol for your friend such that the probability in question is zero.
27. Eddington’s controversy. Four witnesses, A, B, C, and D, at a trial each speak the truth with
probability + independently of each other. In their testimonies, A claimed that B denied that C declared
that D lied. What is the (conditional) probability that D told the truth? [This problem seems to have
appeared first as a parody in a university magazine of the ‘typical’ Cambridge Philosophy Tripos
question.)
28. The probabilistic method. 10 per cent of the surface of a sphere is coloured blue, the rest is red.
Show that, irrespective of the manner in which the colours are distributed, it is possible to inscribe a
cube in S with all its vertices red.
29, Repulsion. The event A is said to be repelled by the event B if P(A | B) < P(A), and to be
attracted by B if P(A | B) > P(A). Show that if B attracts A, then A attracts B, and B° repels A.
If A attracts B, and B attracts C, does A attract C?
30. Birthdays. If m students born on independent days in 1991 are attending a lecture, show that the
probability that at least two of them share a birthday is p = 1 — (365)!/{(365 — m)!365""}. Show
that p > } when m = 23
31, Lottery. You choose r of the first positive integers, and a lottery chooses a random subset L of
the same size. What is the probability that:
(a) L includes no consecutive integers?
(b) L includes exactly one pair of consecutive integers?
(c) the numbers in L are drawn in increasing order?
(d) your choice of numbers is the same as L?
(e) there are exactly k of your numbers matching members of L?
32. Bridge. During a game of bridge, you are dealt at random a hand of thirteen cards. With an
obvious notation, show that P(4S, 3H, 3D, 3C) ~ 0.026 and F(4S, 4H, 3D, 2C) ~ 0.018. However
if suits are not specified, so numbers denote the shape of your hand, show that (4, 3, 3, 3) ~ 0.11
and P(4, 4, 3,2) ~ 0.22.
33. Poker. During a game of poker, you are dealt a five-card hand at random. With the convention
that aces may count high or low, show that:
PA pair) = 0.423, P(2 pairs) ~ 0.0475, P(3 of akind) ~ 0.021,
P(straight) ~ 0.0039, P(flush) ~ 0.0020, (full house) ~: 0.0014,
PG of akind) ~ 0.00024, (straight flush) ~ 0.000015.18 Problems 25
34, Poker dice. There are five dice each displaying 9, 10, J, Q, K, A. Show that, when rolled:
FC pair) ~ 0.46, P(2 pairs) ~ 0.23, P(3 of akind) ~ 0.15,
(no 2 alike) ~ 0.093, (full house) ~ 0.039, P(4 of akind) ~ 0.019,
PG of akind) ~ 0.0008.
35. You are lost in the National Park of Bandrika+. Tourists comprise two-thirds of the visitors to
the park, and give a correct answer to requests for directions with probability }. (Answers to repeated
questions are independent, even if the question and the person are the same.) If you ask a Bandrikan
for directions, the answer is always false.
(a) You ask a passer-by whether the exit from the Park is East or West. The answer is East. What is
the probability this is correct?
(b) You ask the same person again, and receive the same reply. Show the probability that it is correct,
is }
(© You ask the same person again, and receive the same reply. What is the probability that itis
correct?
(@) You ask forthe fourth time, and receive the answer East. Show thatthe probability its corect
is 3
(©) Show that, had the fourth answer been West instead, the probability that East is nevertheless
correct is 7.
36. Mr Bayes goes to Bandrika. Tom is in the same position as you were in the previous problem,
but he has reason to believe that, with probability €, East is the correct answer. Show that:
(a) whatever answer first received, Tom continues to believe that East is correct with probability «,
(b) if the first two replies are the same (that is, either WW or EE), Tom continues to believe that East
is correct with probability €,
(©) after three like answers, Tom will calculate as follows, in the obvious notation:
9 lle
[rage Past comect | WWW) = 55
(East correct | EEE) =
Evaluate these when
37. Bonferroni’s inequality. Show that
°(U 4) = OP) — DPA Ay.
yan ah
rk
38. Kounias’s inequality. Show that
o(U 4) < min {Sea = Arn vo}.
rire
39. Then passengers for a Bell-Air flight in an airplane with n seats have been told their seat numbers.
‘They get on the plane one by one. The first person sits in the wrong seat. Subsequent passengers sit
in their assigned seats whenever they find them available, or otherwise in a randomly chosen empty
seat. What is the probability that the last passenger finds his seat free?
1A fictional country made famous in the Hitchcock film ‘The Lady Vanishes’2
Random variables and their distributions
Summary. Quantities governed by randomness correspond to functions on the
probability space called random variables. The value taken by a random vari-
able is subject to chance, and the associated likelihoods are described by a
function called the distribution function. ‘Two important classes of random
variables are discussed, namely discrete variables and continuous variables.
‘The law of averages, known also as the law of large numbers, states that the
proportion of successes in a long run of independent trials converges to the
probability of success in any one trial. This result provides a mathematical
basis for a philosophical view of probability based on repeated experimenta-
tion. Worked examples involving random variables and their distributions are
included, and the chapter terminates with sections on random vectors and on
Monte Carlo simulation.
2.1 Random variables
We shall not always be interested in an experiment itself, but rather in some consequence
of its random outcome. For example, many gamblers are more concerned with their losses
than with the games which give rise to them. Such consequences, when real valued, may
be thought of as functions which map @ into the real line IR, and these functions are called
‘random? variables’.
(1) Example. A fair coin is tossed twice: 2 = (HH, HT, TH, TT}. For @ € Q, let X(w) be
the number of heads, so that
X(HH)
1, X(T) =0.
X(HT) = X(TH)
Now suppose that a gambler wagers his fortune of £1 on the result of this experiment. He
gambles cumulatively so that his fortune is doubled each time ahead appears, and is annihilated
on the appearance of a tail. His subsequent fortune W is a random variable given by
W(HH) =
W(HT) = W(TH) = W(TT) = 0 e
{Derived from the Old French word randon meaning ‘haste’2.1 Random variables 27
Figure 2.1. The distribution function Fy of the random variable X of Examples (1) and (5)..
After the experiment is done and the outcome € Q is known, arandom variable X : 2 —
R takes some value. In general this numerical value is more likely to lie in certain subsets
of R than in certain others, depending on the probability space (, ¥ ,P) and the function X
itself, We wish to be able to describe the distribution of the likelihoods of possible values of
X. Example (1) above suggests that we might do this through the function f : R — [0,1]
defined by
f(x) = probability that X is equal to x,
but this turns out to be inappropriate in general. Rather, we use the distribution function
F :R— R defined by
F(x) = probability that X does not exceed x.
More rigorously, this is
(2) F(x) =P(A(x))
where A(x) C is given by A(x) = {o € @ : X(w) < x}. However, P is a function on the
collection ¥ of events; we cannot discuss P(A (x)) unless A(x) belongs to ¥, and so we are
led to the following definition.
(3) Definition, A random variable is a function X : 2 —» R with the property that (o € Q:
X(@)
(0, 1] given by F(x) = P(X = x).
This is the obvious abbreviation of equation (2). Events written as (w € @ : X(w) < x}
are commonly abbreviated to {e : X(w) < x} or {X < x}. We write Fy where it is necessary
to emphasize the role of X.28 2.1 Random variables and their distributions
Figure 2.2. The distribution function Fy of the random variable W of Examples (1) and (5).
(5) Example (1) revisited. ‘The distribution function Fy of X is given by
0 ifx <0,
4 if02,
and is sketched in Figure 2.1. The distribution function Fy of W is given by
0 ifx <0,
Fwa)=4 3 if04,
and is sketched in Figure 2.2. This illustrates the important point that the distribution function.
of a random variable X tells us about the values taken by X and their relative likelihoods,
rather than about the sample space and the collection of events. e
(6) Lemma. A distribution function F has the following properties:
(@) lim FG) =0, tim F(@) = 1,
(b) ifx < y then F(x) < F(y),
(©) F is right-continuous, that is, F(x +h) > F(x) ash | 0.
Proof.
(a) Let By = {w € 2: X(w) < —n) = (X <—n). The sequence By, Bo, ... isdecreasing
with the empty set as limit. Thus, by Lemma (1.3.5), P(B,) —> P(@) = 0. The other
part is similar.
(b) Let A(x) = {X < x}, AG, y) = (x < X < y}. Then A(y) = AQ) U AG, y) isa
disjoint union, and so by Definition (1.3.1),
P(A(y)) = P(A(x)) + P(AQ, y))
giving
F(y) = F(X) +PQ@ < X < y) > FQ).
(©) This
is an exercise. Use Lemma (1.3.5). .2.1 Random variables 29
Actually, this lemma characterizes distribution functions. Thats to say, F isthe distribution
function of some random variable if and only if it satisfies (6a), (6b), and (6c).
For the time being we can forget all about probability spaces and concentrate on random
variables and their distribution functions. The distribution function F of X contains a great
deal of information about X.
(7) Example. Constant variables. The simplest random variable takes a constant value on
the whole domain Q. Let c € R and define X : @ > R by
X()=c forall weQ
‘The distribution function F(x) = P(X < x) is the step function
0 x R be given by
X(H)
1,
X(T) =
Then X is the simplest non-trivial random variable, having two possible values, 0 and 1. Its
distribution function F(x) = P(X < x) is
0 x <0,
FQ@)=)1-p Osx<1,
1 xel
X is said to have the Bernoulli distribution sometimes denoted Bern(p). e
(9) Example. Indicator functions. A particular class of Bernoulli variables is very useful
in probability theory. Let A be an event and let /4 : 2 —> R be the indicator function of A;
that is,
1 ifweAd,
USO { 0 ifwe a.
Then [4 is a Bernoulli random variable taking the values 1 and 0 with probabilities P(A) and
P(A) respectively. Suppose {B; : i € 1} is a family of disjoint events with A © Uje, Bi
Then
(10) In =D lane,
an identity which is often useful. e30 2.2 Random variables and their distributions
(11) Lemma. Let F be the distribution function of X. Then
(a) P(X > x) =1- FQ),
(b) Pa x)= F(x), T2(x) = P(X < x) = F(-x),
where x is large and positive, We shall see later that the rates at which the T; decay to zero
as x —> oo have a substantial effect on the existence or non-existence of certain associated
quantities called the ‘moments’ of the distribution.
Exercises for Section 2.1
1. Let X be arandom variable on a given probability space, and let a € R. Show that
(i) aX is a random variable,
(ii) X — X =0, the random variable taking the value 0 always, and X + X = 2X.
2. Arandom variable X has distribution function F. Whats the distribution function of Y = aX+b,
where a and b are real constants?
3. A fair coin is tossed n times. Show that, under reasonable assumptions, the probability of exactly
heads is ('2)(4)". What is the corresponding quantity when heads appears with probability p on
each toss?
4, Show thatif F and G are distribution functions and < A < 1 then AF +(1—A)G isa distribution
function. Is the product FG a distribution function?
5. Let F be a distribution function and r a positive integer. Show that the following are distribution
functions:
(@) Foy’,
(b) 1- (1 F@®)}",
(©) F(x) + (1 — F(x)}log{l — F(x)},
(@) (F(x) = Ne +expll = FO).
2.2. The law of averages
We may recall the discussion in Section 1.3 of repeated experimentation. In each of N
repetitions of an experiment, we observe whether or not a given event A occurs, and we write
N(A) for the total number of occurrences of A. One possible philosophical underpinning of
probability theory requires that the proportion N(A)/N settles down as N — oo to some
limit interpretable as the ‘probability of A’. Is our theory to date consistent with such a
requirement?
With this question in mind, let us suppose that Ay, A2,... is a sequence of independent
events having equal probability P(A;) = p, where 0 < p < 1; such an assumption requires of2.2 The law of averages 31
course the existence of a corresponding probability space (@, F , P), but we do not plan to get
bogged down in such matters here. We think of 4; as being the event ‘that A occurs on the ith
experiment’, We write 5, = "7, [4,, the sum of the indicator functions of Ay, A2,-.. Ans
S; isarandom variable which counts the number of occurrences of A; for 1 0,
P(p-e1 as n> 0,
There are certain technicalities involved in the study of the convergence of random variables
(see Chapter 7), and this is the reason for the careful statement of the theorem. For the time
being, we encourage the reader to interpret the theorem as asserting simply that the proportion
n~'S,, of times that the events Ay, A2,..., An Occur converges as n — 00 to their common
probability p. We shall see later how important itis to be careful when making such statements.
Interpreted in terms of tosses of a fair coin, the theorem implies that the proportion of heads
is (with large probability) near to 4. As a caveat regarding the difficulties inherent in studying
the convergence of random variables, we remark that itis nor true that, in a ‘typical’ sequence
of tosses of a fair coin, heads outnumber tails about one-half of the time.
Proof. Suppose that we toss a coin repeatedly, and heads occurs on each toss with probability
p. The random variable S, has the same probability distribution as the number H, of heads
which occur during the first n tosses, which is to say that P(S, = k) = P(Hy = k) for all k.
It follows that, for small positive values of ,
P(1s.= pte) = > P(A, =k).
7 ken(p+e)
We have from Exercise (2.1.3) that
n
P(Hy =k) (;
Jota = py * for O» (Dota or
where m = [n(p + )], the least integer not less than n(p + €). The following argument is
standard in probability theory. Let 2 > 0 and note that e > (P+ if k > m. Writing
q = 1— p, we have that
P( Se =p +) z Se ett-nwo(t) ptr
0,
an inequality that is known as ‘Bernstein’s inequality’. It follows immediately that P(a~!S, >
p+e)— 0asn > oo. Anexactly analogous argument shows that P(n7!S, < p—€) > 0
as n —> 00, and thus the theorem is proved. .
Bernstein’s inequality (4) is rather powerful, asserting that the chance that S, exceeds its
mean by a quantity of order n tends to zero exponentially fast as n > oo; such an inequality
is known as a ‘large-deviation estimate’. We may use the inequality to prove rather more than
the conclusion of the theorem. Instead of estimating the chance that, for a specific value of
1n, Sy lies between n(p — €) and n(p + €), let us estimate the chance that this occurs for all
large n. Writing An = {p — € ei" 50 as m— oo,
giving that, as required,
1
(6) P(p-e< 18, < pte foraltn > m) +1 asm — 00.
n
Exercises for Section 2.2
1. You wish to ask each of a large number of people a question to which the answer “yes” is
‘embarrassing. The following procedure is proposed in order to determine the embarrassed fraction of
the population. As the question is asked, a coin is tossed out of sight of the questioner. If the answer
would have been “no” and the coin shows heads, then the answer “yes” is given. Otherwise people
respond truthfully, What do you think of this procedure?
2. Accoin is tossed repeatedly and heads turns up on each toss with probability p. Let Hy and Ty be
the numbers of heads and tails in n tosses. Show that, for ¢ > 0,
1
P(2p-1-«< (ty =p) <2p-1 46) = 1 asin 00.
n
3, Let {X; : r > 1} be observations which are independent and identically distributed with unknown
distribution function F. Describe and justify a method for estimating F (x).2.3. Discrete and continuous variables 33
2.3 Discrete and continuous variables
Much of the study of random variables is devoted to distribution functions, characterized by
Lemma (2.1.6). The general theory of distribution functions and their applications is quite
difficult and abstract and is best omitted at this stage. It relies on a rigorous treatment of
the construction of the Lebesgue-Stieltjes integral; this is sketched in Section 5.6. However,
things become much easier if we are prepared to restrict our attention to certain subclasses
of random variables specified by properties which make them tractable. We shall consider in
depth the collection of ‘discrete’ random variables and the collection of ‘continuous’ random
variables.
(1) Definition. The random variable X is called discrete if it takes values in some countable
subset {x1,2,...}, only, of R. ‘The discrete random variable X has (probability) mass
function f : R > [0,1] given by f(x) = P(X =x).
We shall see that the distribution function of a discrete variable has jump discontinuities
at the values x;,x2,... and is constant in between; such a distribution is called atomic. this
contrasts sharply with the other important class of distribution functions considered here.
(2) Definition. The random variable X is called continuous if its distribution function can
be expressed as
Foy= f° frau rer,
for some integrable function f : R +> [0, 00) called the (probability) density function of X.
The distribution function of a continuous random variable is certainly continuous (actually
itis ‘absolutely continuous’). For the moment we are concerned only with discrete variables
and continuous variables. There is another sort of random variable, called ‘singular’, for a
discussion of which the reader should look elsewhere. A common example of this phenomenon
is based upon the Cantor ternary set (see Grimmett and Welsh 1986, or Billingsley 1995). Other
variables are ‘mixtures’ of discrete, continuous, and singular variables. Note that the word
‘continuous’ is a misnomer when used in this regard: in describing X as continuous, we are
referring to a property of its distribution function rather than of the random variable (function)
X itself.
(3) Example. Discrete variables. The variables X and W of Example (2.1.1) take values in
the sets (0, 1,2} and (0, 4) respectively; they are both discrete. e
(4) Example. Continuous variables. A straight rod is flung down at random ontoa horizontal
plane and the angle w between the rod and true north is measured. The result is a number
in Q = [0,2r). Never mind about ¥ for the moment; we can suppose that F contains
all nice subsets of , including the collection of open subintervals such as (a,b), where
0 4m
To see this, let 0 R be given by
X(T) . X((H, x)) =
‘The random variable X takes values in (—1} U [0, 2s) (see Figure 2.3 for a sketch of its
distribution function). We say that X is continuous with the exception of a ‘point mass (or
atom) at -1. e2.4 Worked examples 35
Exercises for Section 2.3
1. Let X be a random variable with distribution function F, and let a = (dm : -00 < m < 00)
be a strictly increasing sequence of real numbers satisfying a—y, > —co and dy, —> 00 as m —> 00.
Define G(x) = F(X < dm) when d,—1 ER be continuous and strictly increasing. Show that
¥ = g(X) isa random variable.
3. Let X be a random variable with distribution function
0 ifx 1.
Let F be a distribution function which is continuous and strictly increasing. Show that ¥ = F~'(X)
is arandom variable having distribution function F. Is it necessary that F be continuous and/or strictly
increasing? .
4. Show that, if f and g are density functions, and 0 < A <1, then Af + (1 —A)g isa density. Is
the product fg a density function?
5. Which of the following are density functions? Find c and the corresponding distribution function
F for those that are.
ex? x>1,
a) f(x) =
@ fo) { 0 otherwise.
(b) f(x) =ce*(1+e*)?, xER,
2.4 Worked examples
(1) Example. Darts. A dart is flung at a circular target of radius 3. We can think of the
hitting point as the outcome of a random experiment; we shall suppose for simplicity that the
player is guaranteed to hit the target somewhere. Setting the centre of the target at the origin
of IR?, we see that the sample space of this experiment is
Qa (ix, yx? +y? <9.
Never mind about the collection ¥ of events. Let us suppose that, roughly speaking, the
probability that the dart lands in some region A is proportional to its area |A|. Thus
2) P(A) = |Al/(9z).
The scoring system is as follows. The target is partitioned by three concentric circles C1, C2,
and C3, centered at the origin with radii 1, 2, and 3. These circles divide the target into three
annuli Ay, Ap, and A3, where
A= {(x,y):k-1< yx? + y? b,
Indicate how these distribution functions behave as a —* —00, b > 00,
2.5 Random vectors
Suppose that X and Y are random variables on the probability space (2, FP). Their dis-
tribution functions, Fy and Fy, contain information about their associated probabilities. But
how may we encapsulate information about their properties relative to each other? The key
is to think of X and Y as being the components of a ‘random vector’ (X, ¥) taking values in
iR?, rather than being unrelated random variables each taking values in R.
(1) Example. Tontine is a scheme wherein subscribers to a common fund each receive an
annuity from the fund during his or her lifetime, this annuity increasing as the other subscribers
die. When all the subscribers are dead, the fund passes to the French government (this was
the case in the first such scheme designed by Lorenzo Tonti around 1653). The performance
of the fund depends on the lifetimes L, L2,..., Ln of the subscribers (as well as on their
wealths), and we may record these as a vector (L.1, L2,..., Ln) of random variables. @
(2) Example. Darts, A dart is flung at a conventional dartboard. The point of striking
determines a distance R from the centre, an angle © with the upward vertical (measured
clockwise, say), and a score S. With this experiment we may associate the random vector
(R, ®, S), and we note that S is a function of the pair (R, ). e
(3) Example. Coin tossing. Suppose that we toss a coin n times, and set X; equal to 0
or 1 depending on whether the ith toss results in a tail or a head. We think of the vector
X = (Xi, X2,..., Xp) as describing the result of this composite experiment. The total
number of heads is the sum of the entries in X. e
An individual random variable X has a distribution function Fy defined by Fy (x) =
P(X < x) for.x € R. The corresponding ‘joint’ distribution function of a random vector
(X1, Xo,..-, Xn) is the quantity P(X, < x1, Xz < x2,...,Xn [0, 1] given by Fx(x) = P(X < x)
forx eR",
As before, the expression {X_< x} is an abbreviation for the event {o € @ : X(w) < x}.
Joint distribution functions have properties similar to those of ordinary distribution functions.
For example, Lemma (2.1.6) becomes the following.
(3) Lemma. The joint distribution function Fy of the random vector (X, Y) has the follow-
ing properties:
(@) limy,y co Fx.v(x, y) = 0, limy,y-+00 Fx, (ty) = 1,
(b) if (41, 91) S (x2, Y2) then Fy.y (x1, 1) < Fx.y (x2, 92),
(©) x,y is continuous from above, in that
Fry(xtuyy+v) > Fxy@y) as u,v} 0.
We state this lemma for a random vector with only two components X and Y, but the
corresponding result for n components is valid also. The proof of the lemma is left as an
exercise. Rather more is true. It may be seen without great difficulty that
(6) yim, Fxy(, y) = Fx(x) (= P(X = x)
and similarly
a dim, Fx.v@, 9) = FrQ) (= PY sy).
This more refined version of part (a) of the lemma tells us that we may recapture the individual
distribution functions of X and Y from a knowledge of their joint distribution function. The
converse is false: it is not generally possible to calculate Fx,y from a knowledge of Fy and
Fy alone. The functions Fy and Fy are called the ‘marginal’ distribution functions of Fx,y
(8) Example. A schoolteacher asks each member of his or her class to flip a fair coin twice
and to record the outcomes. The diligent pupil D does this and records a pair (Xp, Yp) of
outcomes. The lazy pupil L flips the coin only once and writes down the result twice, recording
thusa pair (Xz, Yi) where X, = Yz. Clearly Xp, Yp, X, and Y,, are random variables with
the same distribution functions, However, the pairs (Xp, ¥p) and (Xz, Yz.) have different
joins distibution fonctions. In particular, POXn = Yo = heads) = sinee only one ofthe
four possible pairs of outeomes contains heads only, whereas P(X, = Y; = heads) = }. @
Once again there are two classes of random vectors which are particularly interesting: the
‘discrete’ and the ‘continuous’.
(9) Definition. The random variables X and ¥ on the probability space (@, ¥ , P) are called
(jointly) discrete if the vector (X, Y) takes values in some countable subset of R? only. The
jointly discrete random variables X, Y have joint (probability) mass function f : R? —
10, 1) given by f(x, y) = P(X =x, Y=y).