2.
PROBABILITY THEORY FUNDAMENTALS
2.1 Introductory Concepts
Statistics is the scientific discipline concerned with developing and studying methods for
collecting, analyzing, interpreting and presenting empirical data. Statistics is a highly
interdisciplinary field; statistical analysis has applications in virtually all scientific fields
and used to answer research questions in various disciplines. In effect, it can motivate the
development of new statistical methods and theory. In developing methods and studying
the theory that underlies the methods statisticians draw on a variety of mathematical and
computational tools.
Two fundamental ideas in the field of statistics are uncertainty and variation. There are
many situations encountered in science and more generally in life in which the outcome is
uncertain. In some cases the uncertainty is because the outcome in question is not
determined yet, e.g. we may not know whether it will snow tomorrow, while in other cases
the uncertainty is because although the outcome has been determined already we are not
aware of it, e.g. it may not be known whether a military target has been hit.
Probability is the mathematical language used to discuss uncertain events and probability
plays a key role in statistics. Any measurement or data collection effort is subject to a
number of sources of variation. By this we mean that if the same measurement were
repeated, then the answer would likely change. Probability, then, is exactly this area of
study which involves predicting the relative likelihood of various outcomes. It is a
mathematical area which has developed over the past three or four centuries. One of the
early uses was to calculate the odds of various gambling games. Its usefulness for
describing errors of scientific and engineering measurements was soon realized. Reliability
engineers study probability for its many practical uses, ranging from quality control and
quality assurance to communication theory in electrical engineering. Engineering
measurements are often analyzed using statistics and a good knowledge of probability is
needed in order to understand statistics. Statisticians attempt to understand and control,
where possible, the sources of variation in any situation.
17
18
Statistics is a word with a variety of meanings. To most laypersons it most often means
simply a collection of numbers, such as the number of people living in a country or city, a
stock exchange index, or the rate of inflation. These all come under the heading of
descriptive statistics, in which items are counted or measured and the results are combined
in various ways to give useful results. That type of statistics certainly has its uses in
engineering, and we will deal with it later, but another type of statistics is applicable in
reliability to a much greater extent. That is inferential statistics or statistical inference. For
example, it is often not practical to measure all the items produced by a process. Instead,
we very frequently take a sample and measure the relevant quantity on each member of the
sample. We infer something about all the items of interest from our knowledge of the
sample. A particular characteristic of all the items we are interested in constitutes a
population. Measurements of the diameter of all possible bolts as they come off a
production process would make up a particular population. A sample is a chosen part of the
population in question, say the measured diameters of twelve bolts chosen to be
representative of all the bolts made under certain conditions. We need to know how
reliable is the information inferred about the population on the basis of our measurements
of the sample.
Probability and statistics are two related but separate academic disciplines. Statistical
analysis often uses probability distributions, and the two topics are often studied together.
However, probability theory contains much that is mostly of mathematical interest and not
directly relevant to statistics. Indeed, probability theory is the branch of mathematics
concerned with probability. Although there are several different probability interpretations,
modern probability theory treats the concept in a rigorous mathematical manner by
expressing it through a set of axioms. Typically these axioms formalize probability in
terms of a probability space, which assigns a measure taking values between 0 and 1,
termed the probability measure, to a set of outcomes called the sample space. Any
specified subset of these outcomes is called an event.
Although it is not possible to perfectly predict random events, much can be said about their
behavior. Two major results in probability theory describing such behavior are the law of
large numbers and the central limit theorem. As a mathematical foundation for statistics,
probability theory is essential to many human activities that involve quantitative analysis
of data. Methods of probability theory also apply to descriptions of complex systems given
only partial knowledge of their state.
The main concepts in probability theory include random variables and stochastic processes,
which provide mathematical abstractions of non-deterministic or uncertain processes or
measured quantities that may either be single occurrences or evolve over time in a random
fashion. Specifically, a random variable is described informally as a variable whose values
depend on outcomes of a random phenomenon. A stochastic (or random) process is a
mathematical object usually defined as a family of random variables. Historically, the
random variables are associated with or indexed by a set of numbers, commonly viewed as
19
points in time, giving the interpretation of a stochastic process representing numerical
values of some system randomly changing over time.
An example of the concepts presented so far and inspired from reliability is manufacturing
of standard weights say 1 lb. pieces made from a certain metal or alloy. Given the process,
the impurities or inconsistencies of the material used etc. one can reasonably expect that
neither all pieces manufactured will be identical nor will they be exactly 1 lb. with
unlimited precision; in effect, the actual weight of each piece manufactured is a random
variable. Assume that the manufacturing facility delivers one piece every 10 minutes, i.e.
six in an hour; then, the family of the actual weights manufactured in the interval of one
hour is a stochastic process, since each weight piece which is a random variable on its own
right, can be indexed by its order of manufacturing precedence, i.e. the first, second,
third,… and sixth in the hour considered.
Alternatively, consider another example where a person is throwing once a day for a week
a six-faced cubical die with each and every number between one and six etched on one and
only one face of the die. The result of each daily throw is recorded for the week under
consideration. Here the random variable is the outcome of each individual throw; capital
letters, e.g. X, are normally used to denote random variables. Here it might be convenient
to name the random variable corresponding to each weekday as X1, X2, X3, X4, X5, X6, X7,
for Monday, Tuesday, Wednesday, Thursday, Friday, Saturday and Sunday respectively.
The sample space Ω for all random variables X1, X2, X3, X4, X5, X6, X7 is the set
{1,2,3,4,5,6}, i.e. all possible outcomes of a single die throw. All these outcomes are
equally probable with probability 1/6 each; note that the sum of all these probability values
is equal to 1. The family {X(t)}t∊T of random variables X1, X2, X3, X4, X5, X6, X7 is a
stochastic process; here T is vector (1,2,3,4,5,6,7) with the numerals corresponding the
weekdays as before. One question that immediately arises is what the possible values for
stochastic process X are. It is easy to verify that any permutation of the six possible
outcomes in Ω and of size (length) 7 is a possible value for X. For example, 1234561 or
6543216 or 1111111 or 5555555 etc. are all possible values for stochastic process X. This
means that there are in total 67 = 279,936 possible values for stochastic process X. Here we
used the result from Eq. (1.31) i.e. the number of permutations of six possibilities (as in
sample space Ω) taken seven (one for each day in the week) at a time with repetition.
Notice that without the fundamental result from combinatorics, determining the number of
possible values for X would be practically intractable. Now, even though it may seem
counterintuitive all those possible outcomes for X are all of equal probability since each
one of the seven daily throws in the week can be any possible value in Ω. So the outcome
of the throw e.g. on Thursday, X4, is entirely independent of that that came before (e.g. on
Monday, X1) or those coming after (e.g. on Saturday, X6). So the probability for each one
of the 279,936 possible values for stochastic process X (i.e. the permutations discussed
previously) will be 6–7 = 1/ 279,936.
In conclusion, the outcome according to which all throws in the week come the same
number, e.g. 2, and some other “more random-looking” one, e.g. 2631546, are equally
20
probable! Notice how this may counter your intuitive perception where more “orderly”
outcomes like 1234561 or 6543211 may appear as less probable or more “suspicious” to a
human observer.
2.2 Historical Definitions of Probability
Probability as a specific term is a measure of the likelihood that a particular event will
occur. Just how likely is it that the outcome of a trial will meet a particular requirement? If
we are certain that an event will occur, its probability is 1 or 100%. If it certainly will not
occur, its probability is zero. The first situation corresponds to an event which occurs in
every trial, whereas the second corresponds to an event which never occurs. At this point
we might be tempted to say that probability is given by relative frequency, the fraction of
all the trials in a particular experiment that give an outcome meeting the stated
requirements. But in general that would not be right because the outcome of each trial is
determined by chance. Say we toss a fair coin, one which is just as likely to give heads as
tails. It is entirely possible that six tosses of the coin would give six heads or six tails, or
anything in between, so the relative frequency of heads would vary from zero to one. If it
is just as likely that an event will occur as that it will not occur, its true probability is 0.5 or
50%. But the experiment might well result in relative frequencies all the way from zero to
one. Then the relative frequency from a small number of trials gives a very unreliable
indication of probability.
If we were able to make an infinite number of trials, then probability would indeed be
given by the relative frequency of the event. As an illustration, suppose the weather man
on TV says that for a particular region the probability of precipitation tomorrow is 40%.
Let us consider 100 days which have the same set of relevant conditions as prevailed at the
time of the forecast. According to the prediction, precipitation the next day would occur at
any point in the region in about 40 of the 100 trials. This is what the weather man predicts,
but we all know that the weather man may not always be right!
Although we cannot make an infinite number of trials, in practice we can make a moderate
number of trials, and that will give some useful information. The relative frequency of a
particular event, or the proportion of trials giving outcomes which meet certain
requirements, will give an estimate of the probability of that event. The larger the number
of trials, the more reliable that estimate will be. This is the empirical (i.e. based on
observation or experience) or frequency approach to probability. In effect, if after n
repetitions of an experiment, where n is very large, an event is observed to occur in h of
these, then the probability of the event is h/n.
An alternative approach is possible in certain cases. This includes various gambling games,
such as tossing an unbiased coin; drawing a colored ball from a number of balls, identical
except for color, which are put into a bag and thoroughly mixed; throwing an unbiased die;
or drawing a card from a well-shuffled deck of cards. In each of these cases we can say
before the trial that a number of possible results are equally likely. In effect, if an event can
21
occur in h different ways out of a total of n possible ways, all of which are equally likely,
then the probability of the event is h/n. This is the classical or “a priori” approach. The
phrase “a priori” comes from Latin words meaning coming from what was known before.
This approach is often simple to visualize, so giving a better understanding of probability
and can be applied directly in engineering in certain cases.
Relevant to the classical approach to probability is the geometric probability. This arises in
cases where the sample space Ω can be represented as a continuous set (i.e. open set in the
sense of mathematical topology) on the plane ℝ2 or even the Euclidean space ℝ3 and the
event A as a continuous subset (or a collection of disjoint such subsets) then the probability
of A can be defined to be the ratio between the area of A and that of Ω.
Both the classical and frequency approaches have serious drawbacks, the first because the
words “equally likely” are vague and the second because the “large number” involved is
vague. Because of these difficulties, mathematicians have been led to an axiomatic
approach to probability.
Before we go on to give the modern axiomatic definition of probability, a brief mention of
another type of probability approach and definition, the subjective estimate. This is based
on a person’s experience. To illustrate this, say an engineer examines extensive
geotechnical information on a particular property. He chooses the best site to drill an oil
well, and he states that on the basis of his previous experience he estimates that the
probability the well will be successful is 30%. However, note that some other experienced
engineer using the same information might well come to a different estimate. This, then, is
a subjective estimate of probability. The company administration can use this estimate to
decide whether to take the risk and drill the well or not.
2.3 The Axiomatic Definition of Probability
As introduced in 1933 by A.N. Kolmogorov, the mathematical model of an experiment
with random outcome is a triplet [Ω, F, Pr ], also called probability space. Ω is the sample
space, F the event field, and Pr (or Prob or P or p) the probability of each element of F. Ω
is a set containing as elements all possible outcomes of the experiment considered. So as
we have seen Ω = {1, 2, 3, 4, 5, 6} if the experiment consists of a single throw of a die; in
the case of time-dependent reliability Ω = (0,∞) is the sample space for the failure-free
lifespan of an item, e.g. a bolt or a transistor; indeed, the time the item will need to break
down or fail after put in service starts right after it is installed and theoretically at least has
no upper bound. The elements of sample space Ω are called elementary events. If the
logical statement “the outcome of the experiment is a subset A of Ω’’ is identified with the
subset A itself, combinations of statements become equivalent to set operations with
subsets of Ω. This is why events are commonly enclosed in braces {}. If the sample space
Ω is finite or countable, a probability can be assigned to every subset of Ω. In this case, the
event field F contains all subsets of Ω and all combinations of them. If Ω is continuous,
restrictions are necessary. The event field F is thus a system of subsets of Ω to each of
22
which a probability has been assigned according to the situation considered. Such a field
has the following properties:
1. Ω is an element of F.
2. If A is an element of F, its complement A′ (or A∁) is also an element of F.
3. If A1, A2,… are elements of F, the countable union A1∪A2∪… is also an element of F.
It is evident that concepts from set theory, introduced earlier, will be used here.
Specifically, we introduce the certain event Ω which is the event that occurs in every trial;
for example in a die throw, the certain event is the event occurring whenever anyone of its
six faces shows. In set theory terms Ω = {1,2,3,4,5,6}. The union A + B of two events A
and B in F is the event that occurs when A or B or both occur. In a die throw, the union of
events ‘odd’ and ‘greater than 4’ is the event ‘1 or 3 or 5 or 6’. The intersection AB of two
events A and B in F is the event occurring when both A and B occur. In a die throw, the
intersection of events ‘odd’ and ‘greater than 4’ is event ‘5’. Finally, events A and B are
mutually exclusive is the occurrence of one excludes occurrence of the other; for example,
in a die throw events ‘even’ and ‘odd’ are mutually exclusive. In effect, A and B are
mutually exclusive if AB = ∅. In probability theory the empty set ∅ corresponds to the
impossible event, i.e. one that cannot occur. For example, a single die throw that comes 2.7
or 31.
So now we are ready to present the axiomatic approach to probability, which is based on
the following three axioms (postulates).
1. Probability of any event A is a positive number assigned to the event i.e.
Pr(A) ≥ 0 (2.1)
2. Probability of certain event is 1, i.e.
Pr(Ω) = 1 (2.2)
3. The probability of the union between two mutually exclusive events A, B is equal to the
sum of the two individual probabilities, i.e.
A∩B = ∅ ⇒ Pr(A∪B) = Pr(A) + Pr(B) (2.3)
The last axiom above can be generalized to unions involving any enumerable, even
infinite, collection of mutually exclusive events, A1, A2, … as follows.
∞ ∞
Pr ∪ A k = ∑ Pr ( A k ) (2.4)
k =1 k =1
For any finite collection of mutually exclusive events A1, A2, … , An the above becomes in
simplified notation as follows.
23
Pr ( A1 ∪ A 2 ∪ … ∪ A n ) = Pr ( A1 ) + Pr ( A 2 ) + + Pr ( A n ) (2.5)
If events A1, A2, … are increasing i.e. if:
∞
A n ⊆ A n +1 , n = 1, 2, 3,… ⇒ lim Pr ( A n ) = Pr ∪ A k (2.6)
n →∞
k =1
Using the same axioms, we can further derive the following fundamental properties:
Pr(∅) = 0 (2.7)
If A⊆B then Pr(A) ≤ Pr(B) (2.8)
Pr(A∁) = 1 – Pr(Α) (2.9)
0 ≤ Pr(A) ≤ 1 (2.10)
If A⊂B then Pr(A) ≤ Pr(B) and Pr(B\A) = Pr(B) – Pr(A) (2.11)
For any events A and B, Pr(A) = Pr(A∩B) + Pr(A∩B∁) (2.12)
If A1 and A2 are two arbitrary events then
Pr(A1∪A2) = Pr(A1) + Pr(A2) – Pr(A1∩A2) (2.13)
For three arbitrary, the above needs to be modified as follows.
Pr(A1∪A2∪A3) = Pr(A1) + Pr(A2) + Pr(A3) –
– Pr(A1∩A2) – Pr(A1∩A3) – Pr(A3∩A2) + Pr(A1∩A2∩A3) (2.14)
If A1, A2,…, An are n arbitrary events, the above equation for the probability of their union
A1∪A2∪…∪An can be straightforwardly generalized. The derivations of properties (2.7) to
(2.14) can be derived on the basis of Axioms
2.4 Conditional Probability
Let A and B be two events such that P(A) > 0. Denote Pr(B|A) the probability of B given
that A has occurred. Since A is known to have occurred, it becomes the new sample space
replacing the original Ω. From this we are led to the following definition.
Pr(A ∩ B)
Pr (B|A) (2.15)
Pr(A)
Alternatively, the following definition of the conditional probability Pr(B|A) of an event B
under the condition A (i. e., assuming that A has occurred) can be given.
24
Pr(A∩B) = Pr(B|A) Pr(A) = Pr(A|B) Pr(B) (2.16)
This alternative avoids the pitfall A = {} which inescapably requires Pr(A) = 0 and renders
equation (2.15) meaningless. When this alternative definition, with symmetry in A and B is
used condition Pr(A) > 0 is not required. Equations (2.16) are also known as Bayes’ law,
especially if presented in the following form.
Pr(B|A) Pr(A)
Pr (A|B) = (2.17)
Pr(B)
Probabilities Pr(B|A) can now be defined for all B∊F, where F is the event field introduced
in the previous section. Pr(B|A) is a function of B which satisfies Axioms 1 to 3 given
earlier, obviously with Pr(A|A) = 1. The information “event A has occurred” thus leads to
a new probability space [A, FA, PrA], where field FA consists of events of the form A∩B,
with B∊F and PrA(B) = Pr(B|A). Now A is the sample space and in lieu of probability the
conditional probability PrA is introduced.
Independent Events
It is then only natural to introduce independent events. Two events are independent if the
occurrence of one does not affect the probability of occurrence of the other i.e. it does not
affect the odds. In effect, two events A and B are independent if and only if Pr(A|B) =
Pr(A) or equivalently if Pr(B|A) = Pr(B). In effect, for two independent events the
following hold.
Pr(A∩B) = Pr(B|A) Pr(A) = Pr(A|B) Pr(B) = Pr(A) Pr(B) (2.18)
Pr(A∪B) = Pr(A) + Pr(B) – Pr(A) Pr(B) (2.19)
When dealing with collections of more than two events, a weak and a strong notion of
independence need to be distinguished. The events are called pairwise independent if any
two events in the collection are independent of each other as follows.
n
{Ai }i =1 ; ∀m, k : Pr ( A m ∩ A k ) = Pr ( A m ) ⋅ Pr ( A k ) ,1 ≤ m < k ≤ n (2.20)
On the other hand, saying that the events are mutually independent (not to be confused
with mutually exclusive events!) intuitively means that each event is independent of any
combination of other events in the collection.
k
k
{Ai }i =1 ; ∀k : Pr ∩ A
n
m
( )
= ∏ Pr A m ,1 < k ≤ n,1 ≤ 1 < 2 < < k ≤ n (2.21)
m =1 m =1
25
Example – Conditional Probability & Independent Events
To illustrate the abstract concepts introduced above, the following example will be
considered. A metallurgist is presented with 100 apparently identical samples made from a
new alloy. (S)he proceeds to apply a certain thermal treatment to 40 of these samples that
are randomly selected; the 50 samples that are thermally treated are branded with “T” for
identification. Then, (s)he randomly selects 60 samples to be treated with a certain
mechanical treatment; the selection for mechanical treatment is entirely random and
unrelated to whether a sample has been treated thermally too or not. The 60 mechanically
treated samples are branded with “M” for further identification.
In effect, the probability of thermal treatment is given by:
τ = Pr(T) = 50% = 0.5 = 1/2
The probability of subsequent mechanical treatment is given by:
m = Pr(M) = 60% = 0.6 = 3/5
Per the way the samples are selected for thermal or mechanical treatment, it is easy to
conclude that the two selection processes are independent of each other. In effect, the
probability of a sample being treated both thermally and mechanically is given by
P(TM) = Pr(T∩M) = τm = 0.3 = 30%.
This means that one expects that 30 samples have been branded with both T and M.
Since P(TM) is not zero, it follows that T and M are not mutually exclusive since they are
(pairwise) independent. Indeed, one can verify that P(T|M) = P(T) & P(M|T) = P(M). It
follows then, that the probability of a sample being treated thermally or mechanically (or
both) is given by
P(T+M) = Pr(T∪M) = Pr(T) + Pr(M) – P(TM) = τ + m – τm = 0.8 = 80%.
In effect, one expects that 20 samples have been branded neither with T nor with M, i.e.
Pr(T′∩M′) = Pr( (T∪M) ∁ ) = 1/5.
Then, the metallurgist proceeds to test the strength of all 100 samples, regardless of
whether or what treatments have been applied to them. Consequently, 40 of the samples
develop cracks and branded with “K” for identification. In effect, k = Pr(K) = .4 = 2/5.
Further, examination reveals that 24 samples are branded with both M and K (and possibly
with T, too); 25 samples are branded with both T and K (and possibly with M, too); 12
samples are branded with all three identifiers i.e. T, M & K. Per these results, one can
derive the following.
P(MK) = Pr(M∩K) = 0.24 ; P(TK) = Pr(T∩K) = 0.25 .
26
The following conditional probabilities can be obtained.
Pr(Τ|Κ) = P(TΚ) / k = .25 / .4 = 0.625 > 0.5 = Pr(Τ)
Pr(K|T) = Pr(K∩T) / τ = .25 / .5 = 0.5 > 0.4 = Pr(K)
This means that per the metallurgist’s experimental results, T and K are not independent.
As it turns out, the thermal treatment deteriorates a sample’s strength since Pr(T|K) and
Pr(K|T) are greater than t and k, respectively.
Furthermore:
Pr(M|K) = P(MK) / k = .24 / .4 = 0.6 = Pr(M)
Pr(K|M) = Pr(K∩M) / m = .24 / .6 = 0.4 = Pr(K)
This means that per the metallurgist’s experimental results, M and K are (pairwise)
independent.
The final question would be whether T, M & K are mutually independent. It is clear,
though, that they cannot be mutually independent since T and K are not even pairwise
independent, so condition (2.21) cannot be satisfied at least in the case of T and K.
However, per the experiments, it turns out that T and M as well as M and K are
independent. Also, notice that
P(TMK) = Pr(T∩M∩K) = 0.12 = τmk
So, mutual independence fails just because T and K are not independent. Some additional
results follow.
P(K|TM) = Pr(K|T∩M) = P(TMK) / P(TM) = 0.12 / 0.3 = 0.4 = k
In contrast:
P(K|TM′) = Pr( K|(T∩M∁) ) = Pr(K∩T∩M′) / Pr(T∩M∁)
But:
Pr(K∩T∩M′) = P(TK) – P(TMK) = 0.25 – 0.12 = 0.13
Pr(T∩M∁) = P(T) – P(TM) = 0.5 – 0.3 = 0.2
In effect:
P(K|TM′) = Pr(K∩T∩M′) / Pr(T∩M∁) = 0.13 / 0.2 = 0.65
27
The above results mean that if a sample is treated thermally only then the probability to
develop cracks is 65%! In contrast if it treated thermally and then mechanically the
probability to develop cracks is 40%, i.e. equal to k.
Finally, we calculate P(K|(M+T)′) = Pr(K|(T∪M)∁). In this end, we first need:
Pr( K∩(T∪M) ) = Pr( (K∩T) ∪ (K∩M) ) = P(TK) + P(MK) – Pr((K∩T)∩(K∩M)) =
= P(TK) + P(MK) – P(TMK) = 0.25 + 0.24 – 0.12 = 0.37
Then:
Pr(K∩(T∪M)∁) = k – Pr( K∩(T∪M) ) = 0.4 – 0.37 = 0.03
In effect:
P(K|(M+T)′) = Pr(K∩(T∪M)∁) / Pr( (T∪M) ∁ ) = 0.03 / 0.2 = 0.15
The above means that given that a sample that has been neither thermally nor mechanically
treated will develop cracks with a probability of 15%! Compare that to P(K|M) = 40% and
P(K|T) = 50%!
Before we jump to conclusions though and in order to see how intertwined these results
are, let us consider the following modified scenario that involves an alternative thermal
treatment, identified with “H” instead of “T”. Raw data are given below.
# of samples undergoing H: 50, i.e. h = Pr(H) = .5 = τ
# of samples undergoing M: 60, i.e. Pr(M) = .6 = m
# of samples developing cracks: 40, i.e. Pr(K) = .4 = k
# of samples branded with M and K: 24, i.e. P(MK) = .24 = mk
# of samples branded with H and K: 20, i.e. P(HK) = .20 = hk = τk < .25 = P(TK)
# of samples branded with H, M & K: 12, i.e. P(HMK) = hmk = τmk = P(TMK)
So, the only difference from before is that H and K are pairwise independent and in effect
H, M and K are mutually independent. Goes without saying that H & M are independent
(and in effect P(HM) = hm) and M & K are independent, since P(MK) = mk. In this case
we also obtain like before,
P(K|HM) = Pr(K|H∩M) = P(HMK) / P(HM) = 0.12 / 0.3 = 0.4 = k.
Let us now calculate:
P(K|HM′) = Pr( K|(H∩M∁) ) = Pr(K∩H∩M′) / Pr(H∩M∁)
28
In this end:
Pr(K∩H∩M′) = P(HK) – P(HMK) = 0.2 – 0.12 = 0.08
Pr(H∩M∁) = P(H) – P(HM) = 0.5 – 0.3 = 0.2
In effect:
P(K|HM′) = Pr(K∩H∩M′) / Pr(H∩M∁) = 0.08 / 0.2 = 0.4 = k
So in this case, since H, M & K are mutually independent no matter whether is treated
thermally or mechanically or both, the chance for it to develop cracks is 40%. If a sample
is not treated then the probability to develop cracks is given by the following.
P(K|(M+H)′) = Pr(K|(H∪M)∁)
In this end, we first need:
Pr( K∩(H∪M) ) = Pr( (K∩H) ∪ (K∩M) ) = P(HK) + P(MK) – Pr((K∩H)∩(K∩M)) =
= P(HK) + P(MK) – P(HMK) = 0.2 + 0.24 – 0.12 = 0.32
Then:
P(K|(M+H)′) = Pr(K|(H∪M)∁)
Pr(K∩(H∪M)∁) = k – Pr( K∩(H∪M) ) = 0.4 – 0.32 = 0.08
In effect:
P(K|(M+H)′) = Pr(K∩(H∪M)∁) / Pr( (H∪M) ∁ ) = 0.08 / 0.2 = 0.4
So, when H, M & K are mutually independent, the probability to develop cracks is k =
40% and regardless of whether treatment H is applied alone, M is applied alone, H & M
are applied both or whether no treatment is applied at all.