0% found this document useful (0 votes)
5 views68 pages

Introduction to Probability Theory

Uploaded by

Aneek Saha
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
5 views68 pages

Introduction to Probability Theory

Uploaded by

Aneek Saha
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Probability and Statistics

Dr. Sovik Roy1 *


1
Department of Mathematics, Techno Main Salt Lake Engg. College

May 14, 2025

Abstract
Disclaimer: This Statistics course is a basic level course designed for students of any branch
of Science and Engineering. Starting from designing probability space it goes upto the basic
understanding of the concept of hypothesis testing. This note is a summary of what is taught
in the class and in no way can be considered complete. The students are encouraged to study
any good and convenient book/s on probability theory.

1 Probability Model:
To describe probability model, we start with explaining a few mathematical concept. We consider
a class A containing subsets of some universal set S. Now we define a function f : A → R, where R
is the set of real numbers. This type of function is called set function, e.g. Area of a region, Mass
of a system of particles, Length of a curve etc. are all examples of set functions. A set function
f : A → R is called finitely additive if

f (A ∪ B) = f (A) + f (B), (1)

where A and B are considered to be disjoint subsets of S contained in A such that A ∪ B is also
present in A. This class A is expected to be closed under the operations of union, intersection,
complementation and difference. To make certain of these restrictions, this class A must be of
special type. In this case it is better to opt for a class which is Boolean Algebra.

A non-empty class A of subsets of a given universal set S is called Boolean Algebra if for every A
and B in A, we have

A ∪ B ∈ A, Ac ∈ A, (2)

where Ac = S − A, the complement of A relative to S. Using eq.(2)one can easily show, using De
Morgan’s laws, A ∩ B ∈ A and A − B ∈ A (check!). The smallest Boolean Algebra is the class
A0 = {ϕ, S}. The largest Boolean algebra is the power set of S, which (say) we denote by A1 . It
is pretty obvious that A0 ⊂ A ⊂ A1 .

The set functions which represent area, mass, length etc. must have to be non-negative. So
for a set function f : A → R to be finitely additive measure, f (A) ≥ 0 for each A in the class A
*
[Link]@[Link]

1
under consideration.

In the language of set function, probability is a special kind of measure. In probability theory
we call the universal set S as sample space (the set of all possible outcomes). We also consider a
Boolean Algebra B and a finitely additive measure P. This P is called probability measure defined
on B satisfying the following conditions:

ˆ P is finitely additive.

ˆ P is non-negative.

ˆ P(S) = 1.

The triplet (S, B, P) thus constructed is called Probability Space. For now we shall take S, the
sample space, as finite with respect to which P is defined. A complete knowledge on the sample
space S is important for a probability to be defined on the events. These events are subsets of the
sample space S or members of Boolean Algebra B. The members of S are known as outcomes or
samples. The conditions in bullets above are axioms of probability theory.

Problem: Toss a coin. Construct (a) the sample space S, (b) a suitable Boolean Algebra B.
By taking probability measure P and considering condition P(S) = 1, identify all possible proba-
bilities of the members of B. Build a condition under which your coin is unbiased.

If S = {a1 , a2 , · · · , an }, and if B consists of all subsets of S, the probability function P is completely


determined if we know its values on the one-element subsets i.e. P({a1 }), P({a2 }), · · · , P({an }) are
known. These probabilities are called point probabilities. When A = {a1 } ∪ {a2 } ∪ · · · ∪ {ak }, where
A ⊂ S, then
k
X
P(A) = P({ai }). (3)
i=1

Two events A and B are said to be equally likely if P(A) = P(B). The event A is called more likely
than B if P(A) > P(B), and at least as likely as B if P(A) < P(B).

Example: Prove the following: (a) P(Φ) = 0, (b) 0 ≤ P(A) ≤ 1, (c) if A ⊂ B, then P(A) < P(B),
(d) P(A ∪ B) = P(A) + P(B) − P(A ∩ B). [Use axiomatic approach].

Sol: (a) Φ represents an impossible event. Now it is pretty obvious that S ∪ Φ = S. Apply P
on both sides. Since S and Φ are mutually exclusive (i.e. chance of occurrence of one excludes
the chance of occurrence of the other), so P(S ∪ Φ) = P(S) + P(Φ) (by finite additivity axiom)
and as P(S) = 1 from axiom of probability defined on sample space (as stated above), hence the
result. [Proved].

Naturally a question comes to our mind. It has been shown that probabilities of impossible events
are always zero. But is the converse true? That is, if we have an event A whose probability comes
out to be zero, can we always say, that A is always impossible? The answer is in negative. We shall
explain this in later sections with a counter example.

Do the rest of the parts i.e. (b), (c) and (d) using axioms.

2
You have solved many problems of probability in your school days, using permutation and combi-
nation. Those problems on probability ultimately became problems of combinatorial techniques.
We summarize below the results of combinatorial techniques for your ready reference.

Problem: Let A and B denote events. Show that P(A∩B) ≤ P(A) ≤ P(A∪B) ≤ P(A)+P(B).[Use
axiomatic approach].

Combinatorial Techniques:
In many problems of statistics we must list all the alternatives that are possible in a given situation,
or at least determine how many different possibilities there are. Basic principle of counting helps
us in this regard. Below we list a few theorems on counting rule.

Theorem: If an operation consists of two steps, of which the first can be done in m ways and for
each of these, the second can be done in n ways, then the whole operation can be done in mn ways.

Theorem: If an operation consists of k steps, of which the first can be done in n1 ways, for
each of these the second step can be done in n2 ways, for each of the first two the third can be
done in n3 ways, and so forth, then the whole operation can be done in n1 n2 · · · nk .

Theorem: The number of permutations of n distinct objects is n!.

n!
Theorem: The number of permutations of n distinct objects taken r at a time is nPr = (n−r)! , for
r = 0, 1, 2, · · · , n.

Theorem: The number of permutations of n distinct objects arranged in a circle is (n − 1)!.

Theorem: The number of permutations of n objects of which n1 are of one kind, n2 are of a
second kind, · · · nk are of a k th kind, and n1 + n2 + · · · + nk = n is n1 !n2n!!···nk ! .

n!
Theorem: The number of combinations of n distinct objects taken r at a time is nCr = (n−r)!r!
for r = 0, 1, 2, · · · , n.

Theorem: The number of ways in which a set of n distinct objects can be partitioned into k
subsets with n1 objects in the first subset, n2 objects in the second subset, · · · , and nk objects in
the k th subset is nCn1 ,n2 ,··· ,nk = n1 !n2n!!···nk ! .

Now for further discussion, we shall not always consider the triplet (S, B, P) as discussed in class
1, rather its relevance can be understood from the context itself. So if we talk about tossing a coin
whose sample space is S = {h, t} and from which a Boolean algebra can be constructed by taking
subsets of S as {h}, or {t}, we will simply consider H and T and denote them as events of the
sample space. Also upto certain amount of time we will consider probability as measure defined on
finite sample spaces.

Example: Let S be a finite sample space consisting of n elements. Suppose we assign equal
probabilities to each of the point in S. Let A be a subset of S consisting of k− elements. Prove
that P (A) = nk .

3
Sol: Let O1 , O2 , · · · , On represent individual outcomes in S, each with probability n1 . If A is
the union of k these mutually exclusive outcomes and it does not matter which ones, then

P(A) = P(O1 ∪ O2 ∪ · · · ∪ Ok )
= P(O1 ) + P(O2 ) + · · · + P(Ok )
1 1 1
= + + ··· +
n n n
k
= . (4)
n
Eq.(4) is sometimes treated as classical definition of probability and which has its own limita-
tions.

Conditional Probability:
Let A and B are any two events in a sample space S and P(A) ̸= 0, then P(A ∩ B) = P(A).P(B|A).
While calculating conditional probability there is a change in the size of the sample space. The
sample space actually gets reduced. In the language of probability space we can elaborate our
definition of conditional probability as follows:

Let (S, B, P) be a probability space and let B be an event such that P(B) > 0. The conditional
probability that an event A will occur, given that B has occurred, is denoted by the sample P(A|B)
and is given by the equation
P(A ∩ B)
P(A|B) = . (5)
P(B)

The conditional probability is not defined if P(A) = 0. Hence we need to define another Boolean
algebra B ′ of subsets of B so that the new probability space is (B, B ′ , P).

Two events A and B are independent when P(A|B) = P(A) or P(B|A) = P(B), such that
for independent events we have

P(A ∩ B) = P(A)P(B). (6)

The definition (6) can be generalized to any arbitrary k number of events.

Theorem: If the events B1 , B2 , · · · , Bk constitute a partition of the sample space S and P(Bi ) ̸= 0
for i = 1, 2, · · · , k, then for any event A in S
k
X
P(A) = P(Bi )P(A|Bi ). (7)
i=1

In a situation when cause is there we summarize the effects of that cause which are probable, in a
sample space and we calculate the chance of occurrences of these effects. When occurrence of one
effect is dependent upon the occurrences of some previous events, we deal the system with eq.(5).
But sometimes there may be situation where effects are already counted and we wanted to find out
the probable causes. That is dealt by Bayes’ theorem or theorem of inverse probability.

4
Bayes’ theorem: If B1 , B2 , · · · , Bk constitute a partition of the sample space S and P(Bi ) ̸= 0
for i = 1, 2, · · · , k, then for any event A in S such that P(A) =
̸ 0
P(Br )P(A|Br )
P(Br |A) = Pk , (8)
i=1 P(Br )P(A|Br )

for r = 1, 2, · · · , k.

Although Bayes’ theorem follows from the postulates of probability and the definition of condi-
tional probability, it has been the subject of extensive controversy. There can be no question about
the validity of Bayes’ theorem, but considerable arguments have been raised about the assignment
of the prior probabilities as it goes from effects to cause.

Example (Psychology department): In a T-maze, a rat is given food if it turns left and
an electric shock if it turns right. On the first trial there is a 50–50 chance that a rat will turn
either way; then, if it receives food on the first trial, the probability is 0.68 that it will turn left on
the next trial, and if it receives a shock on the first trial, the probability is 0.84 that it will turn
left on the next trial. What is the probability that a rat will turn left on the second trial?

Sol: Consider the following figure. We define the events as follows. Let L denote the event that the

Figure 1: T-Maze.

rat will turn left and L̄ be the event that the rat will turn right. It is given that p(L) = p(L̄) = 12 .
Agian let F denote the event that on the first trial the rat receives food and F̄ is the event that
the rat receives electric shock. So there is also 50-50 chance that the rat either receives food or
electric shock. As per the given information, we have P(L|F ) = 0.68 and P(L|F̄ ) = 0.84. Thus
the probability that a rat will turn left on the second trial is P(L) = P[(F ∩ L) ∪ (F̄ ∩ L)] =
P(F ∩ L) + P(F̄ ∩ L) = P(F )P(L|F ) + P(F̄ )P(L|F̄ ) = 12 × 0.68 + 12 × 0.84 = 0.76.

Example (Management problem): It is known from experience that in a certain industry


60 percent of all labor management disputes are over wages, 15 percent are over working condi-
tions, and 25 percent are over fringe issues. Also, 45 percent of the disputes over wages are resolved
without strikes, 70 percent of the disputes over working conditions are resolved without strikes, and
40 percent of the disputes over fringe issues are resolved without strikes. What is the probability
that a labor management dispute in this industry will be resolved without a strike? What is the
probability that if a labor management dispute in this industry is resolved without a strike, it was
over wages?

5
Example (Health Dept. Problem): In a certain community, 8 percent of all adults over
50 have diabetes. If a health service in this community correctly diagnoses 95 percent of all per-
sons with diabetes as having the disease and incorrectly diagnoses 2 percent of all persons without
diabetes as having the disease, find the probabilities that (a) the community health service will
diagnose an adult over 50 as having diabetes; (b) a person over 50 diagnosed by the health service
as having diabetes actually has the disease.

Example (Management problem): An explosion at a construction site could have occurred


as the result of static electricity, malfunctioning of equipment, carelessness, or sabotage. Inter-
views with construction engineers analyzing the risks involved led to the estimates that such an
explosion would occur with probability 0.25 as a result of static electricity, 0.20 as a result of mal-
functioning of equipment, 0.40 as a result of carelessness, and 0.75 as a result of sabotage. It is also
felt that the prior probabilities of the four causes of the explosion are 0.20, 0.40, 0.25, and 0.15.
Based on all this information, what is (a) the most likely cause of the explosion; (b) the least likely
cause of the explosion? [Apply Bayes’ theorem].

Example (DRDO1 problem): Radar detection: If an aircraft is present in a certain area, a


radar correctly registers its presence with probability 0.99. If it is not present, the radar falsely
registers an aircraft presence with probability 0.10. We assume that an aircraft is present with
probability 0.05. What is the probability of false alarm (a false indication of aircraft presence), and
the probability of missed detection (nothing registers, even though an aircraft is present)?

Sol: Let us define the events as follows. Let

A : Presence of aircraft.
B : The radar registers an aircraft presence.
Ā : Absence of aircraft.
B̄ : The radar does not register an aircraft presence.

Now,

P(A) = 0.05, P(Ā) = 0.95, P(B|A) = 0.99 and P(B|Ā) = 0.1.

Then P(false alarm) = P(false indication of aircraft presence) = P(Ā ∩ B) = P(Ā)P(B|Ā) =


0.95×0.1 = 0.095. Now we calculate P(mixed detection) = P(nothing registers, even though aircraft
is present) = P(A ∩ B̄) = P(A)P(B̄|A) = 0.05 × 0.01 = 0.0005 (since P(B̄|A) = 1 − P(B|A) = 0.01).

In solving above problems, two things have to be kept in mind, one is we need to identify the
events and define them clearly so that we can apply axiomatic approach to probability and the
second important point is to identify whether a particular problem is of conditional probability or
the problem is that which requires the concept of inverse probability (i.e. Bayes’ theorem).

Exercise: A letter known to have come either from T AT AN AGAR or from CALCU T T A. On
the envelop just two consecutive letters T A are visible. What is the probability that the letters
4
come from CALCU T T A? [Apply Baye’s theorem : 11 ]
1
DRDO stands for Defence Research Development Organization.

6
2 Probability distribution:
Definition: If S is a sample space with a probability measure and X is a real-valued function
defined over the elements of S, then X is called a Random Variable.

Toss a pair of coins. The sample space is {(H, H), (H, T ), (T, H), (T, T )}. If the tossing is ran-
dom, each outcomes are assigned equal probabilities. Suppose we are interested in finding the
number of heads present in each outcome. Let us denote this count of head by X so that
X(H, H) = 2, X(H, T ) = 1, X(T, H) = 1, X(T, T ) = 0, then X is a random variable.
The probabilities are now shown as
Such a function P (X = x) defined in the above table is called probability distribution.

X =x P(X = x) = px
X =2 P(X = 2) = 14
1
X =1 P(X = 1) = P[(H, T ) ∪ (T, H)] = 2
X =0 P(X = 0) = 14

Definition: If X is a discrete random variable, the function given by px = P (X = x) for each x


within the range of X is called probability distribution.

Theorem: A function can serve as the probability distribution of a discrete random variable
X ifPand only if its values, px , satisfy the conditions (a) px ≥ 0 for each value within its domain;
(b) x px = 1, where the summation extends over all the values within its domain.2

Definition: A function with values f (x), defined over the set of all real numbers, is called a
probability density function of the continuous random variable X if and only if P (a ≤ X ≤
Rb
b) = a f (x)dx for any real constants a and b with a ≤ b.

Problem: The key concept of classical information theory is the Shannon entropy which (with
respect to the random variable X) quantifies how much information we gain on average when we
learn the value of X or amount of uncertainty about X before we learn its value. Thus the entropy
of the two outcome random variable is defined to be

H(p) = −p log2 (p) − (1 − p) log2 (1 − p). (9)

We now toss a coin and label the events H and T by 0 and 1 respectively with p = 0 and p + q = 1.
How much information do we gain? (Assume 0 log 0 = 0).

Sol: From the information given in the above question it is clear that we are dealing with a
biased coin. Now if we assume p = 0 (the probability of getting head) then probability of getting
tail is obviously 1. Using eq.(9) we get H(p) = 0. So amount of uncertainty before we learn the
outcome is zero. This means the biased coin gives less information to us.3

Theorem: A function can serve as a probability density of a continuous random variable X if


2
Discrete random variables are those where range is finite or countably infinite while continuous random variable
are those when the range is uncountable.
3
Take p = 12 and give your opinion!

7
R∞
its values, f (x), satisfy the conditions, (a) f (x) ≥ 0 for ∞ < x < ∞; (b) −∞ f (x)dx = 1.

Definition: If X is a discrete random variable, the function given by


X
F (x) = P(X ≤ x) = f (t), f or − ∞ < x < ∞, (10)
t≤x

where f (t) is the value of the probability distribution of X at t, is called the distribution function
or the cumulative distribution, of X.

Theorem: The values F (x) of the distribution function of a discrete random variable X satisfy
the conditions

ˆ F (−∞) = 0, and F (∞) = 1;

ˆ If a < b, then F (a) ≤ F (b) for any real numbers a and b.

It is to be noted that here above we treat −∞ and ∞ as the abstraction of large number and
included in the real number system, thus making the system as extended real number system.

Definition: If X is a continuous random variable, the function given by


Z x
F (x) = P(X ≤ x) = f (t)dt, f or∞ < x∞, (11)
−∞

where f (t) is the value of the probability distribution of X at t, is called the distribution function
or the cumulative distribution, of X.

Theorem: If f (x) and F (x) are the values of the probability density and the distribution function
of X at x, then

P(a ≤ X ≤ b) = F (b) − F (a), (12)

for any real constants a and b with a ≤ b, and

dF (x)
f (x) = , (13)
dx
where the derivative exists.

Definition: If X is a random variable (discrete/continuous), by expectation of a random variable


X we mean the average value of the random variable and is defined as
(P
xpx if x is discrete.
E(X) = R x (14)
x xf (x) if x is continuous.

Moreover we know that in a distribution variance measures scatter of the data from mean. Variance
of a random variable X is thus defined as (and the expression is true for both discrete and continuous
case)

V(X) = E[X − E(X)]2 = E(X 2 ) − [E(X)]2 . (15)

8
Probability distribution are of two types, viz. (a) discrete and (b) continuous. Two special discrete
type probability distributions are Binomial distribution and Poisson distribution. The most signif-
icant continuous probability distribution is Normal distribution.

Example: A random variable X has the following probability density function, f (x) = k(x −
1)(2 − x); 1 < x < 2. Find the value of k and the distribution function of X. Also calculate E(X)
and V(X).

Sol: For completeness, the given p.d.f is defined as


(
k(x-1)(2-x)if 1 < x < 2.
f(x) =
0 if either x ≤ 1 or x ≥ 2.
R
Now using total probability law of the p.d.f we have (−∞, ∞), f (x)dx = 1. We split the range
R2
and using definition, 1 k(x − 1)(2 − x)dx = 1, solving which we get k = 6. Hence the p.d.f is
(
6(x-1)(2-x)if 1 < x < 2.
f(x) =
0 if either x ≤ 1 or x ≥ 2.

Now using eqs.(14) and (15) we get the expectation and variance of the random variable X. There-
fore we get E(X) = 23 and V(X) = 20 1
. We now plot this probability density function f (x) below.
Look at the bulging point at the center. It occurs when x = 32 , so that we can immediately say

Figure 2: Plot of pdf k(x-1)(2-x)

that the data points are scattered around this value (which is nothing but the expectation value).
Also if we take, for instance, average of the sum of the squares of the deviations taken from mean
1
which is of course nothing but variance, the value would come out as 20 . Now using eq.(10), we
can easily calculate F (x). The distributive function is thus given by

0,
 if x ≤ 1.
3 2
F(x) = −2x + 9x − 12x + 5, if 1 < x < 2.

1, if x ≥ 1.

9
Figure 3: Plot of F(x)

The plot of F (x) is shown below and it is way different from the plot of f (x). We now discuss a
few properties of expectation of a random variable and variance of the random variable.

Theorem: If X is a discrete random variable and f (x) is the value of its probability distribu-
tion at x, the expected value of g(X) is given by
X
E[g(X)] = g(x)f (x). (16)

Correspondingly, if X is a continuous random variable and f (x) is the value of its probability
density at x, the expected value of g(X) is given by
Z ∞
E[g(X)] = g(x)f (x). (17)
−∞

Some properties of Expectation and Variance:


Theorem: If a and b are constants, then
E(aX + b) = aE(X) + b,
E(aX) = aE(X)
E(b) = b. (18)

Theorem: If c1 , c2 , · · · , cn are constants, then


n
hX n
X h i
E ci gi (X)] = ci E gi (X) . (19)
i=1 i=1
Pn
Theorem: If X1 , X2 , · · · , Xn are random variables and Y = i=1 ai Xi , where a1 , a2 , · · · , an are
constants, then
n
X
E(Y ) = ai E(Xi )
i=1
Xn XX
V(Y ) = a2i V(Xi ) + 2 ai aj Cov(Xi , Xj ), (20)
i=1 i<j

10
where the double summation extends over all values of i and j, from 1 to n, for which i < j.
Pn
Theorem: If the random variables X1 , X2 , · · · , Xn are independent and Y = i=1 ai Xi , then
n
X
V(Y ) = a2i V(Xi ). (21)
i=1

Pn Pn
Theorem: If X1 , X2 , · · · , Xn are random variables and Y1 = i=1 ai Xi and Y2 = i=1 bi Xi ,
where a1 , a2 , · · · , an , b1 , b2 , · · · , bn are constants, then
n
X XX
Cov(Y1 , Y2 ) = ai bi V(Xi ) + (ai bj + aj bi )Cov(Xi , Xj ). (22)
i=1 i<j

Pn
If however,Pthe random variables X1 , X2 , · · · , Xn are independent, while as before Y1 = i=1 ai Xi
and Y2 = ni=1 bi Xi , then
n
X
Cov(Y1 , Y2 ) = ai bi V(Xi ). (23)
i=1

Example: If the random variables X, Y, Z have the means µX = 3, µY = 5, µZ = 2, the


2 = 8, σ 2 = 12, σ 2 = 18, and Cov(X, Y ) = 1, Cov(X, Z) = −3, Cov(Y, Z) = 2, find
variances σX Y Z
the covariance of U = X + 4Y + 2Z and V = 3X − Y − Z.

Sol: By eq.(22) we have


Cov(U, V ) = Cov(X + 4Y + 2Z, 3X − Y − Z)
= (1 × 3)V(X) + (4 × (−1))V(Y ) + (2 × (−1))V(Z) + (1 × (−1) + 4 × 3)Cov(X, Y ) + (1 × (−1) +
2 × 3)Cov(X, Z) + (4 × (−1) + 2 × (−1))Cov(Y, Z)
= (3 × 8) − (4 × 12) − (2 × 18) + 11 × 1 + 5 × (−3) − (6 × 2) = −76.

We will elaborately discuss the meaning of the notation Cov, when we will discuss bivariate dis-
tribution.

We are now in a position to answer an old question which we raised earlier. That is if any event
A has probability zero, can we say that A is always impossible? The answer is no. We can get a
counter system in the continuous random variable case. Suppose that we are concerned with the
possibility that an accident will occur on a free-way that is 200 kilometres long and that we are
interested in the probability that it will occur at a given location, or perhaps on a given stretch
of the road. The sample space of this experiment consists of a continuum of points, those on the
interval from 0 to 200, and we shall assume, for the sake of argument, that the probability that an
d
accident will occur on any interval of length d is 200 , with d measured in kilometres. By axioms,
d 200
probabilities 200 are all non-negative and P(S) = 200 = 1. So far this assignment of probabilities
applies only to intervals on the line segment from 0 to 200, but if we use additivity axiom, we can
also obtain probabilities for the union of any finite or countably infinite sequence of nonoverlapping
intervals. For instance, the probability that an accident will occur on either of two nonoverlap-
ping intervals of lengths d1 and d2 is d1200
+d2
and the probability that it will occur on any one of

11
a countably infinite sequence of non-overlapping intervals of lengths d1 , d2 , · · · , is d1 +d 2 +···
200 . Thus,
the probability of the accident occurring on a very short interval, say, an interval of 1 centimetre, is
only 0.00000005, which is very small. As the length of the interval approaches zero, the probability
that an accident will occur on it also approaches zero; indeed, in the continuous case we always
assign zero probability to individual points. This does not mean that the corresponding events
cannot occur; after all, when an accident occurs on the 200-kilometre stretch of road, it has to
occur at some point even though each point has zero probability.

Example: Find the mean and variance of a random variable having probability density func-
tion f (x) = 1 − |1 − x| where 0 < x < 2.

Sol: We cannot integrate modulus function directly. So we need to write the function given
explicitly. We shall use the definition of |x|. By this definition we have

1 − (1 − x), if
 1 − x > 0.
f(x) = 1, if 1 − x = 1.

1 − (−(1 − x)), if 1 − x < 0.

i.e.

x, if
 1 > x.
f(x) = 1, if 1 − x = 1.

2 − x, if 1 < x.

Now using eqs.(14) and (15), we can easily calculate the expectation and variance of the random
variable X. While doing so, we split the region of integration into four parts (−∞, 0), (0, 1), (1, 2), (2, ∞).
You may think of, where does this x = 1 go to? We can include x = 1 in either (0, 1) or in (1, 2).
How can we do so? This is because of the following theorem.

Theorem: If X is continuous random variable and a and b are real constants with a ≤ b, then

P (a ≤ X ≤ b) = P (a ≤ X < b) = P (a < X ≤ b) = P (a < X < b). (24)

Example: The chance that a bit transmitted through a digital transmission channel is received in
error is 0.5. Also, assume that the transmission trials are independent. Let X be the number of
bits in error in the next four bits transmitted.

Question:
(a) Tabulate all the outcomes and write the respective probabilities. (Show how you calculate one
such probability for a specific outcome in details). Write about the axiom/theorem that has been
used for this calculation.
(b) Construct an appropriate probability distribution with respect to the random variable defined
in the problem. State the nature of the random variable.
(c) Calculate P (X = 1). (Show details of the calculation). Write about the axiom/theorem that
has been used for this calculation.
(d) Take along x− axis the values of the random variable X and along y− axis the values of the
respective probabilities. Plot line diagram.4 .
4
Line diagram represents the vertical lines corresponding to the probabilities which are weighted on the values of
the random variable. For this the probabilities are known as probability mass functions.

12
(e) Can you formulate any specific formula to calculate, in general, P (X = x).
(f) Which factors play important role in your constructed formula? These factors are called the
parameters of the constructed distribution.
(g) Find the distribution function F (x).
(h) Find expectation and variance of the random variable X.
(i) In the problem statement replace 0.5 by 0.1 i.e. consider the system to be biased. Then repeat
(a) - (h).

A few probability distributions and the nature of their probability function along with their means
and variances are tabulated below. The study of probability distribution is important in the sense

Distribution Probability function Mean Variance


Binomial nCxpx q n−x , p + q = 1, x = 0, 1, 2, · · · , n np npq
exp(−λ)λx
Poisson x! , x = 0, 1, 2, · · · λ λ
Normal √1 exp({ x−µ }2 ), ∞ < x < ∞, σ > 0 µ σ
σ 2π σ (
1
if α < x < β. α+β (β−α)2
Uniform u(x; α, β) = β−α 2 12
0 if elsewhere.
(
1 x
exp(− θ ) if x > 0.
Exponential g(x; θ) = θ θ θ2
 0 if elsewhere.
γ−2
 γ 1 x 2 exp(− x ) if x > 0.
γ 2
Chi-square f (x) = 2 2 Γ( 2 ) γ 2γ
0 if elsewhere.

that in the statistics we are mostly interested in how population behave. The behavior of the
population is studied by observing what probability model, the population under consideration is
following. So the above table is very significant in the theory of probability and statistics. Apart
from the distribution tabulated above, there are other probability distributions which have been
intentionally left out, while those are in no way less important.

Example (Food Technology): Shelf life is defined as the length of time a product may be
stored without becoming unsuitable for use or consumption. Shelf life depends on the degradation
mechanism of the specific product.

Q: The shelf life (in hours) of a certain perishable packaged food is a random variable whose
probability density function is given by
(
20000
(x+100)3
, if x > 0.
f(x) =
0, if elsewhere.
Find the probabilities that one of these packages will have a shelf life of (a) at least 200 hours, (b)
at most 100 hours and (c) anywhere from 80 to 120 hours.

Example (Random Walk Problem): Starting from origin, unit steps are taken to the right
with probability p and to the left with probability q = 1 − p. Assuming independent movements,
find the mean and variance of the distance moved from origin after n steps.

Sol: First we define an appropriate random variable and denote it by Xi . This is defined as

13
(
+1, if ith step is taken towards right with probability p
Xi =
−1, if ith step is taken towards lef t with probability q.
We will first calculate the expectation of the random Xi . Using eq.(14), E(Xi ) = (+1) × p + (−1) ×
(1 − p) = p − q. Also E(Xi2 ) = (+1)2 × p + (−1)2 × (1 − p) = p + q = 1, so that using eq.(15)
V(Xi )P= (p + q)2 + (p − q)2 = 4pq. If now Sn denotes the total distance of after n steps, then
Sn = ni=1 Xi .
Pn Pn Pn Pn
Now
Pn E(Sn ) = E(
Pn i=1 X i ) = i=1 E(X i ) = i=1 (p + q) = n(p − q). Also, V(Sn ) = V( i=1 Xi ) =
i=1 V(Xi ) = i=1 (p + q) = 4npq (using properties of expectation and variance).

Exercise: Let the variate X have the distribution P(X = 0) = P(X = 2) = p; P(X = 1) = 1 − 2 p,
where 0 ≤ p ≤ 21 . For what value of p is the V(X) maximum? [Ans: V(X) = 2p, variance is
maximum when p = 21 ].

Definition: The rth moment about the origin of a random variable X, denoted by µ′r , is
the expected value of X r ; symbolically,
X
µ′r = E(X r ) = xr P (X = x), (25)
x

for r = 0, 1, 2, · · · when X is discrete, and


Z
µ′r = E(X r ) = xr f (x), (26)
x
when X is continuous. Here P (X = x) is the probability mass function and f (x) is the probability
density function.

It is pretty obvious from above that µ′0 = 1 and µ′1 = E(X) i.e. the mean of X.

Definition: The rth moment about the mean of a random variable X, denoted by µr , is
the expected value of (X − µ)r ; symbolically,
X
µ′r = E[(X − µ)r )] = (x − µ)r P (X = x), (27)
x

for r = 0, 1, 2, · · · when X is discrete, and


X
µ′r = E[(X − µ)r )] = (x − µ)r f (x), , (28)
x

when X is continuous. Here P (X = x) is the probability mass function and f (x) is the probability
density function.

It is pretty obvious from above that µ0 = 1 and µ1 = E(X − µ) = 0 i.e. the mean of X.
Also µ2 = E[(X − µ)2 ] is the variance of X, often denoted by σ 2 .

Definition: The rth and sth product moment about the origin of the random variable
X and Y , denoted by µ′r,s , is the expected value of X r Y s ; symbolically,
XX
µ′r,s = E[X r Y s ] = xr y s P (X = x, Y = y), (29)
x y

14
for r = 0, 1, 2, · · · and s = 0, 1, 2, · · · when X and Y are discrete, and
Z ∞Z ∞
′ r s
µr,s = E[X Y ] = xr y s f (x, y)dxdy, (30)
−∞ −∞

where X and Y are continuous.

Definition: The rth and sth product moment about the mean of the random variable X and
Y , denoted by µr,s , is the expected value of (X − µX )r (Y − µY )s ; symbolically,
XX
µ′r,s = E[(X − µX )r (Y − µY )s ] = (x − µX )r (y − µY )s P (X = x, Y = y), (31)
x y

for r = 0, 1, 2, · · · and s = 0, 1, 2, · · · when X and Y are discrete, and


Z ∞Z ∞
′ r s
µr,s = E[(X − µX ) (Y − µY ) ] = (x − µX )r (y − µY )s f (x, y), (32)
−∞ −∞

where X and Y are continuous.

If we substitute r = 1 and s = 1 in eq.(32) we get µ1,1 which is called the Covariance of X


and Y and is denoted by Cov(X, Y ) or σXY .

Theorem: If X and Y are independent, then E(XY ) = E(X)E(Y ) such that Cov(X, Y ) = 0.

Moment Generating Function: The moment generating function of a random variable X,


when it exists, is given by,
X
MX (t) = E(etX ) = etx P(X = x), X is discrete
Zx
MX (t) = E(etX ) = etx f (x), X is continuous. (33)
x

If we consider the discrete case, we get


X (tx)2 (tx)r
MX (t) = E(etX ) = {1 + tx +
+ ··· + + · · · }P(X = x)
x
2! r!
X X t2 X 2 tr X r
= P(X = x) + t xP(X = x) + x P(X = x) + · · · x P(X = x) + · · ·
x x
2! x r! x
t2 ′ tr
= 1 + tµ + µ2 + · · · + µ′r + · · · . (34)
2! r!
Similarly we can show the same for continuous case replacing summation by integration.

Theorem: If MX (t) is the moment generating function of the random variable X (discrete or
continuous), we have

dr
MX (t) = µ′r . (35)
dtr

15
Chebychev’s Theorem: If µ and σ are the mean and standard deviation of a random variable
X, then for any positive constant k, the probability is at least 1 − k12 that X will take on a value
within k standard deviations of the mean; symbolically,
1
P(|X − µ| < k σ) ≥ 1 − . (36)
k2
Obviously, the probability given by Chebychev’s theorem is only a lower bound; whether the prob-
ability that a given random variable will take on a value within k standard deviations of the mean
is actually greater than 1 − k12 and, if so, by how much we cannot say, but Chebychev’s theorem
assures us that this probability cannot be less than 1 − k12 . Only, when the distrbution of a random
variable is known, can we calculate the exact probability. We will now prove the above theorem.

Proof: We know,
R∞
σ 2 = E[(X − µ)2 ] = −∞ (x − µ)2 f (x)dx.

Then, dividing the integral into three parts, as shown in figure, We divide the region of integration

Figure 4: Area under the curve.

as per the figure 4, is split. Therefore we get,


R µ−kσ R µ+kσ R∞
σ 2 = E[(X − µ)2 ] = −∞ (x − µ)2 f (x)dx + µ−kσ (x − µ)2 f (x)dx + µ+kσ (x − µ)2 f (x)dx

Since the integrand (x − µ)2 f (x) is non-negative, we can immediately say


R µ−kσ R∞
σ 2 ≥ −∞ (x − µ)2 f (x)dx + µ+kσ (x − µ)2 f (x)dx

Now for the region (−∞, µ − kσ) we have x ≤ µ − kσ and in the region (µ + kσ, ∞) we have
x ≥ µ + kσ, hence it follows that
R µ−kσ R∞
σ 2 ≥ −∞ k 2 σ 2 f (x)dx + µ+kσ k 2 σ 2 f (x)dx

This implies that


1
R µ−kσ R∞ 1
k2
≥ −∞ σ 2 f (x)dx + µ+kσ σ 2 f (x)dx, provided σ 2 ̸= 0. ⇒ k2
≥ P(|X − µ| ≥ kσ)
1
Hence it follows consequently P(|X − µ| < kσ) ≥ 1 − k2
. [Proved]

Note: Another form of Chebychev’s theorem is as follows


1
P(|X − µ| ≥ k σ) < . (37)
k2

16
Chebychev’s theorem is a weak law of probability. It gives either a lower bound or an upper bound
of the probability of the system but not the actual probability. On the basis of the problem state-
ment, we choose between eqs. (36) and (37). We will try to understand it by the following examples.

Problem: A symmetric dice is thrown 600 times. Find the lower bound for the probability of
getting 80 to 120 sixes. [Done in class]

Hints: Any symmetric dice throwing system corresponds to binomial probability law. Here the
random variable X is the count of sixes and it follows binomial distribution. Hence E(X) =
np = 600 × 61 = 100 and V(X) = npq = 600 × 16 × 56 = 500 6 . Using eq.(36) we therefore get
q
500 1
P(|X − 100| < k 6 ) ≥ 1 − k2 . Comparing with the requirement given in the problem we get
19
P(80 < X < 120) = 24 .

Exercise: A fair dice is thrown 720 times. obtain a lower bound for the probability of get-
ting 91 to 149 sixes.

A little thought on this issue would have shown us immediately, that had we solved the proba-
bility P(80 < X < 120) directly using binomial probability law, it would have been a humongous
calculations of the type 600Cx .

Problem: Two unbiased dice are thrown. If X is the sum of the numbers showing up, prove
that P(|X − 7| ≥ 3) < 35
54 . Compare this with actual probability. [Done in class].

Hints: Use eq.(37). Write the probability distribution and calculate the actual probability di-
rectly from the chart.

Exercise: Show that for any positive constant c the probability that the sample mean X̄ cor-
responding to a random sample of size n will take on a value between µ − c and µ + c is at least
σ2
1 − nc2.

Problem: Use Chebychev’s inequality to determine how many times a fair coin must
be tossed in order that the probability will be at least 0.90 that the ratio of the
observed number of heads to the number of tosses will lie between 0.4 and 0.6.

Sol: Consider the Chebychev’s inequality given in eq.(2). Substitute kσ = ϵ (say). Also fair
coin tossing problem follows binomial distribution and let X be the observed number of heads (or
successes), so that E(X) = n × 12 , where we assume n to be the desired number of tossing and
probability of success is half. Also V(X) = n × 21 × 21 = n4 . Considering all the above, eq.(36) takes
the following form (remember however, as per the problem, we are concerned here about the ratio
of observed number of successes to number of trials), we get
V( X )
 
P |Xn − E( X
n )| < ϵ ≥ 1 − ϵ2n
 
⇒ P |X n − 1
2 | < ϵ 1
≥ 1 − 4nϵ 2

It is given in the problem


X
P(0.4 ≤ n ≤ 0.6) ≥ 0.90.

17
1
Comparing, we get ϵ = 0.1 and 1 − 4nϵ2
= 0.90, solving which we get n = 250. [Try to solve!]

Example: Does there exist a random variable X for which P[µ − 2σ < X < µ + 2σ] = 0.6?

Sol: By Chebychev’s theorem (i.e. eq.(36)), we can write the given expression as
1
P[|X − µ| < 2σ] ≥ 1 − 22
= 0.75.

It immediately invalidates the statement given in the question.

Example: What would be the probability of the following: P(|X − 1| < 4)?

Sol: This is a very wonderful problem to study. Look that we can, by using eq.(36), can rewrite
the probability in three different ways such as

P[|X − 1| < 2 × 2] ≥ 1 − 212 = 0.75.


P[|X − µ| < 1 × 4] ≥ 1 − 112 = 0.
P[|X − µ| < 4 × 1] ≥ 1 − 412 = 0.9375.

The question is which one is appropriate! The second probability value i.e. zero does not tell us
much as it is obvious from axiom of probability that the probability always has to be non-negative.
We have to choose between 0.75 or 0.9375. But if we select 0.9375 then as we are not sure about
the system under consideration, then we cannot say that the probability would always be greater
than 0.9375 whereas if we select the choice of 0.75 then it can also be greater than 0.9375. So 0.75
is our best choice.

Chebychev’s inequality has great utility because it can be applied to any probability distribution
in which the mean and variance are defined.

3 Special Probability Distributions (Discrete and Continuous):


To start with, we must recall the well-known experiment of tossing unbiased coin or dice or likewise.
What we do there is to assume the probabilities of the outcomes as all equal and if there are n
outcomes in the sample space and each outcome is equally likely then we assign probability n1 for
each one of them. The associated random variable is obviously discrete. Thus we can define as
follows.

Discrete Uniform distribution: Let X be the discrete random variable. Then its probabil-
ity mass fucntion is given by
1
n
n, x = 1, 2, · · · , n
P(X = x) = (38)
0, elsewhere

Continuous Uniform distribution: Let X be the continuous random variable. Then its proba-
bility density function is given by
1
b−a , a<x<b
n
u(x; a, b) = (39)
0, elsewhere

18
Theorem: The mean and variance of the continuous uniform random variate X is
a+b
E(X) =
2
(b − a)2
V(X) = . (40)
12
Continuous uniform variate is also known as Rectangular variate. The following figures show the
probability density function (f (x)) and distribution function (F (x)) of the continuous uniform vari-
ate X.

Example: If X is uniformly distributed with mean 4 and variance 43 , find P(X < 0).

Sol: Using eqs.(40) and we see that E(X) = 4 and V(X) = 34 , Solving for a and b and using
the condition a < b we get, a == 1 and b = 3. Then we get from eq.(39)
u(x; −1, 3) = 14 .
R0 1
It is obvious that P(X < 0) = −1 4 dx = 14 .

Binomial Distribution: Let X be the discrete random variable. Then its probability mass
fucntion is given by
n nCx px q n−x , x = 0, 1, · · · , n
P(X = x) = (41)
0, elsewhere, p + q = 1.

In binomial distribution, it is to be remembered, that the trials are independent. The trials are
known as Bernoulli trials.

Theorem: The mean and variance of the continuous uniform random variate X is

E(X) = np
V(X) = npq. (42)

Problem: Show using eq.(14) for discrete case, show that if X is a binomial variate, E(X) = np
and V(X) = npq. [Done is class].

Exercise:: Let X be a random variable having binomial distribution with parameter n and p.
p(1−p)
Consider another random variable Y = X
n . Show that E(Y ) = p and V(Y ) = p .
n
Problem: If X follows Binomial probability distribution with parameters n, p, show that E Xn −
o2 n o
p = pq X n−X
n and Cov n , n = − pq
n.

19
Problem: A multiple choice test consists of eight questions and three answers to each ques-
tion (of which only one is correct). If a student answers each question by rolling a balanced dice
and checking the first answer if he gets a 1 or 2, the second answer if he gets a 3 or 4 and the
third answer if he gets a 5 or 6. What is the probability that he will get exactly four correct answers?

Sol: If you see at the structure of the problem, you can identify easily, the values of the two
parameters. 8 questions need to be attempted. So n = 8. Probability of success is calculated as
follows. You see there are 3 alternatives to each question. Since you select options on the basis of
outcomes of an unbiased dice throwing, as per question, hence, probability of choosing option 1 is
P({1} ∪ {2}) = P(1) + P(2) = 61 + 61 = 31 and likewise. Therefore, probability of success is p = 13 .
The system follows binomial distribution. The probability law is given by eq.(41) which implies
x 8−x
8Cx 13 23 , x = 0, 1, 2, · · · , 8. For the required answer put x = 4. The answer is 0.1707. (Here the
random variable X is the count of correct answers made by you).

Can we solve this problem by any other discrete distribution formula? Let us explore this!

Multinomial Distribution: In some cases, we are interested in multiple outcomes for a cer-
tain trial. As for example, we are often asked to rate certain service with good, average, poor etc
or in election season we take opinion polls of the voters whether the will vote for the candidate or
against him or will remain neutral. In these types of situations we see that the random variables
of the system follows multinomial distribution. Suppose there are n independent trials having k
mutually
Pk exclusive outcomes whose respective probabilities are p1 , p2 , · · · , pk with the condition
i=1 pi . The outcomes are referred as the first kind appearing x1Ptimes, second kind appearing x2
times and in this manner k th kind appearing xk times such that ki=1 xi = n.

The random variables X1 , X2 , · · · , Xn have a multinomial distribution and they are referred
to as multinomial random variable if and only if their joint probability distribution is given by

f (x1 , x2 , · · · , xk ; n, p1 , p2 , · · · , pk ) = nCx1 ,x2 ,··· ,xk px1 1 px2 2 · · · pxk k


n!
nCx1 ,x2 ,··· ,xk = . (43)
x1 !x2 !, · · · , xk !

for x = 0, 1, cdots, n for each i, where ki=1 = n and ki=1 pi = 1.


P P

Multinomial distribution is the immediate generalization of binomial distribution, where the pa-
rameters are n, p1 , p2 , · · · , pk and it is easily seen that the probabilities equal corresponding terms
of the multinomial expansion of (p1 + p2 + · · · + pk )n .

Example (Genetics): According to the Mendelian theory of heredity, if plants with round yellow
seeds are crossbred with plants with wrinkled green seeds, the probabilities of getting a plant that
produces round yellow seeds, wrinkled yellow seeds, round green seeds, or wrinkled green seeds are,
9 3 3
respectively, 16 , 16 , 16 , and 16 . What is the probability that among nine plants thus obtained
there will be four that produce round yellow seeds, two that produce wrinkled yellow seeds, three
that produce round green seeds, and none that produce wrinkled green seeds?

Sol: Let Xi , i = 1, 2, 3, 4 denotes random variables denoting counts of round yellow seeds, wrinkled
yellow seeds, round green seeds, and wrinkled green seeds. The system follow multinomial distri-

20
9 3 3 1
bution and we have x1 = 4, x2 = 2, x3 = 3, x4 = 0 while p1 = 16 , p2 = 16 , p3 = 16 , p4 = 16 . Using
eq.(43) we get
9 3 3 1 9! 9 4 3 2 3 3 1 0
f (4, 2, 3, 0; 9, 16 , 16 , 16 , 16 ) = 4!3!2!0! ( 16 ) ( 16 ) ( 16 ) ( 16 ) = 0.02906.
Problem: A certain city has 3 newspapers, A, B, and C. Newspaper A has 50% of the readers in
that city. Newspaper B, has 30% of the readers, and newspaper C has the remaining 20%. Find
the probability that, among 8 randomly-chosen readers in that city, 5 will read newspaper A, 2 will
read newspaper B, and 1 will read newspaper C. (For the purpose of this example, assume that no
one reads more than one newspaper.)

Exercise: Three card players play a series of [Link] probability that player A will win
any game is 20%, the probability that player B will win is 30%, and the probability player C will
win is 50%. If they play 6 games, what is the probability that player A will win 1 game, player B
will win 2 games, and player C will win 3? [Sol: using multinomial distribution result is 0.135].

Poisson Distribution: Let X be the discrete random variable. Then its probability mass fucntion
is given by
e−λ λx
x! , x = 0, 1, · · · ,
n
P(X = x) = (44)
0, elsewhere.

Theorem: The mean and variance of the continuous uniform random variate X is5

E(X) = λ
V(X) = λ. (45)

Problem: Show using eq.(14) for discrete case, show that if X is a binomial variate, E(X) = λ
and V(X) = λ.

Hints: Use eq.(14) [discretePcase]. Replace P(X = x) by eq.(44) and solve for E(X). For cal-
culation of E(X 2 ), which is x x2 P(X = x), write x2 as x(x − 1) + x and solve!

Problem: A car hire firm has two cars which it hires out day by day. The number of demands
for a car on each day is distributed as Poisson variate with mean 1.5. Calculate the proportion of
days on which (i) neither car is used, (ii) some demand is refused. [Ans: 0.2231, 0.19126]

Hints: Let the random variable be X denoting the number of demands. The demand, however,
is Poisson variate as we are unsure as when it is actually come to that car hire firm as there may
be several car hire firms running in the city. The mean demand is 1.5 and by eq.(45) we have
λ = 1.5. (i) No demand stands for X = 0, so you put this value in eq.(44). (ii) The demand will
be refused when there is no other car is left which the car hire firm can allot as both the two cars
are out for service. That is going to happen when X > 2. But hypothetically there are infinitely
many numbers greater than 2. Hence to calculate it, we would apply P(A) + P(Ā) = 1. Thus
P(X > 2) = 1 − P(X ≤ 2). Then use eq.(44).

We know probability mass function of binomial variate is b(x; n, p) = nCx px (1 − p)n−x . We choose
p = nλ , and we can re-write
5
Poisson distribution is one of the discrete probability distribution where mean and variance are same.

21
 x  n−x
b(x; n, p) = nCx ( nλ )x (1 − nλ )n−x = n(n−1)(n−2)···(n−x+1)
x!
λ
n 1 − λ
n
n−x
1 2
)···(1− x−1

1(1− n )(1− n )
= x!
n
λx 1 − nλ .
n
We rearrange (1 − nλ )n−x = [(1 − nλ )− λ ]−λ (1 − nλ )−x . Hence we get
1
1(1− n 2
)(1− n )···(1− x−1 ) x n
b(x; n, p) = x!
n
λ [(1 − nλ )− λ ]−λ (1 − nλ )−x .
Finally we let n → ∞, while x and λ remain fixed, we find that
1
1(1− n 2
)(1− n )···(1− x−1 ) n
x!
n
→ 1, (1 − nλ )−x → 1 and (1 − nλ )− λ → e. Therefore we obtain eq.(44).

Poisson distribution is the limiting case of binomial distribution. When (a) the number of tri-
als (n) is infinitely large, (in hypothetical sense n → ∞, which I like to see as if we are unsure
of the fact that how many times must we repeat the experiment to get the first success!), (b) the
probability of success p as compared to number of trials is indefinitely small (i.e. p → 0) while
(c) n × p is a finite number, binomial variate X tends to follow Poisson distribution. Although
the Poisson distribution has been derived as a limiting form of the binomial distribution, it has
many applications that have no direct connection with binomial distributions. For example, the
Poisson distribution can serve as a model for the number of successes that occur during a given
time interval or in a specified region when (1) the numbers of successes occurring in nonoverlapping
time intervals or regions are independent, (2) the probability of a single success occurring in a very
short time interval or in a very small region is proportional to the length of the time interval or the
size of the region, and (3) the probability of more than one success occurring in such a short time
interval or falling in such a small region is negligible. Hence, a Poisson distribution might describe
the number of telephone calls per hour received by an office, the number of typing errors per page,
or the number of bacteria in a given culture when the average number of successes, λ, for the given
time interval or specified region is known.

Discussion: Let us once again consider the problem on multiple choice test which we discussed in
the light of binomial probability law. If we assume X, the random variable follows Poisson distribu-
tion, then using eq.(45) we have λ = np = 8× 13 = 2.67. (This is binomial approximation to Poisson
e−2.67 (2.67)4
variate). Using eq.(44) we calculate the probability of four correct answers as 4! = 0.14.

Exercise: The average number of trucks arriving on any one day at a truck depot in a cer-
tain city is known to be 12. What is the probability that on a given day fewer than nine trucks
will arrive at this depot? [Using Poisson distribution P (x < 9) = 0.1550]

Note: The above result differs from that obtained by binomial. The question is which one of
the answers is the correct interpretation of the problem. Now the answers obtained by binomial
law is 0.17 and that obtained by Poisson is 0.14. But if you observe carefully, you see, that probabil-
ity of success is 13 is not too small as compared to n = 8, that is the primary requirement of Poisson
distribution. So the problem, although can be solved by Possion law, it would not be a good choice.

We now calculate moment generating function of Poisson variate. If X is Poisson variate, then its
(λet )
probability density function is given byPeq.(44). Using eqs.(33-34) we get MX (t) = e−λ ∞
P
x=0 x! ,
zx
where x = 0, 1, 2, · · · . It is known that ∞
x=0 x! can be recognized as the Mclaurin’s series of e z , this
P∞ (λet )x t t
implies that x=0 x! is the Mclaurin’s series of eλe . Using this fact we get MX (t) = eλ(e −1) .

22
Now using eq.(35) and substituting t = 0 we get

MX (0) = µ′1 = λ = E(X),
′′
MX (0) = µ′2 = λ + λ2 = E(X 2 ),
V (X) = λ. (46)

Normal Distribution: A random variable X has a normal distribution and it is referred to


as a normal random variable if and only if its probability density is given by
1 1 x−µ 2
f (x; µ, σ) = √ e− 2 ( σ ) , −∞ < x < ∞, σ > 0. (47)
σ 2π
The normal curve is shown below.

Figure 5: Normal Curve

As compared to this curve, Fig.4 shown, while discussing chebychev’s theorem is the skewed curve
and deviates from normality. This is almost appearing in every real life scenario. Normal curve is
an ideal situation, a benchmark with respect to which the skewed data are studied.
R∞ √
Recall a gamma function is defined as Γ(n) = 0 e−x xn−1 dx and Γ(1) = 1 while Γ( 12 ) = π 6 .
By suitable change of variable it can be shown that
Z ∞
1 2
Γ(n) = 2 1−n
z 2n−1 e− 2 z dz, n > 0. (48)
0

For n = 12 , we get from eq.(48)


1 √ Z ∞ 1 2
Γ = 2 e− 2 z dz. (49)
2 0

Now we see that, using eq.(49)


√  
− 21 ( x−µ 1 2
R∞ )2
R∞ −2z
√1 √1 √2 Γ 1
−∞ σ 2π e dx = −∞ e dz = = 1.
σ
2π 2π 2

This means the total area under the curve is 1 [See Fig.9].

6
BSM101 course: Calculus part.

23
X−µ
If X is a normal variate and then if we standardize it as Z = σ , its probability density function
is given by
1 1 2
f (z; 0, 1) = √ e− 2 z , −∞ < z < ∞. (50)

Z is called standard normal variate.

Problem: Suppose X is a normal variate with mean µ and variance σ 2 , show that Z , the
standard normal variate has mean 0 and variance 1. [Hints: Show E(Z) = 0 and V(Z) = 1].

A distribution that has the mean 0 and the variance 1 is said to be in standard form and
when we perform the above change of variable, we are said to be standardizing the distribution
of X.

Properties of normal curve:


ˆ Normal curve (see fig.9) is symmetric about X = µ (i.e. mean). This means both sides of the
line X = µ have equal probability (or area).

ˆ Normal curve is Bell shaped.

ˆ X− axis is the asymptote to the curve, this means the curve represented by functions eq.(47)
(or eq.(50)) touches the horizontal axis either at −∞ or at ∞.

ˆ In the representation of the normal curve, mean, median and mode coincide at a single point.

ˆ Mode of the normal variate X is √1


σ 2π
and that of Z is √1 .

ˆ When n → ∞ while p is not very small, binomial variate tends to follow normal distribution.
Indeed, this is known as binomial approximation to normal distribution. In that case
we write Z = X−np

npq .

Often to solve problems, we standardize the normal variate and use area properties to calculate
the required probability. This is done as standardized Z is symmetric about 0. Remember that
using any numerical integration method, one can easily calculate the area under the normal curve
within a specified limit. Such areas (or probabilities) have been numerically computed as well and
the chart can be found in standard probability book7 .

Problem: If two normal universes A and B have the same total frequency but the standard
deviation of universe A is k times that of the universe B, show that maximum frequency of uni-
verse A is k1 times that of universe B.[Try!]

Example: Suppose that the amount of cosmic radiation to which a person is exposed when flying
by jet across India is a random variable having a normal distribution with mean of 4.35 mrem8 and
a standard deviation of 0.59 mrem. What is the probability that a person will be exposed to more
than 5.20 mrem of cosmic radiation on such a flight?

Sol: Suppose X be the random variable denoting cosmic radiation exposure to a person. X is
7
John E. Freund’s Mathematical Statistics/ Sheldon Ross, Introductory Statistics.
8
mrem stands for millirem.

24
a normal variate with mean 4.35 and standard deviation 0.59 i.e. E(X) = 4.35 and V(X) = 0.59.
Let Z be the standard normal variate Z = X−4.35
0.59 . According to the problem we want to find

P(X > 5.20) = P(Z > 5.20−4.35


0.59 ) = P(Z > 1.44)
R∞
Now this P(Z > 1.44) means calculating 1.44 f (z)dz where f (z) is defined in eq.(50). This integra-
tion needs to be done by numerical integration procedures such as Trapezoidal rule, Simpson’s one
third rule [Link] 9 Fortunately, we can get these probabilities from the chart of ‘Area under normal
curve’ (which is available in any Statistics Book). From this chart we get P(0 < Z < 1.44) = 0.4251.

Figure 6: Normal Curve

From the above figure we see that P(Z > 1.44) = 0.5 − P(0 < Z < 1.44) = 0.5 − 0.4251 = 0.0749.

In most cases, it is found that the data deviates from normality i.e. the curve is skewed to the left
or to the right. If the tail is extended to the right then the data is positively skewed and if extended
to the left, is called negatively skewed. The following figure will clarify this point.

Figure 7: Normal Curve

Normal curve, you can say an idealistic situation from which the actual data deviates. This prop-
erty of being skewed to the left or to the right can be quantified by Skewness.

9
These topics are beyond the scope of the syllabus.

25
Definition: The symmetry or skewness (lack of symmetry) of a distribution is often measured
by means of the quantity
µ3
α3 = . (51)
σ3

Definition: The extent to which a distribution is peaked or flat, also called the Kurtosis. This
is quantified by
µ4
α4 = . (52)
σ4
In the above formulae σ is the standard deviation or second moment about mean i.e. µ2 .

Figure 8: Kurtosis

rth moment about mean µr is connected to rth moment about origin µ′r in the following way.

µr = µ′r − rC1 µ′r−1 µ + · · · + (−1)i rCi µ′r−i µi + · · · + (−1)r−1 (r − 1)µr . (53)

Using eq.(53) and using properties of moments10 we get the following.

µ3 = µ′3 − 3µ′2 µ + 2µ3


µ4 = µ′4 − 4µ′3 µ + 6µ′2 µ2 − 3µ4 . (54)

Example (Skewness): : Consider the following probability distribution. f (1) = 0.05, f (2) =
0.15, f (3) = 0.30, f (4) = 0.30, f (5) = 0.15, f (6) = 0.05 Draw the curve. Also find the skewness of
this distribution.

Sol: If we plot the distribution as bar diagram we get the following (shown in Figure 11). Us-
ing eq.(54) we calculate µ3 . Then using eq.(51) we calculate α3 which comes out to be 0. This is
desirable as the given distribution is of symmetric nature.
10
See in the previous pages where I have discussed moments.

26
Figure 9: Plot of example of skewness

Usif eqs.(54) and (52) we also calculate α4 which is approximately 3. This proves that the given
distribution is Mesokurtic.[Students are encouraged to do the detailed calculation.]

Note: (a) For Normal curve or mesokurtic curve the value of measure of Kurtosis will be ex-
actly equal to 3. (b) For leptokurtic curve the measure of the Kurtosis will be greater than 3 and
(c) for platykurtic curve the measure will be less than 3.

Problem: Analyze the two curves in terms of Skewness and Kurtosis.(Note the difference!) (a)
f (−3) = 0.06, f (−2) = 0.09, f (−1) = 0.10, f (0) = 0.50, f (1) = 0.10, f (2) = 0.09 and f (3) = 0.06.
(b) f (−3) = 0.04, f (−2) = 0.11, f (−1) = 0.20, f (0) = 0.30, f (1) = 0.20, f (2) = 0.11 and
f (3) = 0.04.

Percentiles of normal random variable: Let us consider the following example. For any
0 < α < 1, we define zα as follows.

P(Z > zα ) = α, (55)

where Z is the standard normal variate. What can we say about the value of zα ? [See the figure be-
low!] Suppose α = 0.025 so that P(Z > zα ) = 0.025. Now P(Z < zα ) = 1−p(Z > zα ) = 1−0.025 =

Figure 10: Plot of standard normal variate with zα

0.975. Consider now the following chart which is a part of standard normal table (which is indeed
readily available). 11 Locate the value 0.975 from the table which has already been highlighted.
This value appears for zα = 1.96. We can now say that, 1.96 is known as the 97.5 percentile of the
standard normal distribution. In general, since P(Z > zα ) = α → 1 − P(Z < zα ) = 1 − α. We
call zα as the 100 × (1 − α) percentile of the standard normal distribution. Literally, percentile
means a value on a scale of one hundred that indicates the percent of a distribution
11
This chart represents probability P(Z < zα )

27
Figure 11: Plot of standard normal table to locate zα

that is equal to or below it.

Now as we have got P(Z < 1.96) = 0.975, this means P(−∞ < z < 1.96) = 0.975, so that
using properties of standard normal variate we can immediately say P(−1.96 < Z < 1.96) = 0.95
(Students must try out!).

Problem: Take α = 0.005, find zα and analyze what you get!

Problem: Find moment generating function of binomial and Normal variate. Hence find mean
and variance for both the distributions.12

Hints: For binomial variate the moment generating function is (q + pet )n and for normal vari-
1 2 2
ate it is eµ+ 2 σ t .

Problem: Show that if a random variable has the probability density f (x) = 21 e−|x| , for −∞ <
1
x < ∞, its moment generating function is given by MX (t) = 1−t2.

4 Basic Statistics:
Statistics is the science of data, today it is fashionably known as Data Science. The important
aspect of dealing with data is organizing and summarizing the data in ways that facilitate its in-
terpretation and subsequent analysis.

Numerical summaries of the un-grouped data:


Arithmetic Mean: If the n observations in a sample are denoted by x1 , x2 , · · · , xn , the sample
mean is
x1 + x2 + · · · + xn
X =
Pn n
x
i=1 i
= . (56)
n
The sample mean is the average value of all observations in the data set. Usually, these data are a
sample of observations that have been selected from some larger population of observations. For a
finite population however, with N equally likely values, the probability mass function is f (xi ) = N1
12
For solution you can refer any good book on Statistics.

28
and the mean is
Pn
i=1 xi
µ= . (57)
N
The sample mean X is the reasonable estimate of the population mean µ.

Sample Variance: If x1 , x2 , · · · , xn is a sample of n observations, the sample variance is


Pn
2 (xi − X)2
s = i=1 . (58)
n−1
The sample standard deviation is the positive square root of the sample variance. Analogous to
the sample variance s2 , the variability in the population is defined by the population variance σ 2 .
The positive square root of σ 2 or σ will denote the Population Standard Deviation. When the
population is finite and consists of N equally likely values, we may define the population variance
as
Pn
2 (xi − X)2
σ = i=1 . (59)
N
The sample variance of Eq.(58) is an unbiased estimate of the population variance
of Eq.(59).

From Eq.(58) and Eq.(59), it is easy to find a relationship between σ 2 and s2 , which is
n−1 2
σ2 ≡ s . (60)
N
Eq. (60) gives a relationship between population variance and sample variance, of course the re-
lation is an approximate relation. Later we shall see that, s2 estimates population variance in an
unbiased way.
Pn
(x −X)2
Note: In some books, however, the Eq.(58) is shown to be as i=1 ni . But remember this
is just for the sake of calculation. Although there is not much difference between the values of
variance, when it is divided by n or n − 1, but statistically it is incorrect (when you divide by n).

Sample 100p percentile: The sample 100p percentile is that value having the property that
at least 100p percent of the data are less than or equal to it and at least 100(1 − p) percent of
the data values are greater than or equal to it. If two data values satisfy this condition, then the
sample 100p percentile is the arithmetic average of these values.

To find the sample of 100p percentile of a data set of size n we follow the following algorithm.

ˆ Arrange the data in increasing order.

ˆ If np is not an integer, determine the smallest integer greater than np. The data value in that
position is the sample 100p percentile.

ˆ If np is an integer, then the average of the values in positions np and np + 1 is the sample
100p percentile.

29
The sample 25th percentile is called the first quartile (or Q1). The sample 50th percentile is called
the median (Q2) or the second quartile. The sample 75th percentile is called the third quartile
(Q3).

Problem: Consider the following data 6.2, 2.1, 3.0, 4.5, 6.3, 7.1, 1.2. Calculate (i)25th percentile,
(ii) 50th percentile, and (iii) 75th percentile.

Sol: Total number of data points here are n = 7.

ˆ After arranging the data in increasing order we get, 1.2, 2.1, 3.0, 4.5, 6.2, 6.3, 7.1.

ˆ For 25th percentile, p = 0.25. Then np = 7×0.25 = 1.75 and it is not an integer. The smallest
integer greater than 1.75 is 2.

ˆ Hence the value in the second place i.e. 2.1 is the 25th percentile.

ˆ For 50th percentile, p = 0.50. Then np = 7 × 0.50 = 3.5 and it is not an integer. The smallest
integer greater than 3.5 is 4.

ˆ Hence the value in the fourth place i.e. 4.5 is the 50th percentile or MEDIAN.

ˆ For 75th percentile, p = 0.75. Then np = 7×0.75 = 5.25 and it is not an integer. The smallest
integer greater than 5.25 is 6.

ˆ Hence the value in the sixth place i.e. 6.3 is the 75th percentile.

Problem: Consider the following data 100, 100.5, 91.5, 89. Calculate 50th percentile.

Sol: Total number of data points here are n = 4.

ˆ After arranging the data in increasing order we get, 89, 91.5, 100, 100.5.

ˆ For 50th percentile, p = 0.50. Then np = 4 × 0.50 = 2 and it is an integer.

ˆ Hence the value between 2nd and 3rd place is the 50th percentile or MEDIAN.

ˆ The MEDIAN value, however is, the average of the values of 2nd and 3rd place, i.e., 91.5+100
2
i.e. 95.75.

The following diagram will make things clear. This is to be noted that there is a difference between

Figure 12: Percentile plots

30
the words percentage and percentile. A student sat for a competitive examination scored 31 out
of 40. Her progress report shows that she got 77.5% and 91.52 percentile. What do they mean?
They mean that as the student had got 31 out of 40 what she would have got had the examination
be taken on 100 marks. She would have got 77.5. That is percentage. However, 91.52 percentile
means p = 0.9152. Suppose 100 students appeared in the examination. So at least 100 × 0.9152
percent of the candidates are less than or equal to 91.52 and 100 × (1 − 0.9152) i.e. 8.48 percent of
the candidates are greater than or equal to this value.

Note: There is also another important observation. From the above examples it is clear that
mean, median or any type of average may or may not be the part of the data set (Try to under-
stand!).

Sample mode: Another indicator of central tendency is the sample mode, which is the data
value that occurs most frequently in the data set. As for example we can assume a person going
to BATA to buy shoe. Now the size of the shoes there will range from 2 to 13 say. Now if the
shop is meant for kids, then one can expect more shoes of sizes ranging from 2 to 5. If suppose
that particular shop has shoes of above mentioned sizes with some frequencies, such as, there are
hundred shoes of size 2, three hundred shoes of size 3, five hundred and fifty shoes of size 4 and
eight hundred shoes of size 5, then the mode of the data will be 5 as it is available in maximum
number.

All these measures, such as sample mean, sample percentile and sample mode are measures of
central tendency as these measures give a unique value for the data which tells how other values of
the data are scattered around it. But one has to be very careful about the choice of the appropriate
measure of central tendency. Choosing wrong measure will result in garbage value. Think of the
situation of the person buying shoes from the shop, and what would have happened had he used
sample mean or sample median to measure the average shoe size available in that shop! Will it be
right?

All the averages discussed above are meant for un-grouped data. Now we shall discuss these
measures for grouped data.

Numerical summaries of the grouped data:

5 Curve fitting:
We know that when two variables x (independent) and y (dependent) are connected by a well-
defined function y = f (x), we can visualize it easily by plotting y against x. But suppose we have
a pair of variables (x, y) with some observed values for each. It is not always the case that y and
x will be connected by a well defined mathematical relation. As for example, if you take x to be
amount of petrol consumed by your bike each day and y to be the binge watching time on OTT
platform you spend each day, obviously no mathematical relation connects the two variables. But
what you will have with you is a data set representing the the pair of variables (x, y). In this case
the first thing one can do is to plot the data (taking independent variable on x axis and dependent
variable on y axis). Such a plot of points is called scatter plot, also know as diagram of dots. We
consider one such scatter plot as shown below. Anyone who looks at the figure 4 can say that the
data points have tendency to move upwards. Also it reflects that the nature of the data points is
linear. You can always consider a hypothetical straight line passing through these points. Let us

31
Figure 13: Scatter Diagram in which linear curve can be fitted.

consider another figure as shown. It is quite obvious that the nature of the plot in fig 5 is not linear,

Figure 14: Fig: Scatter Diagram in which linear curve can’t be fitted.

rather curvilinear. By analyzing a scatter plot and setting a linear or polynomial or transcendental
curve to data points constitute the subject matter of curve fitting.

How to fit a linear curve?


In the discussion of curve fitting we shall consider more general variable X and Y (other than x and
y). We call these as Random Variable. The pair of data (X, Y ) taken together is called bi-variate
data. If we want to plot a straight line to the set of data points, we can consider a linear equation
of the type

Y = A + BX, (61)

where A is the intercept and B is slope. These two quantities A and B are unknown, any knowl-
edge about them will reveal the nature of the linear curve. To know two unknowns, one needs to
consider two set of equations involving the unknowns. The process through which we will construct
two such equations containing X and Y is called method of least squares. Look at the fig 1 now. If
you fit a straight line passing through the set of data points, not all points will fall on the straight
line, some may fall on it while some will fall either above or below the hypothetical straight line.
Imagine, you drop perpendiculars from these points on the straight line and these represent the
error in fitting (or more practically you may say that not all of your observed values are falling on
your hypothetical straight line).

Suppose there n observations in your data set and each of these n observations for random variables
X and Y are given by xi and yi . Consider figure 6 now, which is indeed a modified version of figure

32
1 where a hypothetical straight line (green) satisfying eq.(61) has been assumed to pass through
the data points. An arbitrary point P is located whose coordinate is (xi , yi ). Drop a perpendicular
from this point and it cuts your hypothetical line at Q and meets x− axis at M . The coordinate
of Q is thus (xi , a + bxi ). The y− coordinate of Q touches your hypothetical straight line. We can
say yi is observed value of point P and its estimated value is yˆi = a + bxi . The difference yi − yˆi
is the error, denoted by ei . In method of least squares, we consider all such arbitraryPn points P
and consider the sum of squares of these error terms, which may be denoted by Z = i=1 ei . Our2

objective is to minimize this quantity Z. Thus our problem is defined as

n
X
M inimize(Z) = (yi − a − bxi )2 , (62)
i=1

where A and B are not known and which need to be estimated so that form of eq.(61) is known.
We differentiate Z of eq.(62) with respect to A and B, equating each to zero subsequently, we get
the following pair of equations,
n
X n
X
yi = nA + B xi
i=1 i=1
n
X n
X n
X
xi yi = A xi + B x2i . (63)
i=1 i=1 i=1

The pair of equations in (63) are known as normal equations. Solving these normal equations one
can easily find A an B and thus a hypothetical form of eq.(61). If you divide both sides of first
equation of (63) by n, we get

Ȳ = A + BX. (64)

This eq.(64) is so beautiful as it reveals that even if your observed set of points may not all fall
on your constructed straight line eq.(61), yet their respective means will satisfy the said equation.
Before we move further with our analysis, let us know about a quantity which is significant in the
study of regression. By a similar approach we can fit other different types of curves (polynomial,
logarithmic, exponential [Link] ) to a given bi-variate set of data.

Polynomial curve fitting:


The least-squares procedure can be readily extended to fit the data to a higher-order polynomial.
For example, suppose that we fit a second-order polynomial or quadratic:

Y = a0 + a1 X + a2 X 2 + e (65)

33
Following the least square methodology described above the normal equations for eq.(65) can be
found as the following,

X
Y = na0 + a1 X + a2 X 2
X X X X
XY = a0 X + a1 X 2 + a3 X3
X X X X
X 2Y = a0 X 2 + a1 X 3 + a3 X4 (66)

(67)

Other curves:
There are some other forms of curves which can be assumed to be fitted on the set of observed
data. These are

Exponential → Y = abX
P ower → Y = aX b
Gompertz → Y = abcX
k
Logistic → Y = ,b < 0
1 + ea+bX
(68)

Important Remarks:
Some basic facts can be kept in mind. For deciding about the type of curve to be fitted to a given
set of data the following points may be helpful.

ˆ When the Y series is found to be increasing by equal absolute amounts, the straight line curve
is used.

ˆ The logarithmic straight line curve (or exponential curve) is used when the series is increasing
or decreasing by a constant percentage rather than a constant absolute amount.

ˆ Second degree curve fitted to logarithms may be tried if the data plotted on a semi-logarithmic
scale is not a straight line graph but shows a curvature, being concave either upward or
downward.

What is correlation coefficient?


Existence of non-zero correlation coefficient between two random variables X and Y reveal that
there may exist linear relationship between the variables. The correlation coefficient between X
and Y , denoted by rXY , is given by

Cov(X, Y )
rXY = p p , (69)
V(X) V(Y )

where Cov(X, Y ) is the measure of mutual deviation between the variables and V ar(X) is deviation
of the set of values from mean X while V ar(Y ) is deviation of the set of values from mean Ȳ .

34
Mathematically we have
1X
µ11 = Cov(X, Y ) = XY − X Ȳ
n
2 1X 2 2
σX = V(X) = X −X
n
1X 2
σY2 = V(Y ) = Y − Ȳ 2 . (70)
n
Remember, −1 ≤ rXY ≤ 1. The following chart will be helpful in this regard.

rXY > 0 X and Y are positively correlated


rXY < 0 X and Y are negatively correlated
rXY = 1 X and Y are perfectly positively correlated
rXY = −1 X and Y are perfectly negatively correlated
rXY = 0 X and Y are not correlated at all

Example: (a) Find correlation coefficient between X and Y where

X 1 2 3 4
Y 0.5 0.7 0.88 0.92

(b) Find correlation coefficient between X and Y where

X 1 2 3 4
Y 3 5 7 9

Sol: (a) Intuitively we can say there will be positive correlation between X and Y . You see that for
increase in the value of X there is increase in the value Y . Using eq.(69), the value of rXY = 0.99219.
(b) Here, also for increase in the value of X there is increase in the value of Y . Eq.(69) reveals the
correlation coefficient to be perfect positive i.e. rXY = 1.

Example: (a) Find correlation coefficient between X and Y where

X 1 2 3 4
Y 4 3 2 1

(b) Find correlation coefficient between X and Y where

X 1 2 3 4
Y 0.88 .65 .48 .28

Sol: (a) Intuitively we can say there will be perfect negative correlation between X and Y . You
see that for increase in the value of X there is decrease in the value Y . Using eq.(69), the value of
rXY = −1. So the variables have positive correlation.
(b) Here, also for increase in the value of X there is decrease in the value of Y . Eq.(69) reveals the
correlation coefficient to be negative i.e. rXY = −0.99838.

Example: Find correlation coefficient between X and Y where

X 1 1 -1 -1
Y 1 -1 1 -1

35
Sol: You plot the scatter diagram. You will see that the four points will respectively fall into the
four quadrants. So there is no correlation between X and Y . Hence rXY = 0. Interpretation of
correlation coefficient: From eq.(69) we have
Cov(X, Y ) = rXY σX σY
⇒ Cov(X, Y ) ∝ σX σY (71)
Hence rXY is the proportionality constant. Non-zero Cov(X, Y ) value actually implies the exis-
tence of linear relationship between X and Y . So if rXY = 0.11 (say), this means if we multiply
the product σX σY by this factor then it becomes equal to Cov(X, Y ) for the given data.

If we take an observational data such as


X -1 -2 -3 1 2 3
Y 1 4 9 1 4 9
, it is easy to verify that in this case rXY = 0, which in turn means there is no linear relationship
between X and Y , but careful observation will reveal that Y and X actually bear a non-linear
relationship given by Y = X 2 (Verify !).

The existence of non-zero correlation between two variables means that the variables are linearly
related. Zero correlation means the variables are not linearly related while non-linear relationship
between the two variables may exist.

Problem: A coefficient of correlation of 0.2 is derived from a random sample of 625 pairs of
observations. Is this value are significant? What is the 95% confidence limit to the correlation
coefficient in the population?


Sol: H0 : ρ = 0; H1 : ρ ̸= 0, Test statistic: t = r√1−r
n−2
2
= 5.09 > t0.05 = 1.96. Hence H0 is
rejected.

95% confidence limits for ρ are r ± 1.96 × S.E.(r) = r ± 1.96 × (1 − r2 )/ n = (0.125, 0.275)

How can we bring correlation coefficient to normal equations?


We rewrite eq.(63) in terms of random variables.
X X
Y = nA + B X
i i
X X X
XY = A X +B X 2. (72)
i i i

Using eq.(70) in second equation of eq.(72) we get


2 2
µ11 + X Ȳ = AX + B(σX + X ). (73)
Eq.(73) is independent of size n of the sample. Now if we multiply eq.(64) by X and subtract from
the second equation of eq.(73), we get
µ11
B = 2
σX
σY
⇒ B = rXY . (74)
σX

36
You see that the correlation coefficient rXY is directly proportional to slope B of your estimated
straight line eq.(61) whereas the proportionality constant is σσX
Y
. Moreover if you want to write an
equation of straight line passing through the arbitrary pair of points (X, Ȳ ) (the means of X and
Y ), then this equation is therefore given by
σY
Y − Ȳ = rXY (X − X). (75)
σX
This equation (75) is known as regression equation of Y on X.

Similarly by starting with an equation of the form X = C + DY one can find, using the method of
least squares, the regression equation of X on Y . This regression equation is given by
σX
X − X = rXY (Y − Ȳ ). (76)
σY
σY σX
The quantities rXY σX and rXY σY are often denoted by bY X and bXY respectively. This is to be
noted that
2
bY X × bXY = rXY . (77)

Example: Fit a second degree polynomial to the following data


X 0 1 2 3 4
Y 1 1.8 1.3 2.5 6.3

Ans: Y = 1.42 − 1.07X + 0.55X 2

Example: Fit an exponential curve to the following data


X 1 2 3 4 5 6 7 8
Y 1.0 1.2 1.8 2.5 3.6 4.7 6.6 9.1

Ans: Y = .6821(1.38X )

Example: Fit a straight line to the following data


X 1 2 3 4 6 8
Y 2.4 3 3.6 4 5 6
Ans: Y = 1.976 + 0.506X.

Problem: (i) Write the formula for the correlation coefficient r(X, Y ). (ii) Replace expressions of
Cov(X, Y ), V ar(X) and V ar(Y ) using summation notation. (iii) In the expression that followed,
replace Xi − X and Yi − Ȳ by ai and bi respectively and then re-write the expression for r(X, Y ).
(iv) Square both the sides of the expression obtainedPin (iii). [Cauchy P Schwartz’sP inequality states
that if ai , bi , i = 1, 2, · · · , n are real quantities then ( ni=1 ai bi )2 ≤ ( ni=1 a2i )( ni=1 b2i ).] (v) Apply
Cauchy Schwartz’s inequality in (iv). (vi) What do you get? Interpret your answer!

Sol: We know correlation coefficient r(X, Y ) between two random variables X and Y is com-
P 1
Cov(X,Y ) n−1 i (Xi −X)(Yi −Ȳ ) 13 .
puted as, r(X, Y ) = σX σY = q
1
q
1
−X)2 2
P P
n−1 i (Xi n−1 i (Yi −Ȳ )

13
It is to be noted that we have divided by n − 1 which will be clarified later.

37
P
ia i bi
Now let us assume that Xi − X = ai and Yi − Ȳ = bi so that r(X, Y ) = √P √ P . Squaring
i ai i bi
ai bi )2
P
(
both sides we get, r2 (X, Y ) = P i2 P
a 2. Using Causchy-Schwartz’s inequality as mentioned in
i i i bi
the question we get r2 (X, Y ) ≤ 1 that is |r(X, Y )| ≤ 1 i.e. −1 ≤ r(X, Y ) ≤ 1. This shows that
correlation coefficient lies between −1 and 1.

Problem:Let two random variables U and V are defined as U = X−a h and V = Y k−b , where X
and Y are another pair of random variables. (i) Determine the forms of X − E(X) and Y − E(Y )
in terms of U and V . (ii) Write the formula of Cov(X, Y ) in terms of expectation. (iii) Replace
X − E(X) and Y − E(Y ) by the expressions obtained in (i). (iv) Write the formulae of V ar(X) and
V ar(Y ) in terms of U and V . (v) Using expressions obtained in (iii) and (iv) in terms of U and V ,
substitute them in the formula for r(X, Y ). (vi) Determine the final expression of the correlation
coefficient and interpret your result.

Y −b
Sol: Since U = X−a
h and V = k , so X = a + hU ,Y = b + kV and consequently E(X) = a + hE(U )
while E(Y ) = b + kE(V ). Thus, X − E(X) = h(U − E(U )) and Y − E(Y ) = h(V − E(V )).
Cov(X,Y ) E(X−E(X))(Y −E(Y )) E[h(U −E(U ))][k(V −E(V ))]
We know, r(X, Y ) = σX σY = √ √ = √ √ =
E(X−E(X))2 E(Y −E(Y ))2 E[h(U −E(U ))2 ] E[k(Y −E(Y ))2 ]
r(U, V ).

This implies that correlation coefficient is independent of change of origin and scale.

Example: If r(X, Y ) = 1 what is the value of r(X + 2, Y + 3)?

Ans: 1.

Rank Correlation:
Let us suppose that a group of n individuals is arranged in order of merit or proficiency in possession
of two characteristics A and B. These ranks in the two characteristics will in general be different.
Let (Xi , Yi ), i = 1, 2, · · · , n be the ranks of the ith individual with respect to the characteristics
A and B respectively. Karl Pearson’s correlation coefficient r(X, Y ) can be then re-modified to
calculate correlation between Xi and Yi . We assume that no two individuals are ranked same in
either classification and each of the variables X and Y are ranked as 1, 2, · · · , n. Then X = Ȳ =
2
n+1 2 1 P 2
Xi = n(n+1)(2n+1)
P 2
2 (Check!). We know σX = n Xi − X , and as (Check!), we have
2 2
n −1 2
P1 6
σX = 12 (Check!) which is same as σY . Also as Cov(X, Y ) = Pi − X)(Yi2− Ȳ ), by letting
n (X
di = Xi −Yi and running simple manipulation one can easily check that d2i = 2σX −2r(X, Y )σX 2 ,

using eq.(69) (Check!). A little simplification further gives the following formula (Check!).

6 d2i
P
r(X, Y ) = 1 − . (78)
n(n2 − 1)

This formula was deduced by Spearman and hence sometimes is also known as Spearman’s rank
correlation coefficient. To differentiate between the two correlation formulae (69) and (78) we
use the symbol ρXY or ρ(X, Y ) to denote the Spearman’s formula. We also immediately have
−1 ≤ ρ(X, Y ) ≤ 1 (Prove!).

38
Example: Suppose five students sat for Statistics and Mathematics examination and their score
out of 25 (for each examination) are denoted pairwise (11, 13.5), (15, 15), (12, 9), (3, 1), (20, 22),
where the first number in the pair denote marks of Statistics and second number denotes marks
in Mathematics. Arrange the students rank-wise for each subject and find the rank correlation
coefficient between their ranks.

Sol: We summarize the details in the following table.


Student Marks (Maths) Marks (Stat) Rank Xi (Maths) Rank Yi (Stat) di = Xi − Yi
1 11 13.5 4 3 1
2 15 15 2 2 0
3 12 9 3 4 1
4 3 1 5 5 0
5 20 22 1 1 0
P 2
Now from the above table di = 2 and using eq.(78) we get ρXY = 0.9. Ranks are highly pos-
itively correlated. This is true as a person having good Mathematics background is expected to
have sound knowledge in Statistics.

Example: Here the continuous assessment marks (CA1 and CA2) of seven students of a class
are taken. The first number in the bracket represents CA1 marks and the second number repre-
sents CA2 marks. The marks are (25, 15), (25, 9), (25, 12), (25, 8), (22, 15), (25, 12), (21, 11). Now
if we want to arrange the students rank-wise we see that there will be many repeated ranks. The
first question is how to arrange them? We consider all the CA1 marks first. We see that five stu-
dents got 25 and theoretically they must have rank 1. This will not give the correct interpretation
of the data and what we do is to take arithmetic mean of 1, 2, 3, 4, 5 (i.e. first position, second
position and so on), then allotting the mean value to each student, which is 3. The next highest
mark is 22 and third highest is 21, so we allot them rank 6 and 7 respectively. Now for CA2 marks
we first observe that the highest mark 15 was achieved by two students whose theoretical rank were
supposed to be 1. But since there are ties in this rank we take the mean of 1 and 2, which is 1.5
and we allot the number to both of these students. The next highest is 12 and we see that there
is also tie between third and fourth position. Hence we take mean of 3 and 4, which is 3.5. We
allot this rank to 3rd and 4th candidate. Rest of the candidates (where there are no more ties) are
ranked serially allocating from 5th position. Thus we get the following table.
Student Marks (CA1) Marks (CA2) Rank Xi (CA1) Rank Yi (CA2) di = Xi − Yi
1 25 15 3 1.5 1.5
2 25 9 3 6 -3
3 25 12 3 3.5 -0.5
4 25 8 3 7 -4
5 22 15 6 1.5 4.5
5 25 12 3 3.5 -0.5
5 21 11 7 5 2
P 2
From table we get di = 52. Using eq.(78) we calculate the rank correlation to be ρXY = 0.0714.
But this is not the end. In case of repeated ranks or sometimes called tied ranks, Spearman’s
2
correlation formula needs to be corrected. A correction factor m(m12−1) needs to be added where
m denotes the number of times a value repeats itself. For CA1 marks the one correction factor is
to be added (as 25 repeats itself five times making m = 5). The added correction factor would be

39
5(52 −1)
12 = 10, while in case of CA2 marks 15 repeats itself twice and also 12 repeats itself twice.
2
So for each of these cases the correction factor would be 2(212−1) = 0.5. Therefore calculate the
Spearman’s rank correlation as

6( d2i + i Ci )
P P
ρ(X, Y ) = 1 − , (79)
n(n2 − 1)

where Ci is the ith correction factor. In this case we have


d2i +10+0.5+0.5)
P
6(
ρ(X, Y ) = 1 − 7(72 −1)
= −0.125.

This value is justified as you can see from the marks that for increase in the value of CA1 marks
there is decrease in the value of CA2 marks.

Exercise: In a contest, two judges ranked seven candidates in order of their preference as in
the following table:
Candidates A B C D E F G
Rank by Judge I 2 1 4 5 3 7 6
Rank by Judge II 3 4 2 5 1 6 7
Calculate the rank correlation coefficient.
6 d2
P
d2 = 20, R = 1 −
P
Sol: n = 7, n3 −n
= .64

Exercise: Fit the equation y = 2 + abx to the following data by the method of least squares:
x 2 3 4 5 6
y 14 25 50 98 194
Sol: Required equation is y = 2 + 3 × 2x

6 Sampling Distribution:
Statistics concerns itself mainly with conclusions and predictions resulting from chance outcomes
that occur in carefully planned experiments or investigations. In the finite case, these chance out-
comes constitute a subset, or sample, of measurements or observations from a larger set of values
called the population. In the continuous case they are usually values of identically distributed
random variables, whose distribution we refer to as the population distribution, or the infinite
population sampled. The word infinite implies that there is, logically speaking, no limit to the
number of values that we could observe.

Problem: Take a population P = {{1}, {2}}. Suppose the events {1}, {2} are equally likely.
Construct the probability distribution and draw a line diagram. Now construct a sample of size 2
independently, without replacement. Define the random variables X1 and X2 which respectively
take the values from the population. Calculate X. Find the probability distribution of X. Draw
the line diagram. Calculate population mean and population standard deviation. Denote them by
µ and σ. Calculate X. Find the probabilities of X. Now calculate the mean of sample mean that
is E(X). Also calculate V(X). Try to figure out what you observe! If you repeat the experiment
using with replacement sampling, will the result hold? Justify!

40
Problem: Take a population P = {{1}, {2}}. Suppose the events {1}, {2} are not equally likely.
Rather P({1}) = 0.7 and P({2}) = 0.3. Construct the probability distribution and draw a line
diagram. Now construct a sample of size 2 independently, without replacement. Define the random
variables X1 and X2 which respectively take the values from the population. Calculate X. Find the
probability distribution of X. Draw the line diagram. Calculate population mean and population
standard deviation. Denote them by µ and σ. Calculate X. Find the probabilities of X. Now
calculate the mean of sample mean that is E(X). Also calculate V(X). Try to figure out what you
observe!

Problem: Take a population P = {{1}, {2}, {3}}. Suppose the events {1}, {2}, {3} are equally
likely. Rather P({1}) = 0.7 and P({2}) = 0.3. Construct the probability distribution and draw a
line diagram. Now construct a sample of size 2 independently, without replacement. Define the
random variables X1 and X2 which respectively take the values from the population. Calculate X.
Find the probability distribution of X. Draw the line diagram. Calculate population mean and
population standard deviation. Denote them by µ and σ. Calculate X. Find the probabilities of
X. Now calculate the mean of sample mean that is E(X). Also calculate V(X). Try to figure out
what you observe!

Theorem (sampling distribution of mean): If X1 , X2 , · · · , Xn constitute a random sample


from an infinite population with the mean µ and the variance σ 2 , then

σ2
E(X) = µ, V(X) = . (80)
n
Central limit theorem: If X1 , X2 , · · · , Xn constitute a random sample from an infinite population
with the mean µ, the variance σ 2 , and the moment generating function MX (t), then the limiting
distribution of
X −µ
Z= , (81)
√σ
n

as n → ∞ is the standard normal distribution.

Another version of the central limit theorem is as follows. Let X1 , X2 , · · · , Xn be the sample
2
Phaving mean E(Xi ) = µ and V(Xi ) = σ (i = 1, 2, · · · , n). ForP
from a population large n, the
randomPvariable i Xi will approximately have a normal distribution with mean E( i Xi ) = nµ
and V( i Xi ) = nσ 2 .

Example: An astronomer is interested in measuring, in units of light years, the distance from
her observatory to a distant star. However, the astronomer knows that due to differing atmo-
spheric conditions and normal errors, each time a measurement is made it will yield not the exact
distance, but an estimate of it. As a result she is planning on making a series of 10 measurements
and using the average of these measurements as her estimated value for the actual distance. If the
values of the measurements constitute a sample from a population having mean d (the actual dis-
tance) and a standard deviation of 3 light years, approximate the probability that the astronomer’s
estimated value of the distance will be within 0.5 light years of the actual distance.

Sol: The probability of interest here is P(−0.5 < X − d < 0.5). Here X is the sample mean

41
Figure 15: Observatory

of 10 measurements. Since E(X) = d and S.D(X) = √3 , then by central limit theorem (eq.(81)
10
we have Z = X−d
√3
. Hence we get, P( −0.5
√3
< X−d
√3
< 0.5
√3
) which is P(−0.527 < Z < 0.527). Using
10 10 10 10
normal distribution table we can easily calculate this probability to be 0.4038. Therefore we see
that with 10 measurements there is a 40.38% chance that the estimated distance will be within
plus or minus 0.5 light years of the actual distance.

Example: An insurance company has 10000 (104 ) auto-mobile policy holders. If the expected
yearly claim per policy holder is dollar 260 with standard deviation of dollar 800, approximately
the probability that the total yearly claim exceeds dollar 2.8 million i.e. 2.8 × 106 .

Figure 16: Policy holder

Sol: Let us number each policy-holder as 1, 2, · · · , i, · · · and let Xi denote the yearlyPclaim4 of
policy-holder i, where in this problem i = 1, 2, · · · , 104 . By central limit theorem, X = 1i=1
√ 0 Xi
4 6
will approximately be normal with mean 260×10 = 2.6×10 and standard deviation 800× 104 =
8 × 104 . Hence, the probability that the total claim will exceed dollar 2.8 million is given as
P(X > 2.8 × 106 ) = P(Z > 2.5) = 0.0062 (using normal distribution chart)14 .

Theorem: If X is the mean of a random sample of size n from a finite population of size N
with the mean µ and the variance σ 2 , then
σ2 N − n
E(X) = µ, V(X) = . (82)
n N −1

Problem: A random sample of size n = 81 is taken from an infinite population with the mean
14
Students are advised to always draw normal curve to mark regions whose area they are calculating.

42
µ = 128 and the standard deviation σ = 6.3. With what probability can we assert that the value
we obtain for X will not fall between 126.6 and 129.4 if we apply (a) Chebychev’s theorem and (b)
the central limit theorem?

Sol: We are supposed to find what is the probability that X < 126.6 or that of X > 129.4.
For this we will calculate P(126.6 < X < 129.4) first. By central limit theorem, we have using
eq.(81), P( 126.6−128
√6.3 < Z < 129.4−128
√6.3 ) which is P(−2 < Z < 2). Using standard normal table we
81 81
get P(−2 < Z < 2) = 0.9544 i.e. P(|Z| < 2) = 0.9544. Now since P(|Z| > 2) = 1 − P(|Z| < 2)
so we get our required probability as 0.0456. Now using eq.(37) (i.e. Chebychev’s inequality) we
get, P(|X − µ| ≥ kσ) < k12 . Using the values of µ and σ from question and choosing k = 2 we get
P(| X−128 1
6.3 | ≥ 2) < 22 i.e. the required probability has to be less than 0.25 which is indeed the case
as has been seen from central limit theorem.

7 Testing of Hypothesis (with large samples):


Basic objective of testing of hypothesis is to draw inference on a given population (finite or infinite).
Finite population is that population where the complete idea about the size of the population is
known and infinite population is that population where we are not sure about its size. Population
measures are the parameters or you can say the parameters define which probability distribution
the Population is following. These parameters cannot be measured or calculated, rather we hypo-
thetically assume their values. The only way to check the hypothesis is to design a sample and to
analyze that sample to draw conclusions on the population from which the sample has been drawn.
The sample may be of two types viz. large sample and small sample. When the sample size
n > 30, it is treated as large sample and when n ≤ 30, the sample is termed as small sample.
Hence there are two types of tests such as large sample test or sometimes also known as Z test
and small sample test. On the other hand the hypothesis (which literally called assumption)
that we make about the population parameters, are of two types such as (i) Null Hypothesis,
often denoted by H0 and (ii) Alternative hypothesis, often denoted as H1 . Null hypothesis is
the hypothesis of no difference and alternative hypothesis is the complementary statement to that
you have made in null hypothesis. How we can design the hypothesis can be readily understood by
the following example which I have personally observed over the past years.15

Before COVID lock-down, the trend in smoking in females was much less which over time and
after lock-down has significantly increased. Now if I claim that over 80% of the female population
of the city are habitual smokers, this claim is completely based upon the intuition. It may be
correct or may not be. Only way to check the validity of the statement is to test this hypothesis on
the basis of the sample drawn from the city’s female population. We could write then H0 : “The
proportion of the female smokers is over 80%” or mathematically if we say P is the population
proportion then H0 : P ≥ 0.80. Instantly the alternative hypothesis will be H1 : “The proportion
of the female smokers is less than 80%” or mathematically, H0 : P < 0.80. The smoking habits
have been acquired, may be due to the sedentary lifestyle during COVID phase or also may be due
to the binge watching and exposure to OTT platforms. We know smoking has many health hazards
and it may cause the infertility among women. So the objective of the above testing may be to let
females know and to make them get rid of this health menace. Another good example may be set
in this regard. Often you have seen, that in cinema, often a disclaimer, smoking and drinking are
15
The example is totally author’s personal opinion.

43
injurious to health comes when such scenes are displayed. Suppose you want to test H0 : “ More
than 60% of the audiences quit smoking after seeing the disclaimer” against H1 : “ They don’t
quit”. Thus so many real life examples can be thought of where hypothesis testing plays a vital
role for the statistician.

We shall first discuss the algorithm to do hypothesis testing.

ˆ Design your H0 in an appropriate way.

ˆ Choose appropriate test statistic16 and calculate.

ˆ Give your opinion on the hypothesis (either about its acceptance or on its rejection).

ˆ Write conclusion.

In testing of hypothesis (for large samples as well as for small samples), the above algorithm always
fits, only change that occurs is in the choice of proper test statistic. Below we shall discuss first large
sample tests and then small sample tests. We will study different cases. But before we proceed,
another aspect that should be kept in mind is the following. While designing hypothesis if we get
following situations such as H0 :≥, H1 :< or H0 :≤, H1 :>, the test is two tailed test and
when we get H0 :≥, H1 :̸= or H0 :≤, H1 :̸=, the test is one tailed test.

7.1 Testing of hypothesis:


Large sample tests (Z test):
The basic rule remains fixed while the only change is the choice of Z− statistic. We will discuss
the steps case-wise.
X−nP
Case I Test of significance of single proportion Z= √
nP Q
(p1 −p2 )−(P1 −P2 )
Case II Test of significance of difference of proportions Z= q
P1 Q1 P Q
n
+ 2n 2
1 2
X−µ
Case III Test of significance of single mean Z= √σ
n 17
(X̄1 −X̄2 )−(µ1 −µ2 )
Case IV Test of significance of difference of mean Z= r
2
σ1 σ2
n1
+ n2
2
(s1 −s2 )−(σ1 −σ2 )
Case V Test of significance of difference of standard deviation Z= r
2
σ1 σ2
2n1
+ 2n2
2

Test of significance of single proportion:


Example: Suppose a coin is tossed 900 times and head is observed 159 times. Can we say that
the coin is unbiased?

Sol: Let H0 : The coin is unbiased or P = 21 where P is proportion of population. against


H1 : The coin is biased i.e. P ̸= 21 . First of all observe that since the number of observations
n = 900 > 30 the sample is a large sample. According to the design of our hypothesis we can
declare the test to be a two-tailed test. On the basis of the null hypothesis we can choose
16
Any sample measurement is statistic.
17
In the chart corresponding to case III we have X and corresponding to case IV we have X̄1 − X̄2 .

44
X
X−nP −P
Z= √
nP Q
= n
q
PQ
,
n

where p = Xn is the sample proportion and X is the random variable denoting number of heads
observed which is 159 and Q = 1 − P . Using the values given we get
0.1766−0.5
Z= q
0.5×0.5
= −19.4
900

We see that Z < −3 and so the hypothesis is rejected. We conclude that the given coin is biased.

Note 1case1 : In the above problem, if we take X = 500, the calculated value of Z would be
3.34. This implies Z > 3 and null hypothesis is rejected. But if X = 700 then Z = −0.4834. This
shows that Z > −3 and hence we accept the null hypothesis.

Note 2case1 : In the formula for Z, the denominator is called Standard error, usually denoted
by S.E, which is similar to standard deviation.

PQ
Note 3case1 : It is easy to observe that E( X X
n ) = P and V( n ) = n (Try!).

The above problem is, you can say gives the objective of studying this course. In any probability
class when often asked about the probability of getting head in a single toss of a coin, students
reluctantly vote for 12 . But they overlook that the answer is only valid if the coin is unbiased. Now,
given a coin, how can one justify that the coin is unbiased or biased. The only way is to repeat
the tossing, generate sample of observations, design a suitable hypothesis, calculate the value of
test statistic and give opinion. In the above example we also find that the sample of observations
has been chosen from a infinite population as we are now sure how many times must we repeat
the process of tossing to get the desired outcome. Hypothetically we can do it as many times as
we can. Also since it is a single tossing coin problem, so there is only one population that has
been considered here. In some books, after calculating the test statistic Z the authors consider
the absolute value of Z. They accept hypothesis if |Z| ≤ 3 and reject when |Z| > 3. It is because
|Z| ≤ 3 implies −3 ≤ Z ≤ 3 whereas |Z| > 3 means either Z < −3 or Z > 3.

Since in the above problem, the hypothesis has been rejected and the coin comes out to be bi-
ased then this implies that P ̸= 21 . This means that either P > 12 or P < 12 . We can find the
p−P
probable limits of P . For this we start with |Z| ≤ 3 such that q PQ
≤ 3. A quick calculation
n
q q
reveals that p − 3 PnQ ≤ P ≤ p + 3 PnQ . But since P is not known, and as E(p) = P , we estimate
q q
P by p (the sample proportion) so that the probable limit is p − 3 pq
n ≤ P ≤ p + 3 pq
n . Applying
this, we get 0.1385 ≤ P ≤ 0.2147.

While taking a course on probability in the batch of CSE in the fall of 2019 − 20 even session,
I discussed the following problem.

Example: 74100 people have now been infected by the Corona-virus [COVID-19] in China. Out of
which 2000 people died in China [Data collected from The Statesman, 20.2.2020]. Will you accept
the hypothesis that the survival rate, if attacked by the disease, is less than 20%? (i) Formulate
the hypothesis, (ii) Apply the appropriate test statistic. [WHO Resolution: If the survival rate is
less than 20%, then the COVID-19 is Pandemic].

45
Sol: It is also a case of test of significance of single proportion. Just like the previous problem,
we calculate the test statistic Z. We design the hypothesis that H0 : P ≤ 0.02 against P > 0.02.
2000
Out of 74100 infected people, 2000 people died so that the proportion of dead people is 74100 . But
2000
since we are interested in survival rate so the proportion of people who survived is 1 − 74100 which
is the sample proportion p. The test is one tailed (why?) and since we do not have the complete
knowledge of the population of infected people, hence the population is considered to be infinite.
The calculated value of Z is

Z= q0.9731−0.02 = 1604.545.
0.9731×0.0269
74100

We see that Z > 3 and the null hypothesis is rejected i.e. survival rate is greater than 20%. Hence
according WHO resolution, the COVID 19 was not going to be pandemic as on 20th February, 2020.
But you see the funny part is by the end of March of that year, the disease took a terrible turn.

Note 4case1 : As per this old and obsolete problem, calculate the then probable limits of P with
respect to the the given conditions.

So far we have considered the cases of infinite population. Let’s discuss the scheme when pop-
ulation size is finite.

Theorem: When the population size N is finite and a sample of size is drawn, while popula-
tion proportion is P and sample proportion is p, then
r
N − n PQ
S.E(p) = , P + Q = 1. (83)
N −1 n
Example: A statistics exam is conducted for class of computer science engineering second semester
students. The size of the class is 126. After evaluating answer scripts of 35 students it was found
that 20 students got less than 15 out of total of 25 marks. The teacher gets angry and he does not
evaluate the remaining answer scripts. It was pre-decided that if more than 60% of the students’
performance is poor (i.e. students getting less than 15), he will quit teaching them. Now he wants
to analyze the situation on the basis of 35 answer scripts he had so far evaluated. Define H0 and
give your conclusion.

Sol: The population of students here is finite and the size N is 126. The sample size n is 20
and if we denote by X the number of students getting less than 15, as per question, X = 20.
The null hypothesis is H0 : The class performance is poor i.e. P ≥ 0.06 against the alternative
hypothesis H1 : P < 0.06. Then we use Z statistic as
X 20
−P −0.06
Z= q n
N −n P Q
= q 35
126−35
= 14.91
N −1 n 126−1
× 0.06×0.94
35

We see that Z > 3 and we reject H0 . We conclude that the average class performance is not poor
and the teacher decides to stay.

Now instead of designing the null and alternative hypothesis the way they are defined above,
if one writes as H0 : The average class performance is not poor i.e. P ≤ 0.06 against P > 0.06,
then under the same assumption Z = 14.91. This implies again that we reject the null hypothesis

46
and teacher decides to leave(!!)18 19

In all the above examples our Z value was either > 3 or Z < −3. However, when |Z| ≤ 3,
we fail to reject the null hypothesis. In that case, an appropriate level of significance is chosen
and H0 is further tested. We will understand the usage of level of significance now. Later we will
discuss the meaning of this term in details.

Test of significance of difference of proportions:


Often we select different samples from different domains. They initially may reflect to differ from
one another. We may say that the two samples are reflections of two different populations. But
sometimes this may not be the case too. May be the two samples are reflections of same population
but apparently they seem to vary from one another. This situation will be studied here. We design
the hypothesis on the difference of two population proportions P1 and P2 . If the sample proportions
are p1 and p2 , the test statistic Z is defined as

(p1 − p2 ) − E(p1 − p2 )
Z= , (84)
S.E(p1 − p2 )
q
where E(p1 − p2 ) = P1 − P2 and V(p1 − p2 ) = P1nQ 1
1
+ P2nQ
2
2
.

Example: Two insect sprays are to be compared. Two rooms of equal size are sprayed, one
with spray 1 and the other with spray 2. Then 100 insects are released in each room, and after 2
hours the dead insects are counted. Suppose the result is 64 dead insects in the room sprayed with
spray 1 and 52 dead insects in the other room. Is the evidence significant enough for us to reject,
at the 5% level, the hypothesis that the two sprays have equal ability to kill insects?20

Sol: We define null and alternative hypothesis are as follows. H0 : There is no significant dif-
ference between the two sprays i.e. P1 = P2 against H1 : P1 ̸= P2 (the test is two tailed). Here
P1 is the population proportion of population 1 and P2 is the population proportion of population
2. What do we mean by these population proportions? Actually we are assuming that the spray
1 is coming from a population of spray 1 manufactured in a factory and spray 5 is coming from a
population of spray 2 manufactured in another factory and that the ingredients of the two sprays
are equally effective to kill the insects. The sample proportions are p1 = X X2
n1 and p2 = n2 , where
1

X1 is the number of dead insects due to spray 1 and the X2 is the number of dead insects due to
spray 2. Here X1 = 64 and X2 = 52, and n1 = n2 = n = 100. As per our hypothesis P1 − P2 = 0
and therefore we get

Z= q p1 −p2 ,
1 1
P Q( n +n )

where Q1 = 1 − P1 and Q1 = 1 − P1 whereas we have assumed that P1 = P2 = P (say). Since the


value P is hypothetical we will estimate it by sample proportion p = P̂ . Hence we get
X1 +X2
P̂ = n+n = 0.58 and accordingly Q̂ = 1 − P̂ = 0.42.

Thus
18
This fallacy is appearing as sufficient data has not been taken into consideration.
19
We will discuss about error in later section.
20
5% level of significance will be discussed later.

47
Z= q p1 −p2 = q 0.64−.52
= 1.719.
1 1 1 1
P̂ Q̂( n +n ) 0.58×0.42×( 100 + 100 )

We see Z < 1.719. Since we need to study at 5% level of significance, and at this level of significance
we are supposed to have |Z| ≤ 1.96 for acceptance of null hypothesis, hence we can immediately
conclude that the here in this problem H0 is accepted at 5% level. Therefore the two sprays have
same effects in killing the insects. In other words the two sprays, though apparently differ in their
appearance, are part of same population.

Test of significance of single mean:


When we want to test a sample to draw conclusion about population mean then, for large sample,
Z test is used. Suppose a sample is drawn from an infinite population having mean µ and standard
deviation σ (provided σ is known), then it is easy to verify that E(X) = µ and SD(X) = √σn (by
central limit theorem, see eq.(81)) so that
X − E(X) X −µ
Z= = , (85)
S.E.(X) √σ
n
provided σ is known.

Example: An ambulance service claims that it takes on the average 8.9 minutes to reach its
destination in emergency calls. To check on this claim, the agency which licenses ambulance ser-
vices has them timed on 50 emergency calls, getting a mean of 9.3 minutes with a standard deviation
of 1.6 minues. What can they conclude at the level of significance α = 0.05?

Sol: The sample mean X = 9.3 and population mean µ = 8.9 while σ = 1.6. We define the
null hypothesis as H0 : µ = 8.9 against µ ̸= 8.9 (two tailed test). Under H0
X−µ 9.3−8.9 0.04
Z= √σ = 1.6

= 0.226 = 0.176.
n 50

We see that Z < 3, we fail to reject the null hypothesis. Now it is known that, at 5% level of
significance, i.e. at α = 0.05, if Z < 1.96 we accept the null hypothesis which is indeed the case
here.

Test of significance of difference of means:


Similar to the testing of difference of population proportions, we can also test the difference popu-
lation means. For large sample, the test statistic is defined as
X̄1 − X̄2 − E(X̄1 − X̄2 )
Z =
S.E(X̄1 − X̄2 )
X̄1 − X̄2 − (µ1 − µ2 )
= q 2 . (86)
σ1 σ22
n1 + n2
Now if we draw samples from two different populations having different means and same standard
deviation (say σ), the eq.(86) reduces to
X̄1 − X̄2
Z =
S.E(X̄1 − X̄2 )
X̄ − X̄2
= q 1 . (87)
σ 2 ( n11 + n12 )

48
Now there may arise two cases, viz. (i) we have idea about population standard deviation σ, (ii)
we do not have idea about population standard deviation σ. In case (i) we can directly use eq.(87)
while in case (ii) we estimate σ by the following expression.

(n1 − 1)S12 + (n2 − 1)S22


σˆ2 = . (88)
n1 + n2 − 2

For large sample n1 − 1 → n1 and n2 − 1 → n2 , while S12 is approximated by s21 and S22 is
approximated by s22 , so that we have

n1 s21 + n2 s22
σˆ2 = . (89)
n1 + n2
21 Again, there may be case where we have an idea about the two population standard deviations
to be different but we do not have any clue about their values.

Example: A sample of heights of 6400 Englishmen has a mean of 67.85 inches and standard
deviation 2.56 inches, while a sample of heights of 1600 Australians has a mean of 68.55 inches and
standard deviation of 2.52 inches. Do the data indicate that Australians are, on the average, taller
than Englishmen?

Sol: In this study two populations have been considered, viz. population of Englishmen and
population of Australians. We are studying that whether there is any significant difference between
the heights of the two populations. The mean heights of the populations of Australians and En-
glishmen are µ1 and µ2 (say). We define H0 : µ1 = µ2 against H1 : µ1 > µ2 (as per question, it is a
one tailed test). About the variability of the two populations, nothing is mentioned. So we assume
first that variability of Australian population (σ1 is not known) and that of Englishmen population
(σ2 is not known) and also we assume σ1 = σ2 = σ (say), so that we use eq.(88) and therefore

Z= q X̄1 −X̄2 = q 68.55−67.85


= 9.81
σˆ2 ( n1 + n1 ) 1
6.512×( 1600 1
+ 6400 )
1 2

We see that Z > 3 and so null hypothesis is rejected. Hence Australians are on an average taller
than the Englishmen.

Now if we assume that σ1 ̸= σ2 , so we can use eq.(89) and we get

Z= q68.55−67.85 = 9.906,
2.522 2
1600
+ 2.56
6400

which also shows that as Z > 3, H0 is accepted.

Test of significance of single standard deviation:


When X1 , X2 , · · · , Xn is a random sample from a normal distribution and n is large, the sam-
σ2
ple standard deviation has approximately a normal distribution with mean σ and variance 2n .
Therefore, a large sample test for H0 : σ = σ0 can be based on the statistic
S − σ0
Z= q 2 (90)
σ0
2n
21
S2 = 1
− X), s2 = 1
P P
n−1 i (Xi n i (Xi − X)

49
22 .Example: The United States Golf Association tests golf balls to ensure that they conform to
the rules of golf. Balls are tested for weight, diameter, roundness and overall distance. The overall
distance test is conducted by hitting balls with a driver swung by a mechanical device nicknamed “
Iron Byron” after the legendary great Byron Nelson, whose swing the machine is said to emulate.
Following are 100 distances (in meters) achieved by a particular brand of golf ball in the overall
distance test. Test H0 : σ = 10 versus H1 : σ < 10.

Sol: Using the above table we see that S = 12.0667 (S is the sample standard deviation). Given
σ0 = 10. Using eq.(90) we get
12−10
Z= 2 = 0.2828.
√ 10
2×100

We see that for the choice of α = 0.05, Z < 1.96, therefore we fail to reject the null hypothesis.

Test of significance of difference of standard deviations:


To test whether there is any significant difference between the population standard deviations we
use the following Z statistic.

(S1 − S2 ) − (σ1 − σ2 )
Z= q 2 . (91)
σ1 σ22
2n1 + 2n2

Here, we have assumed that H0 : There is so significant difference between the population standard
deviations.

In the above problem of Englishmen and Australian, if we assume that there is no significant
difference between the population variabilities i.e H0 : σ1 = σ2 against H1 : σ1 ̸= σ2 , then using
eq.(91) we have

Z= r S1 −S2 ,
S12 S2
2n1
+ 2n2
2

since σ1 and σ2 are not known. Therefore we get


22
This test is not mentioned in most of the Statistics’ books. The said test has later been found in Montgomery’s
“Applied Statistics and Probability for Engineers”.

50
Z= q 2.52−2.56 = −0.566.
2.522 2
1600
+ 2.56
6400

Therefore we get |Z| < 3. We fail to reject the null hypothesis and further test is needed. We
know at 5% level of significance, |Z| ≤ 1.96 for hypothesis to be accepted, which is indeed the case
here. Hence we conclude that there is no variability in the heights of the two populations of En-
glishmen and Australian although the mean heights of Australian is higher than that of Englishmen.

As we have seen in the discussion of the above problems, the working principles of testing of
hypothesis for large sample test, can thus be summarized as follows:

ˆ Design H0 (null hypothesis) properly.

ˆ Choose appropriate test statistic (in this case Z test).

ˆ If |Z| > 3 we reject the null hypothesis, whereas when |Z| ≤ 3 we test the hypothesis for
some specified level of significance (5% and 1% most of the times)23

ˆ Check whether the test is a two tailed test or one tailed test.

ˆ (a) For two tailed test, at 5% level of significance, null hypothesis is accepted when |Z| ≤ 1.96
and at 1% level of significance, hypothesis is accepted when |Z| ≤ 2.58. (b) For one tailed
test, at 5% level of significance, null hypothesis is accepted when |Z| ≤ 1.645 and at 1% level
of significance, hypothesis is accepted when |Z| ≤ 2.33.

In the next section we are going to design another type of hypothesis where we want to conclude
about the hypothetical value of the variance of a single population. This type of testing requires
χ2 as test statistic.

7.2 Errors:
In hypothesis testing two types of errors may creep in. One is called Type-I error and another is
called Type-II error.

Type-I error:
Rejecting the null hypothesis H0 when it is true is called type-I error.

Type-II error:
Failing to reject the null hypothesis H0 when it is false is called type-II error.

In this respect, the following table will be helpful.

Decision H0 is True H0 is False


Fail to reject H0 No error Type - II error
Reject H0 Type - I error No error

The above table can be understood in various scientific communities as follows. False alarms are
a sub-category of errors known as false positives. A false positive is a test result that indicates
that a particular condition or attribute is present when actually it is not. Typically, false positives
23
It is better to calculate p− value.

51
occur in binary tests. These are the tests with two possible outcomes - positive or negative. In the
context of medical tests, false positives result in people who are not sick being told that they are.
In the courtroom drama, false positives are the innocent people convicted of crimes they did not
commit. Let us figure out another table.
Predicted Condition True condition (Positive) True condition (negative)
Positive True positive False positive
Negative True negative True negative
If one compares above two tables, one finds that false positive is equivalent to type-II
error while true negative is equivalent to type-I error. Now, the probabilities of both
these errors can be calculated.

Probability of type-I error:


Probability of committing type-I error i.e. probability of rejecting null hypothesis when it is actually
true is called significance level or level of significance and it is denoted by α. It is also called
size of the test or α error. Mathematically,
α = P{reject H0 |H0 is true}. (92)
A widely used procedure in hypothesis testing is to use a type-I error or significance level of α = 0.05.
This value has evolved through experience and may not be appropriate for all situations.

Probability of type-II error:


Probability of committing type-II error is the probability of failing to reject null hypothesis when
it is false. Mathematically,
β = P{f ail to reject H0 |H0 is f alse}. (93)
The quantity 1 − β is called power of the test and can be interpreted as the probability of
correctly rejecting a false null hypothesis.

P value
The P-value is the smallest level of significance that would lead to rejection of the null hypothesis
H0 with the given data.

It is customary to consider the test statistic (and the data) significant when the null hypothe-
sis H0 is rejected; therefore, we may think of the P − value as the smallest level α at which the
data are significant. In other words, the P value is the observed significance level. Once the P −
value is known, the decision maker can determine how significant the data are without the data
analyst formally imposing a pre-selected level of significance.

χ2 test:
Definition: If X has the standard normal distribution, then X 2 has the chi-square distribution
with ν = 1 degree of freedom. If X1 , X2 , · · · , Xn are independent random variables having standard
normal distributions, then
n
X
Y = Xi2 (94)
i=1

52
has the chi-square distribution with ν = n degrees of freedom.

Theorem: If X1 , X2 , · · · , Xn are independent random variables having chi-square distributions


with ν1 , ν2 , · · · , νn degrees of freedom, then
n
X
Y = Xi (95)
i=1

has the chi-square distribution with ν1 + ν2 + · · · + νn degrees of freedom.

Theorem: If X1 and X2 are independent random variables, X1 has a chi-square distribution


with ν1 degrees of freedom, and X1 + X2 has a chi-square distribution with ν > ν1 degrees of
freedom, then X2 has a chi-square distribution with ν − ν1 degrees of freedom.

Theorem: If X and S 2 are the mean and the variance of a random sample of size n from a
normal population with the mean µ and the standard deviation σ, then

ˆ X and S 2 are independent.


(n−1)S 2
ˆ The random variable σ2
has a chi-square distribution with n − 1 degrees of freedom.

Since the chi-square distribution arises in many important applications, integrals of its density have
been extensively tabulated24 . It can be seen in the table of χ2 , that, it contains values of χ2α,ν for
α = 0.995 , 0.99 , 0.975 , 0.95 , 0.05 , 0.025 , 0.01 , 0.005 , and ν = 1, 2, · · · , 30, where, χ2α,ν is such
that the area to its right under the chi-square curve with ν degrees of freedom is equal to α. That
is, χ2α,ν is such that if X is a random variable having a chi-square distribution with ν degrees of
freedom, then,

P(X ≥ χ2α,ν ) = α. (96)

When ν is greater than 30, Statistical Table of χ2 test cannot be used and probabilities related to chi-
square distributions are usually approximated with normal distributions (Fisher’s approximation)25 ,
the approximation is given by the formula
p √
Z = 2χ2 − 2n − 1. (97)

24
Chi-square table can be found in good book on Statistics.
25
So we can say χ2 is used for small sample and for n > 30, Z test is applicable

53
Example: (i) A sample of 15 observations shows that standard deviation is 6.4. Is this com-
patible with the hypothesis that the sample is from a normal population with standard deviation
5? (ii) A sample of 50 observations shows that standard deviation is 6.4. Is this compatible with
the hypothesis that the sample is from a normal population with standard deviation 5?

Sol: (i) We define H0 : σ = 5 against H1 : σ ̸= 5 (σ is the population standard deviation).


On the basis of this hypothesis, the test statistic chosen is χ2 and we have
(n−1)S 2 (15−1)×6.42
χ2 = σ2
= 52
= 22.93.
Comparing with table χ2α=0.05,ν=15−1 = 23.685 (i.e. the tabulated value of χ2 ). This is greater than
the calculated value and so we accept the null hypothesis.

(ii) Similarly for sample size n = 50, which is large sample, we get
(50−1)×6.42
χ2 = 52
= 80.2816.
We cannot draw conclusion at this stage. So we apply eq.(97) and consequently we get
√ √
Z = 2 × 80.2816 − 2 × 50 − 1 = 2.7214.
The given test is two tailed and choosing 5% level of significance we see |Z| = 2.7214 > 1.96 ,
subsequently rejecting null hypothesis.

Problem: A manufacturer recorded the cut-off bias (volt) of a sample of 10 tubes as follows:
12.1, , 12.3, 11.8, 12.0, 12.4, 12.0, 12.1, 11.9, 12.2, , 12.2. The variability of cut-off bias for tubes of
a standard type as measured by the standard deviation is 0.208 volts. Is the variability of the new
tube with respect to cut-off bias less than that of the standard type?

Goodness of fit test:


The goodness-of-fit test considered here applies to situations in which we want to determine whether
a set of data may be looked upon as a random sample from a population having a given distribution.
To test the null hypothesis that the observed frequencies constitute a random sample from a certain
population, we must judge how good a fit, or how close an agreement, we have between the two
sets of frequencies. In general, to test the null hypothesis H0 that a set of observed data comes
from a population having a specified distribution against the alternative that the population has
some other distribution, we compute
m
X (Oi − Ei )2
χ2 = . (98)
Ei
i=1

and reject H0 at the level of significance α if χ2 > χ2α, m−t−1 , where m is the number of terms in the
summation and t is the number of independent parameters estimated on the basis of the sample
data. Here, Oi is the observed frequency and Ei is the expected frequency. It is to be remembered
that
X X
Oi = Ei . (99)
i i

Example: Self progenies of a cross between pure strains of plant segregated as follows;

54
P rocess
V ariety Early flowering Late flowering
Tall 120 48
Short 36 13

Do the results agree with the theoretical frequencies which specify a 9 : 3 : 3 : 1 ratio?

Sol: We define H0 : There is no significant difference between the observed and expected fre-
quencies, against H1 : there is. Under H0 , we are going to use the χ2 test defined in eq.(98). The
observed frequencies are given in the question which we summarize once again below.
P rocess
V ariety Early flowering Late flowering Total
Tall 120 48 168
Short 36 13 49
Total 156 61 Grand Total = 217

Since theoretical frequencies are in the ratio 9 : 3 : 3 : 1, hence expected frequencies are
9
E(cell1, cell1) = 16 × 217 = 122.0625,
3
E(cell1, cell2) = 16 × 217 = 40.6875,
3
E(cell2, cell1) = 16 × 217 = 40.6875,
1
E(cell2, cell2) = 16 × 217 = 13.5625,

sum total of which is again 217 satisfying eq.(99). We form the table below.
(Oi −Ei )2
Oi Ei (Oi − Ei )2 Ei
120 122.0625 4.2539 0.0348
48 40.6875 53.4726 1.3142
36 40.6875 21.9726 0.5400
13 13.5625 0.3164 0.0233
2
Now using eq.(98), we get i (Oi −E i)
= 1.9123. Now χ2α,ν = χ2α,m−t−1 = χ20.05,4−2−1 = χ20.05,1 .
P
Ei
Here m = 4 since the summation is over 4 terms and t = 2. Now from the table χ20.05,3 = 7.815.
We see that χ2 = 1.9123 < 7.815 and we accept the null hypothesis. Our final conclusion would be
that the results agree with the theoretical frequencies.

Note: Remember, in the above example when we calculated expected frequencies, we multiplied
number of observations by probability which we do in case of binomial distribution. In fact this is
the mean of the binomial distribution. Hence the value of t is 2, since binomial distribution has
two parameters. Therefore, for the above problem, we can redesign our null hypothesis as H0 : The
binomial distribution can be fitted to the given data against H1 : Binomial law cannot be fitted to
the data.

Example: Consider the following frequency table of observations on the random variable X.

Values 0 1 2 3 4
Observed frequency 24 30 31 11 4

Based on these 100 observations, is a Poisson distribution with a mean of 1.5 an appropriate model?
Perform a goodness of fit procedure with α = 0.05.

55
Sol: We assume H0 : The Poisson distribution can be fitted against that H1 : Poisson law cannot be
fitted. Under the assumption of null hypothesis, we calculate the expected frequencies by Poisson
probability law defined in eq.(44). Given λ = 1.5, we can design the following chart.
−λ x (Oi −Ei )2
Oi Ei = 100 e x!λ (Oi − Ei )2 Ei
−1.5 0
24 100 e 0!1.5 = 22.31 2.8561 0.1281
−1.5 1
30 100 e 1!1.5 = 33.46 11.9716 0.3578
−1.5 2
31 100 e 2!1.5 = 25.10 34.81 1.3869
−1.5 3
11 100 e 3!1.5 = 12.55 157.5025 12.55
−1.5 4
4 100 e 4!1.5 = 4.71 22.1841 4.71

See that two calculate the expected frequencies we have multiplied the Poisson probability density
function by 100 (the sample size). Now using eq.(98), we get

χ2 = 19.1328.

Since summation is over 5 terms, so m = 5. As we have assumed that Poisson law can be fitted,
and as there is only one parameter for Poisson distribution, hence t = 1, so that using eq.(99) we
get χ20.05,5−1−1 = χ20.05,3 = 7.815. We conclude that Poisson data cannot be fitted as H0 is rejected
as χ2 > χ20.05,3 .

Note: χ2 -distribution, is a continuous single-parameter distribution derived as a special case of


the gamma distribution; it is used especially to measure goodness of fit, and to test hypotheses and
obtain confidence intervals for the variance of a normally distributed random variable.

Remember, that earlier we wrote the probability density function of chi-square variate, which
we summarize once again below.
t−2 x
x 2 e− 2
χ2 (x) = t .26 (100)
2 2 Γ( 2t )

In eq.(100), t denotes the degrees of freedom. In the following, we plot the chi-square probability
curve, set for several degrees of freedom such as t = 1, 2, 3, 4, 5, 6. We see that the curves are skewed

56
from normality.

Note: If any probability distribution is, thus, found to be fitted to the population from which
the sample has been drawn, then hypothesis testing to be performed is called parametric test,
otherwise non-parametric test 27 .

Small sample test (t test):


If Y and Z are independent random variables, Y has a chi-square distribution with ν degrees of
freedom, and Z has the standard normal distribution, then the distribution of
Z
T =q (101)
Y
ν

is given by

Γ( ν+1 ) t2 ν+1
f (t) = √ 2 ν (1 + )− 2 , −∞ < t < ∞ (102)
πνΓ( 2 ) ν

and it is called the t− distribution28 with ν degrees of freedom. Sometimes, t distribution is also
known as Student’s t distribution. In view of its importance, the t distribution was extensively
tabulated and contains values for tα,ν , for α = 0.10, 0.05, 0.025, 0.01, 0.005 and ν = 1, 2, · · · , 29
where ν symbolizes degrees of freedom and α stands for degrees of freedom. tα,ν is such that the
area to its right under the curve of the t distribution with ν degrees of freedom is equal to α. This
means

P(T ≥ tα,ν = α), (103)

shown in figure below. The table does not contain values of tα,ν for α > 0.50, since the density

Figure 17: t distribution

is symmetrical about t = 0 and hence t1−α,ν = −tα,ν . When ν ≥ 30, probabilities related to the
t distribution are usually approximated with the use of normal distributions. Among the many
applications of the t distribution, its major application (for which it was originally developed) is
based on the following theorem.
27
Some examples of non-parametric tests are sign test, Kruskal-Wallis test [Link].
28
The t distribution was introduced originally by W.S. Gosset, who published his scientific writings under the pen
name student, since the company for which he worked, a brewery, did not permit publication by employees.

57
Theorem: If X and S 2 are the mean and the variance of a random sample of size n from a
normal population with the mean µ and the variance σ 2 , then

X −µ
T = , (104)
√S
n

has the t distribution with n − 1 degrees of freedom.

Case I:
When sample size is small i.e. (n < 30), population variance (σ 2 ) is unknown and we are interested
in testing the hypothetical value of population mean, then we go for t test. The test statistic is
given by

X − µ0
T = , (105)
√S
n

or we can write the above as


x¯0 − µ0
t= s0 , (106)

n

where x¯0 is the value of X and s0 is the value of S. t in eq.(106) is a value of a random variable
having the t distribution with n−1 degrees of freedom. Thus critical regions of size α for testing the
null hypothesis µ = µ0 against µ ̸= µ0 , µ > µ0 or µ < µ0 are respectively |t| ≥ t α2 ,n−1 , t ≥ tα,n−1
and t ≤ −tα,n−1 . Such a test is called one sample t test.

Example: The specifications, for a certain kind of ribbon, call for a mean breaking strength
of 185 pounds. If five pieces randomly selected from different rolls have breaking strengths of
171.6, 191.8, 178.3, 184.9, and 189.1 pounds, test the null hypothesis µ = 185 pounds against the
alternative hypothesis µ < 185 pounds at the 0.05 level of significance.

Sol: We set up the null hypothesis as follows: H0 : µ = 185 against H1 : µ < 185. From the
given data x¯0 = 183.14 and s0 = 8.2191 (for calculation of s0 we have divided the sum of the
squares of the deviations from the mean value by n − 1 (n = 5).) Using eq.(106) we thus get
t = 183.14−185
8.2191

i.e. t = −0.506. From table we find t0.05,5−1 = 2.132. As per rule defined above we
5
therefore see that t = −0.506 > −2.132. Hence we fail to reject the null hypothesis.

Case II:
When we deal with two samples of sizes n1 and n2 , the individual samples being small in size,
and σ1 , σ2 are population variances (from which the two samples are supposed to be respectively
drawn) that are unknown, then two test the null hypothesis µ1 − µ2 = µ against that it is not, we
use the following test statistic (where µ1 and µ2 are population means).

(x¯1 − x¯2 ) − (µ1 − µ2 )


t= q , (107)
sp n11 + n12

58
(n −1)s2 +(n −1)s2
where s2p = 1 n1 +n 1 2
2 −2
2
, where s21 and s22 are the values of S12 and S22 respectively. Under the
given assumptions and the null hypothesis µ1 − µ2 = µ, this expression for t is a value of a random
variable having the t distribution with n1 +n2 −2 degrees of freedom. Thus, the appropriate critical
regions of size α for testing the null hypothesis µ1 − µ2 = µ against the alternatives µ1 − µ2 ̸= µ,
µ1 − µ2 > µ, or µ1 − µ2 < µ under the given assumptions are, respectively, |t| ≥ t α2 ,n1 +n2 −2 ,
t > tα,n1 +n2 −2 or t < −tα,n1 +n2 −2 . Such a test is called two sample t test.

Example: Below are given the gain in weights (in pounds) of pigs fed on two diets A and B.
Diet A 25 32 30 34 24 14 32 24 30 31 35 25 — — —
Diet B 44 34 22 10 47 31 40 30 32 35 18 21 35 29 22
Test if the two diets differ significantly as regards their effects on increase in weight.

Sol: We set up null hypothesis as H0 : µ1 = µ2 i.e. there is no significant difference between


the mean increase in weight due to diets A and B, against H1 : µ1 ̸= µ2 . The test statistic in
eq.(107) thus takes the following form.
x¯ − x¯2
t= q1 . (108)
sp n11 + n12

The sample sizes for diets A and B P


are respectively n1 = 12 and n2 = 15 respectively. We know
1 P
that S = n−1 (X − X) so that (X − X)2 = (n − 1)S 2 . Now as s1 and s2 are values of S1
2 2

and S2 so calculation of i (xi − X)2 for both the data of diets A and B give us 380 (for diets A)
P
and 1410 (for diets B) (Calculate!). Thus putting the values in Sp2 we get s2p as
1
s2p = 12+15−2 (380 + 1410) = 71.6.
Also X̄1 = 28 and X̄2 = 30 are respectively the means of diets A and B data. Putting these values
in eq.(108), we have
28−30
t= q
1 1
= −0.6103.
8.46( 12 + 15 )

As per the rule described above we take |t| = 0.6103 (since the test is two tailed). Consider level
of significance α = 0.05 so that α2 = 0.025. We find from table t0.025,12+15−2 = 2.06. We see that
|t| < t0.025,12+15−2 and hence we accept the null hypothesis.

Case III:
There is one another arena where t test can be applied. In Paired t test for difference of means,
we consider the case (i) when sample sizes are equal i.e. n1 = n2 = n and the two samples are not
independent but the sample observations are paired together, i.e. the pair of observations (xi , yi )
(i = 1, 2, · · · , n) corresponds to the same ith sample unit. The problem is to test if the sample means
differ significantly or not (while the observations were taken twice). Let xi , yi (i = 1, 2, · · · , n) be
the two readings on ith observation, so that di = xi − yi . The test statistic used in this case is

t= , (109)
√S
n

where d¯ = n1 ni=1 di and S 2 = n−1


1 P ¯2
P
i (di − d) , follows t distribution with n − 1 degrees of freedom.
Thus, the appropriate critical regions of size α for testing the null hypothesis µ1 − µ2 = µ against

59
the alternatives µ1 − µ2 = ̸ µ, µ1 − µ2 > µ, or µ1 − µ2 < µ under the given assumptions are,
respectively, |t| ≥ t 2 ,n−1 , t > tα,n−1 or t < −tα,n−1 .
α

Example: Eleven school boys were given a test in Statistics. They were given a month’s tu-
ition and a second test was held at the end of it. Do the marks give evidence that the students
have benefited by the extra coaching?
Boys 1 2 3 4 5 6 7 8 9 10 11
Marks in 1st test 23 20 19 21 18 20 18 17 23 16 19
Marks in second test 24 19 22 18 20 22 20 20 23 20 18
Sol: Let µ1 and µ2 be the mean marks respectively of the students in Statistics examination before
they took coaching and after they went through coaching. We assume that there is no significant
improvement in the marks i.e. H0 : µ1 = µ2 against H1 : µ2 > µ1 , (which implies coaching is
effective).29 . It is a two tailed test. The sample observations on marks in 1st test are denoted by
xi and that of 2nd test are denoted by yi . We thus get the following table:
Boys 1 2 3 4 5 6 7 8 9 10 11
di -1 1 -3 3 -2 -2 -2 -3 0 -4 1
Now d¯ = −12
12 = −1.09.

Boys 1 2 3 4 5 6 7 8 9 10 11
¯2
(di − d) 0.0081 4.3681 3.6481 16.7281 0.8281 0.8281 0.8281 3.6481 1.1881 8.6481 4.3681
Therefore S 2 = 44.9091. Thus using eq.(109),
−12
t= 6.7014

= −5.9390.
11

By the rule defined above, |t| = 5.9390. We select the level of significance α = 0.05. Now |t| ≥
t α2 ,n−1 = 2.228. The null hypothesis is rejected. This implies that there is improvement in the
grades and hence we can conclude that the coaching is effective in this case.

Test of significance of Zero Correlation:


Suppose ρ be the correlation coefficient of bivariate data taken from a population and suppose we
want to test the hypothesis that H0 : ρ = 0 against H1 : ρ ̸= 0. This means, we would like to test
that if a bivariate sample data is collected from the population and we observe that the variables
are correlated with one another then in reality the variables under study in the population are not
correlated against that indeed they are correlated. The test statistic used in this case is

R n−2
T0 = √ , (110)
1 − R2
following t distribution with n − 2 degrees of freedom and R denoting the correlation coefficient in
the sample data. Thus, we would reject the null hypothesis if |t0 | > t α2 ,n−2 .

Example: The following data gave X which is the water content of snow on April 1 and Y
representing the yield from April to July (in inches) on the Snake river watershed in Wyoming
for 1919 to 1935. (This data were taken from an article in Research Notes vol. 61, 1950), Pacific
Northwest Forest Range Experiment Station, Oregon).
29
You cannot write H1 : µ2 < µ1 as every coaching is supposed to improve the grade and we would expect that.

60
x 23.1 32.8 31.8 32.0 30.4 24.0 39.5 24.2 52.5
y 10.5 16.7 18.2 17.0 16.3 10.5 23.1 12.4 24.9
x 37.9 30.5 25.1 12.4 35.1 31.5 21.1 27.6 —
y 22.8 14.1 12.9 8.8 17.4 14.9 10.5 16.1 —

(a) Estimate the correlation coefficient between Y and X. (b) Test the hypothesis that ρ = 0, using
α = 0.05.

Sol: Using eq.(69), we can easily compute rXY = 0.9332 (compute!), which is the sample corre-
lation coefficient between the variables X and Y (the random variables which take values like x
and y as shown in question). In this problem, we denote this sample correlation by R. We set
up the null hypothesis as follows: H0 : The random variables X and Y in the population are not
correlated at all i.e. ρXY = 0 against ρXY ̸= 0. Under the assumption of the null hypothesis, using
eq.(110) we have

0.9332× (17−2)
T0 = t0 = √1−0.93322 = 10.05730 .

Now for α = 0.05 we see that |t| = 10.057 > t α2 ,n−2 = t0.025,15 = 2.131 (from table).

7.3 Level of significance


After all these discussions mentioned above it is now time to understand the meaning of the term
level of significance which is indeed a kind of probability. Actually while doing testing of hypothesis,
two kinds of error may creep in. Suppose you have designed a hypothesis H0 which is actually true
and that you have accepted it, then there is obviously no error that you have committed. Similar
is the case with the situation where H0 is false and you have to reject it, then also there is no error.
But think of the situation where, the null hypothesis is wrong and for some reason you fail to reject
it or the null hypothesis is true but you are unable to accept it. These two situations will result in
two different types of errors.

Hypothesis H0 Status Error


False You fail to reject it Type - II
True Reject it Type - I

Remember, that the phrase fail to reject is not synonymous to the word accept. The probability of
committing Type - I error is called significance level or α error or size of the test. Let us try
to understand this concept by citing an example.

Let us suppose that an engineer is designing an air crew escape system that consists of an ejection
seat and a rocket motor that powers the seat. The rocket motor contains a propellent, and in order
for the ejection seat to function properly, the propellent should have a mean burning rate of 50
cm/sec. If the burning rate is too low, the ejection seat may not function properly, leading to an
unsafe ejection and possible injury to the pilot. Higher burning rates may imply instability in the
propellent or an ejection seat that is too powerful, again leading to possible pilot injury. So, the
practical engineering question would be to ask, does the mean burning rate of the propellent equals
50 cm/sec or is it some other higher or lower value?

It is clear that the above problem is about testing of population mean burning rate, which we
30
t0 is just the notation denoting specific value of t statistic T0 whose value in this problem is 0.9332 here.

61
shall denote by µ = 50. Here we shall learn how to calculate the probability of type - I error.
So we will do certain things. First we assume that we have collected a sample of size n = 30
and the sample mean burning rate (X) lies between 48.5 and 52.5 (which we want as it should
be)31 . This means that if you run the sample survey several times then propellent burning rate
will always fluctuate between these two values. This means that getting exactly the value similar
to the population mean is a fluke, whose happening chance is too feeble in practical sense. Also we
assume that population standard deviation σ is 2.9. This means that the population mean takes
the value of 50 with standard error of ±2.9. If our sample survey reveals that the sample mean is
falling between our defined range, we fail to reject the null hypothesis which is H0 : µ = 50. But if
X < 48.5 or X > 52.5, then we will reject H0 (with the condition that the null hypothesis is true).
For our calculation we have taken a sample of size n > 30 to make it a large sample test. As we
denote the level of significance by α then by definition

α = P(P rob of T ype − I error) = P(reject H0 |H0 is true). (111)

Using the Eq.(111) we are supposed to find,

α = P((X < 48.5|µ = 50) ∪ (X > 52.5|µ = 50)


= P(X < 48.5|µ = 50) + P(X > 52.5|µ = 50). (112)
X−µ
The appropriate Z statistic is Z test for single mean whose formula is Z = √σ and X is a normal
n
variate which will be standardized using central limit theorem. Thus we ultimately get

α = P(Z < 3.13) + P(Z > 5.2). (113)


32 The probability P(Z > 5.2) is zero. and we are supposed to find the probability P(Z < 3.13)
only. That is, you want to calculate the probability of the yellow shaded region (which is the

area under the curve from ∞ to 3.13). From standard normal distribution chart the probability is
0.999126 i.e.α = 0.99 approx. This implies that there is 99% chance of committing type-I error if
you repeat the above mentioned experiment. Moreover, it is clear from the calculation that |Z| > 3
which forces us to reject H0 (with respect to the given data). What does the analysis mean? This
means that on the basis of the value of test statistic the null hypothesis is rejected implying that
we are not accepting the fact that the population mean burning rate is 50 cm/sec. But if indeed
the propellent burning rate is µ = 50, then we might committed the type-I error with probability
31
This is our choice for the sake of calculation.
32
Do it by yourself.

62
0.99. So probably we need to increase the sample size or adjust some other issues.

Exercise: In a random sample of size 400 there are 80 defective items. Test the significance
at 5% level whether the proportion of defective items in the population may be regarded as 16 .
Standard normal value is given by ϕ(1.96) = 0.475.

Sol: Here H0 : P = 1
: P ̸= 1 q p−P
6 , H1 6, Z = P (1−P )
= 1.77 < 1.96. Therefore null hypothesis
n
will be accepted at 5% level of significance.

Problem:
(i) 20 post-graduate students are randomly selected from a college and their average height
comes to 170 cms with a standard deviation of 3.2 cms. On the other side, 22 under-graduate
students are randomly selected and their average height comes to 168 cms. with a standard
deviation of 6.4 cms. Assuming the distribution of height to be normally distributed with
equal population variances, do you think the post-graduate students are taller than under-
graduate students? Use 0.05 level of significance. (Given that t0.05,40 = 1.684)
(ii) Two independent samples of size 8 and 10 are drawn from population. It is found that the
mean of sample 1 is 22.5 with standard deviation 3.9 and mean of sample 2 is 24.8 with
standard deviation 4.6. Test whether the estimated population variances differ significantly
at 0.1 level of significance. (Given that f0.05 (7, 9) = 3.29, f0.05 (9, 7) = 3.68).
Sol:
q
(n −1)s2 +(n −1)s2
1 2
(i) H0 : µ1 = µ2 ; H1 : µ1 > µ2 ; S = 1
n1 +n2 −2
2
= 26.368.
x¯ − x
¯
t= q 1
1
2
1
= 1.26 < tcritical (1.684). Hence we conclude that there is no difference.
S n1
+n
2
s21
(ii) H0 : σ12 = σ22 ; H1 : σ12 ̸= σ22 ; Test statisticf = s22
= 0.719,
df1 = 8 − 1 = 7, df2 = 10 − 1 = 9, f0.05 (7, 9) = 3.29, f0.95 = f0.051(9,7) = 1/3.68 = 0.271
Since 0.271 < 0.719 < 3.29. Since the F-statistic falls within the range of the critical values, we
fail to reject the null hypothesis. Exercise: In a large lot of electric bulbs, the mean life and
standard deviation of the bulbs are 365 hours and 92 hours respectively. A sample of 630 bulbs
is chosen. It is obtained that mean life and standard deviation of the bulbs are 360 and 91 hours
respectively. Can we conclude that the sample is drawn from the given population? Test at 5%
level of significance.

Sol: Test null hypothesis H0 : µ = 365 against the alternate hypothesis H1 : µ ̸= 365, The test
x̄−µ
√ = 360−365
statistic z = σ/ n

92/ 630
= −1.364. Thus |z| = 1.364 < 1.96. Therefore we fail to reject null
hypothesis. So we may conclude that the sample is drawn from the population.

Problem:
(i) Certain pesticide is packed into bags by machine. A random sample of 10 bags are drawn and
weights (in kg.) of the peesticides in these bags are found as follows: 50, 49, 52, 44, 45, 48,
46, 45, 49, 45. Test whether the average packing can be taken to 50kg at the 5% level of sig-
nificance. Given that t0.025,9 = 2.26.

63
(ii) A random sample of 20 families in a city showed the average monthly expenditure towards
entertainment is Rs. 2040/- with a standard deviation of Rs. 200/-. Another random sample
of 16 families in another city showed the average expenditure as Rs. 2165 with a standard
deviation of Rs. 250/-. Test whether the difference between their average expenditure is sig-
nificant or not at 0.01 level of significance. Assume that the population variances are equal.
(Given that t0.005,34 = 2.728.)

(iii) In order to test whether a coin is perfect, the coin is tossed 5 times. The null hy-
pothesis of perfectness is rejected if more than 4 heads are obtained. What is the probability
of type I error? Find the probability of type II error when the corresponding probability of
head is 0.2.

Sol:

(i) Null hypothesis: H0 : µ = 50 Alternate hypothesis H1 : µ ̸= 50. First calculate the sample
x̄−µ
mean x̄ = 47.3 and sample standard deviation s = 2.53. Value the test statistic t = √ =
s/ (n−1)
−9.6 < −2.26. Therefore H0 is rejected.

x¯1 −x¯2 (n1 −1)s21 +(n2 −1)s22


(ii) H0 : µ1 = µ2 , H1 : µ1 ̸= µ2 , t = q where S 2 = n1 +n2 −2 = 49926.47 Hence
S n1 + n1
1 2
S = 223.44. Therefore, t = 1.67 < 2.728. Hence H0 accepted. There is no significance difference.

Problem:

(i) A machine is making parts with scale diameter of 0.7 inch. A random sample of 10 parts
shows mean 0.742 inch with a standard deviation of 0.04 inch. On the basis of this sample
would you say that the work in inferior? Given that t0.025,9 = 2.26.

(ii) The following table gives the number of aircraft accident that occurred during various days
of the week. Test whether the accidents are uniformly distributed over the week.

Days: Sun Mon Tues Wed Thurs Fri Sat


No. of accidents: 13 14 19 12 11 15 14

Table 1: Caption

[Given χ20.05 = 12.59 for d.o.f.6]

(iii) Find the least value of r in a sample of 18 pairs of observations from a bi-variate normal
population, significant at 5% level of significance. (Given that t0.05 = 2.12 for d.f. 16.)

Sol:

(i) Null hypothesis: H0 : µ = 0.7 Alternate hypothesis H1 : µ ̸= 0.7. Value the test statistic
x̄−µ
t= √ = 3.15 > 2.26. Therefore H0 is rejected. Therefore there is a significant difference.
s/ (n−1)
Hence work is inferior.

(ii) See any standard book.

64
(iii) |r| > 0.4682

8 Bivariate distribution:
Discrete case:
Definition: If X and Y are discrete random variables, the function given by f (x, y) = P(X =
x, Y = y) for each pair of values (x, y) within the range of X and Y is called the joint probability
distribution of X and Y .

Theorem: A bivariate function can serve as the joint probability distribution of a pair of dis-
crete random variables X and Y if and only if its values, f (x, y), satisfy the conditions

ˆ f (x, y) ≥ 0 for each pair of values (x, y) within its domain.

ˆ
P P
x y f (x, y) = 1, where the double summation extends over all possible pairs (x, y) within
its domain.

Let us try to understand this with an example. Suppose two caplets33 are selected at random from a
bottle containing three aspirin, two sedative and four laxative caplets. If X and Y are, respectively
the numbers of aspirin and sedative caplets included among the two caplets drawn from the bottle.
Now suppose we want to find the probabilities associated with all possible pairs of values of X and Y .

In total there are 3 + 2 + 4 = 9 caplets in the bottle. From this two caplets can be drawn in
9C2 ways i.e. 36 ways. We are interested in the number of aspirin caplets (X) and number of
sedative caplets (Y ) as per the conditions given in the problem. Hence we get the following chart
for our favourable outcomes.

(X, Y ) Interpretation Count of caplets Count of favourable outcomes


(0, 0) no aspirin, no sedative aspirin = 3, sedative= 2, laxative = 4 3C0 2C0 4C2 = 6
(0, 1) no aspirin, one sedative aspirin = 3, sedative= 2, laxative = 4 3C0 2C1 4C1 = 8
(1, 0) one aspirin, no sedative aspirin = 3, sedative= 2, laxative = 4 3C1 2C0 4C1 = 12
(1, 1) one aspirin, one sedative aspirin = 3, sedative= 2, laxative = 4 3C1 2C1 4C0 = 6
(0, 2) no aspirin, two sedatives aspirin = 3, sedative= 2, laxative = 4 3C0 2C2 4C0 = 1
(2, 0) two aspirin, no sedatives aspirin = 3, sedative= 2, laxative = 4 3C2 2C0 4C0 = 3

Next we show below the probabilities of the outcomes shown in the above table.

Outcome Total outcome Favourable outcome Probability


1
(0, 0) 36 6 6
2
(0, 1) 36 8 9
1
(1, 0) 36 12 3
1
(1, 1) 36 6 6
1
(0, 2) 36 1 36
1
(2, 0) 36 3 12

This bivariate probability chart can be interpreted in the following way too. A careful study reveals
33
a coated oral medicinal tablet is caplet.

65
the following joint probability distribution:
3Cx 2Cy 4C2−x−y
f (x, y) = , x = 0, 1, 2; y = 0, 1, 2; 0 ≤ x + y ≤ 2.
9C2
Observations:
ˆ The sum of all row sums is 1.

ˆ The sum of all column sums is 1.

ˆ The sum of all the probabilities is 1.


Definition: If X and Y are discrete random variables and f (x, y) is the value of their joint
probability distribution at (x, y), the function given by
X
g(x) = f (x, y) (114)
y

for each x within the range of X is called the marginal distribution of X. Correspondingly, the
function given by
X
h(y) = f (x, y) (115)
x

for each y within the range of Y is called the marginal distribution of Y .

With respect to above caplet problem, we write the marginal distribution of X and marginal
distribution of Y in the following tables.
X=x g(x)
7
0 12
7
1 18
1
2 36
P
g(x) 1
and also
Y =y h(y)
15
0 36
3
1 6
1
2 12
P
h(y) 1

66
Definition: If X and Y are discrete random variables, the function given by
XX
F (x, y) = P(X ≤ x, Y ≤ y) = f (s, t), (116)
s≤x t≤y

for −∞ < x < ∞ and −∞ < y < ∞. Here f (s, t) is the value of the joint probability distribution of
X and Y at (s, t), is called the joint distribution function or joint cumulative distribution,
of X and Y .

Now let’s see how to find joint distribution from caplet problem.

How to calculate?
Using eq.(116) we calculate the following probabilities.

F (0, 0) = P(X ≤ 0, Y ≤ 0)
F (0, 1) = P(X ≤ 0, Y ≤ 1)
F (0, 2) = P(X ≤ 0, Y ≤ 2)
F (1, 1) = P(X ≤ 1, Y ≤ 0)
F (0, 2) = P(X ≤ 1, Y ≤ 1)
F (2, 0) = P(X ≤ 2, Y ≤ 0) (117)

Using eq.(116) we get the following:


XX X 1
P(X ≤ 0, Y ≤ 0) = f (s, t) = f (s, 0) = f (0, 0) = .
6
s≤0 t≤0 s≤0
XX X 7
P(X ≤ 0, Y ≤ 1) = f (s, t) = [f (s, 0) + f (s, 1)] = .
18
s≤0 t≤1 s≤0
XX X 15
P(X ≤ 0, Y ≤ 2) = f (s, t) = [f (s, 0) + f (s, 1) + f (s, 2)] = .
36
s≤0 t≤2 s≤0

From eq.(118) we see that the last probability is the sum of column 1 of the bi-variate table of the
caplet problem (shown above).

Again,
XX X 1
P(X ≤ 1, Y ≤ 0) = f (s, t) = f (s, 0) = f (0, 0) + f (1, 0) = .
2
s≤1 t≤0 s≤1
XX X 8
P(X ≤ 1, Y ≤ 1) = f (s, t) = [f (s, 0) + f (s, 1)] = .
9
s≤1 t≤1 s≤1

Also,
XX X 7
P(X ≤ 2, Y ≤ 0) = f (s, t) = [f (s, 0)] = f (0, 0) + f (1, 0) + f (2, 0) = .
12
s≤2 t≤0 s≤2

Problem: An urn contains four white,four red and four black balls. The balls of each color are
marked 0,1,2,3 respectively. A ball is drawn from the urn at random. A random variable X assumes

67
the value 0,1,2 according as the ball is white, red and black respectively. Y denote the number
marked on the ball. Find E(X),E(Y).Hence find the covariance of the variates.

Sol: E(X) = 1, E(Y ) = 32 , Cov(X, Y ) = 0

Problem: Given the joint probability density


(
2
(x + 2y) f or 0 < x < 1, 0 < y < 1
f (x) = 3 (118)
0 otherwise

Identify the marginal densities of X and Y. conditional densities of X given Y = y, and use it to
evaluate P (X ≤ 21 /Y = 12 ).

Sol:
( (
2 1
g(x) = 3 (x + 1) for 0 < x < 1
h(y) = 3 (1
+ 4y) for 0 < y < 1
(119)
0 otherwise 0 otherwise

[Rest you do].

68

You might also like