B K Sem 1 (SCM-MM) QT & MT Binomial- Poisson Probability Distribution
Probability Distribution
In last module, it may be recalled, we discussed frequency distribution. In a similar
manner, we may think of a probability distribution where just like distributing the total
frequency to different class intervals, the total probability (i.e. one) is distributed to different
mass points in case of a discrete random variable or to different class intervals in case of a
continuous random variable. Such a probability distribution is known as Theoretical
Probability Distribution, since such a distribution exists in theory. We need to study
theoretical probability distribution for the following important factors:
(a) An observed frequency distribution, in many a case, may be regarded as a sample i.e. a
representative part of a large, unknown, boundless universe or population and we may
be interested to know the form of such a distribution. By fitting a theoretical probability
distribution to an observed frequency distribution of, say, the lamps produced by a
manufacturer, it may be possible for the manufacturer to specify the length of life of the
lamps produced by him up to a reasonable degree of accuracy. By studying the effect of a
particular type of missiles, it may be possible for our scientist to suggest the number of
such missiles necessary to destroy an army position. By knowing the distribution of
smokers,
a social activist may warn the people of a locality about the nuisance of active and passive
smoking and so on.
(b) Theoretical probability distribution may be profitably employed to make short term
projections for the future.
(c) Statistical analysis is possible only on the basis of theoretical probability distribution.
Setting confidence limits or testing statistical hypothesis about population parameter(s) is
based on the probability distribution of the population under consideration.
A probability distribution also possesses all the characteristics of an observed distribution.
We define mean (µ) , median (µe) , mode (µ0) , standard deviation (σ) etc. exactly same
way we have done earlier. Again a probability distribution may be either a discrete
probability distribution or a Continuous probability distribution depending on the random
variable under study.
Important discrete probability distributions are (a) Binomial Distribution and (b) Poisson
distribution. (c) Hyper geometric Distribution (d) Negative Binomial Distribution (e)
Geometric Distribution etc.
BINOMIAL DISTRIBUTION
One of the most important and frequently used discrete probability distribution is
Binomial Distribution.
It is derived from a particular type of random experiment known as Bernoulli process
named after the famous mathematician Bernoulli.
Noting that a 'trial' is an attempt to produce a particular outcome which is neither
certain nor impossible,
The characteristics of Bernoulli trials are stated below:
Each trial is associated with two mutually exclusive and exhaustive outcomes, the
occurrence of one of which is known as a 'success' and as such its non- occurrence as a
'failure'. As an example, when a coin is tossed, usually occurrence of a head is known as
a success and its non–occurrence i.e. occurrence of a tail is known as a failure.
The trials are independent.
The probability of a success, usually denoted by p, and hence that of a failure, usually
denoted by q = 1–p, remain unchanged throughout the process.
The number of trials is a finite positive integer.
A discrete random variable x is defined to follow binomial distribution with parameters
n and p, to be denoted by x ~ B (n, p), if the probability mass function of x is given by
f (x) = p (X = x) = nc x px qn-x for x = 0, 1, 2, …., n
= 0, otherwise
→ We may note the following important points in connection with binomial distribution:
(a) As n > 0, p, q ≥ 0, it follows that f(x) ≥ 0 for every x
Also ∑f(x) = f(0) + f(1) + f(2) + …..+ f(n) = 1
(b) Binomial distribution is known as bi-parametric distribution as it is characterised
by two parameters n and p. This means that if the values of n and p are known,
then the distribution is known completely.
(c) The mean of the binomial distribution is given by µ = np
(d) Depending on the values of the two parameters, binomial distribution may be unimodal
or bi- modal. µ 0 , the mode of binomial distribution, is given by
µ 0 = the largest integer contained in (n+1)p if (n+1)p is a non-integer
= (n+1)p and (n+1)p – 1 if (n+1)p is an integer
(e) The variance of the binomial distribution is given by σ2 = npq Since p and q are
X 0 1 2 3 4 5
numerical
f 3 6 10 8 3 2
ly less
than or equal to 1, npq < np variance of a binomial variable is always
less than its mean.
Also variance of X attains its maximum value at p = q = 0.5 and this maximum
value is n/4.
(f) Additive property of binomial distribution.
If X and Y are two independent variables such that X~B (n1, P)and Y~B (n2, P) Then
(X+Y) ~B (n 1 + n2 , P)
Applications of Binomial Distribution
Binomial distribution is applicable when the trials are independent and each trial has just
two outcomes success and failure. It is applied in coin tossing experiments, sampling
inspection plan, genetic experiments and so on.
Sum no. 1: A coin is tossed 10 times. Assuming the coin to be unbiased, what is the
probability of getting (i) 4 heads? (ii) at least 4 heads? (iii) at most 3 heads?
Sum no. 2: If 15 dates are selected at random, what is the probability of getting two Sundays?
Sum no. 3: The incidence of occupational disease in an industry is such that the workmen
have a 10% chance of suffering from it. What is the probability that out of 5 workmen, 3or
more will contract the disease?
Sum no. 4: Find the probability of a success for the binomial distribution satisfying the
following relation 4 P (x = 4) = P (x = 2) and having the parameter n as six.
Sum no. 5: Find the binomial distribution for which mean and standard deviation are 6
and 2 respectively.
Sum no. 6: Fit a binomial distribution to the following data:
X 0 1 2 3 4 5
f 3 6 10 8 3 2
Sum no. 7: 6 coins are tossed 512 times. Find the expected frequencies of heads. Also,
compute the mean and SD of the number of heads.
Sum no. 8: An experiment succeeds thrice as after it fails. If the experiment is repeated 5
times, what is the probability of having no success at all ?
Sum no. 9: What is the mode of the distribution for which mean and SD are 10 and √5
respectively.
Sum no. 10: If x and y are 2 independent binomial variables with parameters 6 and 1/2
and 4 and 1/2 respectively, what is P ( x + y ≥ 1 )?
B K Sem 1 (SCM-MM) QT & MT Binomial- Poisson Probability Distribution
POISSON DISTRIBUTION
Poisson distribution is a theoretical discrete probability distribution which can describe
many processes. Simon Denis Poisson of France introduced this distribution way back
in the year 1837.
Poisson Model
Let us think of a random experiment under the following conditions:
I. The probability of finding success in a very small time interval ( t, t + dt ) is kt, where
k (>0) is a constant.
II. The probability of having more than one success in this time interval is very low.
III. The probability of having success in this time interval is independent of t as well as
earlier successes.
The above model is known as Poisson Model. The probability of getting x successes in
a relatively long time interval T containing m small time intervals t i.e. T = mt. is given
by e-kt .(kt) x/x!
for x = 0, 1, 2, ......… ∞
Taking kT = m, the above form is reduced to
−m
e mx
p(x) =
x!
for x = 0, 1, 2, ...... ∞
Definition of Poisson Distribution
A random variable X is defined to follow Poisson distribution with parameter λ, to be denoted
by X ~ P (m) if the probability mass function of x is given by
−m
e mx
f (x) = P (X = x) = for x = 0, 1, 2, ... ∞
x!
= 0 otherwise
Here e is a transcendental quantity with an approximate value as 2.71828.
It is wiser to remember the following important points in connection with Poisson
distribution:
(i) Since e–m = 1/em >0, whatever may be the value of m, m > 0, it follows that f (x) ≥ 0
for every x. Also it can be established that ∑f(x) = 1 i.e. f(0) + f(1) + f(2) +....... = 1
(ii) Poisson distribution is known as a uni-parametric distribution as it is characterised
by only one parameter m.
(iii) The mean of Poisson distribution is given by m i,e µ = m.
(iv) The variance of Poisson distribution is given by σ2 = m
(v) Like binomial distribution, Poisson distribution could be also unimodal or bimodal
depending upon the value of the parameter m.
µ0 = The largest integer contained in m if m is a non-integer
= m and m–1 if m is an integer
(vi) Poisson approximation to Binomial distribution
If n, the number of independent trials of a binomial distribution, tends to infinity and
p, the probability of a success, tends to zero, so that m = np remains finite, then a
binomial distribution with parameters n and p can be approximated by a Poisson
distribution with parameter m (= np).
In other words when n is rather large and p is rather small so that m = np is moderate
then β (n, p) ≅ P (m).
(vii) Additive property of Poisson distribution
If X and y are two independent variables following Poisson distribution with
Parameters m1 and m2 respectively, then Z = X + Y also follows Poisson distribution
with parameter (m1 + m2 ).
i.e. if X ~ P (m1)
and Y ~ P (m2)
and X and Y are independent, then
Z = X + Y ~ P (m1 + m2 )
Application of Poisson distribution
Poisson distribution is applied when the total number of events is pretty large but the
probability of occurrence is very small. Thus we can apply Poisson distribution, rather
profitably, for the following cases:
a) The distribution of the no. of printing mistakes per page of a large book.
b) The distribution of the no. of road accidents on a busy road per minute.
c) The distribution of the no. of radio-active elements per minute in a fusion process.
d) The distribution of the no. of demands per minute for health centre and so on.
Sum no. 11: Find the mean and standard deviation of x where x is a Poisson variate
satisfying the condition P (x = 2) = P ( x = 3).
Sum no. 12: The probability that a random variable x following Poisson distribution would
assume a positive value is (1 – e–2.7). What is the mode of the distribution?
Sum no. 13: The standard deviation of a Poisson variate is 1.732. What is the probability
that the variate lies between –2.3 to 3.68?
Sum no. 14: X is a Poisson variate satisfying the following relation:
P (X = 2) = 9P (X = 4) + 90P (X = 6).
What is the standard deviation of X?
Sum no. 15: Between 9 and 10 AM, the average number of phone calls per minute coming
into the switchboard of a company is 4. Find the probability that during one particular
minute, there will be,
1. no phone calls
2. at most 3 phone calls (given e–4 = 0.018316)
Sum no. 16: If 2 per cent of electric bulbs manufactured by a company are known to be
defectives, what is the probability that a sample of 150 electric bulbs taken from the
production process of that company would contain
1. exactly one defective bulb?
2. more than 2 defective bulbs?
Sum no. 17: The manufacturer of a certain electronic component is certain that two per
cent of his product is defective. He sells the components in boxes of 120 and guarantees
that not more than two per cent in any box will be defective. Find the probability that a box,
selected at random, would fail to meet the guarantee? Given that e–2.40 = 0.0907.
Sum no. 18: A discrete random variable x follows Poisson distribution. Find the values of
(i) P (X = at least 1)
(ii) P (X ≤ 2/ X ≥ 1)
You are given E (x) = 2.20 and e–2.20 = 0.1108.
Fitting a Poisson distribution
As explained earlier, we can apply the method of moments to fit a Poisson distribution to an
observed frequency distribution. Since Poisson distribution is uni-parametric, we equate m,
the parameter of Poisson distribution, to the arithmetic mean of the observed distribution
and get the estimate of m. i.e. mˆ = X
B K Sem 1 (SCM-MM) QT & MT Binomial- Poisson Probability Distribution
The fitted Poisson distribution is then given by
−m
e mx
f (x) = P (X = x) = for x = 0, 1, 2, ... ∞
x!
Sum no. 19: Fit a Poisson distribution to the following data :
Number of death: 0 1 2 3 4
f 122 46 23 8 1
(Given that e–0.6 = 0.5488)
FC105 - Quantitative Analysis & Modeling Techniques (QA&MT)
26
NORMAL DISTRIBUTION
The Normal Distribution (N.D.) was first discovered by De-Moivre as the limiting
form of the binomial model in 1733, later independently worked Laplace and Gauss.
It is a probability distribution of a continuous random variable and is often used to
model the distribution of discrete random variable as well as the distribution of other
continuous random variables. The basic from of normal distribution is that of a bell,
it has single mode and is symmetric about its central values. The flexibility of
using normal distribution is due to the fact that the curve may be centered over any
number on the real line and it may be flat or peaked to correspond to the amount of
dispersion in the values of random variable.
Definition: A random variable X is said to follow a Normal Distribution with parameter
2
and and if its density function is given by the probability law
(x )2
1 2
f(x) = e 2
- <x< ; - < < ; >0
2
where = a mathematical constant equality = 22/7
e = Naperian base equaling 2.7183
= population mean
= population standard deviation
x = a given value of the random variable in the range - < x <
Characteristics of Normal distribution and normal curve:
The normal probability curve with mean and standard deviation is given by the
equation
(x )2
1 2
f(x) = e 2
;- <x<
2
and has the following properties
i. The curve is bell shaped and symmetrical, about the mean
ii. The height of normal curve is at its maximum at the mean. Hence the mean
and mode of normal distribution coincides. Also the number of observations
below the mean in a normal distribution is equal to the number of observations
about the mean. Hence mean and median of N.D. coincides. Thus, N.D. has
Mean = median = mode
FC105 - Quantitative Analysis & Modeling Techniques (QA&MT)
27
iii.
occurring at the point x = , and given by
1
p[(x)]max =
2
2
3
iv. Skewness 1 = 3
=0
2
4
v. Kurtosis = 2 = 2
= 3 ( i , 2 , 3 and 4 are called central moments)
2
vi. All odd central
i.e. 1 3 5 .............. 0
vii. The first and third quartiles are equidistant from the median
viii. Linear combination of independent normal variates is also a normal
variate
ix. The points of inflexion of the curve is given by
1
1
x , f ( x) e2
2
x. If f (x)dx 1 then
the area under the normal curve is distributed as follows
i) - <x< + covers 68.26% of area
ii) -2 < x < +2 covers 95.44% of area
iii) -3 < x < +3 coves 99.73% of area
Area under Normal curve
FC105 - Quantitative Analysis & Modeling Techniques (QA&MT)
28
The Normal Curve: The graph of the normal distribution depends on two factors - the
mean and the standard deviation. The mean of the distribution determines the location
of the center of the graph, and the standard deviation determines the height and width of
the graph. When the standard deviation is large, the curve is short and wide; when
the standard deviation is small, the curve is tall and narrow. All normal distributions look
like a symmetric, bell-shaped curve, as shown below.
The curve on the left is shorter and wider than the curve on the right, because the curve
on the left has a bigger standard deviation.
Standard Normal Distribution and
X
standard deviation , then Z is a standard normal variate with zero mean and
=
standard deviation = 1.
The probability dens z2
1
f(z) = e 2
and f ( z )dz =1
2
A graph representing the density function of the Normal probability distribution is
also known as a Normal Curve or a Bell Curve (see Figure below). To draw such a curve,
one needs to specify two parameters, the mean and the standard deviation. The graph
below has a mean of zero and a standard deviation of 1, i.e., (m =0, s =1). A Normal
distribution with a mean of zero and a standard deviation of 1 is also known as the
Standard Normal Distribution.
Standard Normal Distribution
Sum no. 20: For a random variable x, the probability density function is given by
2
e (x 4)
f(x)=
for – ∞ < x < ∞ .
Identify the distribution and find its mean and variance.
Sum no. 21: If the two quartiles of a normal distribution are 47.30 and 52.70
52.70 res
respectively, what is the
mode of the distribution? Also find the mean deviation about median of this distribution.
Sum no. 22: Find the points of inflexion of the normal curve
1 2
f (x) . e -(x-10) / 32
4 2
for – ∞ < x < ∞
Sum no. 23: X follows normal distribution with mean as 50 and variance as 100.
What is P(x ≥ 60)? Given p( 1 ) = 0.3413
Sum no. 24: If a random variable x follows normal distribution with mean as 120 and standard
deviation as 40, what is the probability that P (x ≤ 150 / x > 120)?
Given that the area of the normal curve between z = 0 to z = 0.75 is 0.2734.
Sum no. 25: X is a normal variable with mean = 25 and SD 10. Find the value of b such that the
probability of the interval [25, b] is 0.4772 given p(2) = 0.4772.
Sum no. 26: In a sample of 500 workers of a factory, the mean wage and SD of wages are found to
be ` 500 and ` 48 respectively. Find the number of workers having wages:
(i) more than ` 600
(ii) less than ` 450
(iii) between ` 548 and ` 600.
Sum no. 27: The distribution of wages of a group of workers is known to be normal with
mean ` 500 and SD ` 100. If the wages of 100 workers in the group are less than ` 430, what is the
total number of workers in the group?
Sum no. 28: The mean of a normal distribution is 500 and 16 per cent of the values are greater
than 600. What is the standard deviation of the distribution?
(Given that the area between z = 0 to z = 1 is 0.34)
Sum no. 29: In a business, it is assumed that the average daily sales expressed in Rupees follows
normal distribution.
Find the coefficient of variation of sales given that the probability that the average daily sales is less
than Rs.124 is 0.0287 and the probability that the average daily sales is more than Rs. 270 is 0.4599.
Sum no. 30: x and y are independent normal variables with mean 100 and 80 respectively and
standard deviation as 4 and 3 respectively. What is the distribution of (x + y)?
B K Sem 1 (SCM-MM) QT & MT Module 2 Chi Square Test
Chi-Square (χ²) Test –
1. Introduction
The Chi-Square (χ²) test is a non-parametric statistical test used to determine whether there is a
significant association between categorical variables or whether observed data fit expected data. It
was developed by Karl Pearson in 1900.
2. Types of Chi-Square Tests
1. Test of Goodness of Fit – To check whether an observed frequency distribution fits an
expected/theoretical distribution (e.g., checking if dice is fair).
2. Test of Independence (Association) – To check if two categorical variables are independent or
associated (e.g., gender and preference for a product).
3. Test of Homogeneity – To check if distributions of a categorical variable are the same across
different populations (e.g., comparing voting preferences across regions).
3. Chi-Square Statistic Formula
χ² = Σ((O - E)² / E)
Where:
O = Observed frequency
E = Expected frequency
4. Calculation Steps
A. For Goodness of Fit
1. State H₀ and H₁.
2. Calculate Expected Frequencies.
3. Compute χ² = Σ((O - E)² / E).
4. Degrees of freedom (df) = (n - 1).
5. Compare χ² calculated with χ² critical.
6. Decision: Reject or Fail to Reject H₀.
B. For Test of Independence
1. H₀: Variables are independent.
2. Expected frequency: Eij = (Row Total × Column Total) / Grand Total.
3. χ² = Σ((Oij - Eij)² / Eij).
4. df = (r - 1)(c - 1).
5. Assumptions of the Chi-Square Test
1. Data are in frequency form (counts).
2. Observations are independent.
3. Expected frequency in each cell ≥ 5.
4. Random sampling.
6. Applications
• Testing fairness of dice or coins.
• Studying relationship between demographic variables.
• Genetic inheritance studies.
• Market research and social science analysis.
7. Interpretation
Higher χ² → greater difference → likely reject H₀.
Lower χ² → smaller difference → fail to reject H₀.
9. Limitations
• Not suitable for small samples (expected < 5).
• Does not indicate strength or direction of association.
• Sensitive to sample size (large samples may give significant χ² for small differences).
Sum no 1.
Three Hundred digits were chosen at random from set of tables. The frequencies of digits
were as follows
Digits 0 1 2 3 4 5 6 7 8 9
Frequency 28 29 33 31 26 35 32 30 31 25
Using x – test assess the hypothesis that the digits were distributed in equal numbers in the
2
table. (The 5% value of χ2 for 9 degree of freedom is 16.92)
Sum no 2.
The number of automobile accidents per week in a certain community were as follows :
12, 8, 20, 2, 14, 10, 15, 6, 9, 4
are these frequencies in agreement with belief that accident conditions were same during this 10
week period ?
Sum no 3.
The theory predicts the proportion of beans, in the four groups A, B, C and D should be 9 : 3 :
3 : 1. In an experiment with 1600 beans the numbers in the four groups were 882, 313, 287
and 118. Does the experimental result support the theory ? (The table Value of χ2 for 3 d.f. at
5% level of significance 7.81)
Sum no 4.
The following tables gives the number of aircraft accidents that occurred during the various
days of week. Find whether the accidents are uniformly distributed over the week.
Days : Sun Mon Tue Wed Thu Fri Sat Total
14 16 8 12 11 9 14 84
[Link] accidents
Sum 5:
The table value for different degree of freedom is given below :
Degree Freedom 1 2 3 4 5 6 7 8 9
5% value 3.84 5.99 7.82 9.49 11.07 12.59 14.07 15.51 16.92
Sum no 6.
A sample analysis of examination results of 500 students was made. It was found that 220
students had failed, 170 had secured a third class, 90 were placed in second class and 20 got a
first class. Are there figures commensurate with the general examination result which is in the
ratio of 4 : 3 : 2 : 1 for the various categories respectively.
(The table value of χ2 for 3 d.f. at 5 % level of significance is 7.81).
Sum no 7.
Records taken of the number of male and female births in 800 families having four children
are given below :
B K Sem 1 (SCM-MM) QT & MT Module 2 Chi Square Test
Number of births frequency
0 4 32
1 3 178
2 2 290
3 1 236
4 0 64
800
Test whether the data are consistent with the hypothesis that the binomial law holds and the
chance of the male birth is equal to that of a female birth.
Sum no 8.
12 dice were thrown 4096 times and a throw of 6 was reckoned as success the observed
frequencies were as given below :
Success 0 1 2 3 4 5 6 7 & over
Frequencies 447 1145 1181 796 380 115 24 8
Find the value of x2 on the hypothesis that the dice were unbiased and hence show that the
data are consistent with this hypothesis so far as the χ2 test is concerned.
Sum no 9.
A set of 5 coins is tossed 3200 times & number of heads appearing each time is noted. The
results are given below :
No. of Heads 0 1 2 3 4 5
Frequency 80 570 1100 900 500 50
Test the hypothesis that the coins are unbiased.
Sum no 10.
The No. of males in each 106 eight pig litters was found & they are given by the following
frequency distribution :
[Link] males per 0 1 2 3 4 5 6 7 8 Total
litter
Frequency 0 5 9 22 25 26 14 4 1 106
Assuming that the probability of animal being male or female is even (i.e. p = q = ½ ) and the
frequency distribution follows the binomial law calculate the expected frequencies of the nine
classes. Find also the value of χ2 to test the goodness of fit.