0% found this document useful (0 votes)
22 views19 pages

Stat Part1

Chapter 10 discusses parametric families of univariate probability distributions, including both discrete and continuous types, which are essential for decision-making. It introduces key discrete distributions such as Bernoulli, Binomial, and Poisson, as well as continuous distributions like Normal and Exponential. The chapter further elaborates on the Bernoulli distribution, its properties, and the concept of binomial experiments, providing definitions and examples to illustrate these statistical concepts.

Uploaded by

sanjaysajib
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF or read online on Scribd
0% found this document useful (0 votes)
22 views19 pages

Stat Part1

Chapter 10 discusses parametric families of univariate probability distributions, including both discrete and continuous types, which are essential for decision-making. It introduces key discrete distributions such as Bernoulli, Binomial, and Poisson, as well as continuous distributions like Normal and Exponential. The chapter further elaborates on the Bernoulli distribution, its properties, and the concept of binomial experiments, providing definitions and examples to illustrate these statistical concepts.

Uploaded by

sanjaysajib
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF or read online on Scribd
Chapter 10 PROBABILITY DISTRIBUTIONS 10.1 INTRODUCTION In this chapter we shall define and discuss certain parametric families of univariate probability distributions that have been widely applied in a variety of decision-making situations. A parametric family of distributions is a collection of distributions that is indexed by a quantity called parameter. These distributions have standard names and can be derived under certain plausible conditions about the random, variables involved. In addition, they can be expressed by algebraic functions and hence are capable of further mathematical treatment. The distributions that will be presented here include both discrete and continuous distributions. The most frequently encountered discrete distributions, among others, are the following: (a) Bernoulli distribution (b) Binomial distribution (c) Hypergeometric distribution (d) Poisson distribution “ (e) Geometric distribution (f) Negative binomial distribution (g) Multinomial distribution 558 AN INTRODUCTION TO STATISTICS AND PROBABILITY The following are some of the commonly discussed continuous distributions: a) Normal distribution b) Exponential distribution c) Beta distribution d) Gamma distribution e) Rectangular distribution 10.2 BERNOULLI DISTRIBUTION A considerable group of random phenomena, known as Bernoulli Process, named after James Bernoulli (1654-1705), follows the simplest probability. distribution — involving only two possible outcomes, such as head or tail, success or failure, defective or non- defective, smoker or non-smoker, married or unmarried. To standardize the terminology describing these and many other similar processes, we call one of the possible outcomes a success and the other a failure. These names are used only to identify the outcomes and bear no connotation of success or failure in real life. Customarily, the outcome of primary interest in a study is labeled a success (even it is a disastrous event). In a study of prevalence of marriage dissolution or rate of electricity failure, the marital status “divorced” or “load shading” may be attributed the statistical name “success”. The events “success” and “ failure” may be viewed as outcomes of experiments composed of repetitions of independent trials. When the probabilities of the two outcomés remain unchanged from one trial to another, we name the trials Bernoulli trials, in honor of James Bernoulli. In all applications, it is convenient to designate the two possible outcomes of such an experiment as 0 and 1. Each of the following examples demonstrates a Bernoulli trial: (i) Suppose a fair coin is tossed tepeatedly. Let X;=1 if a head is obtained on the ith toss and let X;=0 if a tail is obtained (i=1, PROBABILITY DISTRIBUTIONS 559 2......). Then the random variables X;, Xo, ...... form a sequence of Bernoulli trials with parameter p=0.5. (ii) Suppose that 20 percent of the items produced by a machine of a manufacturing industry are defective and that n items are selected at random and inspected. Asigning X;=0 if the ith item is non- defective and X;=1, if it is defective (i=1, 2, .....), then the variables X;, Xz, .....form n Bernoulli trials with parameter p=0.20. The following definition can then be applied to any experiment of this type. Definition: It is said that a random variable X has a Bernoulli distribution with parameter p (0 S pS) if X can take only the values 0 ‘and land the probabilities are P(X=1)=p and P(X=0)=1-p. (10.1) Since the distribution involves only two classes of events, it is also known as the two-point probability distribution. If we let g=1—p, then the probability function of X can be written as follows: xal-x ba | f(x;p)={P V foe Ot (10.2) 0, otherwise Figure 10.1: Bernoulli function with

p | p q | q Pe bs 1 To verify that the probability function (10.2) actually represents the | Bernoulli distribution specified by the probabilities (10.1), it is 560 AN INTRODUCTION TO STATISTICS AND PROBABILITY necessary to note that f (1, p)=p and (0, p)=q. The tabular presentation of the random variable X is as follows x. P(X =x) 0 gzip ct : P tt 10.2.1 Properties of Bernoulli distribution Similar to other distributions, the Bernoulli distribution has its mean, variance and other descriptive measures. These are obtained as follows: Mean: E(X)= Yi x f(a p)=Ox(- p)+b

£0, p) -[ECOP x (10.4) =0°x(- p)+xp-p” ~ p> =p(l- p)= Pq Moment generating function: Mx (= Ele*)= Dye" fu) =Se"p'l- pe (10.5) =¥(pe'}a- py =pe+q Moments: It is easy to verify that for the Bernoulli distribution, the rth raw PROBABILITY DISTRIBUTIONS 561 moment p/=p. Hence, My = Hl = Ms so Ds Thus, Hy =H - BY = p- P= Pa By = Wh - 3M M, +2? = p-3p" + 2p? = pa(a~ P) my = Hi 4s + OpyMy? -34y" = p—4p* + 6p* ~3p* = 3p’q’ + pql-6pq) Skewness and kurtosis: aes 2 fot B= (q-p) atl Bye 6pq Pq My Pq The value of B; indicates that Bernoulli distribution is skewed unless p=q. If qp it is positively skewed. The value of B> is an indication that Bernoulli distribution is slightly leptokurtic since B, >3 for all values of p except when.p=q. 10.3 BINOMIAL DISTRIBUTION 10.3.1: Binomial experiment Many experiments are composed of repetitions of independent trials, each with two possible outcomes: success and failure. When the probabilities of these two outcomes remain unchanged from one trial to another, we named these trials Bernoulli trials. A binomial experiment consists of a fixed number ‘n’ of Bernoulli trials. Definition: When an experiment has two possible outcomes, success and failure and the experiment is repeated n times independently, and the probability p of success of any given trial remains constant from trial to trial, the experiment is known as binomial experiment. The above definition leads to suggest that a binomial experiment is one that possesses the following properties: 562 4. AN INTRODUCTION TO STATISTICS AND PROBABILITY The overall experiment can be described in terms of a sequence of n identical experiments, which are called trials. Each trial results in an outcome that may be classified as a success or a failure. The probability of success, denoted by p, remains constant from one trial to the next. The repeated trials are independent. 10.3.2 Examples of binomial experiment 4 An insurance salesman contacts ten different families. The outcome associated with visiting each family can be referred to as a success if the family purchases an insurance policy and a failure, if not. If the probability of selling a policy is assumed to be the same for each family, and the decision to purchase or not-a policy by one family is not influenced by the decision of any other family, then we have a situation analogous to the binomial experiment. A customer enters the Arong department store and makes a purchase. For convenience, let us restrict our attention to the next ten customers who enter the shop. The event of interest here is to ascertain whether or not any one of the customers makes a purchase. The process can be viewed as a binomial experiment because of the rpason that (i) the experiment can be described as a sequence of ten identical trials, orje trial for each of the ten customers that will enter the store, (ii) the probability of the purchase (p) may be assumed to be the same for all customers (iii) the purchase decision of each customer is independent of the decision of the other customers. An experiment consists in selecting five radios at random from a‘ lot and inspecting and then classifying them as defective and non- defective. A defective radio is labeled a success, while a non-defective radio is labeled a failure. Suppose 20% of the radios aré defective. If this probability is assumed to remain constant from trial to wial, then X, the number of successes is a binomial random variable assuming PROBABILITY DISTRIBUTIONS 563 the values 0, 1, 2,3 , 4, and 5. Clearly the experiment is a binomial experiment with n=5 and p=0.20. 4. If five cards‘are drawn in succession from an ordinary deck and each trial is labeled a success or failure depending on whether the card is a ted or black. If each catd is replaced and the deck is well shuffled before the next drawing is made,:then the repeated drawings (trials) are independent and the probability of success (drawing a red ball here) remains constant from trial to trial. This is’ a binomial experiment. 5. An each of eight faculties of Dhaka University, 25% of the departments are equipped with computer facilities and 75% are not. A researcher decides to conduct a sample survey, on the basis of a sample formed by taking one department at random from each faculty. We may consider the sampling procedure to consist of n=8 independent trials, each trial being the selection of a department from a faculty. We may classify “computerized” as success and “not-computerized” as failure. Here p=25%=0.25, which is assumed to remain constant from trial to trial ‘This experiment also falls under the category of binomial experiment. In each of the above experiments, there exists a variable that denotes the number of successes in.» Bernoulli trials. This variable is known as the binomial random variable. A binomial random variable may thus be viewed as a set of values representing the number of successes generated through some binomial experiment. Let X), X2, ...-...Xq denote the outcomes of the experiment. Since X takes on the value. ‘1’ if the trial results in success, and ‘0’ if the outcome is a failure, then the number of Is in 7 trials is simply Xp tXatennt X_ lf n X=EX, i=l then X is a binomial random variable. 564 AN INTRODUCTION TO STATISTICS AND PROBABILITY Definition: If X is a random variable designating the number of successes in n Bernoulli trials having the set of values (0,11, ...... n), then X is called the binomial random variable. The distribution associated with the random variable X defined above is said to have a binomial distribution with parameter 7 and.p. The result can be restated as follows: Definition: If the random variables X1, Xz ..-.Xn form n Bernoulli trials with p as probability of success in a single trial and if X= X)+ X54 uu..+X, then X has a binomial distribution with parameters n and p [Link] probability function is as follows: "C.p'- Pp) b(x,n, P)=1.°* mp) fe otherwise (10.6) Obviously, X is discrete because the possible values it may assume are counts of successes out of n trials, which are integers from 0 through n. In this distribution, n must be a positive integer and p must lie in the interval 0

,, X;,.and for p= q= 0.5, we have P(X =0)=P(TIT )= P(X, =0)x P( Xz =0)x P(X3 =0) =qxqxq=@ =(4Pato | P(X =1)= P(HIT )+ P(THT )+ P(TTH ) = P(X, =1)x P(X =0)x P(X; =0) + P(X, =0)x P(X =1)x P(X; =0) + P(X, =0) P(X _ =0)x P(X3 =1) =(pxqxq)t(qxPXq)+(qxqXP) = pq? + pq? + pq? =3pq" HWANG) =F P(X=2)=P(HHT)+P(HTH)+P(THH) 6 =POG=I)xPOG=1)xP(%=0) 566 AN INIRODUCTION TO STATISTICS AND PROBABILITY + P(X, =1)x P(X2 =0)* P(X; =1) + P(X, =0)xP(X2 =1)xP(X3 =!) (px pxq)+( PXI¥ PAC TX PRP) = pat p'qt pa=3P 4 awh P(h)=39 P(X =3)= P( HHH Ja P(X, =1)x P(X 2 =1)K P(X =) 2 pxpxp=pra(4)=3 You can easily demonstrate that the above experiment is a binomial experiment. The number of heads X, in which we are interested, is a binomial random variable. Clearly, X stands for the number of successes, which can assume values 0, 1, 2, and 3. We demonstrate below how the above probabilities can be presented jn a more general way. i aE =0)=*Cop'a” For X=0, the sequence of outcomes is {TTT} Since the trials are independent, we can multiply all the probabilities associated with the different outcomes. Each success (H) occurs with probability p=05 and each failure (T) also occurs with -probability q=0.5. Therefore, the probability for this sequence is P(TET)=P(T) X P(T)x P(T)=q X9 * q=q’ Now the question is: how many ways can we obtain ‘0’ .success and ‘3-0’ ="3° failures from an experiment consisting of 3 trials? Note that in the sequence {TTT} there is one, and only one way to arrange these outcomes since the order of the outcome is of no interest to us. Hence the number of ways in which ‘0° successes and hence ‘3° failures can be obtained from 3 trials is simply 3Co=},Thus the probability of getting 0 success in 3 trials is given by 7 8 2. For X=1, the sequence ‘of outcomes is {HTT}. {THT} and {TTH} By similar arguments, the probabilities of these PROBABILITY DISTRIBUTIONS 567 outcomes are P(HTT)=pxqx4=P» P(THT)= qxpxq= pq and P(TTH)= gxq>p= pq. Note that for one success and hence 2 failures out of 3 trials, we have three different arrangements. Here the order {HHT } is different from the orders {THT} and {TTH}. The question we put now is: in how many ways can We obtain ‘1’ success and ‘3~1" ='2° failures from an experiment consisting of 3 trials? The answer is simply *C,=3. Because the - airangements {HTT}, {THT} and {TTH} are all mutually exclusive, we add the probabilities of all these arrangements to obtain the probability of the. entire sequence or simply multiply pa by °C. Hence 2 LY 4 P(X =="C,p'g* =3PT =4 515 = CP 4 Pq Ble e 8 3. For X<2, the arrangements are {HHT}, {HTH} and {THH} and the associated probability for each of ther is p’q. Since there are 1c; =3 arrangements of two successes and one failure in three trials, the probability of getting such an arrangement is 2 2 1 1). 3 P(X =D="Cyp'q? =3P- =4-)|—|==— ( y= CLP a Pd “y(5) 3 4, For X=3, we have exactly an analogous situation as we had for X=0. Here the outcomes are {HHH}. There is only one way to arrange the outcomes. Thus to get ‘3’ successes and ‘0’ failure, the probability will be 1 ; 1 P(X =3=C:p'g?? = 3 =| =| ae 3P 4 P 3 8 In the general case, if the sample of 1 observations, Xi, Xx. ..%n is taken from a Bernoulli distribution, and the probability of a single trial js p and that of a failure is q with q=l—p, then the probability of x successes inn trials is given by 568 AN INTRODUCTION TO STATISTICS AND PROBABILITY PUSS 3) = BG ni pp) "Cpa" x 0, Le 107) The variable X aboye is a binomial random variable and its probability distribution shown in (10.7) is referred to as the binomial distribution. We now formally state the following theorem regarding binomial random variable: Theorem 10.1 If a Bernoulli trial can result in a success with probability p and a failure with probability q=1-p, then the probability distribution of the binomial random variable X, the number of successes in n independent trials is b(x,n, p)="C,p*q"", ¥=0,12.. n Proof: Let us suppose that a sequence of n Bernoulli trials results in x successes and n-—x failures. If we denote a success by S and a failure by F, then one possible outcome of the sequence is of the form SSSicceS, EF ail. (10.8) The above arrangement shows that x successes will be followed by n—x failures. Since we have assumed that trials are independent, we can multiply all the probabilities corresponding to the different outcomes. If further the probability of success in a single trail is p and that of a failure is g, the probability for the specified order (10.8) is P(SSS....SFF...F) = P(S)P(S)P(S)....P(S). P(F)P(F)...P(F) RTE ELE NSA AS = PPP---P 99-4 x terms =x Yerms x gnnx =P"q This is one possible arrangement of x successes and n—x failures in n trials. We now need to determine the total number of different PROBABILITY DISTRIBUTIONS. 569 arrangements of x successes and n-x failures in n trials. This number is equal to "C, and probability of each such arrangement is p*g" *. Because these arrangements are all mutually exclusive, the probability that there are exactly x successes and n-x failures, is obtained by adding the probabilities of all the different arrangements or multiplying p*g”* by "C,. In other words, the probability of x successes and n—x failures for a binomial experiment of n trials with probability p as success and q as failure in a single trial, is . b(x;n, p)="C, pq", RAO AD vccsced n, ptq=1 (10.9) The binomial distribution derives its name from the fact that the n+1 successive terms in the expansion of (q+p)" correspond to the successive values of b(x;n, p) for x=0, 1, 2........ n. That is n (q+ py" =¥,"C, pq" x=0 a "Gup'g'” oe "C pq" =b(0;n, p) + b(sn, p)+ = Soon, P) x0 +" Capt ig! +"C,p"a"" + b(n -1;n, p) + b(n;n, p) ° ua Since p+q=1, we. note that bx, n, p)=1, a condition that must be = 3 satisfied by any probability distribution. . It is important to note that, given n and p, the probabilities in a binomial distribution are identical with the quantities of the terms of the binomial expansion, Note that each of the terms of the above equation, gives the probability of a specified value of x for x = 0, 1, .....n-1, n. The quantity "C, is known as the binomial coefficient, since it is used to evaluating the terms of the binomial expansion. The choice of n and p determines the binomial distribution uniquely, and the different choices produce different distributions (except when p=0). The set of all binomial distributions is called the family of 570 “AN INTRODUCTION TO STATISTICS AND PROBABILITY binomial distributions. The n and p are called the parameters of the binomial distributions. 10.3.3 Properties of the binomial distribution Mean: The mean of a binomial random variable X, designated 1 or E(X), is the theoretical expected number of successes in n trials. In other words, the mean of a binomial variable is the average number of successes when the number of samples of size n taken from Bernoulli distribution approaches infinity. Symbolically, My, = E(X) = ¥ xdein, P) x0 We will prove that the mean of the binomial distribution is np. That is E(X)=np. The proof is as follows: By definition E(X)= YX, i=l = E(X, +X, + = E(X,) + E(Xy)+ Dt Pt ve =np Variance: Let X be a binomial variate. Then x= 9x, where X; is a Bernoulli random variable and Xjs are all independent. Then by definition, the variance of Xis vane $x] f i=l BV(X, +X toa. 5V(X) + V(X) + mT ae PROBABILITY DISTRIBUTIONS: 571 This is so, since variances of independent random variables are additive. But by (10.4) VAR) + Veet Thus V(X)=npq. Moment generating function: The’ moment generating function: of a binomial variate X definition M y(t) = Ble") = Soe" ban, p) x=0 is, by 2 Sie "C piq’! Fa =Y'C (pe yg" x0 = (9+ pe’) respect to ¢ and set =0. Thus d me Mi = 80x) =| 5+ | . dt 0 = bape! (q+ pe)" ],-0 =np @ uw -0)-| Seas pe)" dt i ; = brea + pe’) 4 thn De*e* at pe)" Jn =np+n(n-1p? Similarly you-can show that 572 AN INTRODUCTION TO STATISTICS AND PROBABILITY r axa Wi, =X y=[seeare) | 10 =np+3n(a—Dp* +n(n—D(n-2)P" and ; dt anr'| Stare] =np+Tn(n—Dp? +6n(n—n— Dp" +m) (n-2(n-3)p"* 1-0 The raw moments obtained above can now be used to obtain the central moments as follows: u,=0 by = Hy WE emp tn 1p? —(npy° = np(1+ np —p— mp) =npQ- p)="pq my =H Bein + 2? = [np +3n(n-Dp? tn - DO —2)p') —3nplnp + nu — Ip?) + AnpyY° =np(l-3p+ 2p”) [on simplification] =np(— pyl-2p) =npaqq— P) tg =H Aph + 6410 oS 3H =[npt+7n(n— Dp’ +6n(n—N(n— 2p +nn-No -2)n- 3)p') —4np|np+3n(n—l)p* +n(n-)(n—2)p"] +6(np) [ap+n(n—Dp"}-3(p)" =3n' p’(1 p)’ +np(l—p)-6p + 6") =3n’ p°q? +npq (1-6p9) = npgll+3(n~2)pql _{on simplifiction] | PROBABILITY DISTRIBUTIONS 573 ‘An alternative way of deriving raw moments and hence central moments is to expand the moment generating function Mx(tiof the tyr! in the binomial distribution and collecting the coefficients of expansion of Mx(t) Thus expressing Mx) = (q+pe'" as Gere) bok at in | A 5 ; i allt p ttc to tenccet a 3! 2 at Po | n(n—1 { Po | =ltnp teatytee lt Pp teatootee | te aw 3! +. BE or 3? Collecting the coefficients of t, 7/21, P/31.....-. , the first, second, third and higher order moments can be obtained. Skewness and Kurtosis: Having calculated the moments, to measure the skewness and kurtosis we can find expressions for Bi and Bo of a binomial distribution as follows: 2 fapatg— iF _ (a= py. 0-20" 3, -2,- ee Hy (npqy np4q npa 3n2p2q? +npq(1=6P4) _ 4, 1-6P4 npq (npq)” e binomial distribution is skewed and The value of B; indicates that th since Bo is greater than 3, it is slightly leptokurtic. When p < 0.5, the _ skewness is positiv hecomes symmetrical for p=0.5. Moreover, as " the number of trials {creases indefinitely, By 0 and B23. — Cumulant generating function: ‘Phe cumulant generating function of Ba 62=—7* Bo binomial distribution is given by and when p>0.5, it is negative. The distribution - 574 AN INTRODUCTION TO STATISTICS AND PROBABILITY K(t)=logeM x (t)=logl(q+ pe! )" ] =nf[log(q+ pe! ] =n [log (1— p+ pe )] 2 3 +44 te =nlog [1+ BCT arta gee 2 3 * Ce pi—+ past Par Pay 2 aul = nlog [1+(pt+p Expanding the log series in the right-hand side and equating the r coefficients ‘of Fin the expansion, we can have the first four iH rt cumulants in terms of the central moments. Example 10.1: Show that for binomial distribution, the (r+1)th cumulant is expressible in the form Ky = mae r2l dp and hence find the first four cumulants. Proof: We have by definition a i K, =|“ togM x(t)| =| los (q+ pe! )" a" me at” a’ t =n ser entah Pe ) 1=0 10 Differentiating with respect to p, we have ak, _ [at (et-1 dp. dt’ | qu pe! ae [since q+ pe’ =1+ p(e' ~1) and hence plat rene -1) PROBABILITY DISTRIBUTIONS Again replacing r bv rel, qe / Kyi an i log(q+ pe ) rH it 10 lat =H ae 4 gar ve) dt’ dt t=0 whi 2h =wiht at at pel Meo Multiplying . by pq and subtracting it from Kies, ‘ dK, a’ aol Ca — pq = p| ——| r+i7 PI da an’ \ ax pel q+ pel - d’ {q+ pe! =np| =| a \ gt Pe jh i9 r =| 40 Lar t=0 =0 from which eh dk K...=pq-— ge = PQ ap ‘ Putting r=1, we get d Ky =pq-—— 2 payin) [ =npq Putting r=2, we get d Ky = paz, (00a) 575° 516 AN INTRODUCTION TO STATISTICS AND PROBABILITY .d = npg & wt p)=npa-2p)= nea a P?) ip Putting r=3, we get d Ka= pq [np q- P= pq—[np\- p 1-2?) dp dp d =npq—I( P -3p?+2p°)] dp = npq{\-6p +6?" 1 = npq(\-6pq) 10.3.4 Binomial recursion relation Tt is usually a laborious job in calculating binomial probabilities directly from the binomial function fix, ”, P)- A much simpler approach is to use the following recursion formula for such cofhputation: n=* P pon, p) (10.10) xt q The recursion formula is arrived at as follows: b(x+hn, p)= By definition ! a ="C FE n x AARX (a) b( x; 1m, PJ CxP 7 qla-a q and. atl noel ni ani helt = b(xthn, p)="CouP Gapin-x-D! Dividing (b) by (a), the result follows. piigh=! @) Thus b (x+1; 1, p) can be obtained from b(x; 7, p) knowing n and p. Example 10.2: Suppose @ biased coin is tossed 5 times. If. the probability of obtaining head in a single toss is p=0.398, find the probability of obtaining 0, 1, and 2 heads. Solution: The binomial probability function in this case is b(x,5,0.398)="C, (0.398) (0.602), x=0,1,2,3,4.5. PROBABILITY DISTI RIBUTIONS 577 Hence P(X =0) = b(0,5,0.398)="Cy (0.398) (0.602)°° = (0.602)° = 0.07907 To obtain the probabilities for x=1, 2, we can use the recursion formula (10.10). Setting x=0, n=5, p=0.398 and q=0.602 in (10.10), we obtain the probability of one head: P(X =1) = B05, 0.398) = n=.b0, 5, 0.398) q 0.398 =5x 5 x0.07907 = 0.26137 Similarly, for two heads, put x=1 in (10.10) P(X =2)=(2; 5, 0.398) _ 5-1 0.398 5, 5, 0.398) 1+1 0.602 =2x0.6611x 0.26137 = 0.34558 Putting x=2, 3, 4 successively, we can similarly obtain the probabilities for three, four and five heads. 10.3.5 Binomial probability table It is almost always a cumbersome task to compute the probability of every outcome for a large set of binomial trials. Extensive tables are available for computation of binomial probabilities. The tables are put in many textbooks in two different forms. One is available to enable you to compute the probability of observing exactly x successes ina binomial experiment consisting of n trials with probability of success in a single trial equal to p. This is of the form b(x; ”, p). The other table gives for the same binomial distribution the probability of observing r cor less successes. This table gives the cumulative probability upto and including the point x rather than the probability of a single number of successes. This is of the form P( Xr), where 578 AN INTRODUCTION TO STATISTICS AND PROBABILITY f PCX 1) and (v) P(X23). Solution: To evaluate the above probabilities, we directly use the binomial probability table included in Appendix IJ. 4 @P(X <4) = THOS 5,0.3) x=0 =0.9977. (ii) P(X =2)=6(2,5,0.3) = P(x <2)- P(X <1) = 0.8370-0.5283 = 0.3087 (iii) P(X <3)= P(X <2) = $(x;5,0.3) =o =0.8370 (iv) P(X >1)=1-P(X SI) =1-0.5283 = 0.4717 P(X $2) — 0.8370 (v) P(X 23) = PROBABILITY DISTRIBUTIONS 579 10.3.6 Binomial frequency distribution If n independent trials constitute one experiment and if this experiment is repeated N times, then we obtain what is known as the binomial frequency distribution. Thus we expect x successes to occur N"C,p*q" “times. This is known as the expected frequency of x successes in N experiments and the possible number of successes together with the expected (theoretical) frequencies will be said to constitute the binomial frequency distribution. It is in fact a theoretical distribution and in practice, the observed frequencies will not always coincide with the expected frequencies of this theoretical distribution because of sampling variations. Nevertheless, this is a useful device that permits a comparison of the experimental results with the theoretical frequencies. For N experiments, each of n trials, the expected frequencies of 0, 1, 2, .n successes are given by the successive terms in the binomial expansion of NV (q+p)", where p+q=1. Example 10.4: A biased coin is tossed 4 times and the number of heads is observed. The experiment is repeated 500 times in all. The results of the experiment were as follows: Number of heads | 0 i | 2 3 4 Frequency 12 | 50 [ 151 | 200 | 87 (i) Find the probability of obtaining a head when the coin is tossed. (ii) Calculate the expected frequencies fitting a theoretical frequency distribution Solution: For the observed distribution gop xf 0x12) + (1x50) + (2x15) + (3x 200) + (487) 500 580 AN INTRODUCTION TO STATISTICS AND PROBABILITY (i) Let X be the random variable ‘the number of heads obtained in 4 tosses’. Then X is a binomial variate with n=4 so that the mean, E(X)= np =2.6 and hence p=0.6S. Therefore the probability that the coin will show head is 0.65. Thus X is b(x; 4, 0.65). (ii) To evaluate the expected frequencies, we need to compute P(X=x) for x=0, 1, 2, 3, 4. Since P(X =x) =0(x:4,0.65 =4C, (0..65)*(0.35)4* P(X =0) = (0.35)* = 0.01500625 P(X =1) = 4(0.65)(0.35)° = 0.111475 P(X = 2) = 6(0.65)? (0.35) = 0.3105375 P(X =3) = 4(0.65)’ (0.35) = 0.384475 P(X = 4) = (0.65)* = 0.1785062 The theoretical distribution is now obtained by multiplying each of the above probabilities by the total frequency 500. The resulting distribution is thus Number of heads 0 | 1 | 2 a 4 Frequency s |s6 [| 155 | 192 | 89 Is this fit good? Apparently a comparison of the frequencies of this fitted distribution tends to demonstrate that the observed frequencies are reasonably close to the expected frequencies. A more valid conclusion can be drawn through a statistical test known as chi-squared test, which is beyond the scope of this text. 10.3.7 Shape of a binomial distribution The shape or pattern of the binomial distribution depends on the values of p and n. If p=q=0.5, the distribution will be symmetrical regardless of the values of n. If p #q, the distribution will be asymmetrical. Given a particular n, the more the difference between p and q, the greater the skewness of the distribution will be. When p < q, the distribution will be positively skewed; when p>q, it will be negatively skewed. PROBABILITY DISTRIBUTIONS 581 However, as the value of n increases, the distribution will become less and less skewed. When n becomes infinitely large, the distribution will approach symmetry irrespective of the difference between p and q. The effects of increases in n on the shape of the binomial distribution are shown in Figures 10.2 (a) and (b) The first set of figures is drawn on the assumption that p=q. The distribution is always symmetrical. As the value of n increases, the bars become narrower and more numerous. As n approaches infinity, the bars become vertical lines with no space in between, and the distribution becomes a bell-shaped smooth curve. The second set of figures is drawn on the assumption that the probability of a success on a single trial is 0.1. It shows that as x becomes larger and larger, the skewness of the distribution disappears and in the long run the distribution becomes continuous. Thus it is apparent that as the value of n increases, the binomial distribution becomes a continuous and symmetrical distribution whether or not p ~ and q are equal. Example 10.5: Obtain the first four raw moments of the binomial distribution and hence derive the corresponding central moments. Solution: This example is designed to demonstrate that the moments of ‘a binomial distribution can be obtained without computing the moment generating function or cumulant generating function. Let X be a binomial random variable having the probability function b(x, n, p). The rth moment about the origin is py, =E(X’). (i) When r=1 Hi = EX) n “ Sx'C,p'a" a) = 0.9" +1."Cyq"! p+2."Caq" p+... mp" =np(q + py =np 582 AN INTRODUCTION TO STATISTICS AND PROBABILITY Figure 10.2: Binomial distribution for varying values of n and p (a): p=q=0.5 (b) : p=0.1 n=l PROBABILITY DISTRIBUTIONS 583 (ii) When r=2 B3 = E(X? ) 2 Ag px nex = dx" "Cyp'q x=0 n = LDi{xtax-1)j] "C,p*q"™ x=0 = Sx "Cpptgt + Saxo) "Cyp%q" x=0 x0 n =p +ufu-l)p* S"[Link]* a" x=2 =np+n(n—1)p?(qt py"? =np[1+(n-1)p] =np(q+np) 2.2 =npq+n'p Hence pt. =p ae =npq (iii) When r=3 wy = EX?) S30 x n-x x Cyp'¢ xed n B(x + 3x(x—1) +x x- 1) x-2)) "Cy p*q"™ or) ‘i “ $e "C,piqt! +33 a0x-1) "Cyp%q ror =) nox 4 x Hex # Dal x-1)(x-2) "Cy p*q 0 enplg > pi +3n(n=1)p?(q4 p) yn n-2 Hine lin=Up\q-=p np + 3n(n= Lp? + n(n=1\(n-2)p? 584 AN INTRODUCTION TO STATISTICS AND PROBABILITY Hence vis =H — 3a + 2H? = np+3nn—1)p> 4 n(n 1) n—2)p? —3npl n? p? + npq)+ 20 p> =np[1-3p+2p] =npdq~P) (iv) When r=4 Expressing x‘ as follows xf ext Tx(x—1) + 6x x— Ix - 2) + xe DE - 2—3) And substituting as before i, =np + Tn(n— lp? +6n(n Dn —2)p> +n(n—D(n—2)(n-3)p* Using the relation Mg = Wy AMI + OME HS ~ 3ux(* , and substituting and simplifying Hy =npqil + 3(n - 2)pq) Example 10.6: A fair coin is tossed 5 times. Find the probability of obtaining (i) exactly 4 heads (ii) fewer than 3 heads. Solution: The experiment is a binomial one. Here p=0.5 and n=5. eet 5 G) P(X =4)=B(5 5, osc(5) (3) as DJ: 32 2 Gi) PX <3) = YBa 5,4) = BO; 5.4) + BU: 545) + BO: 5.2) =o 0 5-0 I ae 2 5-2 1y(1 1\(1 1\(1 sci—l|—-| +*al=|is sc) = || = 4.22 Aa }l2} 7 Gla) \2 iS 404 2 Ee ae 32 32 32 2 Example 10.7: The probability that a patient recovers from a delicate heart operation is 0.9. What is the probability that exactly five of the next seven patients undergoing this operation survive? PROBABILITY DISTRIBUTIONS 585 Solution: Assuming that the operations are made independently and p= (0.9 for each of the seven patients, b(5,7,0.9)="C5(0.9) 0.1" = 21x0.5905x0.01 =0.1240 Example 10.8: A traffic control officer reports that 75% of the trucks passing through a check post are from within Dhaka city. What is the probability that at least three of the next five trucks are from out of the city? Solution: Let X be the number of trucks that pass through are from out of Dhaka city. The probability of such an event is then p=i-.75=1/4. Hence 5. Px 23) =D bG5, 4) =bG: 5.4) + BS.) FBG S. 2) x=} a 2 4 1 3 0 jg 2 co) es c(t a 4)\4 4)\4 4)\4 ee 106 _ 9.1035 1024 1024 1024 1024 The problem can be solved with the help of binomial table also using 2 the relationship P(X23)=1-P(XS2)=1— Ses 5,0.25) x=0 Example 10.9: It is known that 75% of the mice inoculated with a serum are protected from a certain disease. If three mice are inoculated, what is the probability that at most two of the mice contract the disease? Solution: Let X be the number of mice inoculated. Assuming that the inoculation of one mouse is independent of inoculation of the other mice, p=3/4, Since n=3, 586 AN INTRODUCTION TO STATISTICS AND PROBABILITY 2. P(X $2)= YO 3.9) x=0 0 3 I 2 2 I 372 =*C) —||> wal S & wef2 z 4}\4 4}\4 4)\4 1.9.27 _37 s+ t+ eS 64 64 64 64 Example 10.10: Five coins are tossed and the experiment is repeated 200 times. The following table gives the frequency distribution of- the number of heads obtained. No. of heads 0 1 2 3 4 5_| Total 2 |s6 [74 139 118 | 200. Frequency Assuming that the distribution of heads conforms to a binomial experiment, we estimate the probability of obtaining head in a single toss and hence the expected number of heads. Solution: The mean value of the distribution is fx, _0+56+ 1484117 +7245 _| 99 Shi 200 , x= Equating this mean to the binomial mean np, we have np=1.99. Hence p=0.398. Thus the estimated function of the binomial distribution is b(n, p)=b0s 5, 0.398 = 5¢,(0399' (0.603, x=0,1234,5. The expected frequencies can be estimated from Nb(x, 5, 0.398) = 200 5C,(0.398)" (0.602) , x=0,1,2,3, 4,5. We use the binomial recursion formula (10.10) to estimate the probabilities for x=l, 2, 3,4 and 5. For x=0, direct computation yields b(O, 5, 0.398) =(0.602)* = 0.07907 Putting x=1, 2, 3, 4,5 successively in (10.10), PROBABILITY DISTRIBUTIONS 587 5-0Y 0.398 -5,0.398) =| ——— | = | (0.07907 = 0.26136 peso) (53) Ser) ) 5-1) 0.398 5,0.398) =| -— | | (0.26136 = 0.34559 naan) (24\ Geen ) 5-27 0.398 3,5,0.398) =| —— | <= | (0.34559. = 0.22847 mk ) (| can ) 5-3 0.398 4,5,0.398) =( 2 | 2 | (0.22847)= 0.07553 seen (3) San) ) 5-4 0.398 5,5,0.398) =| ——— |< - | (0.07553 = 0.00998 eee (| Son ) Multiplying each of the above probabilities by N=200, the expected frequencies for 30,1, 2,354.5 ae obtained. These are shown in the table below: The number of successes *+ together with the expected frequencies constitutes what we call the binomial frequency distribution. Example 10.11; (i) In a binomial distribution, the mean and the standard deviation are 36 and 4.8. Find n and p. (ii) Is it possible to have a binomial distribution with mean 5 and standard deviation 3? Solution: (i) Given np=36 and. ympq = 4.8 Squaring the second term, mpq = 23.04 or 36q=23.04. This gives q=0.64 and hence p=0.36. From np=36, we obtain n=100, when p=0.36. ——O2RR——— 589 PROBABILITY DISTRIBUTIONS 588 AN INTRODUCTION TO STATISTICS AND PROBABILITY : (ii) When standard deviation is 3, the variance is npq=9. Further the =r, Para Le 9 9 or Pq : fave mean of the distribution np=5. Hence q= eo = = 1.8, which is Multiplying both sides by pq and rearranging the terms, we impossible, since P+q=1. Hence it is not Possible to have a binomial 0.11) distribution with mean 5 and standard deviation 3. It is important to _ i du, 7 note that for a binomial distribution, the variance can not exceed the Hr = PY OTH, dp mean, since p+q=1. i ‘ion Putting r=1, 2, 3 successively in the above equatiot = Ny +— |= pdn+0)=n since W, =0 andy = 1] . Ly = PG My Pe ) pa ,=0 0 ; P ) Example 10.12: Prove that for a binomial distribution with Parameters n and p, the following relation holds: du = 4 hy, Hyat m4 nr, =e } dp, Hence find the first four moments. HM; = pa| 2nw, + dp Solution: By definition, the rth central moment of the binomial variate (4 Xis =p 04 on. 40% A= Xe —npy b(x,n, p) =npq(q- Pp) x=0 . “ dts = >» "C.p* (I= p)"*(x— npy My = pq 3nH> rs x20 iffeienti 2 pq + Linpatq- P| Differentiating 1, with respect to P = pq| 3n’ pq + pla q an 2, oe 2" PPPS amp! + SCCx—npy 4 pia pyr= = pq(3n? pq + n(l- 6p + 6p ») Pp . ip = dp =3n? pig? + npg(l —6pa) SME "CeP (= pl *(x—np = npll +n -2) Pa “ on rs from a ral . ility that a person recovel : n 10.13: The probability # tacted this i Srcta-ney| ot Zapp tao") Ss aicare is 0.4. If 15 people are known to have ae iy ° " fi se, what is the probability that (ji) at least 10 peop! < jisease, i =p + DC x -my Lp n—x)\l-p)™ +x(1- py xp] from 3 to 8 survive (iii) exactly 5 survive? =0 le surviving. ion: Let X be the number of peop =-rm,++ $ "Cp'l=pY(x—npy" fon cme = ™ Then 590 AN INTRODUCTION TO STATISTICS AND PROBABILITY (i) P(X 210) =1-P/X <10) 9 =1~ Y(x;15,0.4) x=0 =1-0.9662 [using tabular value] = 0.0338 8 (ii) PBS X <8)= YH x,15, 0.4) x=3 8 2 = Yb( x15, 0.4)- Sb(x:15, 0.4) x=0 x=0 = 0.9050-0.0271 = 0.8779 (ili) P(X =5)=b(5;15,0.4) 5 4 = Yb x:15,0.4)- ¥b(x:15,0.4) x=0 x=0 = 0.4032-0.2173 = 0.1859 Example 10.14: A binomial random variable X has a mean 4 and variance 3. Determine its probability function. Solution: Assume that the variable has n and p as its parameters. Then np=4 and npq=3. Solving for n, p and q, we have, p=tandg=+ n=16. Hence the probability function of X is , bOn16, = *c(3] (3) s . =012 Example 10.15: Twenty percent of the TVs produced in an industry are defective. If 4 TVs are put in a box for marketing, in how many boxes do you expect to have (i) one defective TV (ii) two defective TVs (iii) at most 2 defective TVs in a consignment of 2000 such boxes? Solution: Here p=20%=1/5 and g=1— p =4/5. If X stands for the number of defective TVs, then X can assume values 0, 1, 2, 3, 4. Hence PROBABILITY DISTRIBUTIONS 3 1\(4) _ 256 i) P(X =1)=“*C,| =| — | === =0.4096 () PX =) (5\§) Be Hence number of boxes having one defective TV is N xP(X=1)=2000 x 0.4096 = 819 1) (4) _ 96 ii) P(X =2)=4C,| =] | — | =~ =0.1536 (ii) P(X =2) (3) (5) 535 Hence number of boxes having two defective TVs is N xP(X=2)=2000 x 0.1536 =307 (iti) P(X $2)= P(X =0)+ P(X =1)+ P(X =2) 07 4x4 174s Bi 052) a (iV (ay an f1Vf4y afl) (4 ="Ol=)lo| + Glelle Cl=| |< (3) (3) G5}(3)* (5) F _ 256 256, 96 _ 608 = = = 09728 625 625 625 625 Hence number of boxes having at most two defective TVs is N xP(X <2)=2000 x 0.9728=1946. Example 10.16: Seventy percent of the passengers who travel on Mohanagar Provati to Chiattagong from Dhaka buy “Daily Star’ at the bookstall before they board the train. The train is full and each compartment holds eight passengers. 4) What is the probability that all the passengers in a compartment have bought the ‘Daily Star’ b) What is the probability that none of the passengers in a compartment has bought the ‘Daily Star’ ©) What is the probability that exactly three passengers in a compartment have bought the ‘Daily Star” dd) What is the most likely number of passengers in a compartment to have bought the “Daily Star’? Solution: Let p be the probability of buying a ‘Daily Star’. Then p=0.7 wnd qe0.3, Let X be the random variable denoting the number of 592 AN INTRODUCTION TO STATISTICS AND PROBABILITY passengers who bought the ‘Daily Star’. Then X is a binomial random variable with n=8 and p=0.7. Consequently, P(X = )=*C,(0.7)' 0.3)" 5 x=0,L..... Hence (a) P(X =8)=(0.7)' = 0.0576 (b) P(X =0)=(0.3) = 6.561x10~ (c) P(X =3)=*C;(0.79 0.3)’ = 0.0467 (d) To find the most likely number of passengers who have bought the ‘Daily Star’, we consider E(X). Here E(X)=np=8x0.7=5.6. Therefore we consider x=4, 5, 6...--. to find the value with the highest probability. Computation shows that P(X =4) = 0.1361 P(X =5) = 0.2541 P(X = 0.2964 P(X =7) = 0.1976 So the most likely number of passengers to have bought the ‘Daily Star’ is 6, since P(X=6) is the highest. 10.3.8 Binomial proportions The last two distributions were concerned with the number of successes in n trials. Quite frequently, however, we are interested not in the number of successes but rather on the proportion of successes. If P (read as p hat) is used to designate the proportion of successes in n trials, then p is also a binomial random variable defined as ' eae p= + n where, as before, X is the number of successes in n trials. Being a proportion, p takes on fractional values such as I/n, 2/n. PROBABILITY DISTRIBUTIONS 593 Similar to the binomial distribution, the distribution of the proportion has its mean and standard deviation as below. x)= 4% = 200)» and via)=v[ x= V0 10.4 HYPERGEOMETRIC DISTRIBUTION The binomial distribution is based on the assumption that the population is infinite and that the random sample is taken with replacement so that the values of the observations are independent of one another. The probability thus remains unchanged for each successive observation. When the population is finite and the random sample is taken without replacement, the probability will change for each additional observation. Under such circumstances, we shall have a probability distribution known as the hypergeometric distribution. In a hypergeometric distribution, we are interested in the probability of selecting x successes and n—x failures from the N-k items labeled ‘failures’ when a random sample of size n is selected from N items. ‘This is known as the hypergeometric experiment. A hypergeometric experiment is one that possesses the following two properties: 1. Arandom sample of size n selected from N items 2. k of the N items may be classified as successes and N-k are classified as failures. If X represents a hypergeometric random variable, then the probability distribution of the X will be called the hypergeometric distribution and will be denoted by h(x; N, n, k), since its values depend on the number ‘of successes k in the set N from which we select n items. a

You might also like