0% found this document useful (0 votes)
3 views72 pages

COSM Unit 3 Computer Oriented Statistical Methods

This document provides an introduction to correlation and regression analysis, outlining the methods for determining relationships between two variables and predicting outcomes. It explains the concepts of positive and negative correlation, the correlation coefficient, and includes practical examples and problems for calculating correlation. Additionally, it discusses properties of the correlation coefficient and its independence from changes in origin and scale.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF or read online on Scribd
0% found this document useful (0 votes)
3 views72 pages

COSM Unit 3 Computer Oriented Statistical Methods

This document provides an introduction to correlation and regression analysis, outlining the methods for determining relationships between two variables and predicting outcomes. It explains the concepts of positive and negative correlation, the correlation coefficient, and includes practical examples and problems for calculating correlation. Additionally, it discusses properties of the correlation coefficient and its independence from changes in origin and scale.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF or read online on Scribd
Correlation 1.1. Introduction to Correlation and Regression Analysis Some of the most useful statistical tools which allow us to make comparisons between sets of data. In this chapter, we consider situations in which there are two distinct quantities but possibly related. In particular, we consider ways in which we can determine whether the two different quantities are related to each other and, if so, how much we find a relationship and use it in a predictive manner. The methods we use in solving such problems is known as correlation and regression analysis. In any correlation and regression analysis situation, there are three stages : 1. Determine whether the two quantitiés are related and the degree of relationship, if any. 2. Assuming that there is a high degree of relationship, find a formula expressing the relationship. 3. Use this relationship to predict the most likely value of the one variable corresponding to any given value of the other variable. In practical situations, we have to study correlation between two variables and some are listed below. 1. Heights and weights of a group of persons in a competition. Demand and supply of a commodity. Supply and price of commodities. Volume and pressure of a gas. Income and expenditure of a group of persons in a locality. ap en 1.2. Definition of Correlation ‘The relationship between two random variables X and Y is called correlation: There are two types of correlation. 1. Positive correlation 1. Positive Correlation the other variable or decrease in ong able leads to incre alled positive correlation. Increase in one vari i t the other variable is ca} variable leads to decrease Examples : ; 1. Heights and weights of a group of persons. 2. Income and expenditure of a group of persons» 3. Demand and supply of a commodity. in a locality. 2. Negative Correlation Increase in one variable leads to decrease in variable leads to increase in the other variable is calle Examples : 1. Volume and pressure of a gas. 2. Supply and price of a commodity. Uncorrelated Variables. If there is no rel variables are called uncorrelated variables. Example : Income of father does not depend on the age of the son. the other variable or decrease in one ed negative correlation. Jation between two variables, the 1.3. Correlation Coefficient Prof. Karl Pearson discovered the measure of correlation i.e. linear relationship between the variables X and Y, called correlation coefficient. It is also known as product moment correlation coefficient. Correlation coefficient is denoted by r(X, Y) or simply r and is defined as — oe yy CKD rer Ys Ox Gy where Cov (&, ¥) = B(IK = BOO) [¥ - EC) = EO) - BOE) (OR) Cov, =2¥ -DO-=* Y, x- FT im ms Variance (K) = oj = E((IX-E()]} = ECK®) - (CK)? (oR) of =2 Y G-2=1Y x2-@P isl Rie 1 Standard Deviation of X = ox = \Var (X) Variance (Y) = of = E[Y - E(Y)]? = ECy2) — (Eq)? 2ly, -, 1% - (OR) oF Py Or-TR= 5D v2 GF Standard Deviation of ¥ = oy = \Var (%) Scatter diagram ‘The diagrammatic representation of a bivariate data (e. dis tter di is a si 1s Ia)s Ca Yo) oo+ Ons Yon called scatter diagram. It is a simple graphical representation by plotting values of the random variable X on X-axis and values of the random variable ¥ on Y-axis. Scatter diagram for positive Scatter diagram for Scatter diagram for no correlation (r > 0) negative correlation (r < 0) correlation (r = 0) Scatter diagram for perfect positive | Scatter diagram for perfect negative correlation (r = 1) correlation (r = —1) Properties of Correlation Coefficient 1. Correlation coefficient is a numerical value or constant and it does not depend on any units of the given data. 2. Limits of correlation coefficient : Correlation coefficient lies between —1 and +1. Proof: We know cw ¥) "Ox Oy 1 x = a BED Xi -H0i-7) fs Ft BLOT" A * Bete mio, 57. ae zy EOD D 2 [3 (@j-%) v9] i cee e+ = ~ DX @-2 YO -7P A it Let xX =a;, y;-¥ = b; a 2 [3 ats] 2s it wn) (2,7) (24) i i By Schwartz's inequality, if), ag, ...a, and by, by, ... b, are real quantities, then [Soo = (Sl aie atest n° 4a (2 ait] ie. at ey v2) (3,07) (+) ma From equations (1) and (2) My <1 = 1 n 5 i 04 5g yon = 19 = 56 ox = [2337 -c@? = [1878 _ Go? 61919 ni 10 = [A%.2_,2_ [B6562 a oy = nD OW = |p -- 66.4)" = 218 in — - 69 Cov KY) = 2D iB. = os ~ 10x 56.4 = 180.5 jel Cov & ¥) 130.5 = 61919 O71 m= Sx oy.” (6.1319) (21.8) aes There is a highly positive correlation between X and Y. Hence if a student works more, he will get more marks, _, PROBLEM 2, Calculate correlation coefficient to th e following data : 10 15 12 17 13 SOLUTION s % = rn 10 30 100 900 300 15 42 225 inte4 aaa 12 45 144 21025 540 46 289 2116 782 18 33 169 1089 429 16 34 256 1156 544 a4 40 576 16100 960 14 35 196 1225 490 22 39 484 1521 858 1444 760 a 7 y » pe 14840 x xj = 6293 5 Ox = (Eb? -@)P = (oe niet oy = 2 So? -G= (Ba «40 niet 4 Cov (X, ¥) 135 xi- -# 5 = Tp a6.9) (88. 2)= nit Cov (% ¥) ___6.64 T= Gay ~ 4.27) (4.98) “. The correlation between X and Y is 0.3123. i.e. X and Y ar'v positively chrrelated. = 0.3123 | Method I! of calculating correlation coefficient | Ny PROBLEM 3: The marks obtained by 10 students in Mathem ties and Statistics are given below. Find the correlation between the two subjects. Marks in Mathematics Marks in Statistics 1% 30 60 80 53 35 | 15 40 38S. 91 58 ained in Mathematics - the obté hen ON a a mathe ined in Statistics ¥ be marks the obt Let us shift the origin and seale of X and Y a8 X-53 uss =. 3, va Bey- 63 m= xi- 58 2 22 484 a 3 529 = z= jy; = 3280} im a terre To ~ $6.7 = 17.66 7 12 Cov (U, V) Bat 2 Pee £0 85.6) 6.7) 0 oan = oo c3,%9 vy) Tuv Sy oy 290.48 “19-14 x 17-68 = 0.8593 Since ray = ruv, therefore rey = 0.8593 PROBLEM 4. The heights of fathers and song are given below in inches. Find the correlation coefficient between them. Heightsof | 65 63 67 64 68 62 70 66 68 67 69 71 Fathers 68 66 68 65 69 69 68 65 71 67 G8 70 SOLUTION. Let X be the heights of fathers Y be the heights of sons Let us shift the origin and scale of X and ¥ as \ X-66 \ v =X—66, vet ay 68 \ — x | oy | up=x,-66| v= 9-68 uf ve ups 65 | 68 -1 0 1 0 0 63 | 66 -3 -2 9 4 6 67 | 68 a 0 1 0 0 64 | 65 ~2 -3 4 9 6 68 | 69 2 1 4 1 2 62 | 69 “4 1 , 16 1 -4 70 | 68 4 0 16 0 10 66 | 65 0 3 1nd, 9 ° 68 | 71 2 3 4 9 6 67 | 67 1 -1 1 1 -1 69 | 68 3 0 9 0 0 7 | 70 5 2 25 4 10 iat (2e-ent = 2.66 “Et -@P= e-¢ 0.10? 2177 : Cov (U,V) =2¥ uv Bd Nie = 3 _ (0.67) (-0.17) = 2.20 Cov (U,V) 220g. 47 = ou ay (2.66) (1.77) Since rxy = rpy, therefore rxy = 0.47. i.e., heights of related. fathers and sons are Positively PROBLEM 6. To calculate the correlation coefficient the following totals were 00, Ly? = 250, xy = 356, obtained for 30 pairs of observations Ex = 120, Ly = 90, te But, at the time of checking, it was found two pairs of observations are noted wrongly, Wrongly noted observations Correct observations = 8 10 12 7 / Obtain correlation coefficient for correct observations, SOLUTION. Given that n = 30, 2x = 120, Sy = 90, Dx? = 600, Dy? = 250, Dey = 356 These value are to be calculated again by subtracting wrong pairs and adding new ‘ pairs, New values are : (Ze)noy = 120-8 - 12 + 18+ 10 = 128 (Zy)new = 90-10-74 12 +17 = 102 (Zx)new = 600 — 8?— 12? + 18? + 102 = 816 (By now = 250 - 102-72 + 12? + 172 = 534 (Ze new = 356 — (8 x 10) - (12 x 7) + (18 x 12) + (10x 17) = 578 1, 198 erat u o4 Ble sie 3] g " | 8 u o =-_ 8" Cov &X, ¥) -35 =F aan iad=ars 1 = ox = Ene (5)? BE (4.27)? = 2.99 1 <2! Vn? 3)? = [84 (3.4)? 2259 rey = SC OGY) 475 Ox dy "2:99 x 2.60 = 06854 oy = 1. Sales and advertisement expenditure of a commodity is given below. Obtain correlation coefficient between them. Advertisement expenses(In [39 65 62 90 82 75 25 98 36 78 54 48 thousands of Rupees) Sales (In lacks of 47 53 58 84 65 68 GO 89 51 84 66 55 Rupees) _ | Ans. r = 0.797 2. Find correlation coefficient to the following data. Ans.r=1 3. Obtain correlation coefficient to the following data : ix. 11.1 10:3. 12 15.1. °13.7 -18.5 .17.3 142 148 15.3 y 10.9 14.2 13.8 21.5 13.2 21.1 164 193 174 19.0 S$ Ans. r = 0.7582 4. Age of wives and-husband of 14 pairs are given below. -Obtain correlation coefficient between the age of wives and husbands. Age of wife 21 25 26 24 22 30 17 24 28 32 31 29 21 28 Age of husband 19 20 24 21 21 24 18 22 19 30 27 26 19 18 5. The following table gives number of blind people per one lakh population in different age groups. Find correlation coefficient between age and blindness. Age in 0-10 10-20 20-30 30-40 40-50 50-60 60-70 70-80 years ‘ Number of blind 55 67 100 111 150 200 300 500 people (per lakh) ics arene meee Ans. r = 0.89 6, A computer while calculating correlation coefficient between two variables X and Y from 25 pairs of observations obtained the following results. n=26, Dx = 15, Dx? = 650, Ly = 100, Dy? = 460, Dey = 508 It was, however, later discovered at the time of checking that he had copied down two pairs as x | y x | 6 | 14 |, whilecorrect values | 8 | 12 | a | 6 6 | 6 | Obtain the correct value of correlation coefficient. Ans. 0.67 7. If X and Y are two random variables and a, 6, c are constants, then obtain the correlation coefficient between aX + bY and cY. Ans. (aro + boy/V a? 0% + 60% + 2abr oxoy 8. If U = oX + bY and V = bX -aY, where X and Y are measured from their respective means and if U and V are uncorrelated, r is correlation coefficient between X and Y, then show that Oy Gy = (a? + B®) ay ay (1 — r2)U2 9.. If U = GX + bY, and V = aX - bY,.X and Y are measured from their respective means and if [Link] V are uncorrelated, r is correlation coefficient between X and Y, then show that Oy Gy = 2ab ox oy (1— r2)12 10. Means of the random variables X, and X, are 5, 10 and variances are 4,9 respectively. If U = 3X; + 4X, V = 3X; — Xp, then obtain correlation coefficient between U and V.. Ans. r= 0. 11. The random variables X and Y are uncorrelated variables with means zero and variances 0,?, 0:2 respectively. If U = X cos 0 + Y sina and V = Xsin a-Y cos a, then show that correlation coefficient between U and V is rt « 0,2 i 22 Vo?= ‘O2) + 40,7 0,2 cosec? 20 12, ies = @X - bY and ris the correlation coefficient between X and Y, then show a @) 0,2 = ao}, + b20% - 2ab r oxy (i) Show = Shay 9k - 9% : 2axoy Hint. Consider @ = 1, 5 = -1) r v Rank Correlation EE 2.1. Rank Correlation “ Rank correlation was introduced by Spearman. Let x4, X2, --+) %, aNd 1, Yo, «+» Yn be the ranks of ‘n’ individuals of the characteristics A and B respectively. The correlation obtained for the ranks of individuals of two characteristics A and B is called rank correlation. Ranks of the individuals may or may not be same. 2.2. Derivation of Spearman Rank Correlation Coefficient Possibility 1. Ranks of the individuals are different (not same). Let the ranks of the individuals due to the characteristics A and B be 1, 2, 3, ... n (which are need not be in the same order) and in general x;#;. _nm+l) ntl a ae) Similarly 2D) bE 2 252424. +n2)— 2) a+ Qn+) (+1? - 6n. 4 (n+1)(2n+1) (n+1)? 6 7 4 Similarly ; ok = 03 42) Let p be the rank correlation coefficient Cov (KY) = oe 8) Cov (X,Y) =p oxy Let us consider deviation of the ranks as d; | i— 2) —(;-9) Squaring, summing over i from 1 ton and dividing with n, we get ahaa, le mong 2)-Q:-Y)? 15 @- +23 o- yP- 22S 2) ;-9) Rind = 0% + 0} -2 Cov (X, ¥) =20%~2 poxoy (from equations (2), (4)] = 20% (1p) * - Ld? 6d a? ei os ft n?—-1\ 2 (21) ea 6 x d? pda PROBLEMS . PROBLEM 1. Compute Spearman’s rank correlation coefficient for the following data. xX 20 14 36 29 5 il y | wi | 9 25 10 2 | SOLUTION n=6 Ranks of X | Ranks of Y x Y { 1 vt 20 19 3 : 14 9 4 4 1 1 0 0 29 10 2 a -1 1 5 2 6 6 0 0 1 6 5 5 0 0 — n x d?=2 isl 7 6S ap zt 6x2 Bet) = 0:9429 P= Gea PROBLEM 2. The ranks of 15 students in Mathematics and Statistics are given below. Obtain rank correlation coefficient between them. Ranks of 1 2 6 9 11 15 10 8 4 7 5 14°13 12 3 | Mathematics Ranks of 10 7 8 11 9 183 15 1 6 8 4 «12 «144«5 2 Statistics SOLUTION Let us denote x; be ranks in Mathematics y; be ranks in Statistics 1 mM qd 1 10 -9 20 7 -5 6. 8 2 9 11 2 PE Y 9 2 15 13 2 10 15 5 8 1 7 4 6 2 7 3 4 5 4 1 14 12 2 13 14 -1 12 5 7 3 2 1 Rank correlation coefficient n 6 d? & ‘ 6x 272 ~n@ay =~ 15 ae ~ 05148 p=l PROBLEM 8. 12 students were participated in a music competition, three judges have given the ranks. Obtain Spearman’s rank correlation coefficient and compare them. Ranks of Judge 1} 1 5 6 10 3 2 9 4 il 7 12 8 (eam of Judge 2| 3 8 5 7 4 6 il 1 9 2 10 #12 Ranks of Judge 3} 4 6 9 8 11 3 12 2 10 5 1 7) SOLUTION. Let us denote x, y and z are ranks given by three judges in the music competition. a ’ z dey dy. dy, a, a a, a 3 4 —2 =3 =1 4 9 2 5 8 6 -3 -1 2 9 1 4 6 5 9 1 -3 4 3) 9 16 10 7 8 3 2 -1 9 4 1 3 4 1 =i -8 -7 1 64 49 2 6 3 -4 -1 3 16 ES 9 9 11 12 -2 -3 -1 4 9 2 4 1 2 3 2 -1 9 4 1 11 9 10 2 1 -1 4 1 1 7 2 5 5 2 -3 25 4 9 12 10 1 2 11 “9 4 121 81 8 12 7 -4 1 5 16 a 25 Xd, | Edz | Lay =102 = 228 =198 n=12 Rank correlation coefficient of ranks given by judges 1 and 2. 6 Za; 6x 102 Ps = 1-7 Ge y=} Fp aa y= 884 Rank correlation coefficient of ranks given by judges 1, 3. 6xrdz 6 x 228 Prs =} 71) =} ye age 1) 02028 Rank correlation coefficient of ranks given by judges 2, 3. 6S dy, 6x 198 pos = 1 Te y= Fa age— yD 7 0.9077 The rank correlation coefficient is high for the ranks of judges 1 and 2, which 0.6434. ie, their judgements are near. Tied Ranks. Ranks of some individuals are same. If some of individuals receive the same rank, the ranks of them are said to be tied. Let us suppose that ranks of m individuals are tied after K ranks, hence they receive the ranks K + 1, K+ 2, ... K+ m. Since ranks are tied for m individuals, they should assign a common rank for all m individuals as the average of K+ 1,K+2,...K+m. For example five individuals got 7th rank after 6 (= K) persons, ‘Then all 6 porsong T+ $+ 94104101 will have to assign a common rank with average of them as" OSU 9. Let p be the Spearman's rank correlation coefficient of X and Y. Cov (X, Y) * Ox Oy 1s a ae YO -HO,-7P) is (1) 1< nat We know o% aod (j- EP =o Y uz Me &. a = S I ic iz. 12 a Similarly >, isl n a Weknow Yd? = Y @-¥? 0 ta -2)-i-y)? = z (uj- vp? isl 2-23 uy ist x wong |B ae s ue 1 § Sa 1 Me g “ie int Now we have to see what significant change tale place ind uf, x uf ond s uy, , U4 * a in case of tied ranks. Since x and ¥ values do not change for tied ranks alao, 2 2 Let S? and S% denote sum of squares of untied and tied ranks respectively. = (K+ 124 (K+ 2)? +... 4 (K =m)" = mK? + (12 + 224... m2) + 2K (L424... bm) = m(m + 1) (21 = m2 AV OM*D stom (mn +) (= 1) (K+2)+..4(K+ my. (2 (K+ 2) +..4+(K+ my m Ss} a : (Genet? 2) ++ (K+ my wm (Mt DEGe De +54 mil en (meaenen tmp m mint] oe mK tog Kemet i a m|K+ sh mj(m?- 1) -e terms With the above adjustments In case of tied vanks, ae sum of the squart are as follows. . _n-1) 1 4 Ty+Ty+ x a) =1 Substitute these values in the equation (1), 2-1 af nae 3 (teeTr+ 2 z 2) Pxy = n(n2— 1) n(n? 1) (“ar -1y) (“S-s) 7 2 Bee, e x ate tee i=l (: ce D_on,) (=D any) which is the formula of rank correlation coefficient if tie occurs. Simple mathematical formulae for numerical calculations : By adjusting only covariance term, but not the variance terms, we have n(n?-1) 1/< nGed 2} (2, dP+Tx+ "| 1 Pxy In(n?—1) n(n?) 12 12 2 * nora) Yds T+] i=l DX d2+Tx+ | ist n(n?—-1) 1 » ! PROBLEM 1. Marks of 11 stu rank correlation coefficient. dents in two subjects A and B are given below. Obtain Student Number 1 2 3 4 5 6 7 8 9 10 i Marks in Subject A | 25 36 20 36 48 52 25 65 385 45 60 Marks in SubjectB | 35 49 30 42 56 68 45 50 42° 55 68 Students Marks in SOLUTION, Let us denote ranks of two subjects A and B be x and y. 1 Marks in | Ranks of | Ranks of| ee a number | Subject A| SubjectB| A=.x; Bay, |P*%n% 1 25 35 9.5 10 0.5 0.25 2 36 42 6.5 8 -15 2.25 3 20 30 11 ll 0 0 4 36 42 6.5 8 -15 2.25 5 48 56 4 3 1 1 6 52 68 3 15 0.5 0.25 7 25 45 9.5 6 3.5 12.25 8 65 50 1 5 4 16 9 35 42 8 8 0 0 10 45 55 5 4 1 1 il 60 68 2 15 0.5 0.25 X d2=35.5 a n=11 In subject A, 25 repeated twice and 36 repeated twice i.e, m, = 2, mg = 2 (r = 2) Hct 1 T= d mj (m?— 1) = 75 lm, (my2— 1) + mg (22-1) =ip (2?—1) + 222-1] =1 In subject B, 68 repeated twice and 42 repeated thrice i.e, m, = 2, mg = 3 (s = 2) FF Rank correlation coefficient with tied ranks is °|, ¥ dz t+ "| im p=1-— oD 6 (35.5 +1425) a ao = 0.8187 2.3. Limits of Rank Correlation Coefficient Limits of rank correlation coefficient are + 1. -l ay, isl es =D iaH +1), +2 x xii where From equations (2) and (3), n(u+DQn+)) o< (3) 3 77% n(n +a? z Rn)? n(n+Q2nt+)_nn+ Dn +2) A 6 = 6 n Cov (X,Y) => DY iyi FF (e142) nel n+l “Cov & ¥) _ “Ox oy p =-1 minimum value of p is -1. From equations (1) and (4), limits of rank correlation coefficient are + 1. -l nD mor aE + bE: & xi =1 is From the equations (4), (5) and (6) Bun +Zy =az +b (0% +z? Multiply equation (3) by Z, EY =az+bz? ~ Subtract equations (8) from (7), nu =6 of a1) +-(2) 4) +5) --(6) 7) +-(8) Since b is the slope of the regression line and regression line passes through @, 9). the regression equation is y-¥ =b(e—z) =X Ge 2). Regression line of ¥ on X is soy, y-y ox (x-x) Regression line of X on Y varinble and y is independent variable. Let the regression line of X on ¥ be xeatby By using principle of least squares, normal equation are * * YL xenatd Dy; ie. ist DX ayi=a Znsod xa fet ist Divide equation a byn, E=a+tby ie. regression line passes through (&, 7) ly =< {We know nix = Cov (X,Y) = 7+ x xni-2F im 12 nh Xs miyi= Wat Ey We know Var (Y) = of “3d P-GP 12 ne at +O Now divide equation (2) by n_ 2 From the equations a, (5) and (6) . By +2¥ =ay +b (0% +9?) Multiply equation (3) by y, EY =ay+by Subtract equations (8) from (7) Bu =6 of Cov (X, TO; Since b is slope and it passes through (%, y), the regression equation is x-k =by- n= Beg -y) .. Regression line of X on Y is be . Ox, * = 1, wok =r oeGsy) Let us consider a bivariate distribution (1 > <1 xr 1 > Zo < AD) ‘We know that <1 = byxbyy $1 1 > byy ta «-(2) From equations (1) and (2), i ber Spo <1 byy <1 3. Regression coefficients are independent of change of origin but not of scale. ie X-a Proof:Let -U ==>", vy. X =a+hu, Y¥ =b4+kv Ox =hoy, Gy = koy and ryy=Tuv Oy ko bye =r ar Fee =£. byy (r do not change) 4. Arithmetic mean of the regression coefficients in greater than the correlation coefficient, provided r > 0. byx + by Proof : We have to prove —™5—"¥zr byx +b Consider “Mer = 2 [- oy, Slo, 2[' ox oy of + o% - 2 lec Ox oy > 0% + 0% 2 2ax ay ' => OY + 0% — 2oxay20 => (ox- oy)? 20 which is always true. Arithmetic mean of regression coefficients is greater than correlation coefficient. 5. Both regression coefficients must have the same sign. ie. if byx is positive, then byy will also be positive. if byy is positive, then byx will also be positive. The reason for this is sign of both regression coefficients depend on sign of the correlation coefficient only. Correlation coefficient and regression coefficients have the same sign. 4.5. Correlation vs Regression The given below are some of basic differences between correlation and regression. Correlation 1. Correlation means the relationship between two or more variables. It measures effect of one variable due the change in the other variables. Correlation need not imply cause and effect relationship between the variables under study. Correlation analysis measures linear relationship only between the variables. Hence practical applications are very less. Sometimes, correlation may be non- sense. Example. There may be a correlation between temperature and height of the persons, which docs not means that real relationship. Regression Regression means stepping back to the means value. It expresses average relationship between two or more variables. But regression analysis indicates the cause and effect relationship. The variable constituting cause is taken as independent variable and the variable constituting the effect is taken as dependent variables. Regression analysis measures linear and non-linear relationships of the variables. Hence practical applications are more. There is no non-sense regression. Correlation measures the direction and degree of linear relationship. It is symmetric ie. ryy = ryy and it is immaterial whether variables are independent or dependent. Regression analysis measures the functional relationship between the variables and it also uses to estimate the value of dependent variable for any given value of independent variable. [t identifies the nature of the variable i which is independent and which is dependent variable. Regression coeffi. cients are not symmetric ie. byx # byy The regression coefficients byy and byy are absolute measures. It measures change in value of one variable to change in the other variable. If the function of the regression curve is known, then the value of dependent variable can be obtain for a given value of independent variable. This value is in the units of measurement of the variable. Correlation coefficient is a relative measure of the linear relationship and is independent of the units of measurement. It is pure number between -1 and +1, lines of Y on X and X and Y. Estimate the value of. [_= [36 | 2] 72 36 | 63 | 47 PROBLEM 1. Twelve Pairs of observations are given below. Obtain the regression y ifx = 65. 85 | 49 | 38 [42 | 62] a0] y | 347 | 125 | 165 | 118 | 14g | 128 | 150 | 145 [115 | 132 [152 | 160 | SOLUTION —. x % aie ve Xi 56 147 3136 8232 42 125 1764 5250 2 165 5184 11880 -36 118 1296 4248 63 149 3969 9387 47 “128 2209 6016 55 150 ~~ 3025 8250 49 145 "9401 7105 38 115 1444 4370 42 132 1764 5544 62 152 3844 9424 60 3600 22500 9000 * = 83636 | > y?=236746| 3, xy; = 88706 ist iz1 1 Cov (X,Y) =— DY xy, - 7 ni 88706 = 7 (51.83) (139.67) = 153.07 Correlation coefficients _ Cov (XY) ___ 153.07 ; r= Gory 7 10814877 9-981 Regression line of Yon X — ig toy yaa (x-2) 14.87 y — 139.67. = (0.9531) 10.8 (x - 51.83) _y = 1.3123x + 70.65 Regression line of X on Y — _ox B-k = or o-y) i 10.8 %- 51.88 = (0.9531) 7 97 — 189.67) x = 0.6922 y— 44.85 Estimation of y If x = 65, by using regression line of Y on X } = 1.3123 x 65 + 70.65 = 155.95 _ PROBLEM 2. The sales and profits of a company are given below. Obtain regression lines. Estimate profit if the sales of the product is Rs. 72 lakhs. eo 75 | 70 | 55 | 65 | 60 | 69 | 80 | 65 | 59 | 61 Profit inlakhs Rs. | 69 | 65 | 45 | 52 | 60 | 62 | 70 | 55 | 45 | 49 SOLUTION Lot X bo the sales of a product: Lot. Y be profit of the product Let U =X - 69, V = Y~ 60 St aM jo ou uy Ly uy; 75 59 6 -1 36 1 6 70 65 ) 1 25 5 55 45 -14 -15 196 225 210 65 52 -4 -8 16 64 32 60 60 -9 0 81 0 0 69 62 0 2 0 4 0 80 70 ce 10 121 100 110 65. 55 4 -5 16 25 20 59 45 -10 -15 100 225 150 61 49 38 -11 64 121 88 =-3. 10 7-38 cae ee can? = 731 790 oy = = a 38" = 8.03 Cov (U, V) =2 DL uy-v.d 609 = 2 _ (3.1) (-3.8) = 49.12 = 70 (-3.1) (-3.8) Cov (U, V) Tay = TU” oy oy 49,12 = — = = 0.8368 = (7,81) (8.03) ag =69, b=60, h=1, k=1 Now E =a+hi=69+(-3.1)=65.9 Cov (X, ¥) =A k. Cov (U, V) = 49.12 Regression line of ¥ on X : ay =P Ny _z II RS (xx) (8.03) ¥- 56.2 = 0.8368 “757 (- 65.9) y =0.9192x - 4.3767 Regression line of X on Y ear 2-F = 2X y_5) _ a x-65.9 = (0.8368) ) (y - 56.2) x = 0.7618y + 23.0886 Ifx = 72, we have to estimate y -. By using regression line of ¥ on X y = 0.9192 x - 4.3767 9 = 0.9192 (72) - 4.3767 = 61.8057 If the sales of the product is Rs. 72 lakhs, then estimated profit is Rs. 61.8057 lakhs. ,OBLEM 4. Regression equations are given 8X — 10Y + 66 = 0, 40X - 18Y = 214 PR ‘and variance of X = 9. Obtain (i) mean values of X and Y i) correlation coefficient between X and Y (ii) standard deviation of Y. SOLUTION. Given that Regression lines are 8X-10Y+66=0 40X - 18Y = 214 and Var (X) = of =9 (@ We know both regression lines pass through the point (Z, 7) 8X-10¥+66=0 40X - 18Y- 214 =0 By solving these two equations, we get X=13, Y=17 Gi) Let regression line of ¥ on X be (It is an assumption) 8X-10Y+66=0 10Y =8X+66 Y =08K +66 : ‘The regression coefficient of ¥ on X is byx= 0.8, Let regression line of X on ¥ be 40X—18Y =214 40X = 18Y +214 X =0.45Y +5.35 The regression coefficient of X on Y is byy= 0,45 We know 7? = byybyy= (0.8) x (0.45) = 0.36 (Since r?< 1, our assumption of regression lines is correct) r=+06 But we know the sign of regression. coefficients and correlation coefficients are -same. Le. r=+06 - roy (ii) Weknow byy = —e . r 0.6 i -8= —5 = 0. Find ROBLEM 5. If the two regression are 2X + SY -8 = 0, K + 2Y 5 f (i) ae coefficient, (ii) means of X and Y, (iii) calculate the variance of Y if - variance of X is 12. byx. 0.8) . (3) oy = 2K- Ox (0.8). (3) _ SOLUTION (i) Let regression line of ¥ on X be 2X + 3Y—8 = 0 (It is an assumption) 3Y =-2X+8 Y =-0.67X + 2.67 byx =-0.67 Let regression line of X on ¥ be X+2¥-5=0 We know 7? = Byy byx = (-2) (0.67) = 1.34>1 Since r?< 1, this assumption of regression lines is not correct. (This is explained for better way of understanding the property) Therefore now consider Regression line of Y on X X+2Y-5=0 “X45 +0.5X + 2.5 byx =-0.5 Regression line of X on ¥ 2X +8Y-8=0 . 2X =-3Y+8 XK =-15Y+4 a byy =-1.5 : We know 72 = byx byy = (-0.5) (-1.5) = 0.75 $1 Since r2< 1, the assumption is correct ” 7 = + 0.8660 We know that the signs of regression coefficients and correlation coefficients are | same. o r =-0,8660 (ii) We know that both the regression lines pass through the point (X, ¥) X+2X-5 =0 2X +3¥-8 =0 By solving these two equations, we get X=1, Y=2. ii) Given that =12 => ox=3.4641 As we know roy OK byx x Ox _ (0.5) (3.4641) ore"; = 0.8660) oy =2 of =4 PROBLEM 6. The price details of dij ifferent commodi Visakhapatnam and Hyderabad are given b ties of two elow. Estimate most likely price ®t Hyderebad if the price at Visakhapatnam is Rs. 195, Visakhapatnam Hyderabad Average Price 165 158. Standard deviation 22.5 13.5 Correlation coefficient between the prices of commodities in the two cities is 0.81, SOLUTION Let X be price of commodity at Visakhapatnam Y be price of commodity at Hyderabad, Given that E % = 165, 7 = 158, ox = 22.5, oy =18.5, r=.0.81 To estimate most likely price at Hyderabad, it requires regressi ion line of ¥ on X, Fare, = = J =); se (w-z) 13.5 9-158 = (0.81) 355 (x — 165) y = 0,486x +77.81 If the price at Visakhapatnam is Rs. 195, then Y = 0.486 x 195 + 77.81 = Rs. 172.58, Hyderabad is Rs. 172.58, PROBLEM 7. The data is given below of marks in two subjects Mathematies an Statistics of [Link]. students, i.e. most likely price at oe Mathematics Statistics Average marks 39.5 49.5 Standard deviation 10.8 16.8 ‘The correlation coefficient between the marks in two subjects is 0.42. @ Estimate the marks in Statistics if the marks in Mathematics is 52. (i) Find angle between two regression lines. SOLUTION. Let X be marks in Mathematics. : Y be marks in Statistics.” @) Given that % = 39.5, ¥ = 49.5, ox= 10.8, oy=16.8, r= 0.42 _ —— To estimate marks in statistics (¥), it requires regression line of Y on X. aE ee y-F Ge (w-x) y—-49.5 =0.42 x88 yy 39.5) : y = 0.6533x + 23,6933 If marks in mathematics (x) is 52 then > = (0.6533) x 52 + 23.6933 57.6649 ~ 58 marks -. Estimated marks in Statistics is 58 marks. (ii) Angle between two regression lines 6 =Tan (Fe, rT ok + of etanct (2210-42 (10.8) (16.8) ) = ten 0.42 BP = Tan7! (0.892) © =41.73° 1. Obtain regression line and estimate the value of X for a given Y = 70 for the following data. x 65 66 67 67 68 69 70 72 | Yy 67 68 65 68 72 72 69 | Ans. y = 0.665x + 23.78, x = 0.54y + 30.74, 3 = 68.54. 2. Obtain correlation coefficient and regression lines to the following data. x 21) 25 | 26 |.24-|.22 | 30 | 17 | 24 | 28 | 32 | 31] 29 | 21] 28 Y 19 | 20 | 24 | 21 | 21| 24 | 18 | 22 | 19| 30| 27 | 26} 19| 18 Ans. r = 0.8534, y = 0.7007x + 4.4875, x = 1.0393y + 2.135 3. The given below are birth and death rates per 1000 population in different years of a country. Obtain regression lines. Estimate death rate of the country if birth rate is 10.5 and also estimate birth rate if death rate is 10.25. Birth rate Death [tate ‘Ans. y = 0.00415x + 8.9364, x = 1275y + 12,3298, 9 = 8.98, Z = 13.63 17.1) 16.5) 154 184] 14.3] 13.6] 12.9] 12.3] 11.7] 11.5] 11.3] 11.3] “ 8.9] 8.9] 8.5] 9.7] 9.0] 8.7] 9.1] 9.0} 9.2 9.3 | 9.3 2a 8.2 5. Estimate the price in Chennai corresponding to the price of Rs. 70 at Calcutta from the following data. Calcutta Chennai . Average price 65 ‘67 Standard deviation 2.5 3.5 Correlation coefficient between the prices of commodities in the two cities 0.8. Ans.) = Rs. 72.6. Regression equations are given 4y = 9x + 15, 25x = 6y +7: Find correlation coefficient and mean values of X and Y. Ans. r = 0.735, % 5 2.565, = 9.52. ‘Two regression lines are given’ below 4x — By =5, 2x-y=3 Find (i) Correlation coefficient (@) Find standard deviation of X when variance of Y is 9 Ans. (i) r= 0.8165 @) ox = 1.8871. PQS TT Se Nhe et Testing of Hypothesis g is a process of testing the significance regarding the population parameter on the basis of sample. A statistic is computed based on the sample observations. Based on statistic, we verify whether the sample has drawn from the parent population with certain specified characteristiés. The computed value of the statistic may differ from the hypothetical value of the parameter. If the difference is small, it can be considered that it has arisen due to sampling fluctuations and this difference is accepted. If the difference is considerable, it may be considered that difference has not arisen due to the sampling fluctuations but it may be due to some other reasons. In this case the difference is said to be significant and the hypothesis is rejected. Testing of hypothesis is the procedure of examining whether difference between the computed statistic (from sample) and the hypothetical parameter (from population) is significant or not. Another definition : Testing of hypothesis is a process of decision making by using various statistical methods and theory of modern probability. DEFINITIONS OF STATISTICAL HYPOTHES| : SS Statistical hypothesis is a certain statement about probability distribution of a random variable (or) it is certain statement about population. Statistical hypothesis is denoted by H. Examples : 1. Average life time of an electrical bulb manufacturing by a company is 1200 hrs. ie. H: p = 1200 hrs. 2. The average height of competitors in a game is 160 ems. i.e. H: 1 = 160 ems. The average marks in the subject mathematics (,) is more than that of marks in computers (19) i.e. H: 11> p12. Machine X has an effective life period of 20 years less than any other machine in the company. i.e. H: 2 < 20. ‘The variation between the product of two companies is 10 hrs in their life time. ie. H: 02-02 =2 Simple Statistical Hypothesis _If a statistical hypothesis specifies the population completely (i.e. probability distribution is known), then it is called simple statistical hypothesis. In a simple hypothesis all the parameters of the population are given clear values. Examples : 1. LetX~N (y, 0%) H: = 1200 kms, o? =2. 2. LetX~B(n, P)H:P =i Composite Statistical hypothesis If a statistical hypothesis do not specify the population completely, then it is called a composite statistical hypothesis. Examples : There are two kinds of essential hypothesis in conducting the test procedures. 1, Null Hypothesis 2. Alternative Hypothesis 1. Null Hypothesis A statistical hypothesis with no difference or with null attitude is called null hypothesis. It is denoted by Hy. According to R.A. Fisher “Null hypothesis is the hypothesis which is tested for possible rejection under the assumption that it is true.” Examples : 1, The average height of the competitors in a game is 160 cms. Ho: = 160 cms. 2. Average life time of an electrical bulbs manufacturing by a company is 1800 hours. Ho: = 1800 hrs. 2. Alternative Hypothesis ‘A statistical hypothesis which is complementary to the null hypothesis is called an alternative hypothesis. It is denoted by Hj. It is clear that null hypothesis is meaningful when we formulate an alternative hypothesis. Examples: 1, Ifthe null hypothesis is the average height of the competitors in a game is 160 cms i.¢., Ho : # = 160 cms, then alternative hypothesis may be formulated as (@) Hy: 160 cms (ii) Hy: 2 < 160 cms (ii) Hy: 42> 160 cms. The alternative hypothesis in (i) referred to two-tailed test and in (éi) and (iii) referred to one-tailed tests. 2, Ifthe null hypothesis is average life time of electrical bulbs in a company is 1800 hours, then alternative hypothesis may be considered as follows : (@) Hy:2# 1800 — (two-tailed test) (ii) Hy:>1800 (one tailed test) (iii) Hy:~<1800 (one tailed test) Oe! b 7 Critical Region 2 ‘ ' Let x1, 2, --- %n be sample obs sample space $ into two disjoint parts W and W. The region W consists of the sample points for which the null hypothesis is rejected when it is true. ervations in the sample space S. Let us divide the Acceptence Region Ww W (Rejection Region/Critical Region) ‘The region of the sample points for which the null hypothesis is rejected when it is true is called critical region. Two types of errors . . Adetision ie. whether the null hypothesis Ho is to be accepted or rejected is made = the batts of the information supplied by the sample data, There may be a chance 0 taking good decision or an error. 5 ssed There are four possible situations that arise in testing the hypothesis 2° expresse in the following table. Statement ‘Avcept Ho Ho true Correct decision [HoFalse | Wrong (Type Tf error) [Wrong (Type Herron) | There are two possible errors in testing the hypothesis. Type I error : The error of rejecting the null hypothesis Hy when His true is calleg e I error. fa IL error : The error of accepting Hy when Hy is false is called Type II error, The probability of type I error is denoted by ‘a’ and the probability of type II error ig denoted by a =P (Type I error) = P (Rejecting Hy when Hy is true) = Pl € WAH) = [ Ide w where Lois the likelihood function of the sample observations #1, 29, ... 2, under Hp. B =P (Type II error) = P (Accepting Hp when Hpis false) “PG EWG = [ Lyte > Suen v where Ly is likelihood function under Hy. Level of Significance The probability of type I error (a), is known as level of significance. This is also called size of the critical region. ais also called producer's risk and f is called consumer's risk in sac. = Power of the test : 1—B = Pe © W/H,) i. probability of rejecting Hy when His false is called power function of testing the hypothesis. The value of the power function is called the power of the test. 2 Most Powerful Test Let the problem of testing a simple null hypothesis Hy : 0 = 6) against a simple alternative hypothesis Hy : 6 = 0. The critical region W is the most powerful eritieal region of size a for testing Hy: 0 = 09 against H, : 0 = 0, if Pe € WH) : f Iodx =a el) Ww and P@ € W/H)) = P(x € W,/Hy) 42) for every other critical region W; satisfying (1). The corresponding test is called most powerful test. : Large Sample Tests Large Sample Entire statistical theory for conducting the test of significance is based on the sample size. Generally in practical situations if the sample size is greater than or equal to 30, then the sample is known as large sample. = OF NORMAL DISTRIBUTION IN LAI We have already studied that, for large values of n, almost all distributions viz., Binomial, Poisson, Negative Binomial, Exponential etc. tends to Normal distribution. In this case, we apply Normal distribution to test the hypothesis. Area’s property of standard normal distribution is used to test the given hypothesis. IfX~N(, 0?)- then z =k ~ NO, X-EQ) sual; z="~——— ~ NO, D “may VVar ®) From the standard normal tables, we have P(C3 8) =1-Pl|z | <3]=0.0027 We expect the standard normal variate lies between + $ ie. if | z | > 3, the null hypothesis Ho is always rejected. Otherwise i.e. if | z | <3, the null hypothesis may be accepted. - The value of a test statistic which separates the rejection (critical) region and the acceptance region is called the critical value or significant value. ~ Critical values depend on the level of significance and alternative hyp ‘ 3, 7 Othes;, | Alternative hypothesis decides whether the test is two tailed or one tailed test "8. Critical values for two tailed test Let critical value at a level of significance for two tailed test he z,. By the definitio, ofa, Pilz |>z) =a => PE >z,) + Pe <-z,) =a => Pe>z,)+Pe>z,) =a (By the property of symmetry) = : 2P(z >zq) = 0 = Pe>z) =$° Similarly P(e <-z,)=$ The normal probability curve can be shown as follows : w = 2, z=0 Za The critical value for two tailed test at o. level of significance can be calculated by P(2,z,) =a Normal probability curve can be shown as follows : “0 z=0 Zu fa The critical value is calculated for right tailed test at a level of significance by Pe <2g)=1-a For left tailed test Plz <-z,) =a Normal probability curve can be shown as follows : | oe =, z=0 a The critical value is calculate for left tailed test at a level of significance by P@ ‘The critical vale is calculated for both right and left tailed tests by using the same formula Pe 1.96) = 0.05 ie. if | z | >-1:96;'the null hypothesis Hy may be rejected at 5% level of si Otherwise i.e. | 2 | < 1.96, Hy may be accepted. The Critical value at 1% level of significance : From the standard normal tables, PC-2.58 2.58) = 0.01 ie-if | 2 | > 2.58, the null hypothesis Hy may be rejected at 1% level of significance. Otherwise ie. | z | = 2.58, Homay be accepted. ignificance. Critical Values for one-tailed test The critical value at 5% level of significance : From the standard normal tables, P(g > 1.645) = 1- P(e 1.645, Hymay be rejected. Otherwise Hy may be accepted. The critical value at 1% level of significance + From the standard normal tables, P(e > 2.33) = 1-P(-@ <2 < 2.33) = 10.99 = 0.01 If | 2 | > 2.83, Hy may be rejected. Otherwise Hp may be accepted. R TEST OTHESIS testing of hypothesis may be outlined as follows The major steps in the 1, Null Hypothesis. Formulate the null hypothesis Hy which is suitable to the problem under study. 2. Alternative Hypothesis. Formulate a meaningful alternative hypothesis H, against to the null hypothesis Ho. 3. Level of Significance. Fix the level of significance o. in advance. 4. Testa statistic Compute the test a statistic ¢—E©) _ (0, 1) under Ho. 2=SE.(@ Where t is a statistic which depends on the sample drawn from the population. 5. Comparison and conclusion Compare calculated value of 2 in step 4 with the significant value (table value) zqat specified level of significance 0. If | z [> za) we may reject the null hypothesis Ho at a% Lo.s. If|z | <2u we may accept the null hypothesis Ho at 0% L.o.s. NGLE PROPORTION. Let A be the attribute having n persons and classify presence of an attribute as a success and absence as a failure. Let x be the number of successes in n independent tiels with probability of success P for each trial. Obviously x follows a Binomial distribution. ; E(x) =nP, Var (x) =n PQ For large values of 7, binomial distribution tends to normal distribution. Null Hypothesis Hy: P = Pyi.e. sample proportion is coming from the population proportion. ‘Alternative Hypothesis Hy: P « Pp or P>Poor P za, we may reject Ho, Other wise, if | z | = 2a, Hp may be accepted. Test whether the dice is unbiased. SOLUTION. Given that n = 900, x = 335 P = Probability of getting 3 or 5 = ze PROBLEM 1. A dice is thrown 900 times and a face of 3 or 5 is observed 335 times. i 3 12 Q=1-3=3 (1) Null Hypothesis : Ho: =i fen dies ietunbined! (2) Alternative Hypothesis : F ie. dice is not unbiased (Two tailed test) (8) Test statistic under Hy: x _ 335 pte Hy:Pe (A) Inference + |z| =2.328 oo Tabulated value (critical value) of z at 5% level of significance for two tailed test is z= 1.96 lz] >2a => — Hyis rejected ie. — dice is not unbiased PROBLEM 2. A coin was thrown 400 times and head resulted 240 timos. Ton, whether coin is unbiased at 1% level of significance. SOLUTION. Given that n = 400, x = 240 P = Probability of getting a head i _ m2) @=2 (1) Null Hypothesis : 1 Ho: P= 2 i.e. coin is unbiased (2) Alternative Hypothesis : Hy: Pe ; ie. coin is biased (Two tailed test) (8) Test statistic under Ho: _P=Po 40072 | fee fig n 2°2 400 (4) Inference : lz|=4 Zq = 2.58 at 1% level of significance and for two tailed test from standard normal tables. |z| >2q > Hyis rejected. |z|>3 => We always reject Hy, -. The coin is not unbiased. Bo FROBLEM 8. a fe eae of 1000 people in Maharashtra, 540 are rice eaters and t rs, Can we assume that b i in Maharashtra at 1% level of significance, RD eee SOLUTION. Given that n = 1000, x = 540 540 P =7o99 = 0-54 P = Probability of rice eaters in Maharashtra = 1 1 . y= 05 pagr0s Q=1-P=1- (1) Null Hypothesis : Hy :P=0.5 ie, both rice and wheat are equally popular in Maharashtra. (2) Alternative Hypothesis : 121253 Hy: P #05 ci ie. rice and wheat are equally popular in Maharashtra (Two tailed test), @ =o Zy= 2-58 (8) Test statistic under Hp : p= rs - z= PaP0 _ Noo, 1) = 054-05 ~NO, 1) =="? 53 PQ (0.5)(0.5) n ~ 1000 2.53 2.58 at 1% level of significance, for two tailed test from standard normal tables, |z| 0.85 Zu, = 1-645 (2=0) 2gp = 1.645 ie. survival rate is more than 85% P= 085 (One tailed test). (8) Test statistic under Ay: P-Po 0.9 - 0.85, ~NO, ) == — _ 0.633 PrQo (0.85X0.15) o Von 20 (4) Inference : |z| =0.633 2q = 1.645 at 5% level of significance and for one tailed test from standard normal tables. - |2| <2q = His accepted. ie. survival rate attacked by the disease is 85% but not more than 85%. PROBLEM 5. A manufacturer claims that 2% of the product is defective. In one day’s production of 200 items, only 8 are defectives. Test his claim of production of defective items are 2% or more at 5% level. Find 95% confidence limits for proportion of defective items. . SOLUTION. Given that n = 200, x =8 =%.%. P =7,= 300 = 0.04 2 P =2%= i007 0.02 Q =1-P=0.98 (1) Null Hypothesis : Hy: P = 0.02 ie. proportion of defective items are 2%. (2) Alternative Hypothesis : 1z1=202 Hy: P> 0.02 - i.e. proportion of defective items | 7, =-1-64 2.) = 1.64 are more than 2% (One tailed | % (@=0) % test). ps2 (3) Test statistic under Ho: ae P-Po 0.04 — 0.02 =2.02 z= ~N(O, 1) =e P,Qo (0.02)(0.98) n 200 (4) Inference : | z| = 2.02 ms Zq= 1.645 at 5% level of significance and for one tailed test from standard normal tables. |z|> za=> Hois rejected. i.e. manufacture claim is not correct. Proportion of defective items are more than 2% .. 95% confidence limits for proportion of defective items (P) are (>100 [#2 ps0 [#8 | A A where P=p,Q=1-p are estimated values. B= 0.04, Q=1-0.04=0.96 «. limits are _ (0.04)(0.96) [(0.04)(0.96) (0.08 1.96 (SoG 0.04 + 1.96 {22ae28 )=c.o12s, 0.0672) 95% confidence limits for the proportion of defective items is (1.28, 6.72) per cent, ERORORTION =u Here, we compare two different populations with respect i 7 . and x be the number of persons having the attribute ofeives oe ciy ioe two populations. Obviously, x, and x2 follows binomial distribution with the parameters ny, P; and nz, Po, Here P; and P, are population proportions. Let the sample proportions be p; and py owt Xa Plans? Poa E@) =Pi, E(p) = Pz P, ~ “Var (py) 278s Var (p2) = 2s Null Hypothesis : Hp: P, = Paice. there is no significant difference between two sample proportions. (or) Populations proportions are equal. Alternative Hypothesis Hy: Py*P zor Py> Pp or P,2a we may reject Hy at « level of significance, otherwise ie. if | z | =2,, we may accept Ho. PROBLEM 1. A survey was conducted on TB. patients in India. The data revealed that 1% of the population are suffering from T.B. in the country. A sample date Wag collected in two colleges. In college A, there are 5 T-B. patients out of 400 students and in follege B, there are 10 T.B. patients out of 1200 students. Test the significance difference between the proportion of T.B. patients in two colleges. j SOLUTION. Given that n= 400, ny = 1200 : x1 =5, x2 =10 %__ 10 a = 007 0.0083 1 P =1% =i00 0.01 1zl=0-7311 Q =14P=1-0.01=0.99 Z (1) Null Hypothesis : Ho: P}=Pp=P=0.01 ie. there is no significant | 29 = difference between the proportion of T.B., patients in two colleges. (2) Alternative Hypothesis ; Hy: P,P, (Two tailed test) - - ie, there is a significant difference between the proportion of T.B. patients in two colleges. (8) Test statistic under Hy: = . — 0.008; z Tia = ——eEeEx = 0.7311 11 Pa (2 + 2) (0.0110.99) (755 + 395 (4) Inference + 0.7811 Zq = 1.96 at 5% level of significance and for two tailed test from standard normal tables. [2 | <2a = Hpis accepted. ie, there is no significance difference between the proportion of T.B. patients in two colleges. PROBLEM 2. Random samples of 400 men and 600 women were asked whether they would like to have a flyover near their residence. 200 men and 325 women were infavour of the proposal. Test the hypothesis that proportions of men and women in favour of the proposal are same or not at 5% level of significance, SOLUTION. Given that n,=400, ng =600 x1 = 200, 1 _ 200 _ 95 Pi=7.= 4007 2 _ 825 P2= 5 = G00 = C5417 Pis not known Iz1=1-269 f_ 1 t+X2 200 +325 y+ nq" 400 +600 525 = 30007 0.525, @ =1-p =1-0.525 = 0.475 (1) Null Hypothesis Hy: P,=P2=P ey ie. there is no significant difference between the opinion of men and women about flyover. (2) Alternative Hypothesis : H, : P; « Pp (Two tailed test) i.e. The opinions of men and women about the flyover are not same. (8) Test statistic under Ho: Pi-P2 e 0.5- a 1.269 11 += (0.525)(0.475) 1 vd (2 +g) Yt0-525K0.478) (795 (4) Inference : lz | = 1,269 = 1.96 at 5% level of significance and for two tailed test from standard normal tables. 2 |< zq => His accepted. : ie. proportion of men and women about flyover are same. PROBLEM 3. A survey was conducted on eye sight of the students of Andhra Pradesh. Two random samples of sizes 900 and 1200 students were selected from the cities Visakhapatnam and Tirupathi, out of these 20% and 15{e/have eye sight respectively. Test the significance of eye sight in Visakhapatnam is more than that of Tirupathi at 1% level of significance. SOLUTION. Given that n,=900, n= 1200 p= 20% = 0.2, py 15.5% = 0.155 Pis not known 5 _ Mit Map2 900 (0.2) + 1200 (0. BAN em 900+ Iz00 = OTIR Q =1-B=1-0.1712 = 0.8288 : (1) Null Hypothesis : Hy: Pi=Pa=P ie. there is no significant difference between the eye sight of students in Viasakhapatnam and in Tirupathi. (2) Alternative Hypothesis : H,: P)> P» (one tailed test) ‘The eye sight of the students in Visakhapatnam is more than in Tirupathi. (8) Test statistic under Ho: 0.2 - 0.155 Pi- Pa ~ NO, ) =——— ———. - 2.87 zee] - ety ae) vq tr * a] (0.1712) (0.8288) (st = an) (4) Inference : | 2 | =2.87 Zq = 2.33 at 1% level of significance and for one tailed test from standard normal tables. | 2 | > zq=> Ho is rejected. ie. we conclude the eye sight of the students in Visakhapatnam is more than of the students in Tirupathi. PROBLEM 4, Before an increase in excise duty on tea, 800 persons out of a sample of 1000 persons were found to be tea drinkers, After an increase in excise duty, 800 people were tea drinkers in a sample of 1200 people. Test whether there is a significant decrease in the consumption of tea after the increase in excise duty. SOLUTION. Given that n,= 1000, ng = 1200 x= 800, x, = 800 x _ 800 _ Pt =n, = 1000 98 x2 800 P is not known, s+, __800+800 _ 1600_ P =e ng * 1000 + 1200 ~ 2200 = 9.7273

You might also like