Statistical Inference 179
Statistical Inference 179
STATISTICAL INFERENCE
COMPLEMENTARY COURSE
[Link]. MATHEMATICS
III SEMESTER
(2011 Admission)
UNIVERSITY OF CALICUT
SCHOOL OF DISTANCE EDUCATION
CALICUT UNIVERSITY P.O., MALAPPURAM, KERALA, INDIA - 673 635
417
School of Distance Education
UNIVERSITY OF CALICUT
SCHOOL OF DISTANCE EDUCATION
STUDY MATERIAL
[Link]. MATHEMATICS
III SEMESTER
COMPLEMENTARY COURSE
(STATISTICS)
STATISTICAL INFERENCE
Prepared by:
Scrutinised by:
©
Reserved
3. INTERVAL ESTIMATION 52 – 68
Interval estimation
Confidence interval for the mean of a normal population
Confidence interval for the difference of means
Confidence interval for the variance of a normal population
Confidence interval for large samples
4. TESTING OF HYPOTHESIS 69 – 89
Statistical hypothesis
Testing of hypothesis
Errors in testing of hypothesis
Steps in testing of hypothesis
Most powerful test
Neymaan-Pearson theorem
SAMPLING DISTRIBUTIONS
1.1. Sampling Distribution
Any function of the statistical population values are called population parameter.
For eg., mean, variance, median etc., of the variable considered.
The process of making inference about the population based on samples taken from
the population is known as statistical inference.
Consider a random sample taken from a population then the function of sample
values like sample mean, sample variance sample moments etc., are known as statistic.
X 1 X 2 ... X n
Sample mean X
n
t
mgf of X , M (t ) M X1 X 2 ... X n (t ) M X1 X 2 ... X n ( )
X n
n
t t t
M X1 ( ) M X 2 ( )....M X n ( ) ( X i ' s are ind .)
n n n
t 2 2
t
We have for X~ N ( , ) , M X (t ) e 2
t t 2 2 t t 2 2 t t 2 2
2
2
2
Therefore, M (t ) e n 2n .e n 2n ....e n 2n
X
t t 2 2 n t 2 2
( ) t
e n 2n 2 e 2n , it is the m.g.f. of a normal
population with parameters and . This implies that
n
N ,
√
Problem 1: A random sample of size 25 is taken from a normal population with mean 1 and
variance 9. What is the probability that the sample mean is negative?
Solution:
Given sample size n = 25, 1 and 3 .
We have X ~ N ( , ).
n
=P <
√ √
=P < − / , where Z = N(0, 1)
√
√
1 2 x n 1
n
2 e 2 x2 , 0 x
f ( x) n2
0, otherwise
is said to follow 2 distribution with n degrees of freedom, where n is the parameter of
Mx (t) = ( − )
Mean and Variance:
Observe that,
d
E( X ) M (t )
dt X / t 0
( − ) t=0
=
= ( − ) − t = 0 = n.
Hence, E(X) = n
= − − − ( − ) − t=0
= n(n + 2).
Hence, V(X) = n(n+2) – n2 = 2n
degrees of freedom.
1
n
2
X 2 ~ ( n2 ) M X (t ) 1 2t 2 .
2
2
n n
1 2
Therefore, M X X (t ) M X (t ) M X (t ) 1 2t 2 1 2t 2
1 2 1 2
n n
1 2
1 2t 2
In general if, X1, X2, ……Xk are n independent random variables with n1, n2, …… nk
degrees of freedom respectively, then X1+ X2 + ……Xk follow chi-square distribution with
n1+ n2 + ……nk degrees of freedom.
Theorem: If X ~ N (0, 1), then Y X 2 follow chi-square distribution with one degree of freedom.
Proof:
2
tx2 1 x2
e
2
e dx
x2
tx2
( e 2 is an even function)
1
2 x2 (t )
2 e 2 dx
0 .
du du
put x 2 u dx
2x 2 u
1
2 u (t ) du
M Y (t )
2 e
0
2
2 u
1
1 u ( 1 t ) 1 1 1
2
e 2 u 2 du
2 1
1
. 2
0
( t) 2
2
Hence Y 2 1 .
Theorem: If X 1 , X 2 ,..., X n are n random samples taken from a standard normal population, then
Proof:
Since X i ' s are random samples from N (0, 1), they are independent and
Here X i 2 ~ 2 (1) for all i . Hence, by additive property of chi-square distribution, the
x
n 2
Y i ~ ( n)
2
i 1
xi
n 2
we get,
i 1
~ ( n) .
2
Problem 1: If X 1 , X 2 ,..., X n are n random samples taken from N ( , ) , find the distribution of
sample variance S 2 .
Solution:
For the random samples X 1 , X 2 ,..., X n taken from N ( , ) , we have
x
n 2
i 1
2
n
x X X
Y i , where X is the sample mean
i 1
2
(x X ) ( X )
n
ie., Y i
i 1
( xi X ) 2 ( X ) 2 2( xi X )( X )
n
i 1 2
n
( xi X ) 2 n ( X ) 2 n 2( xi X )( X )
i 1 2 i 1 2 i 1 2
n
( xi X ) 2 n( X ) 2 n
( ( xi X ) 0)
i 1 2 2 i 1
2 2
n
(x X ) 2 X nS 2 X
i 2 2
i 1
n n
X
where ~ N (0,1)
n
2
X
follow chi-square distribution with one degree of freedom. Since X and S 2 are
n
nS 2
independently distributed, then by additive property of chi-square distribution
2
~ 2 (n 1) .
nS 2
Let U = ~ 2 (n 1) (1)
2
Then,
1 n21 u n 1
2
2 2
1
0u
f (u ) n 1 e u ,
2
0, otherwise
u 2 du
S2 , then, f( S 2 )= f(u) in terms of S 2
n dS 2
n1 n 1
nS 2 1
1 2
f (S 2 ) 2
2 2
nS 2 2
du
e 2
n 1 dS 2
2
n1 n 1
nS 2 1
1 2
2
2 2
nS 2 2
n
.
e 2
2
n 1
2
n1
n 2
nS 2 n 1
S2
1
Hence, ) 2 2 2 0 S2
2
2 2
f (S n 1
e ,
2
x n n 2 x n 1
f ' ( x) 0 e 2 1 x 2 e 2 x 2 0
2
x n x n
n 2 2 1
e 2
1 x e 2 x2
2
n n
n 2 2 1 1
1 x x2 x 1
2 n2
x n2
Let
x −.
Y 2 log e ( ) . This can be written as x =
Note that =
dx
f(y) = f(x) in terms of y .
dy
= . .
1 2y
e
2
2 1
12 2
y 2
1
= 21
e 2 y 2
2 1
f ( y)
1
2
2
e
y
2
2
2
y ,
1
0 y ;
21
2
Problem 4: For large n, show that chi-square distribution approximately normally distributed.
Solution:
Consider a random variable X following chi-square distribution with n degrees of
n
freedom. Then M X (t ) 1 2t 2
. More over,
E( X ) n and V(X) =2n.
X n
Consider Z , then
2n
n
t
2n M t
M (t ) M (t ) e ( )
Z X n X 2n
2n
n
t n
e 2n 1 2 t 2 .
2n
n
2n t t 2
n
Therefore, log M (t ) log e 1 2
Z 2n
n t
t n log 1 2
2n 2 2n
n n t 1 t
2
t 2 2 .....
2n 2 2n 2 2n
1
n n t2
t t ... + (many terms involving n 2 and its
2n 2n 2
higher power in denominator)
t2
As n become very large, log M (t ) .
Z 2
t2
That is M (t ) e . This is the m.g.f. of a standard normal random
2
Z
variable. Then by uniqueness theorem of m.g.f., X~N(n, 2n ) for large n.
nS 2 nS 2 nS 2 nS 2
P( 2 ) 0.60, where 2 ~ 2( n1) .
a b
16 16 16 16
That is, P( 2(15) ) 0.60 --------- (1)
b a
From the table of chi-square distribution,
P ( 2 (15) 10.307 ) 0.80 and P ( 2(15) 19.311) 0.20 .
16 16 16 16
Comparing (1) and (2), we get 10.307 and 19.311
b a
16 16 16 16
Hence b 24.84 , and a 13.26
10.307 19.311
ie., P( 13.26 2 24.84 ) 0.60 .
This is the probability distribution which was introduced by W.S Gossset and
known in his pen name ‘student’. A continuous random variable t with density function
n 1 n 1
2 2
t
f (t ) 2
1 , t is said to follow student’s t-distribution with n
n n
n
2
degrees of freedom.
nS 2
sample variance. Then we have X ~ N ( , ) and 2 ~ 2 (n 1) .
n
X
Note that ~ N (0,1) .
n
X
Therefore, t n ~ t( n 1) .
nS 2
2
(n 1)
( X ) n 1
This implies that t ~ t( n 1) .
S
2. Let X 1 and X 2 be the means, S1 and S 2 be the standard deviations of
samples of sizes n1 and n2 taken independently from two normal populations with same
X1 X 2
t ~ t( n n
2 2)
n1S 1 n2 S 22 1 1
2 1
n1 n2 2 n1 n1
Proof: We have X 1 ~ N ( , ) and X 2 ~ N ( , )
n1 n2
2 2
Then, X 1 X 2 ~ N (0, ).
n1 n2
X1 X 2
This implies that ~ N (0,1)
1 1
n1 n2
n1S12 n2 S 2 2 n1S12 n2 S 2 2
2 ~ 2 (n1 n2 2) ( 2 and are ind .)
2
2
X1 X 2
1 1
n1 n2
So, t ~ t( n n
n1S12 n2 S 2 2 1 2 2)
2
2
n1 n2 2
X1 X 2 n1 n2 2
That is, t ~ t( n n
2 2)
1 1
n1S 12 n2 S 22 1
n1 n2
X1 X 2
That is, t ~ t( n n
2 2)
n1S 12 n2 S 22 1 1 1
n1 n2 2 n1 n1
Tables of t-distribution:
Note that t-distribution is symmetric about zero and bell shapped. Tables of t-
distribution gives the values of t for various degrees of freedom and for various value of
n 1 n 1
2 2
Solution: We have, f (t ) 2 1 t , t
n n
n
2
n 1 n 1
2 2 1
2 n 1 t 2t
f ' (t ) 1 .
n 2 n n
n
2
n 3
2 2
t 0
f (t ) 0 2t 1
'
n
t0
n 1 n 3 n 3
2 2 1 2 2
n 1 2 n 3 t t .
f '' (t ) 2
1 t 1
n 2 n 2 n n
n
2
n 1 n 1
2 n
At at t 0, f '' ( x) 0 . Hence Mode of is at t = 0.
n
n
2
Problem 2: If t ~ t( n ) , then as n , prove that t ~ N (0,1) .
Solution:
n 1 n 1
2 2
t
Given t ~ t( n ) , we have f (t ) 2
1 , t
n n
n
2
Using the following results we can prove as n , t ~ N (0,1)
nk
(i) as n becomes very large, nk
n
n
(ii) lim 1 e
n n
Note that as n
Solution:
The graph of t ~ t(5) is symmetric about zero. To find a such that the area under
Problem 4: Prove that the ratio of two independent standard normal random variables is a
student’s t random variable with 1 degree of freedom.
Solution:
Given X 1 ~ N (0,1) , X 2 ~ N (0,1) and they are independent.
X1
t ~ t(1)
X2
2
1
X1
That is, t ~ t(1)
X2
That is, the ratio of two independent standard normal variables follows t(1) .
Problem 5: If X 1 and X 2 are two independent standard normal variables, find the distribution
2X 1
of t .
X X 22
1
2
X1
Then, t ~ t( 2 ) .
X X 22
1
2
2X 1
That is, t follow t-distribution with 2 degrees of freedom.
X 12 X 22
Problem 6: Find the maximum difference that we can expect with probability 0.95 between the
means of samples of sizes 10 and 12 from a normal population, if their standard deviations are
found to be 2 and 3 respectively.
Solution:
Let x1 and x2 be the means of the samples of sizes n1 =10 and n2 =12 taken
randomly from two normal populations. Assume samples are taken independently. The
sample variances S12 = 4 and S 2 2 = 9 respectively.
P x1 x2 k 0.95 .
n1 n2 2 n1 n1
X1 X 2 k
P 0.95
n1S1 n2 S 22 1 1
2
n1S 12 n2 S 22 1 1
n1 n2 2 n1 n1 n1 n2 2 n1 n1
k
That is, P t( 20 ) 0.95 (1)
1.165
k
Comparing (1) and (2), we get, 2.086 . Hence
1.165
k 2.086 1.165 2.431
That is the maximum difference that can expect with 95% probability is 2.431.
f (F ) n2 , 0F
n1 n2
n1 n2 n1 2
( , ) 1 F
2 2 n2
X1
n1
F ~ F (n1 , n2 )
X2
n2
mean and standard deviation . Let S12 and S 22 are the respective sample variance,
n1S12 (n2 1)
then F ~ F (n1 1, n2 1)
n2 S 22 (n1 1)
STATISTICAL INFERENCE Page 21
School of Distance Education
Proof:
For the set of samples taken from normal population, we have,
n1S12 n2 S22
~ 2( n 1) and ~ 2( n 1)
2 1 2 2
n1S12
2
n1 1
Then, F 2
~ F (n1 1, n2 1)
nS
2 2
2
n2 1
n1S12 (n2 1)
Hence F ~ F (n1 1, n2 1)
n2 S 22 (n1 1)
Tables of F-distribution:
Tables of F-distribution gives the values of F for various values of n1 , n2 and ,
such that P( Fn ,n F ) .
1 2
Mode of F-distribution:
Mode is the point F where f(F) attains its maximum. That is the point F, where
log f ( F ) 2 log f ( F )
f ' ( F ) 0 and f '' ( F ) 0 or the F where 0 and 0.
F F 2
n1
n
n1 2 21 1
F
We have, f (F ) n2
n1 n2
n n n 2
( 1 , 2 ) 1 1 F
2 2 n2
n1 n n n n n n n
Therefore, log f ( F ) log 1 1 1 log F log ( 1 , 2 ) 1 2 log 1 1 F
2 n2 2 2 2 2 n2
STATISTICAL INFERENCE Page 22
School of Distance Education
log f ( F ) n1 1 n n 1 n
i.e., 1 1 2 . . 1
F 2 F 2 n1 n2
1 F
n2
n 1 n n n 1
1 1 1 2 1 .
2 F 2 n2 n1F
log f ( F )
0
n 1
1 1
n1 n2 n1 . 1
F 2 F 2 n2 n1F
n2 n1 2
F .
n1 n2 2
2 log f ( F )
at this point it can be verified that 0.
F 2
n2 n1 2
Hence mode of F ~ F (n1 , n2 ) is F = .
n1 n2 2
n2 n1 2 n2 n 2 . Since F 0 , the
Remark: The mode can be expressed as 1
n1 n2 2 n2 2 n1
mode cannot be negative. Hence n1 should not be less than 2. So the mode exists only
unity.
Problem 1: Prove that the ratio of the squares of two independent standard normal random
variables is an F- random variable with (1, 1) degree of freedom.
Solution:
Let X 1 ~ N (0,1) , X 2 ~ N (0,1) and they are independent
X 12 1
then ~ F (1,1)
X 22 1
X 12
ie., ~ F (1,1) .
X 22
1
Find the distribution of Y .
X
Solution:
1
Given Y we have f(y) = f(x) in terms of y . dx
X dy
n1
n1
n1 2 1
x2
Here f ( x) n2 , 0 x
n1 n2
n1 n2 n1 2
( , ) 1 x
2 2 n2
1 1 dx 1
Y , so X 2
X Y dy y
n1 n1
1
n1 2
1 2
so, f ( y ) n2 y
1
n1 n2
y2
n1 n2 n1 1 2
( , ) 1
2 2 n2 y
n1 n1
1
n1 2 1 2
2
n2 y 1
f ( y) n1 n2
y
n n n y n1 2
( 1 , 2 ) 2
2 2 n2 y
n1 n1
1
n1 2
1 2 n n
( 1 2)
n2 y 1 2
n1 n2
y
n1 n2 n 2
( , ) y 1
2 2 n2
n1 n1 n n
1 1 2
n1 2
1 2 2 2
n2 y
n1 n2 n1 n2
n1 n2 yn2 2 n2 2
( , ) 1
2 2 n1 n1
Y ~ F (n2 , n1 )
1
with n2 , n1 degrees of freedom. Prove that P ( X c ) P ( Y ).
c
Solution:
1 1
P ( X c ) P( )
X c
1
But, given, X ~ F (n1 , n2 ) , then, ~ F ( n2 , n1 )
X
Also Y is a variable following F (n2 , n1 ) .
1
Hence, P ( X c ) P (Y )
c
Problem 4: If X following F distribution with n , n degrees of freedom. If , ( ) are
Solution:
n 1 n 1
2 2
t
Given t ~ t( n ) , we have f (t ) 2
1 , t
n n
n
2
dt 1
Let Y= t 2 ; t Y Therefore
dy 2 Y
,
dt
Note that f(Y) = f(t) in terms of y .
dy
n 1 n 1
Y 2 1
That is, f (Y ) 2. 2 1 . , 0Y
n n 2 Y
n
2
n 1 n 1
1
1 2 Y 2 .Y 2
2. . 1
2 1 n n
n
2 2
n 1
n 1
Y 2 1
f (Y ) 2
1 .Y 2
1 n n
n
2 2
1
1
Y2 mn
m , n
n 1
m n
1 n Y 2
n , 1
2 2 n
Solution:
dx 1
Given Y= n1 X, ;
dy n1
dx
f(Y) = f(x) in terms of y .
dy
n1
n1
n1 2 1
x2
X ~ F( n ,n ; f ( x) n2 , 0 x
n1 n2
1 2)
n1 n2 n1 2
( , ) 1 x
2 2 n2
n1
n1
n12 1
x2
f ( y) n2
1
n1 n2
n1
n1 n2 n1 y 2
( , ) 1
2 2 n2 n1
n1
1
n1 n2 n1
y 2
1 2
n
2 n1
1
n1 n2
n1
n1
n1 n2 y 2
n2 2 1
2 2 n2
n1 n2 1 n1
lim y 2 y 2
n 2
y 2
Also note that as n2 , 1 lim 1 lim 1
n2 n2 n2
n2 n2 n2
y y2
n
e 2 1 lim 1 e y
n2 n2
n1
1 y
y 2 n1 1
n1 2 e 2
f ( y) 1
n n1
Hence, as n2 , n1
2 2 n1
2
n1
1 y
1 2 n1 n1
1 2 1
n1 y
2 2 e
1
n n1
n1
n1
22
2
y n1
1 1
Hence, f ( y) n1 e 2 .y 2 , 0 y
2 2 n1
2
EXERCISES
1. Explain what is meant by sampling distribution. State the relationship between
normal and chi-square distribution.
2. Define chi-square distribution with n degrees of freedom. Derive its mean and
variance.
3. State and prove the reproductive property of chi-square distribution.
4. Show that for Students t-distribution with n degrees of freedom, the mean
n 1
n 2
deviation is given by .
n
2
5. Define F-distribution. Explain its use in statistical inference.
6. State the inter-relationship of t, chi-square and F distributions. A random variable
1
X has F- distribution with (n, m) degrees of freedom. Find the distribution of Y
X
.
7. Derive Student’s t- distribution and establish its relation with F- distribution.
8. If F has F-distribution with (n, m) degrees of freedom, prove that as n , nF tends
to be distributed as chi-square with n degrees of freedom.
9. If X and Y are independent standard normal variables, find the distribution of
X2
Z and write down its p.d.f.
Y2
10. X 1 , X 2 , and X 3 are independent N(0,1) variables. Find the distribution of
X2 X 12 X 2 2
(i ) X 1 2 X 2 2 (ii ) and (iii )
X1 2 X 32
****
CHAPTER 2
(i) Estimation
Estimation of parameters:
Testing of hypothesis deals with the method of deciding whether to accept or reject
the hypothesis regarding the unknown aspects of the population, based on the samples
taken from the population.
Let x1 , x2 ,..., xn are random samples taken from a population with unknown parameter .
The statistic tn t ( x1 , x2 ,..., xn ) is said to be an unbiased estimator of , if E (tn ) . tn is an
unbiased estimator of a function of , say f ( ) , if E (tn ) f ( ) .
Problem 1: A random sample x1 , x2 ,..., xn is taken from a population with mean . Show that the
sample mean x is an unbiased estimator of .
Solution:
Since the samples are taken from a population with mean ,
E ( x1 ) E ( x2 ) ... E ( xn )
x1 x2 ... xn
we have x
n
x x ... xn 1
E(x ) E 1 2 E ( x1 x2 ... xn )
n n
1 1
( ... ) = (n )
n n
E(x)
Hence x is an unbiased estimator of
Remark: Unbiased estimator for a parameter need not be unique. For eg. in the above
case consider the first two observations x1 and x2 only. Then
x x 1 1 x x
E 1 2 E ( x1 x2 ) E ( ) . That means 1 2 is also an unbiased
2 2 2 2
estimator of . In similar way we can find many unbiased estimators for .
Problem 2: A random sample x1 , x2 ,..., xn is taken from a normal population with mean and
1 n 2
standard deviation 1. Show that t
n i 1
xi is an unbiased estimator of 2 1 .
Solution:
1 n
E (t ) E xi 2
n i 1
1
E ( x12 ) E ( x2 2 ) ... E ( xn 2 )
n
Given the population variance as 1, and population mean as ,
1 E ( xi ) E ( xi 2 ) for all i
2
E ( xi 2 ) 1
2
for all i
1 2 1 2 ... 1 2
1
Hence, E (t )
n
n 1 2
1
n
1 n 2
E (t ) 1 , 2
ie., t
n i 1
xi is an unbiased estimator of 2 1 .
Problem 3: For the random sample x1 , x2 ,..., xn taken from N ( , ) , show that the sample
variance is a biased estimator of the population variance.
Solution:
Here to show that E ( S 2 ) 2 , where S 2 is the sample variance of the random
samples x1 , x2 ,..., xn taken from N ( , ) .
Note that E ( S ) S 2 f ( S 2 )dS 2
2
n 1
n 2
2
nS 2 n 1
2 1
i.e,
2
f (S )
2
e 2 2 2
(S ) , 0 S2
n 1
2
n 1
n 2
2
nS 2 n 1
2
E (S 2 ) S 2
2 1
e 2 ( S 2 ) 2 dS 2
0 n 1
2
n 1
n 2
2
nS 2 n 1
2
e 2 2 (S 2
)2 dS 2
n 1 0
2
n 1
n 2
2
nS 2 n 1
2 1
e 2 2 (S 2 ) 2 dS 2
n 1 0
2 2
2
n 1
n 2 n 1 n 1
2 1 1
2 2 2
n 1
n 1 1
n 2
2 2
2
n 1 2
2
2 n
n 1
E (S 2 )
2
n
Hence S 2 is not an unbiased estimator of 2 .
That is S 2 is a biased estimator of 2 .
n 1
Here, E (S 2 )
2
n
nS 2
E( ) 2
n 1
nS 2
That is, is an unbiased estimator of 2 .
n 1
Problem 4: For the random sample x1 , x2 ,..., xn taken from Poisson population with parameter ,
obtain an unbiased estimate of e .
Solution:
Consider a statistic t defined as follows,
Solution:
Here x1 , x2 ,..., xn are from B (1, p ) . Then by the additive property,
T x1 x2 ... xn follows B ( n, p )
T (T 1) 1 1
E E T (T 1) E T 2 T
n(n 1) n(n 1) n(n 1)
E (T 2 ) V (T ) E (T )
2
npq n 2 p 2
T (T 1) 1
E E npq n 2 p 2 np
n(n 1) n(n 1)
1
E np (1 p ) n 2 p 2 np
n(n 1)
1
= E n 2 p 2 np 2
n(n 1)
T (T 1) T (T 1)
E p 2 , ie., is an unbiased estimator of p 2 .
n(n 1) n(n 1)
( ii ) Consistency:
Let x1 , x2 ,..., xn are random samples taken from a population with unknown parameter .
The statistic tn t ( x1 , x2 ,..., xn ) is said to be a consistent estimator of , if
P tn 1 as n or tn is a consistent estimator of a function of , say f ( ) , if
P tn f ( ) 1 as n .
P tn [Link] (tn ) 1
1
t2
c
Let [Link](tn ) c , then t
SD (tn )
P tn c 1
1
2
c
SD (tn )
as n , if E (tn ) or , and V (tn ) 0 ; then [Link](tn ) c becomes a small
c
number and, ,
SD (tn )
P tn c 1 ie., tn
p
Problem 2: For the random sample x1 , x2 ,..., xn taken from Poisson population with parameter ,
nx
show that is consistent estimator .
n 1
Solution:
Here x1 , x2 ,..., xn are taken from Poisson population with parameter , so
E ( X i ) and V ( X i ) for all i , then
x1 x2 ... xn
x E ( x ) and V ( x )
n n
Hence, as n ,
1
1 E ( x ) E ( x ) , and
1
n
nx n2 1
V V (x ) 0
n 1 n 1
2 2
1 n
1
n
nx
Here satisfies the sufficient conditions to be satisfied by consistent estimator
n 1
and hence it is a consistent estimator of .
Problem 3: For the random sample x1 , x2 ,..., xn taken from B (1, p ) , show that T (1 T ) is a
n
1
consistent estimator of p (1 p ) , where T xi .
n i 1
Solution:
Note that x1 , x2 ,..., xn are from B (1, p ) . Then by the additive property,
X x1 x2 ... xn follows B ( n, p )
1 n X 1 pq
T xi
n i 1
E (T ) E (
n
) p , V (T ) 2 V ( X )
n n
E (T 2 ) V (T ) E (T )
2
pq
E (T 2 ) p2 --------- (1)
n
pq
E T (1 T ) E (T ) E (T 2 ) p p2
n
np p (1 p ) np 2 np (1 p ) p (1 p )
n n
n 1
p (1 p )
n
as n , E T (1 T ) p(1 p) --------- (2)
V T (1 T ) E T (1 T ) E T (1 T )
2 2
1
similarly , E T 3 4 n(n 1)(n 2) p 3 3n(n 1) p 2 np
n
as n , E T 3 p 3 and E T 2 p 2 (by (1))
V T (1 T )
2
p 2 p 4 2 p 3 p p 2
Problem 4: For the random sample x1 , x2 ,..., xn taken from N ( , ) , show that the sample
variance is a consistent estimator of the population variance.
Solution:
Let S 2 is the sample variance of the random samples x1 , x2 ,..., xn taken from N ( , ) .
Then we have.
n 1
n 2
2
nS 2 n 1
2 2 1
f (S )
2
e 2 2 2
(S ) , 0 S2
n 1
2
It is already found in a problem of last section,
n 1
E (S 2 )
2
n
Then, as n , E (S 2 ) 2 ------- (1)
n 1
n 2
2
nS 2 n 1
2 2 2
1
E S S
2 2 2
e ( S 2 ) 2 dS 2
2
0 n 1
2
STATISTICAL INFERENCE Page 37
School of Distance Education
n 1
n 2
2
nS 2 n 1
2
e 2 2 (S 2 ) 2 dS 2
n 1 0
2
n 1
n 2
2
nS 2 n 3
2 1
e 2 2 (S 2 ) 2 dS 2
n 1 0
2
n 1
n 2 n3
2
2
2
n 3
n 1
n 2
2 2
2
n 1
n 2 n 1 n 1 n 1
2
2
2 2 2
n 3
n 1
n 2
2 2
2
n2 1
4
n2
2
n 2 1 4 n 1 2
2 2
Hence, V ( S ) E S 2 E S 2
2
n
2
n
as n , V (S 2 ) 4 4 0 -------- (2)
From (1) and (2), it can infer that S 2 is a consistent estimator of the population
variance 2
x 2 x2 1 1 x1 2 x2
(ii) E 1 E x1 2 x2 2 , so is an unbiased estimator of .
3 3 3 3
x x x x 1 4 x1 x2 x3 x4
(iii) E 1 2 3 4 E x1 x2 x3 x4 , so also is an
4 4 4 4
unbiased estimator of .
x1 x2 1 2 2 2
V V x1 x2 ( xi ' s are random samples)
2 4 4 2
x 2 x2 1 5 2
V 1
3 9
V x1 2 x2
9
1 2
4 2
9
x x x x 1 4 2 2
V 1 2 3 4 V x1 x2 x3 x4
4 16 16 4
x x x x x x x 2 x2
Among these V 1 2 3 4 V 1 2 V 1
4 2 3
x1 x2 x3 x4
Hence is more efficient.
4
estimate the parameter p. The result of n tosses x1 , x2 ,..., xn contains no other information
about p than that contains in t. Hence the conditional probability of x1 , x2 ,..., xn given
x
i
i is independent of p.
p i (1 p ) i
x n x
1
That is P ( x1 , x2 ,..., xn / xi t ) n .
Ct p i (1 p ) i
x n x
i
n Ct
L( x1 , x2 ,..., xn , ) L1 (t , ) L2 ( x1 , x2 ,..., xn )
where L1 (t , ) is function of t and alone and L2 ( x1 , x2 ,..., xn ) is a function independent of
.
Proof:
If t is a sufficient estimator of , then the conditional distribution of x1 , x2 ,..., xn
given t = r is independent of . That is,
P ( x1 , x2 ,..., xn / t r ) h( x1 , x2 ,..., xn ) , which is independent of ------- (1)
P( x1 , x2 ,..., xn ) L( x1 , x2 ,..., xn , )
Butbut P ( x1 , x2 ,..., xn / t r )
P(t r ) P (t , )
L( x1 , x2 ,..., xn , )
Then by (1), for sufficient estimator t, h( x1 , x2 ,..., xn )
P(t , )
L( x1 , x2 ,..., xn , ) h( x1 , x2 ,..., xn ) P (t , )
e 1 e 2 e xn
x x
. .....
x1 ! x2 ! xn !
e n i
x
x1 ! x2 !...xn !
1
e n
xi
x1 ! x2 !...xn !
1
e n nx
x1 ! x2 !...xn !
L1 ( x , ) L2 ( x1 , x2 ,..., xn )
1
where L1 ( x , ) e n nx and L2 ( x1 , x2 ,..., xn )
x1 ! x2 !...xn !
Hence, by factorization theorem, x is a sufficient estimator of .
Problem 2: Let x1 , x2 ,..., xn be the random sample taken from a population with p.d.f.
f ( x, ) x 1 ; 0 x 1, 0 . Find a sufficient estimator for .
Solution:
Likelihood function L( x1 , x2 ,..., xn , ) f ( x1 , ) f ( x2 , )... f ( xn , )
i 1
i 1
xi i 1
n
L1 ( xi , ) L2 ( x1 , x2 ,..., xn )
i 1
n
n 1
where L1 ( xi , ) xi and L2 ( x1 , x2 ,..., xn )
n
n
, then by
i 1
i 1
x
i 1
i
Problem 3: Obtain a sufficient estimator for p, using samples x1 , x2 ,..., xn taken from B (n, p).
Solution:
The likelihood function of x1 , x2 ,..., xn ,
L( x1 , x2 ,..., xn , ) f ( x1 , p) f ( x2 , p)... f ( xn , p)
L1 ( x , p ) L2 ( x1 , x2 ,..., xn )
where L1 ( x , p) p nx (1 p)nnx , and L2 ( x1 , x2 ,..., xn ) nCx1 . nCx2 ... nCxn , then by factorization
theorem x is a sufficient estimator of p.
log L
Problem 4: If t is sufficient estimator for , prove that is a function of t and only.
Solution:
If t is a sufficient estimator of , then the likely hood function
L( x1 , x2 ,..., xn , ) L1 (t , ) L2 ( x1 , x2 ,..., xn )
log L( x1 , x2 ,..., xn , ) log L1 (t , ) log L2 ( x1 , x2 ,..., xn )
log L( x1 , x2 ,..., xn , ) log L1 (t , ) 0
Since, L1 (t , ) is a function of t and only, log L is also a function of t and only.
Problem 5: If t is sufficient estimator for , then prove that any 1-1 function of t is also sufficient
for .
Solution:
Let h g (t ) , assume h is a 1-1 function of t, then t g 1 (h)
Since t is sufficient for ,
L( x1 , x2 ,..., xn , ) L1 (t , ) L2 ( x1 , x2 ,..., xn )
Let x1 , x2 ,..., xn be the sample taken from the population with p.m.f/p.d.f
f ( x, 1 , 2 ,.. k ) , where 1 , k ,.. k are the parameters involved. The likelihood function of
the sample L( x1 , x2 ,..xk , 1 , 2 ,.. k ) f ( x1 , 1 , k ,.. k ). f ( x2 ,1 , k ,.. k )..... f ( xn ,1 , k ,.. k ) .
The method of maximum likelihood suggests, the best estimators for estimating the
parameters 1 , k ,.. k are the estimators which maximizes the likelihood function. Such
estimators are known as Maximum Likelihood Estimators (M.L.E) of 1 , k ,.. k .
The Principle of M.L.E says that the best estimators of the parameters based on a
sample obtained are, those values of the parameters which make the probability of getting
that sample a maximum.
Using the method of differential calculus, the function of sample values for a
parameter which maximizing the likelihood function- called MLE of that parameter, can
be obtained. Let L( x1 , x2 ,..xn , 1 , 2 ,.. k ) be the likelihood function corresponds to the
sample x1 , x2 ,..xn . The value of 1 , as a function of x1 , x2 ,..., xn , maximizing the likelihood
L 2 L
function can be obtained from 0 , and if for that value of 1 , 0 . But since we
1 12
know the value of 1 , which maximizing the likelihood function also maximizes logL, such
log L 2 log L
value of 1 can also be obtained by using 0 if for that value of 1 , is less
1 12
than zero.
Maximum Likelihood Estimators possess some desirable properties of a good
estimator.
(i) MLE’s are asymptotically unbiased.
(ii) MLE’s are consistent.
(iii) MLE’s are most efficient.
(iv) MLE’s are sufficient, if sufficient statistics exist.
(v) MLE’s are asymptotically normally distributed.
Problem 1: Find the M.L.E. of and , using the random sample x1 , x2 ,..., xn taken from the
normal population N ( , )
Solution:
Given the random sample x1 , x2 ,..., xn from N ( , )
The likelihood function,
2 2
x x x 2
1 2 n
1 2 1 2 1 2
L( x1 , x2 ,..xk , , ) e 2 e 2 ... e 2
2 2 2
n x 2
n i
1 2
e i 1 2
2
n xi
2
1
log L( x1 , x2 ,..xk , , ) n log 2 2
2 i 1
log L
0 2
n
xi (1) 0
i 1 2 2
n
1 n
xi 0
i 1
xi x
n i 1
2 log L 1
2 ( n) 0
2
Hence, x is the MLE of
To obtain the MLE of ,
1 n x
2
log L 1
0 n 2 2 i 3 (2) 0
2 i 1 2
xi
2
n n
i 1 3
0
n
x n 2
2
i
i 1
1 n
2 xi
2
n i 1
xi
2
2 log L n n
2
2
3
i 1 4
1 n
At 2 xi ,
2
n i 1
2 log L n2 n2
3
2 n n
x x
2 2
i i
i 1 i 1
2n 2
n
0
x
2
i
i 1
n i 1
1 n
considered as xi x , which is the sample variance.
2
n i 1
Problem 2: Find the MLE of , based on random samples taken from Poisson population with
parameter .
Solution:
Let x1 , x2 ,..., xn are the random sample taken from P ( ) , then
L( x1 , x2 ,..., xn , ) f ( x1 , ). f ( x2 , )... f ( xn , )
e 1 e 2 e xn e n i
x x x
. .....
x1 ! x2 ! xn ! x1 ! x2 !...xn !
log L 1
0 n x i
i
(0) 0
1
xi x
n i
2 log L 1
2
x i
i 2
2 log L
x i
at x ; i
0 ( samples xi ' s from
2 x
2
This happens when is the median of the random sample x1 , x2 ,..., xn . So MLE of
is the median of x1 , x2 ,..., xn .
Problem 4: Obtain the MLE of and using the random samples x1 , x2 ,..., xn taken from the
x
1
population with pdf f ( x) e , x , 0 .
Solution:
The likelihood function L (x1, x2, . . . , xn, , ) can be written as
x1 x2 xn
1 1 1
L( x1 , x2 ,..., xn , , ) e . e .... e
n
xi
1
e i
(x ) i
log L n log i
log L n
0 0 0 ------ (1)
log L n (x ) i
0 i
0
2
(x ) i
i
n
1
n
(x )
i
i ----- (2)
Equation (1) cannot imply the MLE of . But we know log L is maximized when
( xi ) a minimum is. This happens when is a maximum. But cannot be greater
i
than Min xi . Hence Min xi is the MLE of . Then by (2) , the value of can be written as
1
( xi Min xi ) and it can be verified that atthis value of ,
n i
Problem 5: Obtain the MLE of a and b using the random samples x1 , x2 ,..., xn taken from a
rectangular population over the interval ( a b , a b ) .
Solution:
Here random samples x1 , x2 ,..., xn taken from a rectangular population over the
interval ( a b , a b ) . Hence f(x) is given by
1 1
f ( x) , a b x ab
a b a b 2b
n
1
In this case the likelihood function L( x1 , x2 ,..., xn , a, b)
a b a b
The method of differential calculus cannot be applied here. The likelihood function L is
maximum when a b a b is minimum. This happens when a b is taking its
minimum and a b is taking its maximum possible value.
But a b cannot be less than the largest value of x1 , x2 ,..., xn and a b cannot be
greater than the smallest value of x1 , x2 ,..., xn .
log L 1
0 10 x
i
i
(0) 0
1
xi x
10 i
2 log L 1
2
x
i
i 2
2 log L
x i
at x ; i
0 ( samples xi ' s from
2 x
2
1 n 2
Second raw moment of the sample is xi
n i 1
1 n 2
Equating these two, we get 2 2 xi
n i 1
1 n 2
2
xi 2 ; But x is the estimator of .
n i 1
1 n 2
Hence moment estimator of 2 is xi x 2 , which is the sample variance.
n i 1
18
Equating these we get, 2 2 , ie., 50 2 25 18 0
25
Solving this quadratic expression, we get = 0.295.
Solution:
Out of 20 samples taken from the population, it is recorded 1 one, 5 twos, 7
threes and the remaining 7 fours.
1 1 5 2 7 3 7 4 60
Then first moment of the sample is = 3,
20 20
First moment of the population,
1 1 1 1
E ( X ) 1 2 3 4
4 4 4 4
10 4
4 4
10 4
Equating these we get, 3,
4 4
solving this quadratic expression, we get = 0.5 .
EXERCISES
1. X 1 , X 2 , X 3 are random samples from population with mean and standard deviation
. T1 , T2 , T3 are defined as T1 X 1 X 2 X 3 ; T2 2 X 1 3 X 3 4 X 2 ; and
1
T3 X 1 X 2 X 3 . Are (i) T1 and T2 are unbiased estimator? (ii) Find , such that
9
T3 is an unbiased estimator of (iii) Which is the most efficient estimator?
2. Define consistent estimator. Obtain the sufficient conditions for consistency.
13. Find an estimator of , based on random samples taken from Poisson population with
parameter by the method of moments.
14. Obtain the moment estimate of , if the probability masses are
X: 1 2 3 4
1 1 1 1 ; 0 1, and the observed
f ( x) :
4 4 4 4
frequencies are 1,5,7 and 7 respectively.
********************
CHAPTER 3
INTERVAL ESTIMATION
confidence coefficient. It is to be noted that there may be many confidence interval for a
particular parameter with same confidence coefficient. Shortness, stability etc., are
some desirable property to identify a good interval.
3.2. Confidence interval for the mean of a normal population with confidence
coefficient 1 :
Let x1 , x2 ,..., xn be the sample taken from N ( , ) and let the sample mean be x . We
use x - the point estimator of for its interval estimation. The mean x follows N ( , ),
n
or t
x n
~ N ( 0,1 ) .
From standard normal table it can observe the value t such that,
2
P( |
x n
| t ) 1
2
P ( t
x n
t ) 1
2 2
P ( x t x t ) 1
2 n 2 n
Multiplying by -1 P ( x t x t ) 1
2 n 2 n
P ( x t x t ) 1
2 n 2 n
x t , x t .
2 n 2 n
Let x1 , x2 ,..., xn be the sample taken from N ( , ) and let the sample mean be x . It
P( | tn 1 | t ) 1
2
P( |
x n 1
| t ) 1
s 2
s s
P ( t x t ) 1
2 n 1 2 n 1
s s
P ( x t x t ) 1
2 n 1 2 n 1
s s
P ( x t x t ) 1
2 n 1 2 n 1
s s
x t , x t .
2 n 1 2 n 1
Solution:
x t , x t
2 n 2 n
17 21 20 18 19 22 20 21 16, 19
We have, x 19.3
10
so that, P ( | t | t ) 0.95 .
2
3 3
19.3 1.96 , 19.3 1.96 = 17.44 , 21.16
10 10
Problem 2: Find the least sample size required if the length of 95% confidence interval for the
mean of a normal population with standard deviation 4 should be less than 5.
Solution:
Let n random samples are taken from the population N ( , 4) . The confidence
interval for the mean is,
x t , x t
2 n 2 n
Given the confidence coefficient is 95 %. Hence from the standard normal table t =
2
1.96.
To find the minimum number of samples such that, the length of the interval of ,
2
x t x t t 5
2 n 2 n 2 n
2 4
2
n 1.96 n 9.83
5
Problem 3: A sample of size 17 taken from N ( , ) . Mean of the sample is 12 and the sample
variance is 4. Using the data, find a 90% confidence interval for .
Solution:
s s
x t , x t
2 n 1 2 n 1
Since the confidence coefficient is 90%, from the table of t distribution for 16 d.f., we
get t = 1.746, so as P( | t16 | t ) 0.90 .
2 2
2 2
12 1.746 , 12 1.746 11.127 , 12.873
16 16
Problem 4: For a N ( ,3) population, construct a 95% confidence interval for 3 5 , on the basis
of the random sample of size 25. The sample mean was found to be 30.
Solution:
y (3 5)
u ~ N (0,1)
9
25
y (3 5)
P ( | t | t ) 0.95 . Hence, P ( | | 1.96 ) 0.95
2
9
25
9 9
y 1.96 , y 1.96
25 25
9 9
95 1.96 ,95 1.96 91.472, 98.528
25 25
Problem 5: Show that the length of the confidence interval for the mean of a normal population
with known variance can be made however small we please by increasing the sample size.
Solution:
The confidence interval for the mean of a normal population when 2 is known is
given by x t , x t .
2 n 2 n
The length of the interval = x t x t
2 n 2 n
2 t ,
2 n
3.3. Confidence interval for the difference of means of two normal populations having
known common variance 2 :
Let n1 and n2 are the number of samples independently drawn from to normal
S 2 be the standard deviations of the samples drawn from the first and second population
respectively.
x1 ~ N 1 , and x2 ~ N 2 ,
n1 n2
2 2
Then, x1 x2 ~ N 1 2 ,
n n
1 1
x1 x2 1 2
t ~ N (0,1)
1 1
n1 n1
From standard normal table it can observe the value t such that,
2
P ( | t | t ) 1
2
x1 x2 1 2
P( | | t ) 1
1 1
2
n1 n1
1 1 1 1
P x1 x2 t 1 2 x1 x2 t 1
2 n1 n1 2 n1 n1
1 1 1 1
P x1 x2 t 1 2 x1 x2 t 1
2 n1 n1 2 n1 n1
1 1 1 1
x1 x2 t , x1 x2 t
2 n1 n1 2 n1 n1
Problem 1: The average mark scored by 32 boys in an examination is 72 with a standard deviation
of 8, while that scored by 32 girls is 70 with a standard deviation of 6. Construct a 99 %
confidence interval for the difference of means. (Assume S.D’s are equal)
Solution:
1 1 1 1
x1 x2 t , x1 x2 t
2 n1 n1 2 n1 n1
n1s12 n2 s2 2
The common value for variance (since we have large samples)
n1 n2
=
32 82 32 62 = 7.07.
32 32
x1 72 ; 1 8; n1 32 and x2 70 ; 2 6; n2 32
1 1 1 1
72 70 2.57 7.07 , 72 70 2.57 7.07
32 32 32 32
= 2 4.543, 2 4.543
= 2.543, 6.543 .
Let x1 , x2 ,..., xn be the sample taken from N ( , ) with sample variance S 2 . Then,
nS 2
2 follow chi-square distribution with (n-1) d.f.
2
P 2 ( n 1) 2 and P 2 ( n 1) 2 1 respectively.
2 1 2
2 2
P 2 2 2 1
1
2 2
2 2
nS 2 1
ie., P 2 2 1 P 2 1 1
2
1 2
nS
2
2 nS 2
2 2
nS 2 nS 2
P 2 2 2 1
1
2 2
nS 2 nS 2
P 2 2 2 1
1
2 2
Problem 1: A sample of size 12 taken from N ( , ) . Mean of the sample is 10 and the sample
Solution:
Given n = 12, S 2 = 9
nS 2 nS 2
Confidence interval for 2 is given by 2 ,
2
1
2 2
For 90% confidence interval, ie., for 0.10 , from table of chi-square distribution for 11
d.f,
2
4.58 and 2 19.68
1
2 2
12 9 12 9
Hence the 90% confidence interval for 2 is,
4.58
,
19.68
5.49 , 23.58
Problem 2: An optical firm purchases glass for making lenses. Assume that the refractive index of
20 pieces of glass have variance of 1.20 X 104 . Construct a 95% confidence interval for the
population variance.
Solution:
Confidence interval for 2 with confidence coefficient 1 is,
nS 2 nS 2
2 , .
2
1
2 2
2 = 32.8523 and 2
= 8.9066.
1
2 2
Problem 3: Construct a 95% confidence interval for the variance 2 of the normal population
with unknown mean using the following sample:
4.5, 10.2, 10.5, 9.8, 13.0, 19.2, 15.5, 13.3, 10.8 and 16.4
Solution:
nS 2 nS 2
(1 )% Confidence interval for 2 is given by 2 ,
2
1
2 2
For (1 )% =95% , 0.05 ; from table of chi-square distribution for n-1 = 9, d.f,
2
2.7004 and 2 19.0228
1
2 2
To find the sample variance of the population, using the given 10 samples
x x2
4.5 20.25
10.2 104.04
10.5 110.25
9.8 96.04
13.0 169
19.2 368.64
15.5 240.25
13.3 176.89
10.8 116.64
16.4 268.96
123.2 1670.96
123.2
x 12.32
10
1 1
x12 x 1670.96 12.32 3.9
2
s
2
10 i 10
10 3.9 2 10 3.9
2
Hence the 90% confidence interval for is, 2
,
19.0228 2.7004
7.996 , 56.325 .
B(n,p). When n becomes very large, X follows normal distribution N (np, npq ) .
X np
Then, t ~ N (0,1)
npq
X
p
t n ~ N (0,1)
pq
n
As an approximation,
p p
t ~ N (0,1) ; q 1 p
pq
n
X
(where, p is the sample proportion)
n
P ( | t | t ) 1
2
p p p p
P( | | t ) 1 P ( t t ) 1
pq 2 2 pq 2
n n
pq pq
P ( p t p p t ) 1
2 n 2 n
pq pq
P ( p t p p t ) 1
2 n 2 n
pq pq
ie., P ( p t p p t ) 1
2 n 2 n
pq pq
p t , p t where, q 1 p
2 n 2 n
Problem 1: Random samples of 120 workers of a factory 40 are dissatisfied with their working
conditions. Form a 95% confidence interval for the proportion of dissatisfied workers of the factory.
Solution:
40 1
Given the sample proportion of dissatisfied workers p = . Also given the
120 3
confidence coefficient = 95%
pq pq
p t , p t
2 n 2 n
For 95% confidence coefficient, from table of standard normal distribution, t 1.96
2
1 2 1 2
1
1 3 0.214 , 0.4526
1.96 3 3 , 1.96 3
3 120 3 120
(n=1492)
Solution:
pq pq
p t , p t
2 n 2 n
The given confidence coefficient is 99%. From standard normal table t 2.57
2
24 4
From the 30 samples taken, p .
30 5
pq
The length of the confidence interval = 2t .
2 n
4 4
1
pq 5 5
2t = 0.05 . That is to get n, such that 2 2.57 0.05
2 n n
2 2.57 4 4
2
n 1 1690.85 1691
0.05 5 5
normal distribution N ( , ) .
The sample mean x ~ N ( , ) for large n.
n
x
t ~ N (0,1), approximately ( E ( x ) )
x
n
P ( | t | t ) 1
2
x x
P ( t x t ) 1
2 n 2 n
x x
P ( x t x t ) 1
2 n 2 n
x x
P ( x t x t ) 1
2 n 2 n
x x
x t , x t
2 n 2 n
EXERCISES
1. The mean of a sample of size 30 drawn from a normal with mean and variance
4. Derive 95% confidence interval for the mean of a normal population when (i)
5. Obtain 95% confidence interval for the parameter of Poisson distribution on the
mean and the sample variance are respectively 22 and 16. Find a 90% confidence
interval for .
7. The mean of a sample of size 24 drawn from N ( , ) is 25.5. The sample variance
10. Obtain a large sample 100 (1 ) % confidence interval for the parameter , in
*********************
CHAPTER 4
TESTING OF HYPOTHESIS
For eg., assume x1 , x2 ,..., xn are n random samples taken from a normal population
H : 500 . This hypothesis is to be tested, against the alternatives, (i) is greater than
500 or (ii) is less than 500 or (iii) either >500 or <500; ie., 500 .
H 0 : 500 .
In this situation we are taking a random sample of n bulbs of new type and find the
average life length of the sample item. Based on the sample mean (which is a good
estimator of population mean) we decide whether to accept or reject the null hypothesis.
Roughly speaking, if the alternate hypothesis considered is H1 : 500 , and the sample
mean is much higher than 500, the hypothesis H 0 : 500 is rejected. If H1 : 500 , and
the sample mean is much lesser than 500, H 0 is rejected and if H1 : 500 , and the sample
Here we are making decision based on the sample mean, because we had to make
a decision on population mean and sample mean is a good statistic to say something
about the population mean. As this, in any statistical test we have to find an appropriate
statistic to make decision based on its value. The value of statistic can be calculated by the
value of the samples selected. Such a statistic used in testing of hypothesis is termed as
test statistic.
Type-I error: It is an error due to rejecting the null hypothesis H 0 , when H 0 is true.
Type-II error: It is an error due to accepting the null hypothesis H 0 , when H 0 is false.
Action taken H 0 is H 0 is
Reject TYPE-I NO
H0 ERROR ERROR
Accept NO TYPE-II
H0 ERROR ERROR
Using better
statistical criteria, the possible errors in testing procedure can be minimized.
(iii) Divide the range of variation of the test statistics into two regions, acceptance
region and rejection region, considering some probabilistic restrictions.
(iv) Take a random sample from the given population and calculated the value of
the test statistic and decide whether to accept or reject the hypothesis.
STATISTICAL INFERENCE Page 71
School of Distance Education
In a testing procedure, the null hypothesis is rejected when the value of the test
statistic falls in the pre-decided rejection region. Since the value of the test statistic is
decided by the sample drawn, there is a chance for the value of the test statistic to fall in
the rejection region even though the null hypothesis is true (type-I-error). The probability
of the value of the test statistic to fall in the rejection region, even though the null
hypothesis is true is known as significance level or size of the test, denoted by .
The probability of the value of the test statistic to fall in the rejection region, when
the alternative hypothesis is true is known as power of the test, denoted by .
1
f ( x, ) , 0 x
. Find the size and power of the tests if the
0, otherwise
critical regions are (i) x>0.5 (ii) 1 < x < 1.5
Solution:
Given H 0 : 1 and H1 : 2
1
dx/ 1
0.5
2
1 1
dx 2 dx
0.5
/ 2 0.5
1 2
x 0.75
2 0.5
(ii) When the test is with critical region 1 < x < 1.5,
1.5
1
dx/ 1
1
, for the given population 0 x
0.
1.5 1.5
1 1
dx dx = 0. 25
/ 2 2
1 1
Problem 2: In a coin tossing experiment, let p be the probability of getting a head. The coin is
tossed 10 times to test the hypothesis H 0 : p 0.5 against the alternative H1 : p 0.7 . Reject H 0 , if
6 or more tosses out of 10 result in head. Find significance level and power of the test.
To test, H 0 : p 0.5 against H1 : p 0.7 . 10 tosses of the coin is considered and let
X denote total number of heads obtained. Then X follows binomial distribution B(10,p).
The critical region is X 6 .
= P (Rej H 0 / H 0 is true)
= P ( X 6 / p = 0.5)
10
10C x p x q10 x / p 0.5
x 6
10
386
10C x 0.5 x (1 0.5)10 x
x 6 210
10
10C x p x q10 x / p 0.7
x 6
10
10C x 0.7 x (0.3)10 x 0.8495
x 6
x 1, find the probability of type–I and type- II errors. Also find the power function of the test, if
H1 suggested is H1 : r , where r 2 .
Solution:
e x dx / 2
1
1 2 1
= P ( x 1/ 1 )
e x dx / 1
1
e x
e dx
x
e
1
1 1 1
e x dx / r
1
e rx
re rx
dx r r
e ,r2
1 r 1
Problem 4: In a city the milk consumption of families, x, is assumed following the distribution
1 x
with, p.d.f. f ( x) e , x 0, 0 . To test H 0 : 5 against H1 : 10 . H 0 is rejected ,if a
family selected at random consumes 15 units or more. Obtain the size and power of the test.
Solution:
x
1
e dx / 10
15
1
x
x
10 15
e 10 dx e 10
15
15 3
e 10 e 2
.
Problem 5: x1 , x2 ,..., x9 are 9 random samples drawn from a normal population N ( ,5) . To test
Solution:
= P (Rej H 0 / H 0 is true) = P ( x 7 / = 5)
Since x1 , x2 ,..., x9 are random sample from normal population, the sample mean x
5
is distributed as N ( , ).
9
x 9
Hence U ~ N (0,1)
5
x 9 7 9
P( x 7 / = 5) = P / 5
5 5
STATISTICAL INFERENCE Page 76
School of Distance Education
= P U 1.2 = 0.1151 (from std. normal table)
= P ( x 7 / = 8)
x 9 7 9
= P / 8
5 5
1 x
f ( x) e , x 0 , 0 . To test H 0 : 2 against H1 : 4 . Reject the hypothesis if
x1 x2 9.5 . Obtain significance level and power of the test
Solution:
Here the given critical region is a plane in the first quadrant as given in the graph,
where x1 x2 9.5 , x1 0, x2 0 .
x x2
1 1 1
f ( x1 , x2 ) e e ,x 0, x2 0 ( x1 , x2 are ind .)
1
P ( x1 x2 9.5 / 2 ) = 1 - P ( x1 x2 9.5 / 2 )
Note that,
9.5 9.5 x1 x1 x2
1
P( x1 x2 9.5 / 2 ) =
x1 0 x2 0
2
e dx2 dx1 / 2
9.5 9.5 x1 x1 x2
1
x1 0
x2 0
16
e 4 dx2 dx1
9.5 9.5 x1
x1 x2
1
16
0 0
e 4 dx2 dx1
9.5 x 9.5 x1 x
1 1 2
4 e
0
2
0
e 2 dx2 dx1
9.5 x
1
9.5 x1
1 2 e 0
4 e 2
2 e
dx1
0
9.5 9.5 x
1 1
2 e
0
2 e 2 dx1
9.5
1 9.5
x1
x1e 2 2e 2
2 0
9.5
1
9.5
9.5
9.5e 2 2e 2 2
2 0
= 1 - P ( x1 x2 9.5 / 4 )
9.5 9.5 x1 x1 x2
1
P ( x1 x2 9.5 / 4 ) =
x1 0 x2 0
2
e dx2 dx1 / 4
9.5 9.5 x1 x1 x2
1
x1 0
x2 0
16
e 4 dx2 dx1
9.5 9.5 x1
x1 x2
1
16
0 0
e 4 dx2 dx1
9.5
1
9.5
9.5
9.5e 4 4e 4 4
4 0
9.5 9.5
9.5
1 e 4 e 4
4
9.5 9.5
9.5
Hence, power of the test = 1- 1 e 4 e 4
4
13.5 9.5
= e 4 = 0.31 (approx.)
4
H 0 : 1 against H1 : 2 , using a random sample x1 , x2 of size 2 and define the critical region as
3
C { x1 , x2 ; x2 } . Obtain significance level and power of the test
4 x1
Solution:
3 3
To test H 0 : 1 against H1 : 2 , given critical region is x2 , or x1 x2
4 x1 4
3
= P( x1 x2 / 1)
4
1 1
3 1
Hence, P( x1 x2 / 1) 3 3 x x dx dx / 1
2
1 2 2 1
4
x1 x2
4 4 x1
1 1 1
3
3 3 1 dx2 dx1 3 1 4 x dx1
1
x1 x2 x1
4 4 x1 4
1
3 3 3 3
x1 log x1 1 log
4 3 4 4 4
4
1 3 3
= log
4 4 4
3
= P( x1 x2 / 2 )
4
1 1
3 3 4 x1 x2 dx2 dx1
x1 x2
4 4 x1
2
3
1
1 4 x1
4 x1 ( )dx1
3 2 2
x1
4
1
1 9
4 3 x ( 2 32 x
1
1
2
)dx1
x1
4
1
x12 9
4 log x1
4 32 3
4
1 9 9 3
power 1 4 log
4 64 32 4
7 9 3
log
16 8 4
So far we considered some testing of hypothesis problem with given critical region.
The power and significance level corresponding to a given critical region is calculated.
Now a question arising is, can we find a critical region with maximum power and zero
significance level? But it is not possible. When we are making the probability of type I
error minimum, the probability of type II error increases, thereby power decreases, and
vice versa.
L0
and for each ( x1 , x2 ,..., xn ) belongs to S, c
L1
L0
for each ( x1 , x2 ,..., xn ) NOT belongs to S c
L1
where L0 and L1 are the likelihood of the sample when H 0 is true and H1 is true
Then S is the most powerful critical region with significance level to test the
simple hypothesis H 0 against a simple alternative H1 .
Proof:
n
when H 0 is true is denoted as, L1 f ( x1 , 1 ) f ( x2 , 1 )....... f ( xn , 1 ) f ( xi , )
i 1
1
L0 L0
c , for each ( x1 , x2 ,..., xn ) S, and c for each ( x1 , x2 ,..., xn ) S
L1 L1
Consider another critical region S ' , for the given test of size . be the sample space.
But, S L dx
0 L0 dx ' L dx 0 and L dx
0 L0 dx ' L dx
0
S ' S S S S' S ' S S S
L0 dx L0 dx ----- (1)
S ' S S ' S
= S L dx
1 L1dx ' L dx 1
S ' S S S
L0 L0
' c
dx L1dx (
L1
c for x1 , x2 ,..., xn S )
S S S S '
L0 L0
By (1),
' c
dx
'
c
dx
S S S S
L0
S L dx
1
'
c
dx L1dx
S S S S '
L0
S L dx
1
'
L1dx ' L dx 1 (
L1
c for x1 , x2 ,..., xn S )
S S S S
S L dx
1 ' L dx
1
S
ie., power of the critical region S Power of the critical region S ' , Hence the
critical region satisfying the conditions of Neymaan-Pearson theorem is the best or most
powerful critical region.
Problem 1: Use Neymaan-Pearson Theorem to find a most powerful test with significance level
for testing the hypothesis H 0 : 0 against, H1 : 1 , ( 1 0 ) using a random sample
1
1 ( x )2
x1 , x2 ,..., xn drawn from the population with pdf f ( x) e 18 , x .
18
Solution:
1
1 ( x )2
Given f ( x) e 18 , x . For the random samples x1 , x2 ,..., xn ,
18
n 1 n 2
1 18 i1( xi )
the likelihood function L f ( x1 , x2 ,..., xn , ) e .
18
n n
But
but , ( xi )2
i 1
(x x x )
i 1
i
2
n n
( xi x ) 2 ( x ) 2 nS 2 n( x ) 2 , where x is the sample
i 1 i 1
n
1 18 S 2 ( x )2
n
Therefore, f ( x1 , x2 ,..., xn , ) L e
18
L0
for ( x1 , x2 ,..., xn ) belongs to S, c
L1
n
1 18 ( S 2 ( x 0 )2 )
n
L0 18
e
ie., c
L1 1 n n ( S 2 ( x 1 )2 )
18
e
18
n (( x 0 )2 ( x 1 )2 )
e 18 c
n
(( x 0 ) 2 ( x 1 ) 2 ) log c ---- (1)
18
n
Since 1 0 , dividing both sides of (1) by a negative quantity, ( 1 0 ) , we get,
18
18 log c
(2 x 0 1 )
n( 1 0 )
18 log c
`` for a most powerful critical region, x 1 0 1
2 n( )
1 0
18 log c
Let, 1 0 1 c1 ; Then for most powerful critical region,
2 n( )
1 0
x c1 --------- (2)
the critical region, ie., x c1 , when H 0 is true should be . Using this condition we
ie., P ( x c1 / 0 )
3
Since x1 , x2 ,..., xn are random samples from N ( ,3) , the sample mean x ~ N ( , )
n
t
x n
~ N (0,1)
3
P(
x n
c1 n
/ 0 )
3 3
c1 0 n
);
P( t t ~ N (0,1) ----- (3)
3
From standard normal table a value t can be identified as shown such that,
t
c1 0 n
c1 0
3 t
3 n
3 t
x 0
n
Remark: In the derived most powerful critical region, it can be observed that whatever be
the value of in the alternative hypothesis, keeping the condition ( 1 0 ) , the most
powerful critical region is unchanged. That is the most powerful test using this critical
region is uniform.
Hence such a critical region is called Uniformly Most Powerful Critical region or
the test by using such a critical region is Uniformly Most Powerful Test (UMPT).
Problem 2: Use Neymaan-Pearson Theorem to find a most powerful test with significance
Solution:
Note that x1 , x2 ,..., xn are random sample taken from N ( , 2 ) . Hence the
L0
for ( x1 , x2 ,..., xn ) belongs to S, c
L1
1 n 2
2 02 i1( xi )
n
1
e
L0 0 2
ie., 1 n c
L1 n ( xi )2
1 212 i1
e
1 2
1n 1 n
( xi )2 ( xi )2
L0 n 2
2 0 i1 212 i1
1n e c
L1 0
1 n 1 n
( xi )2 ( xi )2
2
2 0 i1 212 i1 0n
log e log ( c n )
1
1 1 n 0n
i
2
( x ) log ( c )
2 1 2 0 2 i 1 1n
2
Since 12 0 2 ,
1 1 0 2 12
is negative
2 1 2 0 2 2 0 2 12
2
0n
log ( c )
n
1n
(x )
i 1
i
2
1 1
c1 ( say )
2 1 2 0 2
2
(x )
i 1
i
2
c1
n
ie., P ( (x )
i 1
i
2
c1 / 2 0 2 )
n ( x )2 c1
P i 2 / 2 02
i 1 2
( xi ) n
( xi ) 2
but we have
~N (0,1) and
i 1 2
~ 2( n)
n c1
P 2( n) / H 0 ---- (1)
i 1 02
c1
Comparing (1) and (2), 2 :n c1 2 :n 0 2
02
n
Hence for most powerful critical region, (x )
i 1
i
2
2 :n 0 2
EXERCISES
1. Let p be the probability that a coin will fall head in a single toss in order to test
1 3
H 0 : p against H1 : p . The coin is tossed 5 times and H 0 is rejected if more
2 4
than 3 heads obtained. Find the size and power of the test.
2. 10 random samples x1 , x2 ,..., x10 are taken from N ( ,5) . To test H 0 : 0 against
H1 : 2 . The critical region suggested is x1 2 x2 3 x3 ... 10 x10 1 . Obtain the
probability of type-I and type-II errors.
1 3
3. Let p be the proportion of smokers in a city. To test H 0 : p against H1 : p .
2 4
H 0 is rejected if 60 or more persons are found smokers in a sample of 100 persons.
Compute significance level and power of the test.
4. A single value x is drawn from a normal population N ( m, 5) . The null hypothesis
H 0 : 50 is accepted if x 75 . Otherwise H1 : 60 is considered. Evaluate
significance level and power of the test.
5. A sample of size 10 is taken from a normal population with 1 to test H 0 : 5
against H1 : 6 . The critical region is x 5.52 . Find significance level and power
of the test.
6. Obtain the best critical region for testing H 0 : 0 against H1 : 0 in N ( , )
using a random sample of size n. Also find the power function.
7. In testing H 0 : 0 against H1 : 1 ( 0 ) for the distribution with pdf
( x )
f ( x) e , 0 x , 0 . Show that the UMP test is of the form x i
constant and x i constant.
CHAPTER 5
U E (U )
t ~ N (0,1) for large n. This important result may profitably used for the test
SD (U )
construction.
The standard deviation of any statistic is called its standard error. While testing a
hypothesis H 0 : 0 , naturally taking a statistic U with E(U) = . Using standard error of
U, a test statistic t following standard normal distribution can be formed. Then for large
n, the probability that it will fall in any region in its range of variation can be found using
standard normal table.
Case I: is known:
Let x1 , x2 ,..., xn be the samples from the population with sample mean x and
sample variance S 2 . In testing of the population mean the sample mean x (which is an
P ( x c / 0 ) .
Since the samples are taken from a population with mean and standard
deviation , x ~ N ( ,
) for large n. That is, z
x n
~ N (0,1) for large n.
n
P ( x c / 0 )
P(
x n
c n
/ 0 )
P( z
c 0 n
) , z ~ N (0,1) ---- (1)
From the table of standard normal distribution, one can get a t , such that
c t 0
n
Then the test criterion is to reject H 0 : 0 against H1 : 0 , when, x t 0
n
P ( x c / 0 ) .
Therefore, P( z
c 0 n
) , z ~ N (0,1) ----- (2)
c 0 n
t .
Then the test criterion is to reject H 0 : 0 against H1 : 0 , when, x t 0
n
P( z
c n
) , z ~ N (0,1) ---------- (3)
c n
Hence, from (3) t
2
c t
2 n
Then the test criterion is to reject H 0 : 0 against H1 : 0 , when, | x 0 | t
2 n
Solution:
Let be the mean height of the students. To test whether the data consists of 400
Test criterion is t
x 0 n
.
Here, t
4.75 4.48 400
3.6
1.5
Then, here t t . So reject H 0 at 1% level. That is the data contradict the assumption
2
Problem: An insurance agent has claimed that the average age of policy-holders who insure
through him is less than the average for all agents, which is 30.5 years. A random sample of policy-
holders who had insured through him gave the following age distribution:
Age: 16 – 20 21 – 25 26 – 30 31 – 35 36 – 40
No. of persons: 12 22 20 30 16
Let be the mean age of the policy-holders insured by the agent. A sample of 100
x f
16 – 20 12 18 216 3888
21 – 25 22 23 506 11638
26 – 30 20 28 560 15680
31 – 35 30 33 990 32670
36 - 40 16 38 608 23104
2880 86980
1 1
x
N
fx i
i i
100
2880 28.8
1
S .D. f x x 2 2
i i
N i
1
86980 28.8
2
100
Now t
28.8 30.5 100
2.68
6.35
Here, the calculated value of t = - 2.68 < - t = -1.65. Hence reject H 0 : 30.5 at
samples from first population and n2 samples from the second population taken
P( x1 x2 c / H 0 )
t E (t )
If t x1 x2 , z ~ N (0,1) for large n. Then,
SD(t )
x1 x2 1 2 c 1 2
P / H 0 .
1 2
2 2
1 2
2 2
n 1 n2 n 1 n2
This implies that
P z , from standard normal table P ( z t )
c
12 2 2
n1 n2
c 12 2 2
t ; so, c t
12 2 2 n1 n2
n1 n2
12 2 2
Hence reject H 0 : 1 2 against H1 : 1 2 , if x1 x2 t
n1 n2
x1 x2
The critical region is t t
12 2 2
n1 n2
In similar way,
And,
Remark: If 1 , 2 are unknown, since the sample size is large, the values of sample
n1S12 n2 S 2 2
value of is approximated by .
n1 n2
Problem.1: A sample of 400 men from South India has a mean height of 170 cms. and a standard
deviation of 30 cms. while a sample of 200 men from North India has a mean height of 178 cms
with a standard deviation of 32 cms. Do the data indicate that North Indians are on the average
taller than the South Indians?
Solution:
Large sample of sizies 400 and 200 respectively are taken from two populations and
their mean is found to be 170 and 178. The sample standard deviations are 30 and 32
respectively.
test H 0 : 1 2 against H1 : 1 2 .
x1 x2
The test statistic used is t . Since the sample sizes are large, then if
12 2 2
n1 n2
170 178
Hence here, t = -2.94
302 322
400 200
Reject H 0 , if t t
Consider a significance level of 5%. Then, from standard normal table t 1.645 .
Therefore here, t= -2.94 < t 1.645 . Hence reject H 0 . That is the data indicates that
Problem.2: The mean height of 50 male students who showed above average participation in
college athletics was 68.2 inches with a standard deviation of 2.5 inches; while 50 male students
who showed NO interest in such participation had a mean height of 67.5 inches with a standard
deviation of 2.8 inches.
i. Test the hypotheses that male students who participate in college athletics are taller
than other male students.
ii. By how much should the sample size of each of the two groups be increased in
order that the observed difference of 0.7 inches in the mean heights significant
Solution:
i. Let X 1 and X 2 are the variables representing respectively, the heights of the
students who showed and who not showed interest in athletics. Let n1 be the size of
sample with mean x1 and S.D. s1 , taken from the first set. And n2 be that of the second
Let 1 and 2 be the population mean of the sets of students considered. To test
x1 x2
The test statistic used is t . Since the sample sizes are large, then if
12 2 2
n1 n2
Reject H 0 , if t t
But here t t .Hence we accept H 0 . That is, the data is not significant to believe
the first set students are taller than that of the second set.
ii. Let n be the size of the samples taken from the two sets of students.
The difference of 0.7 units in the mean heights of the samples of sizes n, will
become significant if t t , for that value of n.
68.2 67.5
t 1.645
2.52 2.82
n n
0.7
1.645
14.09
n
2
1.645
n 14.09 77.81 .
0.7
Problem 3: A sample of 200 students from college ‘A’ scored mean mark of 65 with standard
deviation 8 for mathematics in a university examination. Another sample of 100 students from
college ‘B’ scored a mean mark of 60 in the same paper with a standard deviation 6. Does the data
indicate any significance difference between the colleges in terms of the performance in mathematics
paper? Assume the S.D’s are same. (sig. level 5%).
Solution:
Let 1 be the mean marks of the students of college ‘A’ and 2 be that of college
The S.D’s of the populations are unknown and assumed same. Then the
n1S12 n2 S 2 2
common standard deviations can be approximated by,
n1 n2
200 64 100 36
= 7.393
200 100
x1 x2
In this case, the test statistic is, t
n1S12 n2 S2 2 1 1
n1 n2 n1 n2
65 60
t =5.49
200 64 100 36 1 1
200 100 200 100
Problem 4: Electric bulbs manufactured by X and Y companies gave the following result:
X 100 1300 82
Y 100 1248 93
Using standard error of the difference between means, state whether there is any significant
difference in the life of the two makes.
Solution:
Let 1 and 2 be the average life of bulbs by company X and company Y respectively.
To test, H 0 : 1 2 against H1 : 1 2 .
12 2 2
Standard error of the difference between means =
n1 n2
x1 x2
The test criterion is to reject t t , where, t .
12 2 2
2
n1 n2
Since the sample sizes are large use corresponding sample standard deviations
instead of 1 and 2
1300 1248 52
Hence, t = 4.19
2
82 93 2 12.4
100 100
x x
, which is the sample proportion regarding the particular characteristic. E( ) p .
n n
x
Then reject H 0 against H1 : p p0 , if c.
n
Assume the significance level is . Then for the critical region, it is to find c, such
x
that P ( c / H0 ) .
n
x x x x x
n E( n ) c E( n ) E( )
P / H 0 , where, z n n ~ N (0,1) for large n
x x x
SD( ) SD( ) SD( )
n n n
x x c p0
E ( ) p , SD ( ) pq P z
n n p0 q0
n
c p0 p0 q0
t c p0 t
p0 q0 n
n
x p0 q0
Hence the critical region with size is, p0 t
n n
t t , and
t t .
2
Problem: A random sample of 500 pineapples was taken from a large consignment and 65 were
found to be bad. Test the hypothesis that the percentage of bad apples is 20% at 5% level of
significance. Also obtain a 95% confidence interval for the percentage of bad pineapples in the
consignment.
Solution:
x
p0
The test statistic is t n , reject H 0 , if t t .
p0 q0 2
n
65
0.20
Given, n = 500, x = 65. Then t 500 = -3.91
0.20 0.80
500
Hence for 0.05 , we get a confidence interval for the proportion of bad
x p0 q0 x p0 q0
pineapples as t , t .
n n n 2 n
2
x
by .
n
= 0.1005 , 0.1595
independently taken n1 samples from first population and n2 samples from the second
x1 x2
population. Let , are the sample proportions regarding the character considered, for
n1 n2
x1 x1 x x
If H1 : p1 p2 , reject H0 when c where E 1 1 p1 p2 and
n1 n1 n1 n1
x x p1q1 p2 q2 x1 x1
SD 1 1 . To find with significance level , P ( c / H0 )
n1 n1 n1 n2 n1 n1
x1 x1
p1 p2
c p1 p2
P 1 / H0
n n1
p1q1 p2 q2 p1q1 p2 q2
n1 n2 n1 n2
c (0)
P z under H 0 : p1 p2 p ( say )
pq pq
n1 n2
c (0) pq pq
t c t
pq pq n1 n2
n1 n2
x1 x1
n1 n1
Then, reject H 0 , against H1 : p1 p2 , if, t t
pq pq
n1 n2
If H1 : p1 p2 , reject H 0 , if t t
If H1 : p1 p2 , reject H 0 , if t t
2
n1 p '1 n2 p '2 x x
If p is unknown, estimate its value by p , where p '1 1 and p '2 2 .
n1 n2 n1 n2
Problem 1: A random sample of 400 men and 600 women were asked whether they should like to
have a flyover near their residence. 200 men and 325 women were in favor of the proposal. Test
the hypothesis that proportions of men and women in favor of the proposal are same against that
they are not, at 5% level of significance.
Let p1 be the proportion of women in favor of the proposal and p2 is that of men.
Given a sample of size n1 600 from the women, and n2 400 from the group of
325
men. From the sample the proportion of women in favor of fly over p1' and that of
600
200
te second city is p2 ' . Assume p1 p2 p , where
400
n1 p '1 n2 p '2
p ,
n1 n2
325 200
This implies p 0.525
600 400
x1 x1
n1 n1
The test statistic used is t
pq pq
n1 n2
325 200
600 400 0.542 0.5
= = 1.304
0.525 0.475 0.525 0.475 0.0322
600 400
That is men and women do not differ significantly regarding the proposal for the
fly over.
Problem 2: In a year there are 956 births in a town A, of which 52.5% were males, while when
towns A and B are combined; this proportion in a total of 1,406 births was 0.496. Is there any
significance difference in the proportion of male births in the two towns? (sig. level 1%)
Let p1 be the proportion male birth in the town A, and p2 be the proportion of
The number of samples from town A n1 956 . Given the sample proportion of
The combined sample proportion of male birth out of 1,406 births from town A and
B is given as 0.496.
n1 p '1 n2 p '2
From the given combined proportion, p 0.496 , the sample
n1 n2
x1 x1
n1 n1
The test statistic used is t
pq pq
n1 n2
0.525 0.434
0.496 0.504 0.496 0.504
956 450
0.091 0.091
3.791
0.25 0.25 0.024
956 450
Hence reject H 0 , that the proportion of male birth are equal for both the town.
Solution:
Given a sample of size n1 400 and n2 300 . From the sample, the proportion of
20
defective production before overhauling p1' and that of after overhauling is
400
10 n1 p '1 n2 p '2
p2
'
. Assume p1 p2 p , where p ,
300 n1 n2
20 10
This implies p 0.0429
400 300
x1 x1
n1 n1
The test statistic used is t
pq pq
n1 n2
20 10
400 300 0.016666
= = 1.0769
0.0429 0.9571 0.0429 0.9571 0.0154762
400 300
The chi-square test is one of the simplest and most commonly used non-parametric
tests of significance by Karl Pearson. It is the most suitable test to compare the obtained
Consider a set of n possible events, arranged in classes or cells. Let these events
occur with frequencies O1 , O2 ,..., On called observed frequencies. Out of N observations the
expected (theoretical) frequency of each possible event can be evaluated from the
knowledge of the probability distribution suggested for the population. Let they are
denoted by E1 , E2 ,..., En . Our problem is to verify whether the suggestion regarding the
Oi Ei
2
n
2
i 1 Ei
, which follows chi-square distribution with (n-1) degrees of
freedom.
If the calculated value of 2 greater than 2 , it is to conclude that the data could
not have possibly come from a population giving rise to theoretical frequencies.
3. The degree of freedom of 2 is one less than the total number of classes. If r
parameters are estimated using the observations for the calculation of the
theoretical frequencies, then the degree of freedom of 2 is n-r-1, (n is the total
Problem 1: When the first proof of 392 pages of a book of 1200 pages were read, the distribution
of printing mistakes were found to be as follows:
Solution:
Let X denote the number of printing mistakes per page. To fit a Poisson distribution
for the given data, first to identify the parameter . For a Poisson random variable X,
E(X) = . Hence, equate the sample mean to the population mean E(X), to get an
estimate of the parameter .
1
For the given data, Mean, x
N
fx
i
i i
fx
i
i i 0 275 1 72 .... 6 1 189 , N = 392
189
x 0.482 .
392
e 0.482 0.482 x
Hence, P ( X x) ; x 0,1, 2.....
x!
e 0.482 0.4823 28
0.0115
3!
3 7
e 0.482 0.4824
0.00138 5
4!
4 5
e 0.482 0.4825
0.00013
5! 0
5 2 e 0.482 0.4826
0.00001
6!
0
6 1
N = 500 500
Now,
2 30 28 4 0.143
3 7 5
5 0
4 17 5 144 28.8
3 0
5 2 0
392 50.753
Oi Ei
2
square distribution with (4-1-1) =2 degrees of freedom. (4 classes are considered after
pooling and 1 d.f. is lost due to the estimation of the parameter for calculating theoretical
frequencies)
From chi-square table for 2 d.f., and for significance level 0.05 , 2 5.99 .
Here, 2 2 , that is the data is not matching with the hypothesis considered.
Problem 2: A survey of 800 families with four children each revealed the following distribution:
Number of Boys: 0 1 2 3 4
Number of Girls: 4 3 2 1 0
Solution:
Here we have to check whether the given data support that male and female births
are equally probable wit probability of male p 0.5 .
Expected number of families with x male children out of 800 families with 4 Childs,
and with probability of male p 0.5 , can be found by 800 P ( X x ) , where P ( X x ) can
0 32 50 324 6.48
4 64 50 196 3.92
Oi Ei
2
From chi-square table for 4 d.f., and for significance level 0.05 , 2 9.49 .
Here, 2 2 , that is the data is not matching with the hypothesis considered.
Hence we reject the hypothesis that male and female births are equally probable.
Problem 3: A sample analysis of examination results of 200 MBA’s was made. It was found that
46 students had failed, 68 secured a third division, 62 secured a second division and the rest were
placed in first division. Are these figures commensurate with the general examination result which
is in the ratio 4 : 3 : 2 : 1 for various categories respectively?
Solution:
Let us consider the null hypothesis as the data commensurate with the general
examination result. Under this hypothesis our expected number of students in each
category can be calculated as follows:
3
Expected number of students with third division out of 200 = 200 60
4 3 2 1
2
Expected number of students with second division out of 200 = 200 40
4 3 2 1
1
Expected number of students with first division out of 200 = 200 20
4 3 2 1
Oi Ei
Ei
Failed 46 80 14.45
200 28.417
Oi Ei
2
f i.
P (An observation to come in Ai th class) = ( where f.. f ij )
f.. i j
f. j
P (An observation to come in B j th class) =
f..
f i . f. j
then, the probability of an observation to come in Ai th and B j th class is .
f.. f..
f i . f. j f i . f. j
= f.. =
f.. f.. f..
Oi Ei
2
n
2
i 1 Ei
f..
close to zero. If 2 > 2 , we reject the hypothesis that the characteristics are
independent. ( 2 is the table value of chi-square distribution for (l 1)(m 1) d.f. such
Problem 1: For a 2 2 contingency table, where the frequencies are a, b, c and d, as given by,
B1 B2
A1 a b ab ( a b c d )( ad bc )2
; N=a + b + c +d. Show that, 2
A2 c d cd ( a b )(c d )(b d )( a c )
ac bd N
Solution:
Oi Ei
2
n
2
i 1 Ei
. The expected frequencies of each cell is calculated as,
( a b )( a c )
Expected frequency of (1,1)th cell =
N
( a b )(b d )
Expected frequency of (1,2)th cell =
N
(c d )( a c )
Expected frequency of (2,1)th cell =
N
(c d )(b d )
Expected frequency of (1,1)th cell =
N
2 2 2 2
( a b )( a c ) ( a b )(b d ) (c d )( a c ) (c d )(b d )
a b c d
N N N N
2
( a b )( a c ) ( a b )(b d ) (c d )( a c ) (c d )(b d )
N N N N
1 ad bc 2
N ( a b )( a c )
ad bc ;
2
1
Similarly, the other terms are becomes,
N ( a b )(b d )
ad bc ad bc .
2 2
1 1
and
N (c d )( a c ) N (c d )(b d )
Hence,
1 ad bc ad bc ad bc ad bc
2 2 2 2
2
N (a b)(a c) (a b)(b d ) (c d )(a c) (c d )(b d )
ad bc
2
1 1 1 1
N (a b)(a c ) (a b)(b d ) (c d )(a c ) (c d )(b d )
ad bc (b d ) (a c) (a c) (b d )
2
(c d )(a c )(b d )
N ( a b )( a c )(b d )
ad bc
2
N N
N (a b)(a c)(b d ) (c d )(a c )(b d )
2 (c d ) ( a b )
ad bc
(a b)(a c)(b d )(c d )
(a b c d )(ad bc) 2
2
(a b)(c d )(b d )(a c)
Problem 1: Two sample polls of votes for two candidates A and B for a public office are taken, one
from among the residents of rural areas and the other from urban. The results are given in the
adjoining table. Examine whether the nature of the area is related to voting preference in this
election.
A B
Solution:
Considering ‘the nature of the area and voting preference are independent’ as the
null hypothesis. Under this hypothesis we get the expected frequency for each cell as
follows.
1000 1170
Expected frequency for the (1,1) cell = 585
2000
1000 830
Expected frequency for the (1,2) cell = 415
2000
1170 1000
Expected frequency for the (2,1) cell = 585
2000
1000 830
Expected frequency for the (2,2) cell = 415
2000
Oi Ei
2
585 415 585 415
= 10.0891.
voting preference are independent. That is the nature of area is related with
Problem 2: The following contingency table gives the classification of 1000 workers in a factory
according to the disciplinary action taken by the management and their promotional experience:
action
Promoted Not promoted
Test whether promotional experience and disciplinary action taken are associated
or not.
Solution:
700 100
Expected frequency for the (1,1) cell = 70
1000
900 700
Expected frequency for the (1,2) cell = 630
1000
100 300
Expected frequency for the (2,1) cell = 30
1000
900 300
Expected frequency for the (2,2) cell = 270
1000
Oi Ei
2
70 630 30 270
= 84.66.
Problem 3: A group of 200 boys and 100 girls are selected for an IQ test and they are classified as
given below. Examine whether there is any dependency between the intelligence levels and the
gender
boys 86 60 44 10 200
girls 40 33 25 2 100
126 93 69 12 300
Solution:
Here to test whether the gender (boys and girls) and the IQ levels are independent.
Under the hypothesis the two characteristics given are independent, find the
expected frequency in each cell. Then perform the chi-square test.
200 126
expected frequency in (1,1) cell= = 84
300
200 93
Expected frequency in (1, 2)th cell = = 62
300
200 69
Expected frequency in (1,3) rd cell = = 46
300
boys 84 62 46 8 200
girls 42 31 23 4 100
126 93 69 12 300
Oi Ei
2
n
2
i 1 Ei
Gender IQcategory Oi Ei Oi Ei Oi Ei
2
Oi Ei
2
Ei
Here 2 < 2 . Hence the hypothesis that the two characteristics are independent
Problem 4: A marketing agency gives you the following information about the age groups of the
sample informants and their liking for a particular model of scooter which a company plans to
introduce:
Total
Below 20 20 – 39 40 – 59
On the basis of the above data, can it be concluded that the model appeal is independent of
the age group of the informants?
Solution:
Considering ‘the model appeal is independent of the age group’ as the null
hypothesis, the expected frequency in each cell is calculated as follows.
605 200
Expected frequency for the (1,1) cell = 121
1000
605 640
Expected frequency for the (1,2) cell = 387.2
1000
605 160
Expected frequency for the (1,3) cell = 96.8
1000
395 200
Expected frequency for the (2,1) cell = 79
1000
STATISTICAL INFERENCE Page 122
School of Distance Education
395 640
Expected frequency for the (2,2) cell = 252.8
1000
395 160
Expected frequency for the (2,3) cell = 63.20
1000
Ei
Oi Ei
60 96.8 13.99
75 79 0.202
42.764
Oi Ei
2
EXERCISES
1. How do you determine the critical region for testing the mean of a population in
large sample case?
2. What is large sample test? How can the equality of two population proportions be
tested?
3. Explain the principle of a test of goodness of fit.
4. What are the conditions for using chi-square test for testing agreement between
theoretical frequencies and observed frequencies?
5. Derive test statistic and test procedure to test the independence of two attributes in
a 2 2 contingency table.
6. An examination was given to 50 students at college A and to 60 students at college
B. At A, the mean grade was 75 with standard deviation of 9 and at B the mean
grade was 79 with standard deviation 7. Is there significant difference between the
performance of the students at A and those at B at 5%
level of significance?
7. The manufacturer of television tubes knows from past experience that the
average life of a tube is 2,000 hours with a standard deviation of 200 hours. A
sample of 100 tubes has an average life of 1,950 hours. Test at the 0.05 level of
significance, if this sample came from a normal population of mean 2,000 hours.
8. A sample of heights of 6,400 Englishmen has a mean of 67.85 inches and S.D. 2.56
inches, while a sample of heights of 1,600 Australians has a mean of 68.55 inches
and S.D. of 2.52 inches. Do the data indicate that Australians are, on the average,
taller than Englishmen?
9. The mean weekly sale of soap bars in departmental stores was 146.3 bars store.
After an advertising campaign the mean sales in 22 stores for a typical week
increased to 153.7 and showed a standard deviation 17.2. Was the advertising
campaign successful?
10. In a locality 110 persons were randomly selected and asked about their educational
achievement. The results are given as follows:
Sex below SSLC SSLC Above SSLC
Male 15 20 25
Female 25 15 10
Can you conclude in light of this sample, education depends on sex?
**********************
CHAPTER 6
When the number of sample is large, by central limit theorem, almost all test
statistics follows normal distribution. Then for testing the hypothesis, critical region can
be obtained with the help of standard normal table. But when the sample is small, it is to
use test statistics with known probability distributions to perform the testing of
hypothesis.
known standard deviation . Let a sample of size n is taken from the normal population.
we reject H 0 , if t
x 0 n
t , where t is from standard normal table such
that P(t t )
we reject H 0 , if t
x 0 n
t , where t is from standard normal table such
that P(t t )
we reject H 0 , if t
x 0 n
t , where t is from standard normal table
2 2
such that P( t t )
2
(ii) To test the equality means of two normal population with known standard
deviations
Let samples of sizes n1 and n2 are taken from two normal populations N ( 1 , 1 )
x1 x2
The test statistic used is, t , following N(0,1), under H 0 .
12 2 2
n1 n2
that, P( t t )
such that, P ( t t )
such that, P( t t ) .
2
Problem 1: A sample of size 10 taken from a normal population. The sample mean is recorded as
47. Another sample of size 15 gives its mean as 41. Can the samples be regarded as drawn from
the same population of standard deviation 4 . 0.05 .
Solution:
Two samples of sizes 10 and 15 were drawn from a normal population. The sample
means are 47 and 41 respectively. Here to test whether the samples are from normal
Hence H 0 : 1 2 against H1 : 1 2
x1 x2 47 41
The test statistic used is t
12 2 2 42 42
n1 n2 10 15
= 3.6742.
For, 0.05 , from std normal table t 1.96 . Hence, here t t . Then reject
2
2
H 0 . That is, the two samples cannot be regarded as coming from same population.
Problem 2: A sample of size 10 of men and another sample of size 12 of women have mean IQ’s
101 and 98 respectively. Assuming that the IQ’s of men and women are independently
and normally distributed with mean 1 and 2 and S.D’s 4 and 3. Examine whether men are on
Solution:
x1 x2 101 98
The test statistic used is t = 1.96.
12 2 2 42 32
n1 n2 10 12
Here t t . Hence reject H 0 . That is accepting that men are on the average
standard deviation is unknown and since the sample size is small, the test based on
normal distribution is not applicable.
Under the assumptions: (i) the population from the sample am taken is normal
the statistic , t
x1 0 n
which follows t-distribution with (n-1)
s
degrees of freedom is considered as the test statistic.
(ii) To test the equality means of two normal population with known standard
deviations: (when population standard deviations 1 and 2 are unknown)
Let samples of sizes n1 and n2 are taken from two normal populations N ( 1 , 1 )
1 1
n1 n2 2 n1 n2
n1 n2 2 ) degrees of freedom.
If to test the equality of mean effect of two different treatments on a population, let
n different pairs of units of the population in such a way that each pair should contain
units of the population which are homogenous in nature is selected. Apply treatments to
each of the pair such that, the first treatment (X) is to the first unit of the pair and the
second treatment (Y) to the second unit.
Collect the data u1 , u2 , ...., un for each pair such that, ui xi yi , where xi is the value
of the effect of treatment (X) on the first unit of the ith pair and yi is the value of the
and 2 is te mean effect of second . If the mean effect of two treatments is same, we
t
u 0 n 1
, which follows t ( n1)
Su
Solution:
Let X denotes the height of the males and assume it follow normal distribution.
Now to test the hypothesis regarding the mean of the population. That is to test
H 0 : 62 against H1 : 62 .
Sample mean x of the 10 samples given can be calculated. Population S.D. is
unknown.
x 70 66 59 68 62 63 61 60 59 58
xi x 7.4 3.4 -3.6 5.4 -0.6 0.4 -1.6 -2.6 -3.6 -4.6
xi x
2
54.76 11.56 12.96 29.16 0.36 0.16 2.56 6.76 12.96 21.16
x 0 n 1 62.2 62 10 1
The test statistic is, t 0.154 .
S 3.9
Problem 2: The The heights of six randomly chosen sailors are in inches : 63,65,68,69,71, and 72.
Those of 10 randomly chosen soldiers are 61,62,65,66,69,69,70,71,72, and 73. Test whether the
data support the claim that the soldiers are on the average taller than sailors.
Solution:
sailors.
Let 1 denote the average height of soldiers and 2 be that of sailors. Now to test
H 0 : 1 2 against H1 : 1 2 .
x1 x2
The test statistics is t
n S n2 S2 2 1 1
2
1 1
n1 n2 2 n1 n2
408
For the group of soldiers x i
i 408 x1
6
68
x 63 65 68 69 71 72
xi x -5 -3 0 1 3 4
xi x 25 9 0 1 9 16 60
2
x x
2
n1S12 i 1 60
i
678
Again, For the group of sailors, x i
i 678 x2
10
67.8
x 61 62 65 66 69 69 70 71 72 73
xi x -6.8 -5.8 -2.8 -1.8 1.2 1.2 2.2 3.2 4.2 5.2
xi x
2 46.24,33.64,7.84,3.24,1.44,1.44,4.84,10.24,17.64,27.04 153.6
x x
2
n2 S 2 2 i 2 153.6
i
68 67.8
t 0.0991
60 153.6 1 1
6 10 2 6 10
Then it can be observed that , t t . Hence we reject H 0 . That is the soldiers are
Problem 3: In a certain experiment to compare two types of animal foods A and B, the following
results of increase in weights are observed in anumals. The same sets of eight animals were used in
both the foods.
In weight 52 55 52 53 50 54 54 53 423
Food B yi
The above problem is to compare average gain in weight with two foods.
Consider 1 and 2 are the average gain in weight by food A and food B. Let ui xi yi .
paired t-test.
animals 1 2 3 4 5 6 7 8 total
xi 49 53 51 52 47 50 52 53
yi 52 55 52 53 50 54 54 53
ui xi yi -3 -2 -1 -1 -3 -4 -2 0 -16
ui 2 xi yi
2 9 4 1 1 9 16 4 0 44
1 16
u
n i
ui =
8
2
1 1
ui u
2
= (44) (2) 2
2
Su
n i 8
= 1.23
2 0 8 1
t = - 4.30
1.23
nS 2
2 ,
02
P ( 2( n 1) 2 ) .
Problem: It is believed that the weight of one of the product of a company is with variance greater
than 0.16 gms. A Sample of eleven items is taken. Their weights (in gms.) are measured as
follows: 2.5, 2.3, 2.4, 2.3, 2.5, 2.7, 2.5, 2.6, 2.6, 2.7, and 2.5. Test the hypothesis
at 1% level of significance.
Solution:
nS 2
The test statistic is , 2 , following chi-square distribution with )n-1)
02
d.f
1 27.6
Here, x
i
i = 27.6 x
n i
xi
11
2.51
x 2.5 2.3 2.4 2.3 2.5 2.7 2.5 2.6 2.6 2.7 2.5
xi x -0.01 -0.21 -0.11 -0.21 -0.01 0.19 -0.01 0.09 0.09 0.19 -0.01
xi x
2
0.0001 0.0441 0.0121 0.0441 0.0001 0.0361 0.0001 0.0081 0.0081 0.0361 0.0001
1 0.1891
xi x
2
S2 0.0172
n i 11
nS 2 11 0.0172
2 = = 1.1825
02 0.16
of two normal populations. Let S1 and S 2 be the standard deviations of the sample of
The test statistics suggest is the ratio of the unbiased estimators of population
variances, that is;
n1S12 nS2
If > 2 2 ,
n1 1 n2 1
n1S12
n 1 n n 1 S12 ,
F 1 2 1 2
n2 S 2 n2 n1 1 S 2 2
n2 1
hypothesis.
of freedom, such that P ( F n 1, n F ) .
1
2 1 2
2
n2 S 2 2 nS2
If < 1 1 ,
n2 1 n1 1
n2 S 2 2
n 1 n n 1 S2 2 , which follows F-distribution with
Then, consider F = 2 2 2 1
n1S1 n1 n2 1 S12
n1 1
n2 1, n1 1 degrees of freedom.
n2 1, n1 1 degrees of freedom, such that P ( F n
F ) .
2 1, n1 1 2
2
1 10 15 9
2 12 14 9
Test whether the population variances are same at 10% level of significance.
Let the two normal populations are N ( 1 , 1 ) and N ( 2 , 2 ) . If the two samples
H 0 : 1 2 .
To perform the test on H 0 : 1 2 , we can use t-test with the small samples taken.
But to perform the t-test, the basic assumption is 1 2 . Hence first we have to
test H 0 : 1 2 against H1 : 1 2 .
n1S12 10 9
= 10 , and
n1 1 10 1
n2 S 2 2 12 9
= 9.82
n2 1 12 1
n1S12
n 1 n n 1 S12 ,
Hence the test statistic is F 1 2 1 2 follow F-distribution with
n2 S 2 n2 n1 1 S 2 2
n2 1
n1 1, n2 1 d.f.
10
Here calculated value of F = 1.018
9.82
For 10% of significance level, at (10-1,12-1)= (9,11) d.f. From F-table, F 2.90 .
2
x1 x2
t .
n1S12 n2 S2 2 1 1
n1 n2 2 n1 n2
Hence we accept H 0 : 1 2 . Therefore it can be concluded that the two samples are
EXERCISES
2. Stating the assumptions, explain t-test for testing mean of a normal population
3. Stating your assumptions, explain t-test to test the equality of means of two
independent normal populations.
7. Describe the procedure for testing the equality of variances of two normal
populations.
8. Two different diets ‘A’ and ‘B’ were administered to two different groups of pigs.
The gains in weight by the diets are given below
Diet ‘A’ : 25, 32, 30, 34, 24, 14, 32, 24, 30, 31, 35, 25
Diet ‘B’ : 44, 34, 22, 10, 47, 31, 40, 30, 32, 35, 18, 21, 35 29 22
Test whether the diets differ significantly as far as their effects on increasing weight
is concerned.
Boys 70 20 250
Girls 75 15 150
Test whether there is significant difference between the average scores of boys
and girls. Obtain 5% confidence interval for the difference in average scores.
10. Two independent groups of 10 children were tested to find how many digits they
could repeat from memory after hearing them. The results are as follows:
Group A: 8 6 5 7 6 8 7 4 5 6
Group B: 10 6 7 8 6 9 7 6 7 7
Is the difference between the mean scores of the two groups significant?
11. Two random samples of size 8 and 11 drawn from two normal populations are
characterized as follows:
11 16.5 73.26
Examine whether the two samples came from populations having same
12. The nicotine content (in mg.) of two samples of tobacco were found to be as
follows:
Sample A: 24 27 26 21 25
Sample B: 27 30 28 31 22 36
Can it be said that the two samples come from the same normal population?
****************
SYLLABUS
Course-III: Statistical Inference
*********