0% found this document useful (0 votes)
10 views13 pages

Normal Distribution Tests for Proportions

The document discusses statistical methods for hypothesis testing based on normal distribution, particularly focusing on tests for single proportions and differences between two proportions. It provides formulas for calculating test statistics, conditions for rejecting null hypotheses, and examples illustrating the application of these methods. Additionally, it covers the significance levels and confidence intervals for proportions in various scenarios.

Uploaded by

safiyawww
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
10 views13 pages

Normal Distribution Tests for Proportions

The document discusses statistical methods for hypothesis testing based on normal distribution, particularly focusing on tests for single proportions and differences between two proportions. It provides formulas for calculating test statistics, conditions for rejecting null hypotheses, and examples illustrating the application of these methods. Additionally, it covers the significance levels and confidence intervals for proportions in various scenarios.

Uploaded by

safiyawww
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

TEST BASED ON NORMAL DISTRIBUTION (FOR LARGE SAMPLE)

Sampling of Attributes: The presence of an attribute in samples may be treated as success


and its absence as failure. In this case, a sample of n observation is identified with that of a
series of n independent Bernoullian trials with constant probability P of success for each
trail. Then the probability of x (success) in n trails, is given by
 n
Pr ( X = x) =   P x (1 − P) n− x , x = 0,1, 2, L , n .
 x
Test of single proportion
If X is the member of successes in n independent trails with constant probability P of
success for each trail, then
E ( X ) = nP , and V ( X ) = nPQ , where P + Q = 1 .
For large n , we know that binomial distribution tends to normal distribution, hence for large
n,
X ~ N (nP, nPQ )
For testing
H 0 : P = P0 against H A : P ≠ P0 (< P0 or > P0 )
The test statistic is given by
X − E ( X ) X − nP
Z= = ~ N (0, 1)
SE ( X ) nPQ
For H A : P > P0 , H 0 is to be rejected if for the given sample
Z > Z α , where Z α is table value from standard normal distribution at α % level of
significance.
For H A : P < P0 , H 0 is to be rejected if for the given sample
Z < −Zα
For H A : P ≠ P0 , H 0 is to be rejected if for the given sample

Z > Zα / 2 .

Remark:
i) In a sample of size n , X be the number of persons possessing the given attribute, then,
X
observed proportion of successes = = p , therefore,
n
X 1 1 X 1 PQ
E ( p) = E   = E ( X ) = (nP) = P , and V ( p) = V   = V (X ) = .
n n n  n  n2 n
Under H 0 , the test statistic reduces to
p − E ( p) p − P
Z= = ~ N (0, 1) .
SE ( p ) PQ
n
86 RU Khan

 N − n  PQ
ii) If we have a sample from a finite population of size n , then V ( p ) =   ,
 N −1  n
NPQ
S2 = .
N −1
pq
iii) The limit for P at α % level of significance are p ± Z α / 2 .
n
Proof: By definition,
 p−P 
Pr [ Z ≤ Z α / 2 ] = 1 − α or Pr − Z α / 2 ≤ ≤ Zα / 2  = 1 − α
 SE ( p ) 
or Pr [− Z α / 2 SE ( p) ≤ p − P ≤ Z α / 2 SE ( p )] = 1 − α
or Pr [− p − Z α / 2 SE ( p ) ≤ − P ≤ − p + Z α / 2 SE ( p )] = 1 − α
or Pr [ p − Z α / 2 SE ( p) ≤ P ≤ p + Z α / 2 SE ( p )] = 1 − α

Pˆ Qˆ pq
or p ± Z α / 2 or p ± Z α / 2 .
n n
Exercise 1) In a sample of 1000 people in a particular state 540 are rice eaters and the rest
are wheat eaters. Can we assume that both rice and wheat are equally popular in the state at
5% level of significance?
Solution: We are given n = 1000 , if X = Number of rice eaters = 540 , therefore, sample
540
proportion of rice eater ( p ) = = 0.54 . We want to test
1000
H 0 : P = 0.5 (both rice and wheat eater are equally popular), against H A : P ≠ 0.5 .
The test statistic
p−P
Z = ~ N (0, 1) .
PQ
n
Under H 0 , it is reduces to
0.54 − 0.5
Z = = 2.532 .
0. 5 × 0. 5
1000
Since Z = 2.532 > Z α / 2 = Z 0.025 = 1.96 at 5% level of significance, we reject H 0 , i.e. rice
and wheat eaters are not equally popular.
Exercise 2) A dice is thrown 9000 times and a throw of 3 or 4 observed 3240 times. Test
whether dice is biased or unbiased; also find the limits between the probability of a throw of
3 or 4 lies.
Solution: We are given n = 9000 , X = 3240 , therefore, sample proportion of a throw of 3
3240
or 4 is, p = = 0.36 .
9000
Statistical Inference 87

We want to test
1 1
H0 : P = ( probability of getting 3 or 4), against H A : P ≠ .
3 3
Under H 0 , the test statistic
p−P 0.36 − 0.33333
Z = = = 5.367 .
PQ 0.33333 × 0.6667
n 9000
Since Z = 5.367 > 3 , H 0 is to be rejected and conclude that the dice is biased.
Limit for P is given by
pq 0.36 × 0.64
p ± Zα / 2 or 0.36 ± 3 = 0.36 ± 0.015 . Thus (0.345 and 0.375)
n 9000
are the limits of the probability of getting 3 or 4.
Alternative solution
Given n = 9000 , and X = 3240 . We want to test
1 1
H0 : P = ( probability of getting 3 or 4), against H A : P ≠ .
3 3
Under H 0 , the test statistic
X − nP 3240 − 3000
Z = = = 5.367 .
nPQ 1 2
9000 × ×
3 3
Since Z = 5.367 > 3 , H 0 is to be rejected and conclude that the dice is biased.
Exercise 3) Fifty people were attacked by a disease and only 45 survived. Will you reject
the hypothesis that the survival rate of attacked by this disease, is 85% in favour of the
hypothesis that it is more at 5% level of significance.
Solution:
H 0 : P = 0.85 against H A : P > 0.85 .
Under H 0 , the test statistic
p−P 45
Z= , where p = = 0. 9
PQ 50
n
0.9 − 0.85
Z= = 0.05 .
0.85 × 0.15
50
Table value Z α = 1.645 at α = 0.05 .
Since Z = 0.05 < 1.645 at α = 0.05 , H 0 is accepted and conclude that the survival rate of
attacked by this disease, is 85%.
88 RU Khan

Exercise 4) An oil company claims that at least 20% of all automobile owners buy brand
A gasoline. Test this claim at α = 0.05 , if a random check indicates that 58 of 200
automobile owners buy brand A gasoline.
Solution:
H 0 : P = 0.20 against H A : P > 0.20 .
Under H 0 , the test statistic
p−P 58
Z= , where p = = 0.29
PQ 200
n
0.29 − 0.20
Z= = 3.18 .
0.20 × 0.80
200
Table value Z α = 1.645 at α = 0.05 .
Since Z = 3.18 > 1.645 at α = 0.05 , H 0 is rejected. We conclude that brand A gasoline is
brought by more than 20% of all automobile owners.
Test for difference of two proportions
Suppose we want to compare two distinct populations with respect to the prevalence of a
certain attribute, A (say) among their members. Let X 1 and X 2 be the number of persons
possessing the given attribute A in random samples of sizes n1 and n2 from the two
populations respectively, then sample proportions are given by
X1 X
p1 = , and p 2 = 2
n1 n2
Let P1 and P2 be the populations proportions, then

X  1 1
E ( p1 ) = E  1  = E ( X 1 ) = n1P1 = P1
 n1  n1 n1

and
X  1 1 PQ
V ( p1 ) = V  1  = V ( X1 ) = n1P1Q1 = 1 1.
 n1  n12 n12 n1

Similarly,
PQ
E ( p 2 ) = P2 and V ( p 2 ) = 2 2 .
n2
Since for large samples, p1 and p 2 are normally distributed, then p1 − p 2 is also normally
distributed.
Thus, the standard variable corresponding to ( p1 − p2 ) is given by
( p1 − p 2 ) − E ( p1 − p 2 )
Z= ~ N (0,1)
SE ( p1 − p 2 )
Statistical Inference 89

Under H 0 : P1 = P2 = P .
p1 − p 2
Z= .
1 1 
PQ  + 
 n1 n2 
The hypothesis of interest is
H 0 : P1 = P2 against
H A : i) P1 > P2
ii) P1 < P2
iii) P1 ≠ P2
i) If Z > Zα , H 0 is to be rejected at α % level of significance.
ii) If Z < − Z α , H 0 is to be rejected at α % level of significance.
iii) If | Z |> Zα / 2 , H 0 is to be rejected at α % level of significance.
Remark: In general, we do not have any information as to the proportion of attributes in the
populations from which the samples have been drawn. Under H 0 : P1 = P2 = P , an unbiased
estimate of population proportion P based on both the samples is given by
n p + n2 p 2 X 1 + X 2
Pˆ = 1 1 = .
n1 + n2 n1 + n 2
Exercise 5) Random samples of 400 men and 600 women were asked whether they
would like to have a flyover near their residence. 200 men and 325 women were infavour of
the proposal. Test the hypothesis that proportions of men and women in favour of the
proposal are same at 5% level of significance.
Solution: We are given
n1 = 400 , X 1 = 200 , men in favour of proposal
n2 = 600 , X 2 = 325 , women in favour of proposal.
Thus,
X1 X
p1 = = 0.50 , p 2 = 2 = 0.541
n1 n2
We want to test
H 0 : P1 = P2 = P , against H A : P1 ≠ P2 .
Under H 0 , the test statistic is
p1 − p 2
Z = ,
1 1 
Pˆ Qˆ  + 
 n1 n2 
where
n p + n2 p2 X 1 + X 2
Pˆ = 1 1 = = 0.525 and Qˆ = 1 − Pˆ = 0.475 .
n1 + n2 n1 + n2
90 RU Khan

Hence,
0.5 − 0.541
Z = = 1.209
 1 1 
0.525 × 0.475  + 
 400 600 
Since Z = 1.269 < 1.96 , H 0 is accepted at 5% level.
Exercise 6) Before an increase in excise duty on tea, 800 persons out of a sample of 1000
persons were found to be tea drinkers. After an increase in duty, 800 people were tea
drinkers in a sample of 1200 . Using SE of proportion, state whether there is a significant
decrease in the consumption of tea after the increase in excise duty?
Solution: We are given
800
n1 = 1000 , p1 = = 0.80 , sample proportion of tea drinkers before increase in excise
1000
duty.
800
n2 = 1200 , p 2 = = 0.67 , sample proportion of tea drinkers after increase in excise
1200
duty.
We want to test
H 0 : P1 = P2 = P or P1 − P2 = 0 , against H A : P1 > P2 .
Under H 0 , the test statistic is
p1 − p 2
Z= ,
1 1 
Pˆ Qˆ  + 
n
 1 n 2

where
n p + n2 p 2 16 6
Pˆ = 1 1 = and Qˆ = 1 − Pˆ = .
n1 + n2 22 22
Hence,
0.80 − 0.67
Z= = 6.842 .
16 6  1 1 
×  + 
22 22  1000 1200 
Since Z = 6.842 > Zα = 1.64 at α = 0.05 level, H 0 is rejected. i.e., there is a significant
decrease in the consumption of tea after increase in the excise duty.
Exercise 7) A cigarette manufacturing firm claims that its brand A of cigarette outsells its
brand B by 8% . If it is found that 42 out of a sample of 200 smokers prefer brand A and
18 out of another sample of 100 smokers prefer brand B , test whether that 8% difference is
a valid claim at 5% level of significance.
Solution: We are given
n1 = 200 , X 1 = 42 ⇒ p1 = 0.21 ,
n2 = 100 , X 2 = 18 ⇒ p 2 = 0.18 ,
Statistical Inference 91

The hypothesis of interest is


H 0 : P1 − P2 = 0.08 , against H A : P1 − P2 ≠ 0.08 .
Under H 0 , the test statistic is
( p1 − p 2 ) − ( P1 − P2 )
Z = ,
Pˆ1Qˆ1 Pˆ2 Qˆ 2
+
n1 n2
where
Pˆ2 = 0.18 ⇒ Qˆ1 = 1 − Pˆ1 = 0.79
and
Pˆ1 = 0.21 ⇒ Qˆ 2 = 1 − Pˆ2 = 0.82 .
Thus,
(0.21 − 0.18) − 0.08
Z = = 1.02
0.21 × 0.79 0.18 × 0.82
+
200 100
Since Z = 1.02 < Zα / 2 = 1.96 , H 0 is accepted at α = 0.05 level. i.e., 8% difference is a
valid claim.
Exercise 8) In a year there are 956 births in a town A , of which 52.5% were males, while
in towns A and B combined, this proportion in a total of 1406 births was 0.496 . Is there
any significant difference in the proportion of male births in the two towns?
Solution: We are given
n1 = 956 , p1 = 0.525 ,
n2 = n − n1 = 450 ,
n p + n2 p 2
p= 1 1 = 0.496 ⇒ p 2 = 0.434
n1 + n2
We want to test
H 0 : P1 = P2 against H A : P1 ≠ P2 .
Under H 0 , the test statistic is
p1 − p 2 0.525 − 0.434
Z= = = 3.368 .
1 1   1 1 
Pˆ Qˆ  +  0.496 × 0.544  + 
 n1 n2   956 450 

Since | Z |= 3.368 > Zα / 2 = 1.96 at α = 0.05 level, H 0 is rejected. i.e., there is a significant
difference in the proportion of male births in the towns A and B .
Test for single mean
Let x1 ,L, xn be random sample of size n from a normal population with mean µ and
variance σ 2 i.e.,
92 RU Khan

X ~ N ( µ ,σ 2 ) , then

x ~ N ( µ , σ 2 / n) .
Thus the standard normal variate corresponding to x is given by
x −µ
Z= ~ N (0,1) .
σ/ n
Suppose we want to test
H0 : µ = µ 0 (specified), against
H A : i) µ > µ 0
ii) µ < µ 0
iii) µ ≠ µ 0
Under H 0 , the test statistic is
x − µ0
Z= .
σ/ n
If i) Z > Zα , H 0 is to be rejected.
ii) Z < − Z α , reject H 0 .
iii) | Z | > Zα / 2 , reject H 0 .

Remark: If the population variance σ 2 is unknown, then we use its estimate, i.e., σˆ 2 = s 2 .
Confidence limits for µ : 100 (1 − α ) confidence limits for µ is given by
Pr [| Z | ≤ Zα / 2 ] = 1 − α ,
where Z α / 2 be obtained from standard normal table for given α . Thus,

 x−µ 
Pr  ≤ Zα / 2  = 1 − α
 σ/ n 
 x−µ 
⇒ Pr − Zα / 2 ≤ ≤ Zα / 2  = 1 − α
 σ/ n 
 σ σ 
⇒ Pr  x − Zα / 2 ≤ µ ≤ x + Zα / 2  = 1−α .
 n n
Hence the confidence limits for µ is
σ
x ± Zα / 2 .
n
Exercise 9) Suppose that 100 tyres of a certain brand on the average 21431 kms with a
standard deviation 1295 kms, using α = 0.05 , test
H 0 : µ = 22000 kms vs H A : µ < 22000 .
Statistical Inference 93

Solution: Given x = 21431 , σ = 1295 , n = 100 and


H 0 : µ = 22000 kms vs H A : µ < 22000 .
Under H 0 , the test statistic
x − µ0 21431 − 22000
Z= = = −4.39 .
σ/ n 1295 / 10
Since Z = −4.39 < −1.645 at α = 0.05 , H 0 is to be rejected at 5% level of significance, i.e.,
we conclude that the tyres are as good as claimed.
Exercise 10) The mean breaking strength of cables supplied by a manufacturer is 1800 with
a standard deviation 100 . By a new technique in the manufacturing process, it is claimed that
the breaking strength of the cables have increased. In order to test this claim, a sample of 50
cables is tested. It is found that the mean breaking strength is 1850 . Can we support the claim
at 1% level of significance?
Solution: Given x = 1850 , σ = 100 , µ = 1800 , n = 50 and the hypothesis of interest is
H 0 : µ = 1800 kms vs H A : µ > 1800 .
Under H 0 , the test statistic
x − µ0
Z= ~ N (0,1)
σ/ n
1850 − 1800
= = 3.535 .
100 / 50
Since Z = 3.535 > Zα = 2.33 at α = 0.01 , H 0 is rejected and conclude that breaking
strength of the cable has increased.
Exercise 11) From a large lot of fresh coins, a random sample of size 50 is taken. The
mean weight of coins in the sample is found to be 28.57 gm. Assuming that the population
standard deviation of weight is 1.25 gm. Will it be reasonable to suppose that the mean is 28
gm? If not, obtain the 99% confidence limits to the mean weight of all coins in the lot.
Solution: Given x = 28.57 , σ = 1.25 , n = 50 and the hypothesis of interest is
H 0 : µ = 28 gm vs H A : µ ≠ 28 gm.
Under H 0 , the test statistic
x − µ0
Z= ~ N (0,1)
σ/ n
28.57 − 28
= = 3.224 .
1.25 / 50
Since Z = 3.224 > Zα / 2 = 1.96 at α = 0.05 , H 0 is to be rejected, i.e., we conclude that
µ ≠ 28 gm.
99% confidence limit of µ are
σ 1.25
x ± Zα / 2 or 28.57 ± 2.58 or (28.11 and 29.03) .
n 50
94 RU Khan

Test for difference of two means

Let x1 be the mean of a sample of size n1 from a population with mean µ1 and variance σ 12
and x 2 be the mean of a sample of size n2 from a population with mean µ 2 and variance
σ 22 . Since sample sizes are large, then

x1 ~ N ( µ1 , σ 12 / n1 ) and x 2 ~ N (µ 2 , σ 22 / n2 ) , so that

 σ2 σ2
x1 − x2 ~ N  µ1 − µ 2 , 1 + 2  ,
 n1 n2 

where
E ( x1 − x 2 ) = E ( x1 ) − E ( x 2 ) = µ1 − µ 2
and
V ( x1 − x 2 ) = V ( x1 ) + V ( x2 ) − 2 Cov ( x1 , x 2 )

σ 12 σ 22
= + − 0.
n1 n2
Thus,
( x1 − x 2 ) − ( µ1 − µ 2 )
Z= .
σ 12 σ 22
+
n1 n2
The hypothesis of interest is
H 0 : µ1 = µ 2 , against
H A : i) µ1 > µ 2
ii) µ1 < µ 2
iii) µ1 ≠ µ 2
Under H 0 , the test statistic is
x1 − x2
Z= .
σ 12 σ2
+ 2
n1 n2

Under i) µ1 > µ 2 , H 0 is rejected if Z > Zα on the basis of given sample values.


Under ii) µ1 < µ 2 , H 0 is rejected if Z < − Z α .
Under iii) | Z | > Zα / 2 , reject H 0 .
Remark:

i) If σ 12 = σ 22 = σ 2 i.e., the samples have been drawn from the populations with common
variance, then under H 0 the test statistic becomes
Statistical Inference 95

x1 − x 2
Z= .
1 1
σ +
n1 n 2

ii) If σ 2 is unknown, then its estimate is used, and is defined as

n1s12 + n 2 s 22 1 n
σˆ 2 =
n1 + n 2
, where s 2 = ∑ ( xi − x ) 2 .
n − 1 i =1

iii) If σ 12 ≠ σ 22 and are unknown, then

x1 − x2
Z= .
s12 s 22
+
n1 n2

Confidence limit for difference of two means


By definition,
Pr [| Z | ≤ Zα / 2 ] = 1 − α

⇒ Pr [− Zα / 2 ≤ Z ≤ Zα / 2 ] = 1 − α ,

where Z α / 2 be obtained from standard normal table for given α . Thus,

 ( x − x 2 ) − ( µ1 − µ 2 ) 
Pr − Zα / 2 ≤ 1 ≤ Zα / 2  = 1 − α
 SE ( x1 − x 2 ) 
⇒ Pr [( x1 − x 2 ) − Zα / 2 SE ( x1 − x 2 ) ≤ ( µ1 − µ 2 ) ≤ ( x1 − x 2 ) + Zα / 2 SE ( x1 − x 2 )] = 1 − α

Hence the confidence limits for ( µ1 − µ 2 ) is


( x1 − x 2 ) ± Zα / 2 SE ( x1 − x 2 ) .
Exercise 12) Suppose that the nicotine contents of two brands of cigarettes are being
measured. If in an experiment, 50 cigarettes of the first brand had an average nicotine
contents 2.61 mg, while 40 cigarettes of the second brand had a average nicotine contents
2.38 mg., test the claim that brand first, had more nicotine contents than the brand second at
5% level of significance. The population standard deviations are given 0.12 mg and 0.14
mg respectively.
Solution: Given x1 = 2.61 mg, x 2 = 2.38 mg, n1 = 50 , n2 = 40 , σ 1 = 0.12 , σ 2 = 0.14
and the hypothesis of interest is
H 0 : µ1 = µ 2 vs H A : µ1 > µ 2
Under H 0 , the test statistic

x1 − x 2
Z= ~ N (0,1)
σ 12 σ 22
+
n1 n2
96 RU Khan

2.51 − 2.38
= = 8.12 .
1/ 2
 (0.12) 2
(0.14) 2 
 + 
 50 40 
 
Since Z = 8.12 > Z α = 1.645 at α = 0.05 level, H 0 is to be rejected, i.e., we conclude that
brand first had more nicotine contents than second.
Exercise 13) Random samples drawn from two populations when standard deviations are
2.50 inch and 2060 inch respectively, gave he following data relating to the heights of adult
males.

Population A Population B
Mean height (in inches) 67.25 67.42
Sample size 1200 1000

Do the data indicates that he member of population B are on average taller than the people of
population A ?
Solution: Given x1 = 67.25 inches, n1 = 1200 , σ 1 = 2.5 , x 2 = 67.42 inches,
n2 = 1000 , σ 2 = 2.6 .
We want to test
H 0 : µ1 = µ 2 vs H A : µ1 < µ 2 .
Under H 0 , the test statistic

x1 − x 2
Z= ~ N (0,1)
σ 12 σ 22
+
n1 n2

67.25 − 67.42
= = −1.56 .
1/ 2
 (2.5) 2 (2.6) 2 
 + 
 1200 1000 
 
Since Z = −1.56 > Zα = −1.645 at α = 0.05 level, H 0 is to be accepted and we conclude
that on average height of peoples of population A and B are same.
Exercise 14) In a survey of buying habits, the following data was chosen
x1 = 250 , n1 = 400 , s1 = 40 , x 2 = 220 , n2 = 400 , s 2 = 55 .
Test whether the average weekly food expenditure of the two populations are same. Find
95% confidence limits.
Solution:
The hypothesis of interest is
H 0 : µ1 = µ 2 vs H A : µ1 ≠ µ 2
Under H 0 , the test statistic
Statistical Inference 97

x1 − x 2
|Z |= ~ N (0,1)
s12 s 22
+
n1 n2

250 − 220
= = 8.82 .
1/ 2
 (40) 2
(55) 2 
 + 
 400 220 
 
Since | Z | = 8.82 > Zα / 2 = 1.96 at α = 0.05 level, H 0 is to be rejected.
95% confidence limits may be

s12 s 22
( x1 − x 2 ) ± Zα / 2 + or 30 ± 1.96 4 + 7.5625 or (23.335 and 36.665) .
n1 n2
Exercise 15) The means of two large samples of 1000 and 2000 members are 67.5 inches
and 68.0 inches respectively. Can the samples be regarded as drawn from the same
population of standard deviation 2.5 inches? Test at 5% level of significance.
Solution: Given x1 = 67.5 inches, n1 = 1000 , x 2 = 68.0 inches, n2 = 2000 , σ = 2.5 .
We want to test
H 0 : µ1 = µ 2 vs H A : µ1 ≠ µ 2
Under H 0 , the test statistic

x1 − x 2
|Z |= ~ N (0,1)
1 1
σ +
n1 n2

67.5 − 68.0
= = 5.1 .
1/ 2
 1 1 
2.5  + 
 1000 2000 
Since | Z | = 5.1 > Zα / 2 = 1.96 at α = 0.05 level, we reject H 0 , and conclude that the
samples are certainly not drawn from the same population with standard deviation 2.5
inches.

You might also like