0% found this document useful (0 votes)
7 views39 pages

Module 4

This document covers the fundamentals of hypothesis testing, including key definitions such as population, sample, and sampling error. It explains the concepts of null and alternative hypotheses, types of errors, critical regions, and the power of tests, along with procedures for testing hypotheses and specific tests for small samples. Additionally, it provides examples of applying the Student's t-test for various scenarios.

Uploaded by

Soumodeep Nayak
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
7 views39 pages

Module 4

This document covers the fundamentals of hypothesis testing, including key definitions such as population, sample, and sampling error. It explains the concepts of null and alternative hypotheses, types of errors, critical regions, and the power of tests, along with procedures for testing hypotheses and specific tests for small samples. Additionally, it provides examples of applying the Student's t-test for various scenarios.

Uploaded by

Soumodeep Nayak
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Module-4-Test of Significance

Testing of Hypothesis

Important Definitions and Results:

Population:
The group of individuals, under study is called population. The population may be finite or
infinite.
Sample:
A finite subset of statistical individuals in a population is called sample.
Take Sample

Inference Population
Sample

Sample size:
The number of individuals in a sample is called the sample size. It is denoted by n when n<30,
then the sample is called small sample when n≥30, then the sample is called large sample.
Sampling error:
On examining a sample of a particular stuff we arrive at a decision of purchasing or rejecting
the stuff. The error involved in such approximation is known as sampling error.
Parameters & statistics:
Parameters : The statistical constants of population( namely mean µ &,varianceσ2 ) parameters.
statistics: The Statistical constants computed from sample observations alone( namely mean
𝑥̅ &,variance𝑠 2 )
Notation:
Population sample
Size 𝑁 𝑛
Mean 𝜇 𝑥̅
Standard Deviation 𝜎 𝑠
Proportion 𝑃 𝑝

Sampling distribution:
If we draw a sample of size ‘n’ from a given finite population of size N then total number of
N N
possible samples is cn .From each of the cn samples if we compute a statistic (mean,
N
variance)and then we form a frequency distribution for these cn values of a statistic. Such a
distribution is called sampling distribution of that statistic.
Standard error:
The Standard deviation of the sampling distribution is called standard error.
Testing of Hypothesis:
Hypothesis:
A Statistical hypothesis is a statement about a population parameter.
Hypothesis testing(or) Test of significance:
The procedure which enable us to decide on the basis of sample result whether a
hypothesis is true or not is called test of hypothesis (or) test of significance there are two type
[Link] Hypothesis
[Link] hypothesis

1
Null hypothesis: Null hypothesis is the hypothesis which is tested for possible rejection under
the assumption that it is true . Null hypothesis is denoted by H0

Alternative hypothesis:
Any hypothesis which is complementary to null hypothesis is called alternative hypothesis
and is denoted by H1 .
For example: To test the null hypothesis that the population has a specified mean 𝜇0
Null Hypothsis 𝐻0 : 𝜇 = 𝜇0
Alternative Hypothesis (𝑖)𝐻1 : 𝜇 ≠ 𝜇0 (𝑇𝑤𝑜 𝑡𝑎𝑖𝑙𝑒𝑑)
(𝑖𝑖) 𝐻1 : 𝜇 < 𝜇0 (𝑙𝑒𝑓𝑡 𝑡𝑎𝑖𝑙𝑒𝑑)
(𝑖𝑖)𝐻1 : 𝜇 > 𝜇0 (𝑅𝑖𝑔ℎ𝑡 𝑡𝑎𝑖𝑙𝑒𝑑)

Errors in Sampling:
1) Type-I error (𝜶): Reject 𝐻0 When 𝐻0 is true
2) Type-II error (𝜷): Accept 𝐻0 When 𝐻0 is false
P(rejecting a good lot)= 𝛼
P(Accepting a bad lot)= 𝛽
Critical region:
A region in which null hypothesis H0 is rejected is called the critical region (or) region of
rejection .

Level of significance (l.o.s):


The maximum probability of making type-I error is called level of significance and is
denoted by  . Where  = P (making Type-I error)
= P (H0 is rejected when it is true)
This can be measured in terms of percentage i.e. 5%, 1%, 10% etc…….
Power of the test:
The probability of rejecting a false hypothesis is called power of the test and is denoted by 1  
Power of the test =P (H0 is rejected when it is true)
= 1- P (H0 is accepted when it is false)
Test statistic:
The test statistic is defined as the difference between the sample statistic value and the
hypothetical value, divided by the standard error of the statistic.
t  E (t )
i.e. test statistic Z 
S .E (t )
= 1- P (Committing Type-II error) = 1- 
One tailed and two tailed tests:
A test with the null hypothesis H 0 :    0 against the alternative hypothesis H 1 :    0 , it is
called a two tailed test. In this case the critical region is located on both the tails of the
distribution.
A test with the null hypothesis H 0 :    0 against the alternative hypothesis H 1 :    0 (right
tailed alternative) or H 1 :    0 (left tailed alternative) is called one tailed test. In this case the
critical region is located on one tail of the distribution.
H 0 :    0 against H 1 :    0 ------- right tailed test
H 0 :    0 against H 1 :    0 ------- left tailed test

2
Utility of standard error:
1. It is a useful instrument in the testing of hypothesis. If we are testing a hypothesis at 5% l.o.s
t  E (t )
and if the test statistic i.e. Z   1.96 then the null hypothesis is rejected at 5% l.o.s
S .E (t )
otherwise it is accepted.
2. With the help of the S.E we can determine the limits with in which the parameter value
expected to lie.
3. S.E provides an idea about the precision of the sample. If S.E increases the precision decreases
1
and vice-versa. The reciprocal of the S.E i.e. is a measure of precision of a sample.
S.E
4. It is used to determine the size of the sample.

Procedure for testing of hypothesis:


1. Set up a null hypothesis i.e. H 0 :    0 .
2. Set up a alternative hypothesis i.e. H 1 :    0 or H 1 :    0 or H 1 :    0
3. Choose the level of significance i.e.  .
4. Select appropriate test statistic Z.
5. Select a random sample and compute the test statistic.
6. Calculate the tabulated value of Z at  % l.o.s i.e. Z  .
7. Compare the test statistic value with the tabulated value at  % l.o.s. and make a decision
whether to accept or to reject the null hypothesis.

Test of significance of small samples:


When the size of the sample n<30, then the sample is a small sample.
The following are some important tests for small samples.
i. Student’s ‘t’ test
ii. F-test
iii. 𝜒2-test
Small sample tests:
Student’s ‘t’ test(For single mean):
The student’s ‘t’ is defined by the statistic,
𝑥̅ − 𝜇
𝑡= 𝑠
√𝑛 − 1

1
Where 𝑥̅ = 𝑛 ∑ 𝑥𝑖 = 𝑠𝑎𝑚𝑝𝑙𝑒 𝑚𝑒𝑎𝑛 𝜇 = 𝑝𝑜𝑝𝑢𝑙𝑎𝑡𝑖𝑜𝑛 𝑚𝑒𝑎𝑛
1
𝑠 2 = 𝑛−1 ∑𝑖 (𝑥𝑖 − 𝑥̅ )2 𝑛 = 𝑠𝑎𝑚𝑝𝑙𝑒 𝑠𝑖𝑧𝑒

3
𝑛 − 1 = 𝑑𝑒𝑔𝑟𝑒𝑒𝑠 𝑜𝑓 𝑓𝑟𝑒𝑒𝑑𝑜𝑚
Note:
̅−𝝁
𝒙
1. If the standard deviation of a sample is given directly then 𝒕 = 𝑺.𝑫
√𝒏−𝟏
2. If the standard deviation of a sample is not given directly (i.e., if we have to find S.D) then 𝒕 =
̅−𝝁
𝒙
𝑺.𝑫
√𝒏
3. If calculated value of |𝑡| > tabulated value of |𝑡| at 5% or 1% l.o.s then the null hypothesis H0 is
rejected.
4. If calculated value of |𝑡| < tabulated value of |𝑡| at 5% or 1% l.o.s then the null hypothesis H0 is
accepted.
5. If 𝑡0.05 is the table value of |𝑡| for n-1 degrees of freedom at 5% l.o.s then 95% confidence limit
𝑺
̅ ± 𝒕𝟎.𝟎𝟓
for 𝜇 is given by 𝒙
√𝒏−𝟏

Applications (uses) of t-distribution:


1. To test if the sample mean (𝑥̅ ) differs significantly from the hypothetical value 𝜇 of population
mean.
2. To test the significance of the difference between two sample means.
3. To test the significance of an observed sample correlation coefficient and sample regression
coefficient.
Assumptions for t-distribution:
1. The parent population from which sample is drawn is normal.
2. The population standard deviation σ is unknown.
3. Sample size n<30.
Student’s ‘t’-test for single mean:
Problems:
Type-I When S.D is given directly:
1. A machinist is making engine parts with axle diameters of 0.700 inch. A random sample of 10
parts shows a mean diameter of 0.742 inch with a S.D of 0.40 inch. Compute the statistic you
would use to test whether the work is meeting specifications.
Solution: Given n = 10<30 (small sample)
Sample mean 𝑥̅ = 0.742 inches.
Population mean 𝜇 = 0.700 inches.
Sample S.D = 0.040 inches.
Null Hypothesis: H0 : The product is meeting the specifications. i.e., 𝜇 = 0.700
Alternative Hypothesis: H1 : The product is not meeting the specifications. i.e., 𝜇 ≠ 0.700(Two
tailed)
Test Statistic:
𝑥̅ − 𝜇 0.742 − 0.700
𝑡= 𝑠 = = 3.15
0.040
√𝑛 − 1 √10 − 1
Calculated |𝑡| = 3.15
Tabulated |𝑡| at 5% level with 10-1 = 9 degrees of freedom is 2.26.
∵ Calculated |𝑡| > Tabulated |𝑡|(i.e., 3.15 > 2.26)
∴ H0 is rejected.

2. The mean weekly sales of soap bars in departmental stores was 146.3 bars per store. After an
advertising campaign the mean weekly sales in 22 stores for a typical week increases to 153.7
and showed a S.D of 172, was advertising campaign successful.
Solution: Given n = 22<30 (small sample)
Sample mean 𝑥̅ = 153.7

4
Population mean 𝜇 = 146.3.
Sample S.D = 17.2
Null Hypothesis: H0 : The advertising campaign was not successful. i.e., 𝜇 = 146.3
Alternative Hypothesis: H1 : The advertising campaign was successful. i.e., 𝜇 > 146.3(Right
tailed)
Test Statistic:
𝑥̅ − 𝜇 153.7 − 146.3
𝑡= 𝑠 = = 1.97
17.2
√𝑛 − 1 √22 − 1
Calculated |𝑡| = 1.97
Tabulated |𝑡| at 5% level with 22-1 = 21 degrees of freedom is 1.72.
∵ Calculated |𝑡| > Tabulated |𝑡|(i.e., 1.97 > 1.72)
∴ H0 is rejected.
i.e., The advertising campaign was successful.

3. A sample of 26 bulbs gives a mean life of 990 hours with a S.D of 20 hours. The manufacture
claims that the mean life of bulbs is 1000 hours. Is the sample not up to the standard?
Solution: Given n = 26<30 (small sample)
Sample mean 𝑥̅ = 990
Population mean 𝜇 = 1000
Sample S.D = 20
Null Hypothesis: H0 : The sample is upto the standard. i.e., 𝜇 = 1000.
Alternative Hypothesis: H1 : The advertising campaign was successful. i.e., 𝜇 < 146.3(Left
tailed)
Test Statistic:
𝑥̅ − 𝜇 990 − 1000
𝑡= 𝑠 = = −2.5
20
√𝑛 − 1 √26 − 1
Calculated |𝑡| = 2.5
Tabulated |𝑡| at 5% level with 26-1 = 25 degrees of freedom is 1.708.
∵ Calculated |𝑡| > Tabulated |𝑡|(i.e., 2.5 > 1.708)
∴ H0 is rejected.
i.e., The sample is not up to the standard.

4. The mean life time of a sample of 25 fluorescent light bulbs produced by a company is
computed to be 1570 hrs with a S.D 120 hrs. The company claims that the average life of the
bulbs produced by the company is 1600 hrs using the l.o.s of 0.05. Is the claim complete.
Solution: Given n = 25<30 (small sample)
Sample mean 𝑥̅ = 1570
Population mean 𝜇 = 1600
Sample S.D = 120
Null Hypothesis: H0 : The claim is acceptable. i.e., 𝜇 = 1600 hrs.
Alternative Hypothesis: H1 : 𝜇 ≠ 1600 hrs (Two tailed).
Test Statistic:
𝑥̅ − 𝜇 1570 − 1600
𝑡= 𝑠 = = −1.22
120
√𝑛 − 1 √25 − 1
Calculated |𝑡| = 1.22
Tabulated |𝑡| at 5% level with 25-1 = 24 degrees of freedom is 2.06.
∵ Calculated |𝑡| > Tabulated |𝑡|(i.e., 1.22 < 2.06)
∴ H0 is accepted.

5
i.e., The claim that the average life of the bulbs produced by the company is 1600 hrs is
acceptable.

5. A soap manufacturing company was distributing a particular brand of soap through a large
number of retail shops. Before a heavy advertisement campaign, the man sales per week per
shop was 140 dozens. After the campaign, a sample of 26 shops was taken and the mean
sales was found to be 147 dozens with standard deviation 16. What conclusion do you draw
on the impact of advertisement on sales? Use 5% significance level.
Solution: It is given S.D = 16, n = 26, 𝑥̅ = 147, μ = 140
Null Hypothesis: H0: The advertisement was not successful. i.e., μ = 140 against
Alternative Hypothesis: H1 = μ ≠ 140
Test Statistic:
𝑥̅ − 𝜇 147 − 140 7 7
𝑡= 𝑠 =
16
=
16
=  2.19
3.2
√𝑛 − 1 √25 5

From the table, for (26-1) = 25 d.o.f., t005 = 2.06. Since Computed value of t > t0.05, we reject the
null hypothesis. That is, the advertisement may be considered to have changed the average
sales volume.

6. A random sample of 25 cups from a certain coffee dispensing machine yields a mean
6.9ounces per cup. Use α = 0.05 level of significance to test, on the average, the machine
dispense μ = 7.0 ounces. Assume that the distribution of ounces per cup is normal, and that
the variance is the known quantity σ2 = 0.01ounces.
Solution:
Given n = 25, x  6.9,   7.0,  2  0.01
S.D  0.01  
Null Hypothesis H0:  =7 (The sample mean x does not differ significantly from the population
mean  .
Alternative Hypothesis H1:   7.
Test statistic:
x 6.9  7
t=   4.89
s 0.1
n 1 24
Calculated Value of t  4.89 and table value = 1.645
Since the calculated value of t > the tabulated value of t, we reject the null hypothesis H 0.
i.e., the population mean  = 7 is unacceptable.

Home Work:
1. A machine is designed to produce insulating washers for electrical devices of average thickness
of 0.025 cm. A random sample of 10 washers was found to have a thickness of 0.024 cm with a
S.D of 0.002 cm. Test the significance of the deviation. Table value for t for 9 degrees of
freedom at 5% l.o.s is 2.262.
(t0.05=2.262, D.f = 9, |𝒕| = 𝟏. 𝟓)
2. The average breaking strength of the steel rods is specified to be 18.5 thousand pounds. To test
this a sample of 14 rods was tested. The mean and standard deviations obtained were 17.85
and 1.955 respectively. Is the result of the experiment significant?
(t0.05=2.16, D.f =13, |𝒕| = 𝟏. 𝟏𝟗𝟗)

Student’s t-test:
Type-II When S.D is not given directly:

6
1. A random sample of 10 boys had the following I.Q’s 70, 120, 110, 101, 88, 83, 95, 98, 107, 100.
Does the data support the assumption of a population mean I.Q of 100? Find a reasonable
range in which most of the mean I.Q values of samples of 10 boys lie.
Solution: Given n = 10<30 (small sample)
Since the standard deviation (S.D) is not given, we have to find out S.D and sample mean.
∑𝑥 2 1
𝑥̅ = ,𝑠 = ∑(𝑥𝑖 − 𝑥̅ )2
𝑛 𝑛−1
x-𝑥̅
x (x-𝑥̅ )2
x-97.2
70 70-97.2 = -27.2 739.84
120 22.8 519.84
110 12.8 163.84
101 3.8 14.44
88 -9.2 84.64
83 -14.2 201.64
95 -2.2 4.84
98 0.8 0.64
107 9.8 96.04
100 2.8 7.84
∑ 𝑥 = 972 1833.60

∑𝑥 972
Sample mean 𝑥̅ = = = 97.2
𝑛 10
1 1833.60
𝑠 2 = 𝑛−1 ∑(𝑥𝑖 − 𝑥̅ )2 = = 203.73
9
S.D = √203.73 = 14.27
Null Hypothesis: H0 : The data support the assumption of a population mean I.Q of 100
i.e., 𝜇 = 100
Alternative Hypothesis: H1 : 𝜇 ≠ 100 (two tailed)
Test Statistic:
𝑥̅ − 𝜇 97.2 − 100
𝑡= 𝑠 = = −0.61
14.72
√𝑛 √10
Calculated |𝑡| = 0.61
Tabulated |𝑡| at 5% level with 10-1 = 9 degrees of freedom is 2.26.
∵ Calculated |𝑡| > Tabulated |𝑡|(i.e., 0.61 < 2.26)
∴ H0 is accepted.
i.e., The data support the assumption of a population mean I.Q of 100.
Next to find the reasonable range in which most of the mean I.Q values of samples of 10 boys
lie.
𝑠
𝑥̅ ± 𝑡0.05 = 97.2 ± 2.26 × 4.514 = 97.2 ± 10.20 = 107.41,86.99
√𝑛 − 1
The 95% confidence limits with in which the mean I.Q values of sample of 10 boys will
lie in (range), [86.99,107.41]

2. The height of 10 males of a given locality are found to be 70, 67, 62, 68, 61, 68, 70, 64, 64, 66
inches. Is it reasonable to believe that the average height is greater than 64 inches?
Solution: Given n = 10<30 (small sample)
Since the standard deviation (S.D) is not given, we have to find out S.D and sample mean.
∑𝑥 2 1
𝑥̅ = ,𝑆 = ∑(𝑥𝑖 − 𝑥̅ )2
𝑛 𝑛−1
x x-𝑥̅ (x-𝑥̅ )2

7
x-66
70 70-66 = 4 16
67 1 1
62 -4 16
68 2 4
61 -5 25
68 2 4
70 4 16
64 -2 4
64 -2 4
66 0 0
∑ 𝑥 = 660 90

∑𝑥 660
Sample mean 𝑥̅ = = = 66
𝑛 10
1 90
𝑠 2 = 𝑛−1 ∑(𝑥𝑖 − 𝑥̅ )2 = = 10
9
S.D = √10 = 3.16
Null Hypothesis: H0 : The average height is not greater than 64 inches
i.e., 𝜇 = 64
Alternative Hypothesis: H1 : 𝜇 > 64 (Right tailed)
Test Statistic:
𝑥̅ − 𝜇 66 − 64
𝑡= 𝑠 = =2
3.16
√𝑛 √10
Calculated |𝒕| = 𝟐
Tabulated |𝒕| at 5% level with 10-1 = 9 degrees of freedom is 1.833.
∵ Calculated |𝒕| > Tabulated |𝒕|(i.e., 2 > 1.833)
∴ H0 is rejected.
i.e., The average height is greater than 64 inches

3. A certain pesticide is packed into bags by a machine. A random sample of 10 bags is drawn
and their contents are found to weight (in Kg) as follows: 50, 49, 52, 44, 45, 48, 46, 45, 49, 45.
Test if the average packing can be taken to be 50 kg. (t0.05 = 2.262 at 9 d.f)
Solution: Given n = 10<30 (small sample)
Since the standard deviation (S.D) is not given, we have to find out S.D and sample mean.
∑𝑥 2 1
𝑥̅ = ,𝑠 = ∑(𝑥𝑖 − 𝑥̅ )2
𝑛 𝑛−1
̅
x-𝒙
x ̅)2
(x-𝒙
x-47.3
50 50-47.3 = 2.7 7.29
49 1.7 2.89
52 4.7 22.09
44 -3.3 10.89
45 -2.3 5.29
48 0.7 0.49
46 -1.3 1.69
45 -2.3 5.29
49 -1.7 2.89
45 -2.3 5.29
∑ 𝑥 = 473 64.1

8
∑𝑥 473
Sample mean 𝑥̅ = = = 47.3
𝑛 10
1 64.1
𝑠 2 = 𝑛−1 ∑(𝑥𝑖 − 𝑥̅ )2 = = 7.12
9
S.D = √7.12 = 2.67
Null Hypothesis: H0 : The average packing is 50 kgs.
i.e., 𝜇 = 50
Alternative Hypothesis: H1 : 𝜇 ≠ 50 (two tailed)
Test Statistic:
𝑥̅ − 𝜇 47.3 − 50
𝑡= 𝑠 = = −3.19
2.67
√𝑛 √10
Calculated |𝒕| = 𝟑. 𝟏𝟗
Tabulated |𝒕| at 5% level with 10-1 = 9 degrees of freedom is 2.262.
∵ Calculated |𝒕| > Tabulated |𝒕|(i.e., 𝟑. 𝟏𝟗> 2.262)
∴ H0 is rejected.
i.e., The average packing is not 50 kgs.

4. The individuals are chosen at random from a population and their highest are found to be in
inches 63, 63, 66, 67, 68, 69, 70, 70, 71 and 71. In the light of these data, decreases the
suggestion that the mean height in the population is 66. Use 5% significance level.
Solution: Given n = 10<30 (small sample)
Since the standard deviation (S.D) is not given, we have to find out S.D and sample mean.
∑𝑥 2 1
𝑥̅ = ,𝑠 = ∑(𝑥𝑖 − 𝑥̅ )2
𝑛 𝑛−1
𝑥 𝑥 − 𝑥̅ (𝑥 − 𝑥̅ )2
63 -4.8 23.04
63 -4.8 23.04
66 -1.8 3.24
67 -0.8 0.64
68 0.2 0.04
69 1.2 1.44
70 2.2 4.84
70 2.2 4.84
71 3.2 10.24
71 3.2 10.24
∑(𝑥 − 𝑥̅ )2
∑ 𝑥 = 678
= 81.60

678
x  67.8
10

s 
2  ( xi  x) 2 81.60
  9.066
n 1 9
s  9.066  3.011
Null Hypothesis: H0 : μ = 66” against
Alternative Hypothesis: H1 : μ ≠ 66”
Test Statistic:
𝑥̅ − 𝜇 67.8 − 66 1.80
𝑡= 𝑠 = = = 1.89
3.01 3.01
√𝑛 √10 3.16
Calculated |𝒕| = 1.89
Two tailed table value of t with 9d. D. f at 5% level is found to be 2.26.

9
As computed value as t < tabulated ‘t’. So computed t does not lie in the critical region. We
accept H0 and conclude that the experiment provides no ground for doubt that the mean height
is 66”.

5. A certain medicine administered to each 10 patients resulted in the following increase in the
blood pressure 8, 8, 7, 5, 4, 1, 0, 0,-1,-1. Can it be concluded that the medicine was
responsible for the increase in blood pressure.
Solution: Given  = 0, n = 10 (Small sample n < 30) .
Since the standard deviation (S.D) is not given, we have to find out S.D and sample mean.
∑𝑥 2 1
𝑥̅ = ,𝑠 = ∑(𝑥𝑖 − 𝑥̅ )2
𝑛 𝑛−1
𝑥 𝑥 − 𝑥̅ (𝑥 − 𝑥̅ )2
8 4.9 24.01
8 4.9 24.01
7 3.9 15.21
5 1.9 3.61
4 0.9 0.81
1 -2.1 4.41
0 -3.1 9.61
0 -3.1 9.61
-1 -4.1 16.81
-1 -4.1 16.81
∑(𝑥 − 𝑥̅ )2
∑ 𝑥 = 31
= 124.9

∑ 𝑥 31
𝑥̅ = = = 3.1
𝑛 10

s 2

(x
i  x) 2

124.9
 13.88
n 1 9
s  13.88  3.72
Null Hypothesis: H0 :  = 0
Alternative Hypothesis: H1 :   0 LOS  = 0.05, d.f = 10 – 1 = 9 = 2.26
Test statistic:
x  3.1  0
t   2.6314 
s 3.5341
n 1 3
Calculated Value of t  2.6314 and table value = 2.26
Calculated Value of t  2.6314 > table value = 2.26
We reject the null Hypothesis and we concluded that there is significant difference; the
medicine is responsible for the increase of blood pressure.

HOME WORK:
1. The average breaking strength of steel rods is specified to be 17.5. To test this, sample of 14
rods tested & gave the following results 15, 18, 16, 21, 19, 17, 17, 15, 17, 20, 19, 17, 18. Is the
result of the experiment significant? Also obtain the 95% confidence limits.
(t0.05=2.16, D.f =13, |𝒕| = 𝟎. 𝟕𝟎𝟗𝟎)

2. A random sample of 16 values from a normal population showed a mean of 53 and a sum of
squares of deviations from the mean equals to 150. Can this sample be regarded as taken from
the population having 56 as mean? obtain the 95% confidence limits of the mean of the
population.

10
1
Hint: Given ∑(𝑥𝑖 − 𝑥̅ )2 = 150, 𝑓𝑖𝑛𝑑 𝑆 2 = ∑(𝑥𝑖 − 𝑥̅ )2
𝑛−1
(t0.05=2.13, D.f =15, |𝒕| = 𝟑. 𝟕𝟗)
Student’s t-test:
Difference of means (independent samples):
̅𝟏 & 𝒙
To test the significant difference between two means 𝒙 ̅𝟐 of sample sizes 𝒏𝟏 and 𝒏𝟐 ,

̅𝟏 − 𝒙
𝒙 ̅𝟐
𝒕=
𝟏 𝟏
𝒔√𝒏 + 𝒏
𝟏 𝟐
𝟐
∑(𝒙𝟏 − 𝒙 ̅𝟐 ) 𝟐
̅𝟏 ) + ∑(𝒙𝟐 − 𝒙
𝒔𝟐 =
𝒏𝟏 +𝒏𝟐 − 𝟐
(or)
𝟏
𝒔𝟐 = [𝒏 𝒔 𝟐 + 𝒏𝟐 𝒔𝟐 𝟐 ]
𝒏𝟏 +𝒏𝟐 − 𝟐 𝟏 𝟏

𝒔𝟏 𝒂𝒏𝒅 𝒔𝟐 are sample standard deviation.


𝒏𝟏 +𝒏𝟐 − 𝟐 = 𝒅𝒆𝒈𝒓𝒆𝒆𝒔 𝒐𝒇 𝒇𝒓𝒆𝒆𝒅𝒐𝒎

Problems:
6. Samples of two types of electric light bulbs were tested for length of life and following data
were obtained
Type-I Type-II
Sample number 𝒏𝟏 = 𝟖 𝒏𝟐 = 𝟕
Sample means ̅𝟏 = 𝟏𝟐𝟑𝟒 𝒉𝒓𝒔
𝒙 ̅𝟐 = 𝟏𝟎𝟑𝟔 𝒉𝒓𝒔
𝒙
Sample S.D 𝒔𝟏 = 𝟑𝟔 𝒉𝒓𝒔 𝒔𝟐 = 𝟒𝟎 𝒉𝒓𝒔
(A.U NOV/DEC 2010)
Is there a difference in the means sufficient to warrant that type I is superior to type II
regarding length of life.
Solution: Given that 𝑛1 = 8, 𝑛2 = 7, 𝑥̅1 = 1234, 𝑥̅2 = 1036, 𝑠1 = 36, 𝑠2 = 40
Null hypothesis H0: The two types I and II of electric bulbs are identical.
i.e., H0: µ1 = µ2
Alternate Hypothesis H1: µ1 > µ2 (right tailed).
To find S:
1 1
𝑠2 = [𝑛1 𝑠1 2 + 𝑛2 𝑠2 2 ] = [8(36)2 + 7(40)2 ] = 1659.08
𝑛1 +𝑛2 − 2 8+7−2
𝑠 = √1659.08 = 40.73
Test Statistic:

𝑥̅1 − 𝑥̅2 1234 − 1036


𝑡= = = 9.39
1 1 1 1
𝑠√𝑛 + 𝑛 40.73√8 + 7
1 2
Calculated |𝒕| = 𝟗. 𝟑𝟗
Degrees of freedom (D.f): 8+7-2 = 13
Tabulated value of t for 13 D.f at 5% l.o.s is 1.77(one-tailed)
∵ Calculated |𝒕| > Tabulated |𝒕|(i.e., 𝟗. 𝟑𝟗> 1.77)
∴ H0 is rejected.

7. Two types of batteries are tested for their length of life and the following data are
obtained:

11
[Link] Mean life in
Battery
samples Hrs Variance
type A 9 600 121
type B 8 640 144

Is there is a significant difference in the two means? (Value of t for 15 degrees of freedom at
5% level is 2.131)
Solution: Given 𝑛1 = 9, 𝑛2 = 8, 𝑥̅1 = 600, 𝑥̅2 = 640, 𝑠1 2 = 121, 𝑠2 2 = 144
We set up Null hypothesis H0: There is no significant difference in the two means
i.e., H 0 : 1   2
Against Alternate Hypothesis H1 : 1  2
The appropriate test statistic is Fisher’s t which under H0 follows t distribution with (n1+n2-2)
d.o.f.
To find S:
1 9  121  8  144
𝑠2 = 𝑛 [𝑛1 𝑠1 2 + 𝑛2 𝑠2 2 ] ⟹ s   149.4  12.2
1 +𝑛2 −2 982
Test Statistic:
600  640  40
We compute t   6.7
1 1 12.2  0.486
12.2 
9 8
d.o.f. = (n1 + n2 - 2) = 9 + 8 - 2 = 15
Computed t = 6.7 which is larger than tabulated t (=2.13 for 15 d.o.f. of and α = 0.05 )
So we reject H0 and accept H1 .
Our conclusion is there is significant difference in the two means.

8. Two random samples gave the following results.

Sum of squares of deviations


Sample Size Sample Mean
from the mean
1 10 15 90
2 12 14 108
Test whether the samples come from the same normal population using t-test
Solution: : Given 𝑛1 = 10, 𝑛2 = 12, 𝑥̅1 = 15, 𝑥̅2 = 14, 𝑠1 2 = 90, 𝑠2 2 = 108
Null hypothesis H0: The two samples have been drawn from the same normal population.
i.e. H0 = 1   2 and  12   22
Against Alternate Hypothesis H 1 : 1   2 and  12   22
The appropriate test statistic is Fisher’s t which under H0 follows t distribution with (n1+n2-2)
d.o.f.
To find S:

s2 
1
n1  n2  2
 
( x1  x1 ) 2   ( x2  x2 ) 2 
1
20
(90  108)  9.9

Test Statistic:
x1  x2 15  14
The test statistic t =  = 0.742.
1 1 1 1
9.9 (  )
s   
 n1 n2  10 12
d.f = (n1 + n2 - 2) = 20.
Tabulated value of t for 20 d.f at 5% level of significance is 2.086.
Since calculated value of t = 0.74 < the tabulated value of t = 2.086, we accept the null
hypothesis H0 = 1   2 and  12   22
Hence the given samples come from the same normal population

12
9. Two independent samples of sizes 8 and 7 contained the following values:
Sample-I: 19 17 15 21 16 18 16 14
Sample-II: 15 14 15 19 15 18 16
Is the difference between the sample means significant? (A.U APR/MAY 2010)
Solution: Null hypothesis H0: There is no significant difference between the means.
i.e., µ1 = µ2
Alternate Hypothesis H1: µ1 ≠ µ2
To find S:
𝒙 𝒙−𝒙 ̅ (𝒙 − 𝒙 ̅) 𝟐 𝒚 𝒚−𝒚 ̅ ̅ )𝟐
(𝒚 − 𝒚
19 19 – 17 = 2 4 15 15 – 16 = -1 1
17 0 0 14 -2 4
15 -2 4 15 -1 1
21 4 16 19 3 9
16 -1 1 15 -1 1
18 1 1 18 -2 4
16 -1 1 16 0 0
14 -3 9
∑𝒙 = 𝟏𝟑𝟔 36 ∑𝒚 = 𝟏𝟏𝟐 20

𝟏𝟑𝟔 𝟏𝟏𝟐
̅=
𝒙 ̅=
= 𝟏𝟕, 𝒚 = 𝟏𝟔, 𝒏𝟏 = 𝟖, 𝒏𝟐 = 𝟕
𝟖 𝟕
̅𝟏 )𝟐 + ∑(𝒙𝟐 − 𝒙
∑(𝒙𝟏 − 𝒙 ̅𝟐 ) 𝟐 𝟏
𝒔𝟐 = = [𝟑𝟔 + 𝟐𝟎] = 𝟒. 𝟑𝟎
𝒏𝟏 +𝒏𝟐 − 𝟐 𝟖+𝟕−𝟐
𝒔 = √𝟒. 𝟑𝟎 = 𝟐. 𝟎𝟕
Test Statistic:
̅−𝒚
𝒙 ̅ 𝟏𝟕 − 𝟏𝟔
𝒕= = = 𝟎. 𝟗𝟑
𝟏 𝟏 𝟏 𝟏
𝑺√𝒏 + 𝒏 𝟐. 𝟎𝟕√𝟖 + 𝟕
𝟏 𝟐
Calculated |𝒕| = 𝟎. 𝟗𝟑
Degrees of freedom (D.f): 8+7-2 = 13
Tabulated value of t for 13 D.f at 5% l.o.s is 2.16(two-tailed)
∵ Calculated |𝒕| < Tabulated |𝒕|(i.e., 𝟏. 𝟑𝟗< 2.16)
∴ H0 is accepted.

10. Below are given the gain of weight (in lbs) of pigs fed on two diets A & B.
Diet
25 32 30 34 24 14 32 24 30 31 35 25 - - -
A
Diet
44 34 22 10 47 31 40 30 32 35 18 21 35 29 22
B

Test if the two diets differ significantly as regards to their effect on increase in weight.
Solution: Null hypothesis H0: There is no significant difference between the mean increase in
weight due to diets A & B. i.e., µ1 = µ2
Alternate Hypothesis H1: µ1 ≠ µ2
To find S:

𝒙 𝒙−𝒙 ̅ ̅) 𝟐
(𝒙 − 𝒙 𝒚 𝒚−𝒚 ̅ ̅ )𝟐
(𝒚 − 𝒚
25 25 – 28 = - 3 9 44 44 – 30 = 14 196
32 4 16 34 4 16

13
30 2 4 22 -8 64
34 6 36 10 -20 400
24 -4 16 47 17 289
14 -14 196 31 1 1
32 4 16 40 10 100
24 -4 16 30 0 0
30 2 4 32 2 4
31 3 9 35 5 25
35 7 49 18 -12 144
25 3 9 21 -9 81
- - - 35 5 25
- - - 29 -1 1
- - - 22 -8 64
∑𝒙 = 𝟑𝟑𝟔 380 ∑𝒚 = 𝟒𝟓𝟎 1410

336 450
𝑥̅ = = 28, 𝑦̅ = = 30, 𝑛1 = 12, 𝑛2 = 15
12 15
∑(𝑥1 − 𝑥̅1 )2 + ∑(𝑥2 − 𝑥̅2 )2 1
𝑆2 = = [380 + 1410] = 71.6
𝑛1 +𝑛2 − 2 12 + 15 − 2
𝑠 = √71.6 = 8.46
Test Statistic:

𝑥̅ − 𝑦̅ 28 − 30
𝑡= = = −0.609
1 1 1 1
𝑆√𝑛 + 𝑛 8.46√12 +
1 2 15
Calculated |𝑡| = 0.609
Degrees of freedom (D.f): 12+15-2 = 25
Tabulated value of t for 25 at 5% l.o.s is 2.06(two-tailed)
∵ Calculated |𝒕| < Tabulated |𝒕|(i.e., 𝟎. 𝟔𝟎𝟗< 2.06)
∴ H0 is accepted.

HOME WORK:
11. The average number of articles produced by two machines per day are 200 & 250 with S.D 20 &
25 respectively on the basis of records of 25 days production. Can you regard both the
machines equally efficient at 1% l.o.s
( s = 23.10, |𝒕| = 𝟕. 𝟔𝟓, t0.01 = 2.58, D.f = 48)
12. The means of two random samples of size 9 and 7 are 196.42 & 198.82 respectively. The sum of
the squares of deviation from mean are 26.94 & 18.73 respectively. Can the sample be
considered to have been drawn from the same normal population.
( s = 1.81, |𝒕| = 𝟐. 𝟔𝟑, t0.05 = 2.15, D.f = 14)
13. The height of six randomly chosen sailors are in inches 63, 65, 68, 69, 71 & 72. Those of 10
randomly chosen soldiers are 61,62, 65, 66, 69, 69, 70, 71, 72 & 73. Discuss the light that these
data throw on the suggestion that sailors are on the average taller than soldiers.
( s = 15.257, |𝒕| = 𝟎. 𝟎𝟗𝟗, t0.05 =1.76, D.f = 14)

Test for equality of two means – paired t test


(Testing of significance of difference of two means – Dependent samples)
This test used when for the same sample values which indicate an increase or decrease inthe
variable is given. Before and after values of the variable are normally given or the differences
may be given directly.
Here d = (Xafter – Xbefore) or (Xbefore – Xafter)

14
 
We take Null Hypothesis H 0 :  0 against H1 : 0
d d
 
Or Alternative Hypothesis H1 :  0 or H1 : 0
d d
Test statistic:
d
t
s
n
Where,

d 
d
and s 
 (d  d ) 2


1 
 d 2

( d ) 2 

n n 1 n  1  n 
and d. o. f is n – 1.

Problems:

1. To verify whether a course in accounting improved performance, a similar test was given
to 12 participant both before and after the course. The marks are
Before
44 40 61 52 32 44 70 41 67 72 53 72
course
After
53 38 69 57 46 39 73 48 73 74 60 78
course
Whether the course is useful?
Solution: Given 𝑛 = 12 and the samples are dependent, so we use paired t-test.
Null hypothesis H0: The course is not useful.
i.e., 𝑑̅ = 0
Alternate Hypothesis H1: 𝑑̅ ≠ 0 . We take d = (Xafter – Xbefore).

Before course After course 𝒅 = 𝑨. 𝑪 − 𝑩. 𝑪 𝑑𝟐

44 53 9 81
40 38 -2 4
61 69 8 64
52 57 5 25
32 46 14 196
44 39 -5 25
70 73 3 9
41 48 7 49
67 73 6 36
72 74 2 4
53 60 7 49
72 78 6 36
𝟐
∑𝑑 = 60 ∑𝑑 =578

∑𝑑 60
𝑑̅ = 𝑛 = 12 = 5,
(∑𝑑)2 3600
∑ 𝑑2 −
𝑠2 = 𝑛 = 578 − 12 = 578 − 300 = 25.273
𝑛−1 12 − 1 11
𝑠 = √25.273 = 5.03
Test Statistic:
15
𝑑̅ 5 √12
𝑡= 𝑠 = = 5× = 3.44
5.03 5.03
√𝑛 √12
Calculated |𝑡| = 3.44
Degrees of freedom (D.f): n-1 = 12 – 1 = 11.
Tabulated value of t for 11 at 5% l.o.s is 2.201(two-tailed)
∵ Calculated |𝒕| >Tabulated |𝒕|
∴ H0 is rejected. i.e., The course is useful.

2. An IQ test was administered to 5 persons before and after they were trained. The results are
given below:

I II III IV V
IQ before training 110 120 123 132 125
IQ after training 120 118 125 136 121

Test whether there is any change in IQ after the training programme. (Given t0.01,4=4.6)

Solution: This problem is a case for paired t test as the scores before training (x before) and
after training (x after) are not independent, but the latter to be affected by the former
i.e. they are correlated (dependent).
We take d = (Xafter – Xbefore).

And set up null hypothesis H 0 :  0
d
i.e, H0: There is no change in the I.Q after the training programmed.

against H1 :  0
d
I.Q = x before I.Q = x after d d2
110 120 10 100
120 118 -2 4
123 125 2 4
132 136 4 16
125 121 -4 16
∑𝑑 = 10 ∑𝑑𝟐 =140

∑𝑑 10
𝑑̅ = = = 2,
𝑛 5
(∑𝑑)2 100
∑ 𝑑2 − 140 −
𝑠2 = 𝑛 = 5 = 140 − 20 = 30
𝑛−1 5−1 4
𝑠 = √30 = 5.48
Test Statistic:
𝑑̅ 2 √5
𝑡= 𝑠 = = 2× = 0.82
5.48 5.48
√𝑛 √5
Calculated |𝑡| = 0.82
Degrees of freedom (D.f): n-1 = 5 – 1 = 4.
Tabulated value of t for 4 at 1% l.o.s is 4.6(two-tailed)
∵ Calculated |𝒕| <Tabulated |𝒕|
∴ H0 is rejected. i.e., The course is useful.

16
3. A manufacturer of shock absorbers is comparing the durability of one of his models with
those of his competitors. To conduct the test the manufacturer installed one of his shock
absorbers, also one of his competitor's shock absorbers on each of ten pairs of cars selected
at random. Each of the cars was driven 20,000 miles. Then, each shock absorber was
measured for strength. The results are shown below:
Car Manufacturer’s Shock Competitor’s Shock
1 10.0 9.6
2 11.7 11.9
3 13.7 13.1
4 9.9 9.4
5 9.8 10.0
6 14.4 14.0
7 15.1 14.6
8 10.6 10.8
9 9.8 9.4
10 12.1 12.3

At the 0.1 level of significance, is there any evidence that the manufacturer's shock absorbers
last longer?
Solution:
Null Hypothesis:
The manufacturer’s shock absorbers do not last significantly longer than the competitor’s shock
absorbers.
Alternative Hypothesis:
The manufacturer’s shock absorbers last significantly longer than the competitor’s shock
absorbers.
Car Manufacturer’s Shock Competitor’s Shock Difference d
1 10.0 9.6 0.4
2 11.7 11.9 -0.2
3 13.7 13.1 0.6

d 
4 9.9 9.4 0.5
2
5 9.8 10.0 -0.2 d  0.2
6 14.4 14.0 0.4 n 10
7 15.1 14.6 0.5
8 10.6 10.8 -0.2
Σx 2 
Σx 2 1.5 
22
9 9.8 9.4 0.4 n 10  0.35
s 
10 12.1 12.3 -0.2 n 1 9

d 0.2
t t  1.807
sd 0.35
n 10

Critical region at 1% level of significance t  1.383


Since the computed t  0.82 is less than 1.383, H0 cannot be rejected and we conclude that
The manufacturer’s shock absorbers do not last significantly longer than the competitor’s shock
absorbers.

F-Test:
To test whether if there is any significant difference between two estimates of population
variance (or) To test if the two samples have come from the same population we use F-test.
In this case we set up null hypothesis H0 = 𝝈𝟏 𝟐 = 𝝈𝟐 𝟐 (i.e., The population variances are same).
Under H0 the test statistic is,

17
𝑺𝟏 𝟐 𝑔𝑟𝑒𝑡𝑎𝑒𝑟 𝑣𝑎𝑟𝑖𝑎𝑛𝑐𝑒
𝑭= ( )
𝑺𝟐 𝟐 𝑠𝑚𝑎𝑙𝑙𝑒𝑟 𝑣𝑎𝑟𝑖𝑎𝑛𝑐𝑒
Where
𝟐
̅) 𝟐
∑(𝒙 − 𝒙
𝒔𝟏 = ; 𝒏𝟏 = 𝑭𝒊𝒓𝒔𝒕 𝒔𝒂𝒎𝒑𝒍𝒆 𝒔𝒊𝒛𝒆
𝒏𝟏 − 𝟏
̅ )𝟐
∑(𝒚 − 𝒚
𝒔𝟐 𝟐 = ; 𝒏𝟐 = 𝑺𝒆𝒄𝒐𝒏𝒅 𝒔𝒂𝒎𝒑𝒍𝒆 𝒔𝒊𝒛𝒆
𝒏𝟐 − 𝟏
Note:
1. Always 𝒔𝟏 𝟐 > 𝒔𝟐 𝟐
2. Degrees of freedom are 𝜸𝟏 = 𝒏𝟏 − 𝟏; 𝜸𝟐 = 𝒏𝟐 − 𝟏
3. If sample variance S2 is given we can obtain population variance S2 by using the relation
n S2 = (n-1) S2

Problems:
1. In one sample of 8 observations the sum of the squares of deviations of the sample values
from the sample mean was 84.4 and in another sample of 10 observations it was 102.6. Test
whether this difference is significant at 5% level.
Solution: Given 𝑛1 = 8 ; 𝑛2 = 10
∑(𝒙 − 𝒙 ̅)𝟐 = 84.4, ∑(𝒚 − 𝒚
̅)𝟐 = 102.6
𝟐
∑(𝒙 − 𝒙̅)𝟐 84.4
𝒔𝟏 = = = 12.057
𝒏𝟏 − 𝟏 7
∑(𝒚 − 𝒚 ̅)𝟐 102.6
𝒔𝟐 𝟐 = = = 11.4
𝒏𝟐 − 𝟏 9
Null hypothesis H0: There is no significant difference in the sample variance.
i.e., 𝒔𝟏 𝟐 = 𝒔𝟐 𝟐
𝑺𝟏 𝟐 12.05
𝑭= 𝟐= = 1.057
𝑺𝟐 11.4

Calculated 𝐹 = 1.057
Degrees of freedom (D.f): 𝜸𝟏 = 𝒏𝟏 − 𝟏 = 𝟖 − 𝟏 = 𝟕; 𝜸𝟐 = 𝒏𝟐 − 𝟏 = 𝟏𝟎 − 𝟏 = 𝟗
Tabulated value of 𝐹 for (7, 9) at 5% l.o.s is 3.29.
∵ Calculated 𝐹 < Tabulated 𝐹 i.e., 1.057< 3.29
∴ H0 is accepted.

2. Two independent samples of eight and seven items respectively had the following values of
the variables.
Sample-1: 9 11 13 11 15 9 12 14
Sample-2: 10 12 10 14 9 8 10
Do the two estimates of population variance differ significantly at 5% level of significance.
(A.U APR/MAY 2010)
Solution:
Given 𝒏𝟏 = 8, 𝒏𝟐 = 7
Null Hypothesis: H0: The two estimates of population variance does not differ significantly.
i.e., H0 : 𝝈𝟏 𝟐 = 𝝈𝟐 𝟐
Alternative Hypothesis: H1: 𝝈𝟏 𝟐 ≠ 𝝈𝟐 𝟐

𝒙 𝒙−𝒙 ̅ ̅) 𝟐
(𝒙 − 𝒙 𝒚 𝒚−𝒚 ̅ ̅ )𝟐
(𝒚 − 𝒚
9 9 – 11.75 7.56 10 10 – 10.43 0.18

18
= -2.75 = -0.43
11 -0.75 0.56 12 1.57 2.46
13 1.25 1.56 10 -0.43 0.18
11 -0.75 0.56 14 3.57 12.74

15 3.25 10.56 9 -1.43 2.04

9 -2.75 7.56 8 -2.43 5.90

12 0.25 0.06 10 -0.43 0.18

14 2.25 5.06 -

∑ 𝒙 = 94 33.48 ∑ 𝒚 = 73 23.68
∑ 𝒙 𝟗𝟒
𝟐
̅)𝟐 𝟑𝟑. 𝟒𝟖
∑(𝒙 − 𝒙
̅=
𝒙 = = 𝟏𝟏. 𝟕𝟓 𝒔𝟏 = = = 𝟒. 𝟕𝟖
𝒏𝟏 𝟖 𝒏𝟏 − 𝟏 𝟖−𝟏
∑ 𝒚 𝟕𝟑 ̅)𝟐 𝟐𝟑. 𝟔𝟖
∑(𝒚 − 𝒚
̅=
𝒚 = = 𝟏𝟎. 𝟒𝟑 𝒔𝟐 𝟐 = = = 𝟑. 𝟗𝟒
𝒏𝟐 𝟕 𝒏𝟐 − 𝟏 𝟕−𝟏
Test statistic:
𝒔𝟏 𝟐 𝟒. 𝟕𝟖
𝑭= = = 𝟏. 𝟐𝟏
𝒔𝟐 𝟐 𝟑. 𝟗𝟒
Calculated 𝑭= 𝟏. 𝟐𝟏
Tabulated value of F at 5% l.o.s is 4.21(two-tailed)
∵ Calculated 𝑭 < Tabulated 𝑭(i.e., 𝟏. 𝟐𝟏< 4.21)
∴ H0 is accepted.

3. The nicotine content in milligrams of two samples of tobacco were found to be as follows:

Sample A 24 27 26 21 25 -
Sample B 27 30 28 31 22 36

(i) Can it be said that two samples come from normal populations having the same
mean.
(ii) Can it be said that two samples come from normal populations having the same
Variance.

Solution:
Given 𝒏𝟏 = 5, 𝒏𝟐 = 6
Equality of means will be tested by t-test and the equality of variances will be tested by F-test.
Since F-test assumes equality of variances, we shall apply t-test first.

𝑥 𝑥 − 𝑥̅ (𝑥 − 𝑥̅ )2 𝑦 𝑦 − 𝑦̅ (𝑦 − 𝑦̅)2

24 24 - 24.6 = -0.6 0.36 27 -2 4


27 2.4 5.76 30 1 1
26 1.4 1.96 28 -1 1
21 -3.6 12.96 31 2 4
25 0.4 0.16 22 -7 49
- - - 36 7 49
123 0 21.2 174 0 108

∑ 𝑥 123 ∑ 𝑦 174
𝑥̅ = = = 24.6 𝑦̅ = = = 29
𝑛1 5 𝑛2 6
∑(𝑥 − 𝑥̅ )2 = 21.2

19
∑(𝑦 − 𝑦̅)2 = 108

For t-test:
1 1
𝑠2 = [∑(𝑥 − 𝑥̅ )2 + ∑(𝑦 − 𝑦̅)2 ] = [21.2 + 108] = 14.35
𝑛1 +𝑛2 − 2 5+6−2
𝑠 = √14.35 = 3.78
For F-test:
2
∑(𝑥 − 𝑥̅ )2 21.2
𝑠1 = = = 5.3
𝑛1 − 1 4
2
∑(𝑦 − 𝑦̅)2 108
𝑠2 = = = 21.6
𝑛2 − 1 5
Here 𝑠2 2 > 𝑠1 2

(i) Using students t-test(mean):


Null Hypothesis: H0 : The two samples have been drawn from the normal populations with the
same mean. i.e., μ1 = μ2 .
Alternative Hypothesis: H1 : 𝜇1 ≠ 𝜇2
Test Statistic:
𝑥̅ − 𝑦̅ 24.6 − 29
𝑡= = = −1.92
1 1 1 1
𝑠√𝑛 + 𝑛 3.78√ + 6
1 2 5
Calculated value of |𝑡| = 1.92
Degrees of freedom (D.f): 5+6-2 = 9
Tabulated value of t for 9 D.f at 5% l.o.s is 2.262(two-tailed)
∵ Calculated |𝑡| < Tabulated |𝑡|(i.e., 2.262 < 1.92)
∴ H0 is accepted.

(ii) Using F-test (variance):


Null Hypothesis: H0 : The two samples have been drawn from the normal populations with the
same variance. i.e., H0 : σ1 2 = σ2 2 .
Alternative Hypothesis: H1 : σ1 2 ≠ σ2 2
Test Statistic:
𝑆2 2 21.6
𝐹= 2 = = 4.07
𝑆1 5.3
Calculated value of 𝐹 = 4.07
Degrees of freedom (D.f): 𝜸𝟏 = 𝒏𝟏 − 𝟏 = 𝟔 − 𝟏 = 𝟓; 𝜸𝟐 = 𝒏𝟐 − 𝟏 = 𝟓 − 𝟏 = 𝟒
Tabulated value of 𝐹 for (5, 4) at 5% l.o.s is 6.26.
∵ Calculated 𝐹 < Tabulated 𝐹 i.e., 4.07< 6.26
∴ H0 is accepted.
i. e., The two samples have been drawn from the normal populations with the same variances.

4. Two random samples gave the following results.


Sum of squares of deviations
Sample Size Sample Mean
from the mean
1 10 15 90
2 12 14 108
Test whether the samples come from the same normal population.
Solution: Given 𝒏𝟏 = 10, 𝒏𝟐 = 12
Null Hypothesis: H0: The two samples have been drawn from the same normal population.
i.e. H0: 1   2 and  12   22
Alternative Hypothesis: H1 : 𝜇1 ≠ 𝜇2 and σ1 2 ≠ σ2 2

20
Equality of means will be tested by t-test and the equality of variances will be tested by F-test.
Since t-test assumes equality of variances, we shall apply F-test first.

(i) Using F-test (variance):


Null Hypothesis: H0:  12   22
Alternative Hypothesis: H1 : σ1 2 ≠ σ2 2
Given n1 = 10, n2 = 12, x1  15, x2  14
(x  x1 )  90,  ( x2  x2 )  108
2 2
1

1 90 1 108
S12 
n1  1
 ( x1  x1 ) 2 
9
 10, S 22 
n2  1
 ( x2  x2 ) 2 
11
 9.82

Test statistic:
S12
F = 2  1.018
S2
Calculated F = 1.018.
F0.05 (9, 11) = 2.90
Since calculated F < tabulated F, we accept the null hypothesis H0.
i.e.,  12   22
Since  12   22 , we can apply t-test now.
(ii) Using students t-test(mean):
Null Hypothesis: H0: 1  2
Alternative Hypothesis: H1 : 𝜇1 ≠ 𝜇2

s2 
1
n1  n2  2
 ( x  x )   ( x
1 1
2
2 
 x2 ) 2 
1
20
(90  108)  9.9

Test statistic:
x1  x2 15  14
t=  = 0.742.
1 1 1 1
9.9 (  )
s   
 n1 n2  10 12
d.f = (n1 + n2 - 2) = 20.
Tabulated value of t for 20 d.f at 5% level of significance is 2.086.
Since calculated value of t = 0.74 < the tabulated value of t = 2.086, we accept the null
hypothesis H0: 1  2
Combining (i) and (ii), we conclude that the samples have come from the same normal
population.

5. In comparing the variability of family income in two areas, a survey fielded the following
data.n1 = 100, n2 = 100, s12 = 25, s22 = 10 where n1 and n2 are the sample sizes and s12, s22 are
the sample variance of incomes for the two areas respectively. Assuming that the populations
are normal test the hypothesis H0: σ12 = σ22 against H1: σ12 > σ22 at 5% level of significance.
Solution:
Given n1 = 100, n2 = 100, s12 = 25, s22 = 10
H0: There is no significant difference between the variances.
s12 25
The test statistic is F =   2 .5
s 22 10
Tabulated value of F for (99, 99) d.f is 1.35.
Since calculated F > tabulated F, we reject the null hypothesis H0.
i.e., the variances differ significantly.

Chi Square Distribution (𝝌2 distribution) – Definition:


If Oi (i = 1, 2, 3, …n) are set of observed (experimental) frequencies & E i (i = 1, 2, 3, …n)

21
are the corresponding set of expected (theoretical or hypothetical) frequencies then the
statistic 𝝌2 is defined by,
(Oi – Ei )2
(𝐶ℎ𝑖𝑠𝑞𝑢𝑎𝑟𝑒)  2 = ∑
Ei
and the degrees of freedom of this statistic is ν = n-1.
Note:
1. The number of degrees of freedom is the total number of observations less the number of
independent constraints imposed on the observations.
2. Chi-Square Distribution has only one parameter called the degrees of freedom. (d.o.f).
3. For fitting binomial distribution v = n – 1, Poisson distribution v = n – 2 and Normal distribution
v = n – 3.
4. The shape of chi-square distribution depends on the number of degrees of freedom.
The important uses (applications) of 𝝌2 test are:
(a) As a test for goodness of fit
(b) As a test for independence of attributes
(c) To test the homogeneity of independent estimates of the population variance.

Conditions:
(a) The sample observations should be independent.
(b) The total frequency should be large, say greater than 50.
(c) Constraints of the cell frequencies must be linear.
(d) No theoretical cell frequency should be less than 5.
Method -1:
χ2 test The Goodness of Fit Test:
It is a very powerful test for testing the significance of the discrepancy between theory
and experiment. It helps us to find if the deviation of the experiment from theory is just by
chance or it is really due to the inadequacy of the theory to fit the observed data.
Pearsonian Chi-Square is
(Oi – Ei )2
(𝐶ℎ𝑖𝑠𝑞𝑢𝑎𝑟𝑒)  2 = ∑
Ei

Where Ei = Expected frequency

Oi = Observed frequency

K = Number of Categories

Problems:
1. A survey of 320 families with five children each revealed the following distribution:
No. of Boys 0 1 2 3 4 5
No. of Girls 5 4 3 2 1 0
No. of Families 12 40 88 110 56 14
Is this result consistent with the hypothesis that male and female births are equally
probable? (A.U APR/MAY 2010)
Solution: Given n = 6 (Small sample n < 30).
Null hypothesis H0: Male and female births are equally probable.
𝟏 𝟏 𝟏
p(male birth) = 𝟐 ; q =1 - 𝟐 = 𝟐 (∵ 𝒑 + 𝒒 = 𝟏 𝒊𝒏 𝒑𝒓𝒐𝒃𝒂𝒃𝒊𝒍𝒊𝒕𝒚 𝒕𝒉𝒆𝒐𝒓𝒚)
𝟏 𝟓
Based on H0, the probability that a family of 5 children has r male children = 𝟓𝒄𝒓 (𝟐)
𝟏 𝒓 𝟏 𝟓−𝒓 𝟏 𝒓+𝟓−𝒓
(by binomial distribution req. probability = 𝒏𝒄𝒓 (𝒑)𝒓 (𝒒)𝒏−𝒓 = 𝟓𝒄𝒓 (𝟐) (𝟐) = 𝟓𝒄𝒓 (𝟐) =
𝟏 𝟓
𝟓𝒄𝒓 (𝟐) )
Alternate Hypothesis H1: Male and female births are not equally probable.
𝟏 𝟓 𝟏
∴ Expected number of families having r male children = 320 × 𝟓𝒄𝒓 (𝟐) = 𝟑𝟐𝟎 × 𝟓𝒄𝒓 𝟐𝟓
𝟏
= 𝟑𝟐𝟎 × 𝟓𝒄𝒓 𝟑𝟐 = 𝟏𝟎 × 𝟓𝒄𝒓 (320 = total no. of families).
(Oi – Ei )2
Test statistic =  2 = ∑ Ei

22
(Oi  Ei ) 2
Oi Ei (Oi  Ei ) (Oi  Ei ) 2
Ei
4
12 10 × 5𝑐0 =10 12 – 10 = 2 4 = 0.4
10
40 10 × 5𝑐1 =50 -10 100 2
88 10 × 5𝑐2 =100 -12 144 1.44
110 10 × 5𝑐3 =100 10 100 1
56 10 × 5𝑐4 =50 6 36 .72
14 10 × 5𝑐5 =10 4 16 1.6
(Oi – Ei )2
∑ = 7.16
Ei

𝜸 = 𝒏−𝟏 =𝟔−𝟏 = 𝟓
At 5% l.o.s table value for 5 D.f chi-square distribution is 𝝌𝟎.𝟎𝟓 𝟐 = 𝟏𝟏. 𝟎𝟕
Since calculated value 𝝌𝟐 = 𝟕. 𝟏𝟔 < 𝟏𝟏. 𝟎𝟕, H0 is accepted.

2. A Personal Manager is interested in trying to determine whether absenteeism is greater on


one day of the week than on another. His records for the past year show the following
sample distribution.

Day of Week :Monday Tuesday Wednesday Thursday Friday


No. of absentees :66 56 54 48 75
Test whether the absentee is uniformly distributed over the week.
Solution:
Given n = 5 (Small sample n < 30)
We set up H0 : Number of absence is uniformly distributed over the week against
H1: Number of absence is not uniformly distributed over the week.
The number of absentees during a week are 300 and if absentism is equally probable on all
300
days, then we should expect  60 absentees on each day of the data as follows:
5
Test statistic =  = ∑ i E i
2
(O – E )2
i

(Oi  Ei ) 2
Category Oi Ei (Oi  Ei ) (Oi  Ei ) 2
Ei
Monday 66 60 6 36 0.60
Tuesday 57 60 -3 9 0.15
Wednesday 54 60 -6 36 0.60
Thursday 48 60 -12 144 2.40
Friday 75 60 -15 225 3.75
(Oi – Ei )2
Total ∑ = 7.50
Ei
The critical value of  = 9.49 at   0.05 and d.o.f. = (5-1) =4
2

As calculated value of  = 7.5 is less than its critical value, the null hypothesis is accepted.
2

3. A die tossed 264 times with the following results.


a) Face 1 2 3 4 5 6
Frequency 40 32 28 58 54 52
Show that the die is biased.
Solution:
Given n = 6 (Small sample n < 30)
We set up H0 : The die is unbiased.
H1: The die is biased.

23
264
The expected frequency of each of the numbers  44
6
Test statistic =  = ∑ i E i
2 (O – E )2
i

(Oi  Ei ) 2
Face Oi Ei (Oi  Ei ) (Oi  Ei ) 2
Ei
1 40 44 -4 16 0.36
2 32 44 -12 144 3.27
3 28 44 -16 256 5.82
4 58 44 14 196 4.45
5 54 44 10 100 2.27
6 52 44 8 64 1.45
(Oi – Ei )2
Total ∑ = 17.62
Ei
The critical value of  = 11.07   0.05 and d.o.f. = (6-1) =5
2

As calculated value of  = 17.62 is > critical value, the null hypothesis is rejected.
2

Hence the die is biased.

4. The theory predicts the proportion of beans in the 4 groups A, B, C and D should be 9:3:3:1. In
an experiment among 1600 beans, the numbers in the 4 groups were 882, 313, 287 and 118.
Does the experimental result support the theory? (A.U MAY/JUN 2012)
Solution:
H0: The experimental result support the theory.
i.e., The four categories are in the ratio 9:3:3:1

H1: The experimental result does not support the theory.

The expected frequencies of the four classes are,


9 3 3 1
 1600  900,  1600  300,  1600  300,  1600  100
16 16 16 16
Test statistic =  = ∑ i E i
2 (O – E )2
i

(Oi  Ei ) 2
Group Oi Ei (Oi  Ei ) (Oi  Ei ) 2
Ei
A 882 900 18 324 0.36
B 313 300 13 169 0.56
C 287 300 13 169 0.56
D 118 100 18 324 3.24
1600 (Oi – Ei )2
∑ =4.72
Ei

d.o.f. = (4-1) =3
i.e, Calculated value of 2 = 4.72. Table value = 7.81 at 5% level.
Calculated value 2 = 4.72 < table value, so we accept H0 at 5% level of significance.
The four categories are in the ratio 9:3:3:1.

4. A sample analysis of examination results of 500 students was made. It was found that 220
students have failed, 170 have secured a third class, and 90 have secured a second class and
the rest, a first class. Do these figures support the general belief that the above categories are
in the ratio 4:3:2:1 respectively?
Solution:
Given n = 4 (Small sample n < 30)

24
H0: the results in the four categories are in the ratio 4:3:2:1
H1: the results in the four categories are not in the ratio 4:3:2:1
The expected frequencies of the four classes are,
4 3 2 1
 500  200,  500  150,  500  100,  500  50
10 10 10 10
Test statistic =  = ∑ i E i
2 (O – E )2
i

(Oi  Ei ) 2
Class Oi Ei (Oi  Ei ) (Oi  Ei ) 2
Ei
Failures 220 200 20 400 2.00
III 170 150 20 400 2.67
II 90 100 - 10 100 1.00
I 20 50 -30 900 18.00
500 (Oi – Ei )2
∑ =23.67
Ei

i.e, Calculated value of 2 = 23.67. Table value = 7.815 at 5% level.


Calculated value 2 = 23.67 > table value, so we reject H0 at 5% level of significance. The four
categories are not in the ratio 4:3:2:1

Method -2:
𝝌2 test - for independence of attributes:
Frequencies related to two attributes may be tested to find whether an association
between the attributes exist or not for the two –way table between the attributes the expected
frequencies are calculated for each cell of the observed frequencies. The Chi-square value is
less than evaluated using the formula,

(Oi  Ei ) 2
 
2

Ei

RowTotal  ColumnTotal
Where Ei 
Totalfrequency
The corresponding null hypothesis is H 0 (attributes are independent) against
alternative hypothesis H1 (attributes are not independent)

Rejection Rule: If the calculated value of Chi-Square is greater than or equal to the
tabulated value for (r-1)(c-1) d.o.f. where r = number of rows, and c=number of
columns, then H 0 rejected otherwise accepted.
Note: for 2 x2 tables, the Chi-square formula for test of independence of attributes becomes
N (ad  bc) 2
simplified and can be used directly as  2 
R1 R2 C1C 2
Where given table explains the notation,

Row
2 X 2 table table
a b R1
c d R2
Column C1C2 N = Total Frequency Total
Rejection Rule: The null hypothesis is rejected when the Chi-square value, is equal to or
exceeds the tabulated value for (2-1)(2-1) =1 d.o.f.
Yates’ Correction: For (2 x 2) table, there is always only one degree of freedom. To get better
result of  2 it is necessary to make a correction to the above formula and corrected formula is
given by

25
2
 N
N  ad  bc  
2  
2
R1 R2 C1C2

Problems:
1. Find the value of 2 for the following 22 contingency table:
6 2
3 5

Solution: Here a = 6, b = 2, c = 3, d = 5
2
 N
N  ad  bc  
2   2
(a  b)(a  c)(b  d )(c  d )
and N= a + b + c + d = 16
2 = 1.0159

2. Find the coefficient of attributes for the 22 contingency table


7 6
10 8
Solution:
ad  bc
Coefficient of attributes =
ad  bc
56  60
= = - 0.0345
56  60

3. Given the following contingency table for hair color and eye color, find the value of Chi
square. Is there good association between the two?
(A.U NOV/DEC 2010)

Hair color
Fair Brown Black Total
Blue 15 5 20 40
Eye Grey 20 10 20 50
color Brown 25 15 20 60
Total 60 30 60 150
Solution: Null hypothesis H0: The given attributes are independent.
Alternate Hypothesis H1: The given attributes are not independent.
Test statistic:
(O  Ei ) 2
2   i
Ei
Given,
Hair color
Fair Brown Black Total
Blue 15 5 20 40
Eye Grey 20 10 20 50
color Brown 25 15 20 60
Total 60 30 60 150

𝐑𝐨𝐰 𝐭𝐨𝐭𝐚𝐥×𝐜𝐨𝐥𝐮𝐦𝐧 𝐭𝐨𝐭𝐚𝐥 (Oi  Ei ) 2


Oi Ei = (Oi  Ei ) (Oi  Ei ) 2
𝐆𝐫𝐚𝐧𝐝 𝐭𝐨𝐭𝐚𝐥 Ei
𝟒𝟎×𝟔𝟎 1
15 =16 15 – 16 = -1 1 = 0.06
𝟏𝟓𝟎 16
𝟒𝟎×𝟑𝟎
5 =8 -3 9 1.13
𝟏𝟓𝟎

26
𝟒𝟎×𝟔𝟎
20 =16 4 16 1.00
𝟏𝟓𝟎
𝟓𝟎×𝟔𝟎
20 =20 0 0 0.00
𝟏𝟓𝟎
𝟓𝟎×𝟑𝟎
10 =10 0 0 0.00
𝟏𝟓𝟎
𝟓𝟎×𝟔𝟎
20 =20 0 0 0.00
𝟏𝟓𝟎
𝟔𝟎×𝟔𝟎
25 =24 1 1 0.04
𝟏𝟓𝟎
𝟔𝟎×𝟑𝟎
15 =12 3 9 0.75
𝟏𝟓𝟎
𝟔𝟎×𝟔𝟎
20 =24 -4 16 0.67
𝟏𝟓𝟎
(O – E)2
∑ = 3.65
E

𝐷. 𝑓 = (𝑟 − 1)(𝑠 − 1) = (3 − 1)(3 − 1) = 4
At 5% l.o.s table value for 4 D.f chi-square distribution is 𝜒0.05 2 = 9.48
Since calculated value 𝜒 2 = 3.65 < 9.48, H0 is accepted.

4. The following data are collected on the two characters


Smokers Non Smokers
literates 83 57
illiterates 45 68
Based on this, can you say that there is no relationship between smoking and literacy?
Solution:
Null hypothesis H0: literacy and smoking habit are independent
Alternate Hypothesis H1: literacy and smoking habit are not independent
Test statistic:
(O  Ei ) 2
2   i
Ei

𝐑𝐨𝐰 𝐭𝐨𝐭𝐚𝐥×𝐜𝐨𝐥𝐮𝐦𝐧 𝐭𝐨𝐭𝐚𝐥 (O i  E i ) 2


Oi E I= 𝐆𝐫𝐚𝐧𝐝 𝐭𝐨𝐭𝐚𝐥 Ei
128  140
83 =70.83 2.09
253
125  140
57 = 69.17 2.14
253
128  113
45 = 57.17 2.59
253
125  113
68 = 55.83 2.65
253
9.48

 = 0.05 d.f (r-1)(c-1) = (2-1)(2-1) = 1 x 1 = 1


 Calculated value = 9.48 and 2 1d.f at 5% table value = 3.84
2

2 Calculated value = 9.48 > 2 table value.


∴We reject the null hypothesis.
There is some relationship between literacy and smoking habit.

Home Work:
1. The following table gives the number of aircraft accidents that occur during various days of the
week. Find whether the accidents are uniformly distributed over the weekdays.
Days : Mon Tues Wed Thurs Fri Sat Sun
No of accident: 14 16 8 12 11 9 14

27
[Ans. χ 2 = 4.16 < 12.59, Yes]

2. Three Hundred digits were chosen at random from a set of tables. The frequencies of the
digits were as follows:
Digits 0 1 2 3 4 5 6 7 8 9
Frequency 28 29 33 31 26 35 32 30 31 25
Using χ 2 test assess the hypothesis that the digits were distributed in equal number in the
table. [Ans. Yes, χ 2 = 2.864 <16.92]

3. A die tossed 120 times with the following results.


b) Face 1 2 3 4 5 6 Total
Frequency 30 25 18 10 22 15 120
Test the hypothesis that the dice is unbiased. [Ans. No, χ 2 = 12.90 > 11.07]
c) Face :1 2 3 4 5 6
Observed freq : 15 22 20 14 18 31
2
Can you regard the die as honest. (Givenx = 13.6 at 5% level of significance for 5 d.o.f.

4. From the table given below, test whether the colour of the son’s eyes is associated with
that of father’s eyes.
Eye colour of sons
Not light Light
Eye colour Not light 230 148
of fathers Light 151 471
2
(Given χ = 3.84 at 1df &5% significance)
[Ans. 133.39 > 3.84, there is relation]
5. The following are collected on two characters.
Cinema goers Non cinema goers
Literate 92 48
Illiterate 208 52
Based on this can you conclude that there is a relation between habit of cinema going and
literacy (Given 5% points of χ 2 distribution are 3.841, 5.991 and 9.488 for degrees of freedom 1,
2 and 4 respectively) [Ans. χ 2 = 9.9 > 3.84, Yes there is relation]

6. Calculate the expected frequencies for the following data presuming the two attributes viz. ,
conditions of home and condition of child are independent.
Condition Condition of home
of child Clean Dirty
Clean 70 50
Fairly clean 80 20
Dirty 35 45
[Ans. χ 2 = 25.633 > 5.99, there is an association between the two]

7. In an examination on immunization of cattle from tuberculosis the following result were


obtained.
Affected Unaffected
Inoculated 12 28
Not inoculated 13 7
Examine the effect of vaccine in controlling the incidence of the disease.
[Ans. χ 2 6.729 > 3.84 , Yes.]
8. Two sample polls of votes for two candidates A and B for public office are taken, one from
among residents of rural areas and the other from urban areas. The results are given in the
table. Examine whether the nature of the area is related to voting preference in this election.
Vote for
Area A B Total
Rural 620 380 1000
Urban 550 450 1000
Total 1170 830 2000
2
[Ans. χ =10.089 > 3.84, Nature of area is related to voting]

28
9. Two groups of 100 people each were taken for testing the use of a vaccine.15 persons
contracted the disease out of the inoculated persons while 25 contracted the disease in the
other group. Test the efficiency of vaccine using χ 2 – test.
[Ans. Not effective, χ 2 =3.125 < 3.84]

10. In a survey of 200 boys, of which 75 were intelligent. 40 had skilled fathers, while 85 of the
unintelligent boys had unskilled fathers. Do these data support the hypothesis that skilled
fathers have intelligent boys.
[Ans. Yes, χ 2 = 8.89 > 3.84]

11. Out of 800 persons 25% were literates and 300 had enjoyed T.V. programmes. 30% of those
who had not enjoyed T.V. Programmes were literate. Test at 5% level of significance whether
T.V. Programmes influence literacy.
Degree of freedom : 1 2 3 4
Chi – Square at 5% level : 3.841 5.999 7.815 9.448
[Ans. χ 2 = 17.78 ( Ho is rejected T.V. program significantly influence literacy]

12. Out of 8000 graduates in a town 800 are females. Out of 1600 graduates employees 120 are
[Link] whether there is any distinction made in the appointment on the basis of female
or male criterion.
[Ans. χ 2 = 13.89. Yes distinction is made]

Large sample tests:


The sample size which is greater than or equal to 30 is called as large sample and the test
depending on large sample is called large sample test.
The assumption made while dealing with the problems relating to large samples are
Assumption-1: The random sampling distribution of the statistic is approximately normal.
Assumption-2: Values given by the sample are sufficiently closed to the population value and
can be used on its place for calculating the standard error of the statistic.

Assumptions: z-test
1. The underlying distribution is normal or the Central Limit Theorem can be assumed to hold
2. The sample has been randomly selected and the sample size is ≥ 30.
3. The population standard deviation is known or the sample size is at least 25.

Critical values of Z for selected levels of significance:

Level of significance (𝛼) 0.10 0.05 0.01 0.005

Two Tail Critical values of |Z| 1.645 1.96 2.58 2.81

Right tail Critical value of Z 1.28 1.645 2.33 2.58

Let tail Critical value of Z -1.28 -1.645 -2.33 -2.58

Large sample test for single mean (or) test for significance of single mean:
Test statistic:
x
Z ~ N 0,1

n
Where 𝜎 is the S.D of the population.
Now calculate Z
Find out the tabulated value of Z at  % l.o.s i.e. Z

29
If Z > Z , reject the null hypothesis H0
If Z < Z , accept the null hypothesis H0
Note:
1. If the population standard deviation is unknown then we can use Test statistic, with sample
S.D s,
x
Z ~ N 0,1
s
n
𝜎 𝜎
2. At 5% level 95% confidence limits are 𝑥̅ − 1.96 < 𝜇 < 𝑥̅ + 1.96
√𝑛 √𝑛
𝜎 𝜎
3. At 1% level 99% confidence limits are 𝑥̅ − 2.58 < 𝜇 < 𝑥̅ + 2.58
√𝑛 √𝑛

Problems:
(1) In order to test whether average weekly maintenance cost of a fleet of buses is more than Rs.
500, a random sample of 50 buses was taken. The mean and the standard deviation were
found to be Rs. 508 and Rs. 40 (Use   0.05 )
Solution:
Given n = 50 > 30 (large sample), 𝑥̅ = 508, 𝑠 = 40
Null hypothesis H 0 :  = 500
Alternate Hypothesis H1: µ > 500.
The significance level is 0.05 and H1 is a right tailed.
The test criterion is the Z test.
Test statistic:
x   508  500 8
Z    1.414
s 40 5.657
n 50
|𝒛| = 𝟏. 𝟒𝟏𝟒
At 𝛼 = 0.05, the critical value of Z = 1.645 (i.e. right tailed) and
The computed value of Z (1.414) < critical value of Z (1.645)
∴ we accept null hypothesis H0 that μ = 500. i.e., The average maintenance cost is not more
than Rs. 500.

(2) The quality control department as processing company specifies that the mean net weight
per pack as its produce must be 20 gms. Experience has shown that the weight are
approximately normally distributed with a S-d, of 1.5 gms. A random sample of 50 packs yield
a mean weight if 19.5 gms as this sufficient evidence to indicate that the true mean weight of
the packs has decreased (use 5% significance levels).
Solution:
Given n = 50 > 30 (large sample), 𝑥̅ = 19.5, 𝜎 = 1.5
Null hypothesis H 0 :  = 20
Alternate Hypothesis H1: µ < 20.
The significance level is 0.05 and H1 is a left tailed.
The test criterion is the Z test.

Test statistic:
x   19.5  20  0.5
Z    2.358
 1.5 0.212
n 50
|𝒛| = 𝟐. 𝟑𝟓𝟖
At 𝛼 = 0.05 level of significance the critical value of Z (Left tailed) is -1.645.
The computed value of Z (2.358) > critical value of Z (-1.645)
∴ we reject null hypothesis H0 that μ = 20.

30
(3) A company is engaged in the packaging of a suppressing quantity tea in jar of 500 gm each.
The company is of the view that as long as jars contains 500 gms of tea, the process is in
control. The standard deviation is 50 gm. A sample of 225 jars is taken at random and the
sample average is found to be 510 gm. Has the process gone out of control at 5% l.o.s.
Solution
Given n = 225 > 30 (large sample), 𝑥̅ = 510, 𝜎 = 50
Null hypothesis H 0 :  = 500 gms
Alternate Hypothesis H1: µ ≠ 500 gms
The significance level is 0.05 and H1 is a two tailed.
The test criterion is the Z test.
Test statistic:
x   510  500 10
Z   3
 50 3.33
n 225
|𝒛| = 𝟑
At 𝛼 = 0.05 level of significance the critical value of Z (two tailed) is 1.96.
The computed value of Z (3) > critical value of Z (1.96)
∴ we reject null hypothesis H0 that μ = 500 gms.
This means process is not in control.

4. The average (mean) live weight of a farmer’s steers prior to slaughter was 390 pounds in past
years. This year his 50 steers were fed on a new diet. Suppose we consider these 50 steers on
the new diet as a random sample taken from a population of all possible steers that may be
fed the diet now or in the future and it S.D is give by 35.2. Use the sample data given below
and α =.01 to test the research hypothesis that the mean live weight for steers on the new
diet is greater than 380.
Solution: Given x =390, n = 50(large sample), s = 35.2.
Null hypothesis H0: µ = 380
Alternate Hypothesis H1: µ > 380
Test Statistic:
x   390  380
z=   2.01
s 35.2
n 50
|𝒛| = 𝟐. 𝟎𝟏
At 𝛼 = 0.01 level of significance the critical value of Z (right tailed) is 2.33.
The computed value of Z (2.01) > critical value of Z (2.33)
∴ we accept null hypothesis H0 that µ = 380
There is not sufficient evidence to conclude that the mean live weight for steers on the new
diet is greater than 380.

5. Chennai Municipality uses thousands of fluorescent light bulbs each year. The brand of bulb it
currently uses has a mean life of 900 hours. A manufacturer claims that its new brand of
bulbs, which cost the same as the brand the university currently uses, has a mean life of more
than 900 hours. The university has decided to purchase the new brand if, when tested, the
test evidence supports the manufacturer’s claim at α = .05. Suppose sixty-four bulbs were
tested with the following results:
x = 930 hours s = 80 hours

Will Chennai Municipality purchase the new brand of fluorescent bulbs? Conduct
hypothesis test.

Solution: Given x =930 hours, n = 64(large sample), s = 80 hours

31
Null hypothesis H0: µ = 900
Alternate Hypothesis H1: µ > 900 (the mean life for the new brand of bulbs is higher than the
mean life for the old brand)
Test Statistic:
x   930  900 30
z=    3.00
s 80 10
n 64
|𝒛| = 𝟑. 𝟎𝟎
At 𝛼 = 0.05 level of significance the critical value of Z (right tailed) is 1.645.
The computed value of Z (3.00) > critical value of Z (1.645)
∴ we reject null hypothesis H0 that µ = 900.
There is sufficient evidence to conclude that the mean life for the new brand of bulbs is greater
than 900.

Large sample test for difference between two means:


Test statistic
( x1  x2 )
Z ~ N 0,1
1 2 2 2

n1 n2
Now calculate Z
Find out the tabulated value of Z at  % l.o.s i.e. Z
If Z > Z , reject the null hypothesis H0
If Z < Z , accept the null hypothesis H0
Note:
1. If  1 2 and  2 2 are unknown then we can consider S1 2 and S 2 2 as the estimate value of  1 2
and  2 2 respectively.
2. Under H0: µ1 = µ2 , if the samples are drawn from the same population where
̅𝑥̅̅1̅− ̅̅̅̅
𝑥2
𝜎1 = 𝜎2 = 𝜎, then 𝑍 = 1 1
𝜎√ +
𝑛1 𝑛2
̅𝑥̅̅1̅− ̅̅̅̅
𝑥2
4. If 𝜎1 , 𝜎2 are not known and 𝜎1 ≠ 𝜎2 , then 𝑍 =
2 2
√𝑆1 +𝑆2
𝑛1 𝑛2

Problem:
1. The means of 2 large samples 1000 & 2000 members are 67.5 inches & 68.0 inches
respectively. Can the samples be regarded as drawn from the same population of standard
deviation 2.5 inches?
(A.U NOV/DEC 2010, A.U MAY/JUNE 2012)
Solution:
Given 𝑛1 = 1000 ; 𝑛2 = 2000 (large sample test)
𝑥1 = 67.5 𝑖𝑛𝑐ℎ𝑒𝑠; ̅̅̅̅
̅̅̅ 𝑥2 = 68 𝑖𝑛𝑐ℎ𝑒𝑠
σ = 2.5 𝑖𝑛𝑐ℎ𝑒𝑠
Null hypothesis H0: The samples have been drawn from the same population of S.D 2.5 inches.
i.e., H0: µ1 = µ2 and σ = 2.5 𝑖𝑛𝑐ℎ𝑒𝑠
Alternate Hypothesis H1: µ1 ≠ µ2
Test Statistic:
𝑥1 − ̅̅̅̅
̅̅̅ 𝑥2 67.5 − 68 −0.5
𝑍= = = = −5.16
σ2 σ2 (2.5) 2 (2.5) 2 0.0968
√ + √
𝑛1 𝑛2 1000 + 2000
Calculated |𝑍| = 5.16
Tabulated value of Z at 5% l.o.s is 1.96.

32
∵ Calculated |𝑍| > Tabulated |𝑍|(i.e., 5.16> 1.96.)
∴ H0 is rejected.
𝑖. 𝑒., The samples are not drawn from the same population of S.D 2.5 inches.

2. The mean yield of two sets of plots and their variability are as given below.
Set of 40 plots Set of 60 plots
Mean yield per plot 1258 Kg 1243 Kg
Standard deviation per plot 34 28
Examine at 5% level whether the difference in mean yields of the two sets of plots is
significant.
(A.U MAY/JUNE-2010)
Solution:
Given 𝑛1 = 40 ; 𝑛2 = 60 (large sample test)
𝑥1 = 1258 𝑘𝑔𝑠; ̅̅̅̅
̅̅̅ 𝑥2 = 1243 𝑘𝑔𝑠
s1 = 34, s2 = 28
Null hypothesis H0: There is no significant difference in mean yields of two sets of plots.
i.e., H0: ̅̅̅
𝑥1 = ̅̅̅̅
𝑥2
Alternate Hypothesis H1: ̅̅̅ 𝑥1 ≠ ̅̅̅̅
𝑥2 (two tailed)
Test Statistic:
𝑥1 − ̅̅̅̅
̅̅̅ 𝑥2 1258 − 1243 15 15 15
𝑍= = = = = = 2.31
2 2 (34) 2 (28) 2 √ 28.9 + 13.067 √41.967 6.48

√𝑆1 + 𝑆2 40 + 60
𝑛1 𝑛2
Calculated |𝑍| = 2.31
Tabulated value of Z at 5% l.o.s is 1.96.
∵ Calculated |𝑍| > Tabulated |𝑍|(i.e., 2.31> 1.96.)
∴ H0 is rejected.

Large sample test for single proportion (or) test for significance of proportion:
𝑝−𝑃
Test statistic 𝑧 = 𝑃𝑄

𝑛

Now calculate Z
Find out the tabulated value of Z at  % l.o.s i.e. Z
If Z > Z , reject the null hypothesis H0
If Z < Z , accept the null hypothesis H0
Large samples
Test of significance for single proportion:
1. In a big city 325 men out of 600 men were found to be smokers. Does this information
support the conclusion that the majority of men in this city are smokers.
(A.U MAY/JUNE 2011,NOV/DEC 2010)
Solution: Given n = 600 (large sample).
𝒙 = Number of smokers = 325
𝑝 = Proportion of smokers in the sample.
𝒙 𝟑𝟐𝟓
= 𝒏 = 𝟔𝟎𝟎 = 0.5417
𝑷 = Proportion of smokers in the Population.
𝟏
= 𝟐 = 0.5
𝑸 = 1 − 𝑃 = 1 − 0.5 = 0.5
Null Hypothesis: H0: The number of smokers and non-smokers are equal in the city.
Alternative Hypothesis: H1: P > 𝟎. 𝟓 (Right tailed)
Test statistic:

33
𝑝−𝑃 0.5417 − 0.5
𝑧= = = 2.04
√𝑃𝑄 √(0.5)(0.5)
𝑛 600
Calculated |𝑍| = 2.04
Tabulated value of Z at 5% l.o.s is 1.64(Right-tailed)
∵ Calculated |𝑍| > Tabulated |𝑍|(i.e., 2.04> 1.64)
∴ H0 is rejected.
𝑖. 𝑒., The majority of men in this city are smokers.

2. Experience has shown that 20% of a manufactured product is of top quality. In one day’s
production of 400 articles, only 50 are of top quality. Show that either the production of the
day chosen was not a representative sample (or) the hypothesis of 20% was wrong.
(A.U APR/MAY 2010)
Solution: Given n = 400 (large sample).
𝑝 = Proportion of top quality products in the sample.
𝟓𝟎 𝟏
= 𝟒𝟎𝟎 = 𝟖
𝑷 = Proportion of top quality products in the Population.
𝟐𝟎 𝟏
= 𝟏𝟎𝟎 = 𝟓
1 4
𝑸=1−𝑃 =1−5=5
1
Null Hypothesis: H0: Proportion of top quality products in the Population P = 5
1
Alternative Hypothesis: H1: P ≠ 5
Test statistic:
1 1 3
𝑝−𝑃 − −
𝑧= = 8 5 = 40 = −3.75
√ 𝑃𝑄 1 4 √ 4
𝑛 √(5) (5) 25 × 400
400
Calculated |𝑍| = 3.75
Tabulated value of Z at 5% l.o.s is 1.96(two-tailed)
∵ Calculated |𝑍| > Tabulated |𝑍|(i.e., 3.75 > 1.96)
∴ H0 is rejected.

3. The product manager wises to determine whether or not to change the package design for
her product. She feels that it will be worth considering only if more than 60% of the non-users
prefer the new box to the old one. She selects or random sample of 100 persons who are non-
users and finds that 73 persons prefer the new box. Should she change the design?
(Significance level  = 0.05)
Solution: Given n = 100 (large sample).
73
𝑝= = 0.73
100
𝟔𝟎
𝑷 = 𝟏𝟎𝟎 = 0.6
𝑸 = 1 − 𝑃 = 1 − 0.6 = 0.4
Null Hypothesis: H0: P = 0.6.
Alternative Hypothesis: H1 : P > 0.6 (right-tailed)
Test statistic:
𝑝 − 𝑃 0.73 − 0.6 0.13
𝑧= = = = 2.65
√ 𝑃𝑄 √ 0.6 × 0.4 √ 0.24
𝑛 100 100
Calculated |𝑍| = 2.65
At  = 0.05, Z0.05 = 1.645 as the test is right-tailed.

34
As computed |𝑍| > Tabulated value as Z (= 1.645), Null hypothesis is rejected
∴ we can conclude that the sample gives sufficient evidence that more than 60% of the non-
users prefer the new box design and the product manger should change the package design of
her product at the specified significance level.

Large sample test for Difference of proportions (or) test for significance of proportion:
p1  p 2
Test statistic Z  ~ N 0,1
1 1
PQ(  )
n1 n2
n1 p1  n2 p2
When 𝑃 is not known 𝑃 can be calculated by P  and Q  1  P
n1  n2
Now calculate Z
Find out the tabulated value of Z at  % l.o.s i.e. Z
If Z > Z , reject the null hypothesis H0
If Z < Z , accept the null hypothesis H0

Problem:
1. A machine puts out 10 defective units in a sample of 200 units. After the overhauling the
machine puts out 4 defective units in a sample of 100 units. Has the machine been improved.
(use 𝜶 = 𝟎. 𝟎𝟓)
Solution:
10 4
Given n1  200, p1   0.05 , n2  100, p2   0.04
200 100
Null Hypothesis: H 0 : p1  p2
i.e., The proportions of defective before and after overhauling are equal.
Alternative Hypothesis: H1 : p1  p2
i.e., The proportion of defectives has decreased after overhauling.
n p  n2 p 2 10  4
Proportion P  1 1   0.047
n1  n2 200  100
𝑄 = 1 – 𝑝 = 0.953
Test statistic:
p1  p2 0.05  0.04
Z=   0.385
1 1
PQ(  )  1 1 
0.047  0.953  
n1 n2  200 100 
Calculated |𝑍| = 0.385
At  = 0.05, Z0.05 = 1.645 as the test is right-tailed.
Calculated Value of | z | = 0.385 < table value = 1.645
Since | z | < 1.645, we accept the null hypothesis H0 at 5% level of significance.
i.e., The proportions of defective before and after overhauling are equal.

3. In a large city A, 20% of a random sample of 900 school boys had a slight physical defect. In
another large city B, 18.5% of a random sample of 1600 school boys had the same defect. Is
the difference between the proportions significant?

Solution:
Given n1 = 900, n2 = 1600,p1 = 0.2, p2 = 0.185
Null Hypothesis: H 0 : p1  p2
i.e., The differences between the two proportions are not significant

35
Alternative Hypothesis: H1 : p1  p2
i.e., The differences between the two proportions are significant.
n p  n2 p 2
Proportion P = 1 1 = 0.1904.
n1  n2
Q = 1 – P = 0.8906.
Test statistic:
p1  p2 0.2  0.185
Z=   0.9375
1 1
PQ(  )  1 1 
0.1904  0.8906   
n1 n2  900 1600 
Calculated |𝑍| = 0.9375
At  = 0.05, Z0.05 = 1.645 as the test is right-tailed.
Calculated Value of | z | = 0.9375 < table value = 1.645
Since | z | < 1.645, we accept the null hypothesis H0 at 5% level of significance.
i.e., the differences between the two proportions are not significant.

36
Solved Two Mark Questions:
TESTING OF HYPOTHESIS
PART-A
1. Define Type I and Type II errors.
Solution:
Type I error: Reject H0 when it is true
Type II error: Accept H0 when it is false
2. If we want to estimate the true proportion of defectives in a very large shipment of adobe
bricks, and that we want to be at least 95% confident that the error is at most 0.04. How large
a sample will we need if we know that the true proportion does not exceed 0.12?
Solution:
Z0.05 = 1.96, E = 0.04 and P = 0.12
( PQ)
n  z2
E2
=253.55
 254
3. In a sample of 100 ceramic pistons made for an experimental diesel engine, 18 were cracked.
Construct a 95% confidence interval for the true proportion of cracked pistons.
Solution:
95% confidence interval for the true proportion is
 pq pq 
 p  z 0.05 , p  z 
 n
0.05
n 
 
 0.18  0.82 0.18  0.82 
 0.18  1.96  ,0 .18  1.96  
 100 100 
 
i.e., (0.1047, 0.2553)
4. The length of certain machine parts looked upon as a random variable having a normal
distribution with a mean of 2cm and a standard deviation of 0.05cm. We want to test the null
hypothesis  = 2 against the alternative hypothesis   2 on the basis of a random sample of
size n = 30 and mean 2.01. If the probability of a Type I error is to be  = 0.05, find whether
the sample mean differ significantly from the population mean.
Solution:
x   2.01  2.00
z   1.095
 0.05
n 30
Since   0.05, z  1.96 .
z  z . Therefore the null hypothesis   2 is accepted. There is no significant
difference between the sample mean and population mean.
5. In a random sample of 1000 people in Maharashtra, 540 are rice eaters and the rest are
wheat eaters. If both rice and wheat are equally popular in the state, find the standard error
of the proportion of wheat eaters.
Solution:
n = 1000, P = ½ and Q = ½
PQ 0.5  0.5
S .E.   = 0.0138
n 1000
6. Write short notes on critical region.
Solution:
A region in the sample space which amounts to the rejection of H0 is known
as the critical region or region of rejection.
7. Distinguish between parameters and statistics.
Solution: Statistical constants of the population namely mean(), variance(2), etc., are usually
referred as parameters. Statistical measures computed from the sample observations alone are
known as Statistic.(eg.)Mean(x), variance(s2), etc.,
8. Define degrees of freedom
Solution:
Number of independent variables which make up the statistic is known as degrees of freedom.

37
9. Explain level of significance
Solution:
The probability ‘’ that a random value of the statistic t belongs to the critical region is
known as the level of significance. ie., level of significance is the size of the Type I error or
maximum producer’s risk. Usually 5% and 1% levels of significance are employed in testing of
hypothesis.
If E(X) be the expected value of X and Z  X  E ( X )
S .E ( X )
If P  Z  1.96   0.05 , we say that H0 is rejected at 5% level of significance.
10. Define critical value
Solution:
The value of the test statistic which separates the critical region and the
acceptance region is called the critical value or significant value. This value is
dependent on (i) the level of significance and (ii) the alternative hypothesis,
whether it is one –tailed or two-tailed.
11. What do you mean by Confidence level and confidence limits
Solution:
If  is the probability level and if the estimated value of a statistic is  and if
P(c1<<c2) =1- then c1 and c2 are called the confidence limits and the interval
(c1, c2) is known as confidence interval at the probability level .
12. A normal population has a mean 0.1 and s.d. 2.1. find the probability that mean of a sample
of size 900 will be negative.
Solution:
 x o 
P x  0  P  
 / n  / n 
  
 P z  
 / n
Given  = 0.1,  = 2.1, n = 900
 
 = Pz  1.43 = 0.0764
0.1
P  x  0  P  z  
 2.1/ 900 
13. In a recent study, 69 of 120 meteorites were observed to enter the earth’s atmosphere with a
velocity of less than 26 miles per second. If we estimate the corresponding true proportion as
o.575, what can we say with 95% confidence about the maximum error?
Solution:
PQ
E  z
n
0.575  0.425
 1.96  = 0.0884
120
14. A random sample of 500 toys was taken from a large consignment and 65 were found to be
defective. Find the percentage of defective toys in the consignment.
Solution:
n= 500 p = 65/500 = 0.13 and q = 1- p = 0.87
Confidence interval for proportion of defective toys is
 pq pq 
 p 3 , p  3 
 n n 
 
(0.0849, 0.1751)
Percentage of defective toys in the consignment lies between 8.5 and 17.51.
15. Test the hypothesis that  = 10 given that s = 15 for a random sample of size 50 from a normal
population.
Solution:
H0:  = 10
Given that n = 50 and s = 15
s  15  10
z = = 5.05
 10
2n 100

38
z0.05 = 1.96
z > z0.05. Hence, H0 is rejected.
16. What are the uses of ‘F’ – test?
Solution:
(i) F - test is used to test whether two independent samples have been drawn from the normal
populations with the same variance
(ii) F – test is used to test whether the two independent estimates of the population variance are
homogeneous or not.
21. Write the applications of ‘2’ test.
Solution:
2 - test is used
(i) to test the goodness of fit.
(ii) to test the independence of attributes.
(iii) to test the homogeneity of independent estimates of the population variance.
17. Define errors in sampling and critical region.
Solution:
Type I error: Reject H0 when it is true
Type II error: Accept H0 when it is false
A region in the sample space which amounts to the rejection of H0 is known as the critical
region or region of rejection.
18. Define null hypothesis and alternative hypothesis.
Solution:
Null hypothesis:
It is a definite statement about the parameter, that there is no difference. It is denoted by H 0
Alternative hypothesis:
Complementary hypothesis to null hypothesis is called the alternative hypothesis & is denoted by
H1.
19. What are the conditions under which chi-square test is valid?
Solution:
(i). The sample observations should be independent.
(ii). Constraints of the cell frequencies must be linear.
(iii) No theoretical cell frequency should be less than 5.

PART-B
1. A cubical die is thrown 9000 times and a throw of three or four is observed 3240 times. Show
that the die cannot be regarded as an unbiased one and find the extreme limits between
which the probability of a throw of three or four lies.
Solution:
1
H0: The die is unbiased. i.e., P = (= the probability of getting 3 or 4)
3
1
H1: P  Two tailed test is used. Level of significance α = 5 % Zα = 1.96
3
 1
3240   9000 x 
X  np  3
Z   5.37
nPQ 1 2
9000 x x
3 3
Z  5.36 > Z = 1.96. We reject H0
1
We concluded that the dice is almost certainly biased. P 
3
3240
P=  0.36 and Q  1 - P  1 - 0.36  0.64
9000
Hence the probable limits for the population proportions of successes may be taken by
 0.36x0.64
P  PQ n  0.36   0.360  0.015 = 0.345 and 0.375
9000
Hence the probability of getting 3 or 4 almost certainly lies between 0.345 and 0. 375.

39

You might also like