0% found this document useful (0 votes)
3 views123 pages

Module 6 (Full)

Uploaded by

sambit.talcher
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views123 pages

Module 6 (Full)

Uploaded by

sambit.talcher
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

MODULE 6

Small sample tests

 Student’s t-test
 F-test
 Chi-square test
Goodness of fit
Independence of attributes
 Design of Experiments - CRD- RBD- LSD
 Analysis of variance
One and Two way classifications
Small Sample Tests
In this study, we discuss test of significance for small samples. The important tests for
small samples are
• t-test
• F-test
Note:
Large Sample: Size of the sample is greater than or equal to 30.
Small Sample: Size of the sample is less than 30.
T-test
When the population standard deviation is not known and the size of the sample is
less than thirty, we use t-test.

t-distribution is also known as “students t-distribution”.

Let be the members of a random sample drawn from a


normal population with mean and variance . We define the test statistic as

(Note: we are using n-1)


2
i
where

2
i
Or one can write , where S=
(students t)
A random variable X is said to follow t−distribution, if its probability density function is
given by

where v is known as the degrees of freedom and k is constant.


The constant value k is chosen in such a way that
After simplification, we get

Note on Degree of Freedom


The number of degrees of freedom can be interpreted as the number of useful bits of information
generated by a sample of given size for estimating a population parameter. Suppose we wish to find the
mean of a sample with observations 1 2 n . We have to use all the values taken by the variable
with full freedom for computing . Hence is said to have degrees of freedom.
Assumption of t-distribution:
1) The population from which the sample is drawn is normal.
2) The sample is random and size n < 30.
3) The population S.D. σ is not known.

Properties of t-distribution:
1) The probability curve of t-distribution is symmetrical.
2) The tails of the curve are asymptotic to x-axis.
3) When n→ ∞, t- distribution tends to normal distribution.
4) The form of the t-dist. varies with the degrees of freedom.

Application of t- distribution: The t- distribution is used to test the significance of


the difference between
1) The mean of a small sample and the mean of the population
2) The means of two small samples.
Test of significance of the difference between single sample
mean and population mean
To test the significance of a mean of a small sample the test statistic is

where standard deviation = s =


𝑠/ 𝑛 = standard error
of the mean
The degrees of freedom is

Then, find out tabular value based on the level of significance and degree of
freedom. To find value follow -distribution table. Compare calculated value
with tabular value and then conclude the result.
Computation of t value by using deviation method:

Consider the deviation Then,

Where

Conclusions (two-tailed):
If the calculated value of |t| is less than the table value , then we accept null
hypothesis for the degrees of freedom and level of significance.
If the calculated value of |t| is greater than the table value , then we reject null
hypothesis for the degrees of freedom and level of significance.

Interval estimate of population mean:


To find confidence interval or critical region for single mean of a small sample we
need critical points. The critical points can be calculated by the formula:
i.e., Estimate±(t critical)×(Standard Error)
Problem 1: A random sample of size 7 from a normal population gave a mean of
977.51 and a standard deviation of 4.42. Find 95% confidence interval for the
population mean.
Problem 2: A random blood sample for test of fasting sugar for 10 boys gave the following
data in mg/dl:

70, 120, 110, 101, 88, 83, 95, 107, 100, 98.
Does this support the assumption of population mean of 100 mg/dl? Find a reasonable
range in which most of the mean fasting sugar test of the 10 boys lie.
Test statistic is

that is, calculated value of t at 5% L.S for 9 d.f is less


than the table value.
Hence we accept null hypothesis.

The 95% confidence limit for are


Test of significance of the difference means of two small samples drawn from the
normal populations having different mean values

Let and be the sample means for two small samples drawn from the normal
population with means and . The test statistics is

or (For pooled)

and the degrees of freedom is

Critical points: Where 2


Test of significance of the difference between means of two small samples drawn
from the same normal population

Let and be the sample means for two small samples drawn from a normal
population. The test statistics is

The degrees of freedom is


Que: Two types of batteries are tested for their length of life (in hours). The following
data is the summary descriptive statistics

Is there any significant difference between the average life of the two batteries at
5% level of significance?
Under null hypotheses 0 :

where is the pooled standard deviation given by,

2 2
X F
p

2 2

The value of is calculated for the given information as

p
Problem 3: The nicotine content in milligrams of the samples of tobacco were found as
follows
Sample A: 24, 27, 26, 21, 25.
Sample B: 27, 30, 28, 31, 22, 36.
Can it be said that the two samples come from normal populations with the same mean?
Paired T-Test
If 𝑛1 = 𝑛2 = 𝑛 and the pairs of values of 𝑥1 and 𝑥2 are associated (correlated),
then the usual two-sample t-test should not be used.

Instead, we test the hypothesis:


𝐻0 : 𝑑ҧ = 𝑥ҧ − 𝑦ത = 0
meaning the average difference between the paired observations is zero, i.e.,
there is no significant difference.
Test Statistic Standard deviation of the differences

𝒅
t= 𝒔/ 𝒏−𝟏

Where 𝑑𝑖 = 𝑥𝑖 − 𝑦𝑖 𝑖 = 1 2…𝑛
Mean difference: 𝑑ҧ = 𝑥ҧ − 𝑦,
ത with 𝑣 = 𝑛 − 1
Example
The following data relate to the marks obtained by 11 students in two tests; one
held at the beginning of the year and the other at the end of the year after
intensive coaching. Do the data indicate that the students have benefited by
coaching?

Test 1 (Beginning of the year): 19, 23, 16, 24, 17, 18, 20, 18, 21, 19, 20

Test 2 (End of the year): 17, 24, 20, 24, 20, 22, 20, 20, 18, 22, 19
Solution:
|Test 1|Test 2| (di = xi-yi) |
| 19 | 17 | 2 |
| 23 | 24 | -1 |
| 16 | 20 | -4 |
| 24 | 24 | 0 |
| 17 | 20 | -3 |
| 18 | 22 | -4 |
| 20 | 20 | 0 |
| 18 | 20 | -2 |
| 21 | 18 | 3 |
| 19 | 22 | -3 |
| 20 | 19 | 1 |
Problem
Memory capacity of 9 students was tested before and after a course of
meditation for one month. State whether the course was effective or not.
Before After
10 12
15 17
9 8
3 5
7 6 ?
12 11
16 18
17 20
4 3
F-test
A random variable X is said to follow F distribution, if its probability density function is
given by

where and are the degrees of freedom of samples.

Use of F distribution:
F distribution is used to test the equality of the variance of the populations from
which two small samples have been drawn.
F test of significance of the difference between population variances
To test the significance of the difference between population variances, we shall first
find their estimates, and based on the sample variances and . We compute
estimates by the following formulas:

σ 𝑥 − 𝑥ҧ 2
2
𝑆 =
𝑛−1

2 2
2 1
Test statistics is 2 , where or 2 , where
1 2

The value of F is greater than 1. Then, compare calculated value F with the table value
Fv1,v2 (α) (in case of ) or Fv2,v1 (α) (in case of ) by using F−table.
If the calculated value F is less than the table value, then fail to reject null hypothesis.
σ 𝑥 − 𝑥ҧ 2
2
𝑆 =
4.95 𝑛−1
Note: After computing two variance estimates and for two different
samples our next job is to compute F-test statistic value. If you look at the F-test
statistic formula and it’s provided condition always in numerator we are taking
highest value and in denominator lowest value of estimates.

Since the calculated value of F is less than the table value of F for [Link] (5, 6) at 5%
L.S we accept H0.
Example: A sample of size 13 gave an estimated population variance of 3.0, while
another sample of size 15 gave an estimate of 2.5. Could both samples be from
populations with the same variance?
-Test
In this study, we introduce chi-square distribution, the measure of which enables us
to find the degree of discrepancy between the observed and expected frequencies is
due to error of sampling or due to chance.
• It is a square of standard normal variate, i.e., z = x-μ/ σ

• The chi-square is denoted by the symbol . It is always positive. The value of


chi-square lies between 0 and ∞.
• Since chi-square is not derived from the observation in a population, it is not a
parameter. The chi-square test is not a parametric test.

• Chi-square is computed on the basis of frequencies in a sample and the


value of chi-square so obtained is a statistic.
χ2-Test is defined as

where
Oi = Observed frequencies &
Ei = Expected frequencies.

Note: It works with categorical data (frequencies), not raw continuous values
-Distribution
Let samples of size n be drawn from a normal population with standard deviation σ. If
for each sample we calculate χ2 a sampling distribution of χ2 can be obtained. It is
given by

where 0 < χ2 < ∞ and v is the number of degrees of freedom.


Conditions for using χ2-Test

• Each of the observations making up the sample for the χ2 test should be independent
of each other.
• The test is wholly dependent on the degrees of freedom.
• The frequencies used in a χ2 test should be absolute and not relative in terms.
• The expected frequency of any item or cell should not be less than 5. If it is less
than 5, then the frequencies from the adjacent items or cells should be pooled
together in order to make it 5 or more than 5
• The observations collected for χ2 must be based on the method of random sampling.

• σ 𝑂𝑖 = σ 𝐸𝑖 = 𝑁
Uses of χ2-Test

χ2-test is an important test. We require only the degrees of freedom for using
this test. It is used as a test of –

• Goodness of fit,
• Independence of attributes, and
• homogeneity. (This concept is not in your syllabus.)
χ2-test as a test of Goodness of fit
χ2-test is applied as a test of goodness of fit to determine whether the actual (i.e.
observed) frequencies are close to the expected (i.e. theoretical) frequencies. The
degrees of freedom in this case are v = n − 1, where n is the number of observations.
Problem 1: In 90 throws of a die, face 1 turned 9 times, face 2 or 3 turned 27 times,
face 4 or 5 turned 36 times and 6 turned 18 times. Test at 10% level if the die is honest.

Expected frequency (E) = Np


χ2-test as a test of goodness of fit (Conti...)
2
i i
Face turned Observed Oi Expected Ei
i

1 9 15 2.4
2 or 3 27 30 0.3
4 or 5 36 30 1.2
6 18 15 0.6
Total 90 90 4.5

χ2
Since the calculated value of χ2 = 4.5 is less than the table value at 10% level of
significance and for 3 degrees of freedom (i.e., 6.25) , we accept the null
hypothesis, and conclude that the die honest/ unbiased.
Problem 2: A sample analysis of examination results of 500 students was made. It was
found that 230 students had failed. 160 had secured a third class, 80 were placed in
second class and 30 got first class. Do these figures commensurate with the general
examination results which is in the ratio 4 : 3 : 2 : 1 for various categories
respectively?

Solution: H0 : The observed results commensurate with the general examination results
H1 : It is not true that the observed results commensurate with the general examination
results.
Level of significance=5%. Degrees of freedom=4 − 1 = 3.
The total frequency=N = 500.

Table value χ2 for 3 df at 5% level of significance = 7.815.


Dividing 500 in the ratio 4 : 3 : 2 : 1, we get 200, 150, 100 and 50.
Therefore, the expected frequencies are 200, 150, 100 and 50 correspond-
ing to the observed frequencies 230, 160, 80 and 30.

2
i i
Class/division Observed Oi Expected Ei i
Failed 230 200 4.500
Third 160 150 0.666
Second 80 100 4.000
First 30 50 8.000

χ2

Since the calculated value of χ2 is (17.066) greater than the table value of χ2
(7.815), the null hypothesis is rejected.
Problem 3:The table below gives the number of aircraft accidents that occurred during the
various days of the week. Test whether the accidents are uniformly distributed over the week.

Days Mon Tue Wed Thu Fri Sat


No. of accidents 14 18 12 11 15 14
χ2-test as a test of goodness of fit (Conti...)

Problem 4: Four coins are tossed 160 times and the following results were obtained.

Numbers of heads: 0 1 2 3 4
Observed frequencies: 17 52 54 31 6

Under the assumption that coins are balanced, find the expected frequencies of
getting 0,1,2,3 or 4 heads and test the goodness of fit.

Hint: Output is either H or T


Degree of Freedom = n-1

Thus, p is calculated through binomial distribution


Observe: n for Binomial distribution is 4, n for df is 5
(since you have 5 categories)
Problem 5: Goodness of Fit with Pooling
Fit a Poisson distribu3tion to the following data and test the goodness of fit

x: 0 1 2 3 4 5 6
f: 275 72 30 7 5 2 1

Note: Degree of freedom v= n-1 –(no. of parameters estimated from data).


In Poisson distribution v= n-2= n-1-1 (due to 𝜆)

Solution:
Expected frequencies are less than 5.
So, we pool adjacent categories to make it greater
than 5
Test for independence of attributes

The chi-square test can also be applied to test the association between attributes such as
honesty, smoking etc, when the sample data is presented in the form of a contingency
table with any number of rows and columns.

Contingency Table: A classification table containing m rows and n columns with observed
3 / 9

frequencies is called a contingency table.


Test for independence attributes (Conti...)
Problem 1:The following table gives the classification of 150 workers according to
gender and nature of work. Test whether the nature of work is independent of the
gender of the work.

Stable Unstable
8
Males 60 30
Females 15 45
Solution: H0 : The nature of the work is independent of the gender of the worker.
H1 : The nature of the work is not independent of the gender of the worker.
Degrees of freedom =(No. of Rows- 1)*(No. of Columns-1)=(m− 1)(n− 1) = (2 − 1)(2− 1) = 1.
Level of significance = 5%
Table value of χ2 = 3.84
Contingency Table:
Stable Unstable Total
Males 60 30 90
Females 15 45 60
Total 75 75 150

Expected frequencies are given in the following table

Stable Unstable
Males 75×90/150=45 75×90/150=45
Females 75×60/150=30 75×60/150=30

Expected Cell Frequency = (Row Total * Column Total)/N.


Calculation of χ2

χ2

The calculated value of χ2 is greater than the table value at 5% LS and 1 df.
Hence, we reject the null hypothesis.
• Problem 2:A tobacco company claims that there is no relationship between smoking and
lung ailments. To investigate the claim, a random sample of 300 persons in the age group
of 40 and 50 are given a medical test. The observed sample results are tabulated below

Lung ailment Non-lung ailment


Smokers 75 105
Non smokers 25 95

On the basis of this information, can it be concluded that smoking and lung ailments
are independent? ( 3.841 for 1 df)
Experiment
• An experiment is a means of getting an answer to the problem under consideration.

• Absolute experiment deals with determining the absolute value of some characteristic like
obtaining the average intelligent quotient of a group of people.

• Comparative experiment are designed to compare the effect of two or more objects on
some population characteristic.
Ex. Comparison of different kinds of varieties of crops.

Design of Experiment
Design of Experiments (DoE) is a statistical framework used to plan, conduct, and
analyze experiments so that we can determine the effect of one or more factors on a
response variable.

Instead of random trial-and-error, DoE uses a systematic and efficient approach.


• It helps to ensure that the experiment is as efficient and effective as possible by carefully
controlling the variables and measuring the results.
• DOEs are commonly used in fields such as manufacturing, engineering,
and pharmaceuticals to improve processes and products.

• Example: We are conducting an agriculture experiment to verify the truth to


claim that the fertilizers increase the yield of wheat.

• Then there two variables involved directly here, the fertilizer and the yield of
wheat. These two variable are called the experimental variables.

• In addition, there are other variables also involved here : quality of seed, climate,
nature of soil etc.,. These variables are called the extraneous variable.
Aim of Design of Experiment

The main aim of Design of experiment is the following:


• To control the extraneous variable,
• To minimize the experimental error.
Basic Components
•Factors: Independent variables (e.g., fertilizer type, temperature)
•Levels: Values of factors (e.g., type A/B , low/high)
•Response Variable: Outcome measured (e.g., crop yield)
• Experimental units: The object on which we make the observations/
measurements under the study is termed as the experimental units.
• Ex. Plot of a land.

• Treatment: Various objects of comparison in a comparative experiment


are termed as a Treatment (Combination of factor levels)
Ex. In a field of experiment: different fertilizer, different varieties of
crops, different method of cultivation.
Basic principles of experimental design

[Link]
Avoid bias by randomly assigning treatments.

2. Replication
Repeat experiments to estimate variability.

3. Local Control (Blocking)


Group similar experimental units to reduce noise.
Types of Experimental Designs (In syllabus)

•Completely Randomized Design (CRD)

•Randomized Block Design (RBD)

•Latin Square Design (LSD)


Completely Randomized Design (CRD)
• CRD is the simplest of all the designs based on the
principles of randomized and replication. In this design,
treatments are allocated at random to the experimental
units over the entire experimental material.
• Let us suppose we have treatments. The being
replicated times ( ), then the whole
experimental material is divided into experimental
units and the treatments are distributed completely at
random over the units subject to the condition that the
treatment occurs times.
Completely Randomized Design (CRD)
• Randomization assures that the extraneous factor does not
continuously influence one factor.
ANOVA
ANOVA, or Analysis of Variance, is a statistical method used to test the differences
between the means of two or more groups. It is used to determine whether there
is a significant difference between the means of the groups, or whether any
observed differences are due to chance.

Instead of comparing means directly, ANOVA compares variances.

ANOVA can be used with data from experiments that have a single factor, or
independent variable, with more than two levels, or with data from experiments
that have multiple factors. It is a powerful and widely used tool for analyzing data
from designed experiments, and is often used in fields such as agriculture, biology,
engineering, and psychology.
Types of ANOVA

One-Way ANOVA
→ One factor (e.g., fertilizer type)

Two-Way ANOVA
→ Two factors (e.g., fertilizer + irrigation)

×Factorial ANOVA
→ Interaction effects included

×: Not in syllabus
Statistical analysis of CRD
• ANOVA is a technique used to test the means of more than
two samples. It divides the total variance in the group into
parts, which are associated to different factors. The
variation is split into two components as the variation
within subgroup and a variation between subgroups.

Total variation is split into:


Total Variation = Between Group Variation
+
Within Group Variation
Assumption and hypothesis of one way ANOVA
• It is assumed that the (treatments) populations are
independent and normally distributed with means
and common variance . We wish to derive
appropriate methods for testing the hypothesis.

• At least two of the means are not equal.
Connection Between DoE and ANOVA

DoE: Designs the experiment


ANOVA: Analyzes the results
Example 1: A researcher wants to study the effect of fertilizer type on crop yield.
Three fertilizers are tested: Ammonium Chloride, Urea , Control (No fertilizer).
The experiment is conducted on 12 identical plots of land.
Plot No Fertilizer Applied Yield (kg)
Ammonium
1 13.4
Chloride
2 Urea 12.0
3 Control 10.5
Ammonium
4 10.9
Chloride
5 Urea 11.7
6 Control 9.8
Ammonium
7 11.2
Chloride
8 Urea 10.7
9 Control 10.2
Ammonium
10 11.8
Chloride
11 Urea 11.2
12 Control 9.9
Component Description
Factor Fertilizer Type
3 (Ammonium Chloride,
Levels
Urea, Control)
Experimental Units 12 plots
Treatments 3 treatments (T1, T2, T3)
Response Variable Yield (kg)

Solution:

Null Hypothesis 𝐻0:


𝜇1 = 𝜇2 = 𝜇3

i.e., All fertilizer means are equal

Alternative Hypothesis 𝐻1:


At least one mean is different
Treatment-wise Arrangement (for One way ANOVA):

Treatment Observations Ti= Sum of i-th row

Ammonium Chloride 13.4, 10.9, 11.2, 11.8 T1=47.3

Urea 12.0, 11.7, 10.7, 11.2 T2=45.6


T3=40.4
Control 10.5, 9.8, 10.2, 9.9
G= T1+T2+T3
=133.3

Number of groups, k=3


Total observations, N=12
ni = 4 for all i
ANOVA TABLE
Use of F test in ANOVA
• Test statistic:

Since, F >1 always


Critical value
• :
• Degrees of freedom :
• If then .
• Otherwise .
Source SS df MS=SS/df F=MS/MS
Treatments 6.46 2 3.23 5.01
Error 5.79 9 0.644 —
Total 12.25 11 — —
Value (F-table)

Level of significance: 5%
Degrees of freedom:
• Numerator 𝑑𝑓1 = 2
• Denominator 𝑑𝑓2 = 9

𝐹𝑐𝑟𝑖𝑡𝑖𝑐𝑎𝑙 = 𝐹0.05 2 9 ≈ 4.26

Decision Rule
•If 𝐹𝑐𝑎𝑙𝑐𝑢𝑙𝑎𝑡𝑒𝑑 > 𝐹𝑐𝑟𝑖𝑡𝑖𝑐𝑎𝑙 →Reject 𝐻0
•If 𝐹𝑐𝑎𝑙𝑐𝑢𝑙𝑎𝑡𝑒𝑑 ≤ 𝐹𝑐𝑟𝑖𝑡𝑖𝑐𝑎𝑙 →Do not reject 𝐻0

Since, 5.01>4.26 , therefore Reject 𝐻0

At 5% level of significance, the calculated F-value exceeds the tabulated


value. Hence, the null hypothesis is rejected.
Therefore, there is a significant difference among the mean yields of the
fertilizers.
Problem 1
• A set of data involving four tropical feedstuffs A, B, C, D
tried on 20 chickens is given below. All the twenty chickens
are treated alike in all respects except the feeding
treatments and each of the feeding treatment is given to 5
chickens. Analyze the data. Weight gain of baby chickens fed
on different feeding materials composed of tropical
feedstuffs.
Problem 2
• A completely randomized design experiment with 10 plots
and 3 treatments gave the following results.
• Plot No. : 1 2 3 4 5 6 7 8 9 10
• Treatment : A B C A C C A B A B
• Yield :5 4 3 7 5 1 3 4 1 7
• Analyze the results for treatment effects.
Problem 3
• Set up an analysis of variance table for the following per
acre production data for three varieties of wheat, each
grown on 4 plots and state if the variety differences are
significant.
Problem 4
• The following table shows the lives in hour for four brands
of electric lamps.
• Brand
•A : 1610, 1610, 1650, 1680, 1700, 1720, 1800
•B : 1580, 1640, 1640, 1700, 1750
•C : 1460, 1550, 1600, 1620, 1640, 1660, 1740, 1820
•D : 1510, 1520, 1520, 1570, 1600, 1680
• Perform the analysis of variance and test the homogeneity
of mean lives of the four brand of lamps.
Randomized Block Design
• If the whole experimental area is not homogenous, then a simple
method of controlling the variability of experimental material consists
in grouping the whole area into relatively homogenous subgroups.

• The treatments can be applied in a random manner to relatively


homogenous units within each subgroups/blocks and replicated over
all the block. This design is known as RBD.
Conclusion

• Make conclusion based on F test.


• F statistic > 1 (always)
• Compare FStatistic and Ftabular value
Problem 1
• Three verities of crops are tested in a randomized block
design with four replications. The layout is given below. The
yields are given in kilo gram analyze the significance
difference.
Problem 2
• Consider the results given in the following table for an experiment
involving six treatments in four randomized blocks. The treatments
are indicated by numbers within parentheses. Analyze whether
there is any significant difference in the treatments and blocks are
homogenous.
Problem 3
• A tea company appoints four salesmen A, B, C, and D and observes
their sales in three seasons. Summer, Winter and Monsoon. The
figures (in lakhs) are given in the following table:
Salesmen
Seasons A B C D
Summer : 36 36 21 35
Winter : 28 29 31 32
Monsoon : 26 28 29 29
Problem 4
• The following data represent the number of units of production per
day turned out by 5 different workmen using different types of
machines.
• i) Test whether the mean
productivity is the same for the
four different machine types.
• ii) Test whether 5 men differ
with respect to mean productivity.
Problem 5
• The following table shows the experiment conducted between
detergents A,B,C and D , and the three Engines 1,2 and 3
• Looking at the detergents as treatments
and the engines as blocks, perform
the two way analysis of variance test
and comment on your results.
Problem 6
• Four doctors each test four treatments for a certain disease and
observe the number of days each persons takes to recover. The
results are as follows.
• Discuses is there any difference between
the (a) doctors (b) treatments.

You might also like