0% found this document useful (0 votes)
20 views105 pages

History and Applications of Statistics

Uploaded by

heyderus9936
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOC, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
20 views105 pages

History and Applications of Statistics

Uploaded by

heyderus9936
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOC, PDF, TXT or read online on Scribd

1.

INTRODUCTION

1.1 Definition and Brief History of Statistics

In biology, including agricultural sciences, “the laws of nature” are not that simple. Biological phenomena
often show variation that obscures the law we want to establish. For instance, if we treat two fields in
identical way, the yields obtained from the same variety of a crop will not be the same. Similarly, two cows
of the same age, breed, and body weight fed with the same type and amount of fodder will produce different
amounts of milk. Thus, variation is a typical feature of biological data. Such variation is systematically
studied, quantified, and interpreted using statistical sciences.

Statistics is defined as study of numerical data based on variation in nature. It is a science, which deals with
collection, classification, tabulation, summarization and analysis of quantitative data or numerical facts.
Statistics was developed to deal with problems in which, for the individual observations, laws of cause and
effect are not apparent to the observer and where an objective approach is needed. In such problems, there
must always be some uncertainty about any inference based on a limited number of observations.

The word statistics also refers to numerical and quantitative data such as statistics of births, deaths,
marriage, production, yield, etc. The application of statistical methods to the solution of biological problems
is biometry, biological statistics or bio-statistics. The word biometry comes from two Greek roots: 'bios'
mean life and metron mean 'to measure'. Thus, biometry literally means the measurement of life.

The term statistics is an old one. Statistics must have started as a state arithmetic technique to assist a ruler
who needed to know the wealth and number of his subjects in order to levy a tax or wage a war. We know
that Caesar Augustus sent out a decree that the entire world should be taxed. Consequently, he required that
all persons report to the nearest statistician-in that day the tax collector.

William, the conqueror ordered a survey of the lands of England for purposes of taxation and military
service. This was called the Domesday Book, which is a manuscript record of the "Great Survey" of much
of England and parts of Wales completed in 1086 by order of King William the Conqueror.

Several centuries after the Domesday Book, we find an application of empirical probability in ship
insurance, which seems to have been available to Flemish shipping in the fourteenth century. This can be
little more than speculation or gambling, but it developed in to the very respectable form of statistics called
insurance.

Gambling in the form of games of chance led to the theory of probability originated by Pascal and Fermat
about the middle of the seventeenth century because of their interest in the gambling experiences of the
Chevalier de Mere. To the statistician and the experimental scientist, the theory contains much of practical
use for the processing of data.

The normal curve or normal curve of error has been very important in the development of statistics. The
equation of this curve was first published in 1733 by de Moivre. De Moivre had no idea of applying his
result to experimental observations and his paper remained unknown until Karl Pearson found it in a library

1
in 1924. However, the same result was later developed by two mathematical astronomers, Laplace (1749-
1827) and Gauss (1777-1855) independently of one another.

Charles Darwin (1809-1882), a biologist, received the second volume of Lyell’s book while on the Beagle.
Darwin formed his theories later and he may have been stimulated by his reading of this book. Darwin’s
work was largely biometrical or statistical in nature and he certainly renewed enthusiasm in biology. Gregor
Mendel (1822-1884) too, with his studies of plant hybrids published in 1866, had a biometrical or statistical
problem.

In the nineteenth century, the need for a sounder basis for statistics became apparent. Karl Pearson (1857-
1936) initially a mathematical physicist applied his mathematics to evolution as a result of the enthusiasm in
biology created by Darwin. Pearson spent nearly half a century in serious statistical research. In addition, he
founded the journal Biometrica and a school of statistics; as a result the study of statistics gained impetus.

While Pearson was concerned with large samples, large-sample theory was proved to be somewhat
inadequate for experimenters with necessarily small samples. Among these was W. S. Gosset (1876-1937),
a student of Karl Pearson and a scientist of the Guinness firm of brewers. Gosset’s mathematics appears to
have been insufficient to the task of finding exact distributions of the sample standard deviation, of the ratio
of the sample mean to the sample standard deviation, and of the correlation coefficient, statistics with which
he was particularly concerned. Consequently, he resorted to drawing shuffled cards, computing and
compiling empirical frequency distributions. Papers on the results appeared in Biometrika in 1908 under the
name of student, Gosset’s pseudonym. Today Student’s t is a basic tool of statisticians and experimenters.
Now that the use of Student’s t distribution is so widespread, it is interesting to note that the German
astronomer, Helmert, had obtained it mathematically as early as 1875.

R.A. Fisher (1890-1962) was influenced by Karl Pearson and Student and made numerous and important
contributions to statistics. He and his students gave considerable impetus to the use of statistical procedures
in many fields, particularly in agriculture, biology and genetics.

Abrahm Wald (1902-1950) has contributed two books on Sequential Analysis and Statistical Decision
Functions. Thus, it is in the 20th century that most of the statistical methods presently used have been
developed.

Currently, statistics is used as an analytical tool in many fields of research.

1.2 Applications and Limitations of Statistics

Statistics is important in the field of social science, agriculture, medical, engineering, etc because it provides
tools to analyze collected data. Scientists frequently use statistics to analyze research data. Statistics
provides scientific methods for appropriate data collection, analysis and summarization of data and
inferential statistical methods for drawing conclusions in the face of uncertainty. Statistical methodologies
have wide applicability to almost any branch of science dealing with the study of uncertain phenomena.

Statistics has already become a very important and useful subject and the various techniques are being used
to analyze and solve the problems in different discipline. Thus, currently statistics is used as an analytical
tool in many fields of research.

2
However, there are some of the limitations, i.e. it is not suited to study the qualitative phenomenon; statistics
does not study individuals as it deals with an aggregate of objects and does not give any special importance
to the individual of a series; statistical laws are not exact like as physical or natural law of sciences,
statistical analysis is only in terms of probability and chance not an exact; and statistics is liable to be
misused as statistical methods are more dangerous tools in the hand of the inexpert.

1.3 Definition of Some Basic Terms


Data: Qualitative or quantitative information taken on a certain character. For example, data of height,
weight, color, etc.

Variable: A property with respect to which individuals in a sample differ in some way, e.g. length, weight,
height, color, etc. Characteristics, which show variations, are called random variables.

Variables can be:


1. Measurement variables: Variables that can be expressed in a numerical order. They are of two types:
a) Continuous variables: Variables which can assume an infinite number of values between any two fixed
points, e.g. height, grain yield, score of students, etc.
b) Discontinuous/Meristic/Discrete variables: Variables that take any certain fixed numerical values with
no intermediate values in between, e.g. number of seeds per pod, number of plants in a quadrat, number
of students taking the course biometry, etc.
2. Ranked variables: Some variables cannot be measured but at least can be ordered or ranked by their
magnitude, e.g. disease score; no infection(0), mild infection(1), high infection(2), severe infection(3)
Here we cannot say that the difference between 1 and 2 is identical or equal to the difference between 2 and
3. Such assumption is made for the measurement variables.

3. Attribute or nominal variables: Variables that cannot be measured but are expressed qualitatively, e.g.
color of common bean seed: white, black, red; sex of animals could be male or female, blood group of
humans could be A, AB, O, etc.
When such attribute data are combined with frequency (number of occurrences), they can be treated
statistically. Such data are called enumeration data. Suppose we have 18 mixed bean seeds, they can be
grouped as: black (3), white (5) and red (10).

Variate (datum): A single reading, score or observation of a given variable. If we measure height of 5
plants from a plot, each of the 5 readings of height will be a variate: 10 cm, 5 cm, 20 cm, 25 cm, 15 cm.

Population and sample


Biologists usually wish to make inferences (draw conclusions) about a population, which is defined as the
collection of all the possible observations of interest. Sample is a subset from the population selected by a
specific procedure. In an experiment, we can rarely include the whole population because it is costly, time
consuming and sometimes impossible to determine. The idea is then to use the sample data to make
statements (decisions) about the population. Thus, collection of observations we take from the population is
called a sample and the number of observations in the sample is called the sample size (usually indicated by
the symbol n).

The basic method of collecting the observations in a sample is called simple random sampling. This is
where any observation has the same probability of being collected, e.g. giving each student in a class equal
3
chance in measuring height. The aim is always to sample in a manner that does not create a bias in favour of
any observation being selected. Nearly all applied statistical procedures that are concerned with using
samples to make inferences (i.e. draw conclusions) about populations assume some form of random
sampling. If the sampling is not random, then we are never sure as to what population is represented by our
sample. When random sampling from clearly defined populations is not possible, then interpretation of
standard methods of estimation becomes more difficult. [Read the other different types of probability and
non-probability sampling methods and their applications from statistical books].

Populations must be defined at the start of any study and this definition should include the
spatial and temporal limits to the population. Our formal statistical inference is restricted to these limits. For
example, if we sample from a population of animals at a certain location in December 2010, then our
inference is restricted to that location in December 2010. We cannot infer what the population might be like
at any other time or in any other place, although we can speculate or make predictions.

Parameter: A population value, which we generally do not know, but would like to infer (estimate) about.
For example, national average yields of maize in 2010 in Ethiopia. Parameters are designated using Greek
letters such as µ, , , etc.

Statistics: Are sample estimates of population value (parameters) and designated using Latin letters such as
, s, p, etc.

The population parameters cannot be measured directly because the populations are usually too large, i.e.
they contain too many observations for practical measurement. It is important to remember that population
parameters are usually considered to be fixed, but unknown, values so they are not random variables and
do not have probability distributions. Sample statistics are random variables, because their values
depend on the outcome of the sampling experiment, and therefore they do have probability distributions,
called sampling distributions.

1.4 Types of Statistics

1. Descriptive (deductive) statistics: Are methods, which are used to describe a set of data without
involving generalization. Deal with the presentation of research data or any numerical information. Help
in summarizing and organizing data so as to make them readable for users, e.g. mean, median, mode,
standard deviation, etc.

2. Inferential (inductive) statistics: It is a statistics, which helps, in drawing conclusion about the whole
(population) based on data from some of its parts (samples).

2. STATISTICAL INFERENCE

2.1 Estimation

Most of the time, we make decisions about population on the basis of sample information because it is
generally difficult and sometimes impossible to consider all the individuals in a population for economic,

4
time and other reasons. For example, we take random sample and calculate the sample mean ( ) as an
estimate of population mean (µ) and sample variance (s2) as an estimate of population variance (2).

Estimation is a technique, which enables us to estimate the value of population parameter based on sample
values. The formula that is used to make estimation is known as estimator whereas the resulting sample
value is called an estimate.

There are two kinds of estimation: point estimation and interval estimation.

Point estimation is a technique by which a single value is obtained as an estimate of a population parameter
like = 10; s2 = 0.6 are point estimates of µ and 2, respectively.
Interval estimation is an estimation technique in which two limits within which a parameter is expected
to be found are determined. It provides a range of values that might include the parameter with a known
probability, e.g. confidence intervals. For example, population mean (µ) lies between two points such that
a<µ<b where a and b are lower and higher limits, respectively, and are obtained from sample
observations, e.g. the average weight of students in a class lies between 50 kg and 60 kg.

2.1.1 Properties of best estimates

 Unbiasedness: the expected value of the sample statistic (the mean of its probability distribution)
should be equal to the parameter. Repeated samples should produce estimates which do not consistently
under or over estimate the population parameter.
 Consistency: as the sample size increases, then the estimate will get closer to the population
parameter. Once the sample includes the whole population, the sample statistic will obviously equals the
population parameter.
 Efficiency: it has the lowest variance among all competing estimators. For example, the sample
mean is a more efficient estimator of the population mean of a variable with a normal probability
distribution than the sample median, despite the two statistics being numerically equivalent.
 Sufficiency: the estimates should have all the required information about the parameter, which is
being estimated. For example, the sample mean is a more sufficient estimator than the range as it includes
all the observations.

2.1.2 Confidence Interval

The sample mean ( ) and variance (s2) are point estimates of population mean (µ) and population variance
(2), respectively. These sample estimates may or may not be equal to the population mean. Thus, it is better
to give our estimation in interval and state that the population mean (µ) is included in the interval with some
measure of confidence.

Confidence interval is a probability statement concerning the limit within which a given parameter lies,
e.g. we can say the probability that µ lies within the limit a and b is 0.95 where a< b; P (a<µ< b) = 0.95

2.1.3 Estimation of a single population mean (µ)

5
Case 1: When population standard deviation () is known (given), the sample size could be large (n 
30) or small (n<30)
A (100- ) % Confidence Interval for population mean (µ) =  Z/2 where
-  = allowable error rate
- - Z/2  = L1 (lower confidence limit)

- + Z/2  = L2 (upper confidence limit)


- Z is the value of standard normal curve

Example: A certain population has standard deviation of 10 and a sample of 100 observations were taken
from a population and the mean of the sample was 4.00
Find the point estimate and construct the 95% confidence interval for the population mean (µ).
Given:  =10; = 4.00; n = 100. Thus,
- Point estimate = 4.00
- A 95% confidence interval for population mean
=  Z/2 = 4.00  Z0.025  = 4.00  1.961.00
- L1 = 4.00-(1.96 x 1.00) = 2.04
- L2= 4.00 + (1.96 x 1.00) = 5.96

Thus, 95% C. I. for population mean is 2.04 < µ < 5.96. This means, the probability that the population
mean lies in between 2.04 and 5.96 is 0.95 or we are 95% confident that the population mean is included in
this range.

If we increase the confidence coefficient from 0.95 to 0.99, the width of the interval increases and the more
certain we can be that the true mean is included in this estimated range. However, as we increase the
confidence level, the estimate will loose some precision (will be more vague).

Case 2: When  (pop standard deviation) is unknown


In many practical situations, it is difficult to determine the population mean and the population standard
deviations. Under such cases, a sample standard deviation (s) can be used to approximate the population
standard deviation ().

A). Estimation with large sample size (n  30)


A (100- ) % Confidence Interval for population mean (µ) =  Z/2 
Example: A poultry scientist is interested in estimating the average weight gain over a month for 100
chicks introduced to a new diet. To estimate the average weight gain, he has taken a random sample of 36
chicks and measured the weight gain, the mean and standard deviation of this sample were = 60 g and s =5
g.
Find the point estimate and construct 99% confidence interval for population mean.
- Point estimate = 60 g

6
- A 99% C. I. =  Z/2  = 60  Z0.005  = 60  2.58  0.83
Thus, 99% C.I. for population mean = 57.86 g<µ<62.14 g. The probability that the mean weight gain per
month is between 57.86 g and 62.14 g is 0.99, or we are 99% confident that the population mean is found in
this interval.

B) Estimation with small sample size (n<30)

The central limit theorem states that as sample size increases, the sampling distribution of means approaches
a normal distribution. However, samples of small size do not follow the normal distribution curve, but t-
distribution.
Thus, A (100- )% C.I. for population mean (µ) =  t/2 (n-1) where n-1 is degree of freedom:

Properties of t-distribution:
- symmetrical in shape and has a mean equal to zero like normal (Z) distribution.
- the shape of t-distribution is flatter than the Z distribution, but as sample size approaches 30, the
flatness associated with the t-distribution disappears and the curve approximates the shape of Z-
distribution.
- there are different distributions for each degree of freedom

t for n = 30

t for n = 20

t for n =10

μ =0

Example: A sample of 25 horses has average age of 24 years with standard deviation of 4 years. Calculate
95% C.I. for population mean.
A 95% C.I. =  t/2 (n-1)  = 24  t0.025 (25-1)  = 24  2.064  0.8. Thus, a 95% C.I. = 22.35
years < < 25.65 years

2.1.4 Estimating the difference between two population means

There are three cases:

Case 1: When 12 & 22 (population variances) are known (n1 and n2) could be large or small)

7
- Point estimate = - if > or - if >

- A (100-)% C.I.= -  Z/2  if > or

-  Z/2  if >

where and are sample mean of 1st and 2nd group; n1 and n2 are sample sizes of 1st and 2nd group; 12 & 22
are population variances of 1st and 2nd group.

Case 2: 12 & 22 are unknown, but with large sample sizes (n1 and n2  30)
For large sample size, s12 is an estimate of 12 and s22 is an estimate of 22
- Point estimate = - if > or - if >

- A(100- ) % C.I. = -  Z/2  if > or -  Z/2  if > >

Example: The performance of two breeds of dairy cattle (Holstein & Newjersy) was tested by taking a
sample of 100 cattle each. The daily milk yield was as follows. Find point estimate and construct a 90% C.
I. for population mean difference.
- Holstein (x): = 15 liters; s1 = 2 liters.
- Newjersy(y): = 10 liters; s2 = 1.5 liters

 Point estimate = - = 15 - 10 = 5 liters

 A 90% C.I. = -  Z/2  = 15-10  Z0.05 

 = 5  1.65  0.25 = 5  0.4125


 Thus, a 90% C.I. for population mean difference is 4.59 lt <1–2< 5.41 lt

Case 3: 12 & 22 are unknown and n1 and n2 are small (< 30)
- Point estimate = - if > or - if >
- A (100-)% C.I. for population mean difference
= -  t/2 (n1+n2-2)  sp

Sp (pooled standard deviation) =

n1, n2 are sample sizes; n1+ n2-2 = degrees of freedom

8
Example: Two groups of nine framers were selected, one group using the new type of plough and another
group using old type of plough. Assume the soil, weather, etc. conditions are the same for both groups
(variation is only due to plough types). The yields and standard deviation were:

= 35.22 qt; s1 = 4.94 qt; = 31.55 qt; s2 = 4.47 qt

Estimate the difference between two plough types and construct a 99% confidence interval for the
difference of the population means:
- Point estimate = - = 35.22-31.55 = 3.67 qt

- A 99% C. I. = 35.22-31.55  t0.005 (9+9-2)  sp

Sp (pooled standard deviation) = = 3.67 6.48

Thus, 99% C.I. for pop mean difference = -2.81qt <1–2< 10.15 qt

2.2 Hypothesis Testing

2.2.1 Types of hypothesis and errors in hypothesis testing

In attempting to reach at decisions, we make assumptions or guesses about the population. Such assumption,
which may or may not be true is called statistical hypothesis.

Examples:
- The ratio of male to female in Ethiopia is 1:1
- The cross of 2 heterozygous varieties for color of tomato will produce plants with red and white
flowers in the ratio of 3:1. Rr  Rr RR: 2Rr: rr where red is dominant over white.

There are two types of hypothesis:


1. Null hypothesis (HO): Is a hypothesis which states that there is no real difference between the true
values of the population from which we sampled, e.g. male to female ratio is 1:1.
2. Alternate hypothesis (HA): Any hypothesis which differs from a given null hypothesis
HA: sex ratio is not 1:1
Male ratio > Female ratio
Male ratio< Female ratio
If we find that results observed in random sample differ markedly from those expected under null
hypothesis, we will say the observed differences are significant and we would reject the null hypothesis.

A procedure that enables us to decide whether to accept or reject the null hypothesis or procedure used to
determine whether observed samples differ significantly from expected results are called test of hypothesis
or rules of decision.

9
Rejection and Acceptance of Hypothesis
A hypothesis is rejected if the probability that it is true is less than some predetermined probability. The
predetermined probability is selected by the investigator before he collects his data based on: his research
experience, the consequences of an incorrect decision and the kind of risk the researcher prepares to
take.

When the null hypothesis is rejected, the finding is said to be statistically significant at that specified level
of significance. When the available evidence does not support the rejection of the null hypothesis, the
finding is said to be non-significant.

Errors in Hypothesis Testing


Two kinds of errors are encountered in hypothesis testing:
 Type I error () is committed when a true null hypothesis is rejected. Concluding there is a
difference between the groups being studied when, in fact, there is no difference. Thus, it would seem
reasonable to select a small value of .
 Type II error () committed when a false null hypothesis is accepted. Concluding there is no
difference between the groups being studied when, in fact, there is a difference.

The two kinds of correct decisions are accepting a true null hypothesis and rejecting a false null
hypothesis.

The probability of type I error is the significance level, i.e. in a given hypothesis, the maximum
probability which we will be willing to risk a type I error is called the level of significance of the test. It is
denoted by  and is generally specified before samples are taken. In practice 5% or 1% levels of
significance are in common use. A 5% level of significance means that there are about 5 chances in 100 that
we would reject null hypothesis (H O) when it should be accepted, i.e. we are 95% sure that we make a
correct decision.

The probability of type II error (accepting a false null hypothesis) is represented by . Type II error usually
occurs when the sample size is small. The more an investigator protects his experiment against type I error,
the greater the likelihood of committing type II error.

Thus, the value of  must be chosen to minimize type I as well as type II error.

Power of Test: This is defined as the probability of rejecting the hypothesis when it is false and
symbolically given as (1-). We have to design our experiment in such away that the power of test is
maximized.

2.2.2 One- tailed and two-tailed tests

In most cases, the HO (null hypothesis) is one of no effect (e.g. no difference between two means) and the
HA (the alternative hypothesis) can be in either direction; the H O is rejected if one mean is bigger than the
other mean or vice versa. This is termed a two-tailed test because large values of the test statistic at either
end of the sampling distribution will result in rejection of H O. In two-tailed test, to do a test with α = 0.05,
then we use critical values of the test statistic at α/2 = 0.025 at each end of the sampling distribution.

10
Sometimes, our HO is more specific than just no difference. We might only be interested in whether one
mean is bigger or smaller than the other mean but not the other way. In one-tailed test, to do a test with α =
0.05, then we use critical values of the test statistic at α = 0.05 at one end of the sampling distribution.

Thus, in one tailed test, the area corresponding to the level of significance is located at one end of the
sampling distribution while in a two tailed test the region of rejection is divided into two equal parts, one at
each end of the distribution.

Most statistical tables either provide critical values for both one- and two-tailed tests but some just have
either one- or two-tailed critical values depending on the statistic, so make sure that you look-up the correct
α values when you use tables. Statistical software usually produces two-tailed α values so you should
compare the α value to 0.10 for a one-tailed test at 0.05.

Acceptance and rejection regions


Acceptance and rejection regions Acceptance and rejection regions
in case of two-tailed test
(with 5% significance level) in case of one-tailed test (right tail) in case of one-tailed test (left-tailed)
with 5% significance level with 5% significance level
Rejection region Acceptance region Rejection region
(Accept H 0 if the sample Acceptance region Rejection region Rejection region Acceptance region
mean (x) falls in this region) (Accept H 0 if the sample (Accept H 0 if the sample
mean (x) falls in this region) mean (x) falls in this region)
Limit

Limit

Limit

0.475 0.475 Limit


of area of area 0.50 0.45 0.45 0.50
of area of area of area of area
0.025 of area 0.025 of area
Both taken together equals 0.05 of area 0.05 of area
0.95 or 95% of area Both taken together equals Both taken together equals
0.95 or 95% of area 0.95 or 95% of area
Z = -1.96 H0 Z = 1.96
H0 Z = -1.645 Z = -1.645 H0
Reject H 0 if the sample
mean (x) falls in either of these Reject H 0 if the sample Reject H 0 if the sample
two region mean (x) falls in this region mean (x) falls in this region

5.2.3 Test of a single population mean ()

Procedure:
i. State the null and alternate hypothesis
ii. Determine the significance level (). The level of significance establishes a criterion for rejection or
acceptance of the null hypothesis (10%, 5%, 1%)
iii. Compute the test statistic: The test statistic is the value used to determine whether the null
hypothesis should be rejected or accepted. The test statistic depends on sample size or whether the
population standard deviation is known or not:
a. when  is known n could be large or small, the test statistic

Zc = (normal distribution)

b. when 2 unknown, but with large sample size (n 30)

11
Zc (test statistic) = (normal distribution)

c. when  unknown and with small sample size (n<30)

tc (test statistic) = (t-distribution)

iv. Determine the critical regions (value). The critical value is the point of demarcation between the
acceptance or rejection regions. To get the critical value, read Z or t-table depending on the test statistic
used.

v. Compare calculated value (test statistic) with table value of Z or t and make decision.
If the absolute value of calculated value (test statistic) is greater than table Z or t-value, reject the null
hypothesis and accept the alternate hypothesis.
 If /Zc/ or /tc/ > Z or t table value at the specified level of significance, reject the H O (null hypothesis)
 If /Zc/ or/tc/  Z or t-table value at the specified level of significance, accept H O

Example 1: A standard examination has been given for several years with mean (µ) score of 80 and
variance (σ2) of 49. A group of 25 students were taught with special emphasis on reading skills. If the 25
students obtained a mean grade of 83 on the examination, is there reason to believe that the special emphasis
changed the result on the test at 5% level of significance assuming that the grades are normally distributed?
µ = 80; σ2 = 49; = 83; n =25; α = 0.05.

1. State the null and alternate hypothesis


HO:  = 80
HA:  ≠ 80 (two-tailed test)
2. Compute the test statistic: we use the normal (Z-distribution) because of known population variance

Zc = = = 2.14

3. Obtain the table Z-value at Zα /2 = Z0.025 = 1.96


4. Make decision. Since /Zc/ (2.14) > Z-table (1.96), we reject the null hypothesis and accept the alternate
hypothesis, i.e. there is sufficient reason to doubt the statement that the mean is 80. That means, the especial
emphasis has significantly changed the result of the test

Example 2: It is known that under good management, Zebu cows give an average milk yield of 6 liters/day.
A cross was made between Zebu and Holstein and the result from a sample of 49 cross breeds gave average
daily milk yield of 6.6 liters with standard deviation of 2 liters. Do the cross breeds perform better than the
pure Zebu cows at 5% level of significance?
Given: =6 liters; n = 49; =6.2; s = 2

1. State the null and alternate hypothesis


HO:  = 6 (No difference between cross breed and pure Zebu)
HA:  > 6 (Cross breeds perform better than the pure zebu). This is one-tailed test.
12
2. Compute the test statistic: we use the normal (Z-distribution) because of large sample size.

Zc = = = 2.1

3. Obtain the table Z-value: Z0.05 = 1.65


4. Make decision. Since /Zc/(2.1) > Z-table(1.65), we reject the null hypothesis and accept the alternate
hypothesis, i.e. cross breeds perform better than the pure Zebu or cross breeds gave statistically
significant higher milk yield than the pure Zebus.

2.2.4 Test of the difference between two means

Procedure:
1. State the null and alternate hypothesis
2. Compute the test statistic
 Case I: When 12 & 22 (population variances) are known (n1 and n2) could be large or small

 Zc = ( - )/

 Case II: When 12 & 22 are unknown, n1 & n2 are large (30)

 Zc = ( - )/

 Case III: When 12 & 22 unknown, n1 & n2 small ( 30);
 tc = ( - )/sp where sp (pooled standard

deviation) =

3. Read Z or t- table depending on the test-statistic used


4. Make decision. If /Zc/ or /tc/ > z table or t table, reject the null hypothesis at the specified level of
significance.

Independent t-test

In this case, the allocation of treatments on experimental units is done completely at random, e.g. varieties A
(filled) & B (open) are allocated to 12 fields, each on 6 fields randomly.

13
Example: A researcher wanted to compare the potentials of 2 different types of fertilizer. He conducted an
experiment at 10 different locations and the two fertilizer types were randomly applied in five of the fields
each on maize. The following yields were obtained in tons/ha.

Fertilizer x: 9.8 10.6 12.3 9.7 8.8


Fertilizer y: 10.2 9.4 9.1 11.8 8.3

Is there any difference in potential yield effect of the fertilizers at 5% level of significance?

Solution
1. Calculate the means and standard deviations: = 10.24; s1 = 1.32; = 9.76; s2 = 1.33; both n1 & n2 = 5

2. State the hypothesis


H0: 1 = 2;
HA: 1  2 (two tailed test)
3. Compute the test statistic (12 & 22 unknown, n1 & n2 small)
tc = ( - )/sp 

= sp 

sp = = 1.32

= 1.32 = 0.58
4. Read t- table value for t(/2) at 5 + 5 – 2 d.f., t0.025 (8) = 2.306.
5. Make decision. Since t-calculated (0.58) < t-table (2.306), we accept the null hypothesis, i.e. the data do
not give sufficient evidence to indicate the difference in potential yield of two fertilizers. Thus, the
difference between the two fertilizer types is not statistically significant.

Matched (paired) t-test


Paired (matched) data occur in pairs, e.g. blood pressure measure of patient before medication compared to
blood pressure measure after medication of the same person. Another example is when each of 10 fields
divided in to two parts and variety A planted on one part and variety B on another part.

If five plots were divided in to two half; one fertilized (Nitrogen - N) and other was control (H)

14
Example: An experiment was done to compare the yields (qt/ha) of two varieties of maize, data were
collected from seven farms. At each farm, variety A was planted on one plot and variety B on the
neighboring plot.

______________________________________________________________
Farm: 1 2 3 4 5 6 7
______________________________________________________________
Variety A: 82 68 109 95 112 76 81
Variety B: 88 66 121 106 116 79 89
______________________________________________________________

Test whether there is a significant difference between two varieties at 5% level of significance.

Solution
1. Calculate the differences, mean of the differences, and variance of the differences as shown below
______________________________________________________________
Farm: 1 2 3 4 5 6 7 
______________________________________________________________
Variety A: 82 68 109 95 112 76 81
Variety B: 88 66 121 106 116 79 89
Deterrence (d) -6 2 -12 -11 -4 -3 -8 -42
(d- ) 2 0 64 36 25 4 9 4 142
______________________________________________________________

(mean of the difference) = =

Sd2(variance of the difference) = = = 23.67

Note that in paired data some farms are poor (2, 6), thus low yields for both varieties are obtained while
others are good (3, 5) and thus high yield for both varieties.

2. State the hypothesis: HO: 1 = 2; HA: 1  2 (two-tailed test)

3. Calculate the test statistic

15
tc = = = = -3.26

4. Read table t value: t/2 (n-1) = t 0.025(n-1) d.f. = t0.025(6) = 2.447

5. Make decision: Since /tc/ (3.26) > t-table (2.447), there is a significant difference between the varieties of
maize at 5% level of significance.

3. PRINCIPLES OF EXPERIMENTAL DESIGN

3.1 Introduction

In research, a scientist identifies solution to problems through experimentation. Research can be broadly
defined as systematic investigation in to a subject to discover new facts or principles or to confirm or
deny the results of previous finding. Such investigation will help in decision making such as
recommending a new procedure, a new fertilizer rate, a new pesticide, etc.

The procedure for research is generally known as the scientific method, which, although difficult to define
precisely, it usually involves, the following steps:
a. Formulation of hypothesis: a tentative explanation or solution
b. Planning an experiment to objectively test the hypothesis
c. Careful observation and collection of the data
d. Analysis and interpretation of the experimental results.

Experiment is an important tool of research.

Remark: Not all researches involve experimentation! Example: observational researches, surveys, etc. in
which no variable is manipulated by the researcher. These are not experiments!!

Experiment is a collection of research designs which use manipulation and controlled testing to understand
causal processes. Generally, one or more variables are manipulated to determine their effect on a dependent
variable.

Some important characteristics of well-planned experiments are:


a. Simplicity: The selection of treatments and the experimental arrangement should be as simple as
possible, and consistent with the objectives of the experiment.
b. Degree of precision: The probability should be high so that the experiment will be able to measure
differences with the degree of precision the experimenter desires. This implies an appropriate design
and sufficient replication.
16
c. Absence of systematic error: The experiment must be planned to ensure that experimental units
receiving one treatment differ from those receiving another treatment in no systematic way so that an
unbiased estimate of each treatment effect can be obtained.- This is achieved through randomization
d. Range of validity of conclusions: Conclusions should have as wide range of validity as possible. An
experiment replicated in time and space (e.g., over years and locations) would increase the range of
validity of the conclusions that could be drawn from it. A factorial set of treatments is another way for
increasing the range of validity of an experiment. In a factorial experiment, the effects of one factor
are evaluated under varying levels of a second factor.
e. Calculation of degree of uncertainty: In any experiment, there is always some degree of uncertainty
as to the validity of the conclusions. The experiment should be designed so that it is possible to
calculate the probability of obtaining the observed results by chance alone.-setting the appropriate
level of significance (α).

3.2 Design of Experiments

The term refers to five interrelated activities required in the investigation. These are:
a. Formulating statistical hypothesis and making plans for laying out, collection and analysis of data
b. Stating the decision rules to be followed in testing statistical hypothesis (e.g., F-test)
c. Collecting data according to plan
d. Analyzing data according to plan
e. Making decisions based on decision rules

Purposes of experimental designs:


a. To provide estimates of a treatment effects or differences among treatment effects
b. To provide an efficient way of testing hypothesis about the response to treatments
c. To assess the reliability of estimates and assumptions
d. To estimate the variability of the experimental material
e. To increase precision by eliminating extraneous external source of variation from the comparisons of
interest
f. To provide a systematic, and efficient pattern of conducting an experiment

3.3 Concepts Commonly Used in Experimental Design

Treatment: It is an amount of material or a method that is to be tested in the experiment such as crop
varieties, insecticides, feedstuffs, fertilizer rates, method of land preparation, irrigation frequency, etc.

Experimental unit: It is an object on which the treatment is applied to observe an effect, e.g. cows, plot of
land, petri-dishes, pots, etc

In the study of the effect of different rations on milk production, ration is a treatment and animal is the
experimental unit; while in the study of different fertilizer rates on yield of maize, the fertilizer rates are
treatment and plot of land is experimental unit.

Experimental error: It is a measure of the variation, which exists among observations on experimental
units treated alike. Variation generally comes from two main sources:
17
1. Inherent variability that exists in the experimental material to which treatments are applied.
2. Lack of uniformity in the physical conduct of an experiment or failure to standardize the
experimental techniques such as lack of accuracy in measurement, recording data on different days,
etc.

Therefore, every possible effort should be made to reduce the experimental error.

Methods aimed at reducing the experimental error


a) Increase the size of experiment either through provision of more replicates or by inclusion of additional
treatments
b) Refine the experimental technique
- have uniformity in the application of treatments such as equally spreading of fertilizers, recording data
on the same day, etc.
- control should be done over external influences so that all treatments produce their effects under
comparable conditions, e.g. protect against diseases, insects, etc. as their effects are not uniform on all
plots.
c) Blocking: Dividing the field into several homogenous parts. Blocks are the levels at which we hold an
extraneous factor fixed, so that we can measure its contribution to the total variability of the data by means
of analysis of variance.

Replication: A situation where a treatment appears more than once in an experiment, it is said to be
replicated. The functions of replication are:
a) It provides an estimate of experimental error because it provides several observations on experimental
units receiving the same treatment. For an experiment on which each treatment appears only once, no
estimate of experimental error is possible and when there is no method of estimating the experimental
error, there is no way to determine whether observed differences indicate the real differences or due to
inherent variability.
b) It improves the precision of an experiment: As the number of replicates increase, the estimates of
population means as observed treatment means become closer to the true value.
c) It increases the scope of inference and conclusion of the experiments: Field experiments are normally
repeated over years and locations because conditions vary from year to year and location to location.
The purpose of replication in space and time is to increase the scope of inference. The results of an
experiment are applicable only to conditions that are similar to that condition.

Factors determining the number of replications are:


a) The degree of precision required: the higher the precision desired the greater is the number of replicates
(the smaller the departure from null hypothesis to be measured or detected, the greater the number of
replicates)
b) Variability of experimental units: certain experimental materials are more variable than others. For the
same precision, less replication is required on uniform experimental units (soils) than on variable
experimental unit (soils).
c) The number of treatments: more number of replications are needed with few treatments than with many
treatments
d) The type of experimental designs also affect the precision of an experiment and the required number of
replications, e.g in Latin Square Design the number of replications should be equal to the number of
treatments while in balanced lattice design the number of replications is one more than the square root of
the number of treatments, i.e. r = + 1 where r is number of replications and t is number of treatments.
18
e) Fund and time availability also determines the number of replicates: If there are adequate fund and time,
more replications can be used.

Usually three replications are taken as the minimum number for standard experiments

Identifying the Number of Replications

Two procedures for calculating the number of replications are described below:

Procedure 1: This method takes into consideration the variability of experimental material and field, which
is measured in terms of coefficient of variation (CV) and standard error of means (SEM). To calculate the
number of replications we can use the following formula:

N = (CV/SEM)2

Where N = Number of replications; CV = Coefficient of variation; SEM = Standard Error of Means.

Example: CV (%) = 12; SEM ± = 6; N = (12/6)2 = 4

Procedure 2: If CV & SEM are not known, then the number of replications can be arrived at using the
principle that the precision of treatment comparisons increases if the experimental error is kept to minimum.
The experimental error can be kept to the minimum by providing more degrees of freedom for the
experimental error. In other words, a lower number of degrees of freedom for experimental error results in
enlarged experimental error. Based on this principle, the number of degrees of freedom for error should not
be less than 15 (not less than 10 in any case).

When ‘t’ treatments are replicated ‘r’ times, the error is based on (t-1) (r-1) degrees of freedom in
Randomized Complete Block Design (RCBD) and t(r-1) in Completely Randomized Design (CRD), which
should not be less than 15.

No. of treatments: 2 3 4 5 6
Minimum no. of replications (CRD): 9 6 5 4 4
Minimum no. of replications (RCBD): 16 9 6 5 4

Increasing either number of replications or plot size can improve precision, but the improvement achieved
by doubling plot size is almost always less than the improvement achieved by doubling replications.

Randomization
Assigning the treatments to the experimental units in such away that any unit has equal chance to receive
any treatment, i.e. every treatment should have an equal chance of being assigned to any experimental units.
Thus, a particular treatment should not be consistently favored or disfavored.
Purposes of randomization
a) To eliminate bias: randomization ensures that no treatment is favored or discriminated against the
systematic assignment to units in a design
b) To ensure independence among the observations. This is necessary to provide valid significance tests.

Randomization is usually done by using tables of random numbers or by drawing cards or lots.
19
Confounding
It occurs when the differences due to experimental treatments, i.e. the contrast specified in your hypothesis,
cannot be separated from other factors that might be causing the observed differences. Example, if you
wished to test the effect of a particular hormone on some behavioral response of sheep. You create two
groups of sheep, males and females, and inject the hormone into the male sheep and leave the females as the
control group. Even if other aspects of the design are ok, differences between the means of the two groups
cannot be definitely attributed to effects of the hormone alone. The two groups are also different in sex and
this may also be, at least partly, determining the behavioral responses of the sheep. In this example, the
effects of hormone are confounded with the effects of sex.

Controls
A control is a part of the experiment that is not affected by the factor or factors studied, but otherwise
encounter exactly the same circumstances as the experimental units treated with the investigated factor(s).
For example, when investigating the effect of spraying micronutrients, the crop being in the control should
also be sprayed with the same amount of water except the micro nutrients. The reason for this way to work
is to make sure that it is only the effect of the substance of interest that is investigated. Otherwise, it may
not be possible to draw conclusions of the reason(s) to the outcome of the experiment.

Local Control
The principle of local control is another important principle of experimental designs. In local control, the
extraneous factor, i.e., the known source of variability, is made to vary deliberately over as wide a range as
necessary; and this needs to be done in such a way that the variability it causes can be measured and hence
eliminated from the experimental error. In other words, according to the principle of local control, we first
divide the field into several homogeneous parts, known as blocks, and then each such block is divided into
parts (units) equal to the number of treatments. Then the treatments are randomly assigned to these parts of
a block. Dividing the field into several homogenous parts is known as ‘blocking’. In general, blocks are the
levels at which we hold an extraneous factor fixed, so that we can measure its contribution to the total
variability of the data by means of analysis of variance. In brief, through the principle of local control we
can eliminate the variability due to extraneous factor(s) from the experimental error.
Remark: Blocking direction should be against the variability gradient!

Degrees of Freedom (D.F.)


Is the number of observations in a sample that are “free to vary” after knowing the mean when determining
the variance. After determining the mean, then only n-1 observations are free to vary because knowing the
mean and n-1 observations, the last observation is fixed. A simple example – say we have a sample of
observations, with values 3, 4 and 5. We know the sample mean (4) and we wish to estimate the variance.
Knowing the mean and one of the observations does not tell us what the other two must be. But if we know
the mean and two of the observations (e.g. 3 and 4), the final observation is fixed (it must be 5). So, knowing
the mean, only two observations (n-1) are free to vary. As a general rule, the d.f. is the number of
observations minus the number of parameters included in the formula for the variance.

3.4 Analysis of Variance

Analysis of variance was first introduced by R.A Fisher in 1930s. Analysis of variance is used in all fields
of research where data are quantitatively measured and it is used: to estimate and test about population
variance and to estimate and test about population means.
20
It is defined as an arithmetic technique whereby the total variation in a set of data is divided (or partitioned)
into meaningful components or parts to have a better idea how a particular object (person, plant, animal,
etc.) reacts to a change in the conditions applied to it and how the environment plays a role in the expression
of this reaction.

The total variance has two basic components:


1. Explained variance: variance due to block, and variance due to treatment or factor combinations. –
known factor
2. Unexplained variance: variance due to unknown or uncontrolled factors. This is also termed as
experimental error. Unknown factors
In conducting experiments, differences among plants (or animals or things) will occur but it is important to
be able to identify what types of differences there are: one is due to the treatment applied (treatment
differences), another is due to the condition in the area of the experiment (area or site differences), and last
is due to conditions that are not known or cannot be explained or controlled. The first two are commonly
referred to as explained variation while the last one is known as unexplained variation. In analysis of
variance, we want to know how much of the total variation is due to known or explained factors (explained
variance) and how much is due to unknown or unexplained factors (unexplained variance or unexplained
error).

The test of significance deals with computing the ratio between the explained and unexplained variances
and comparing with a probability value. If the ratio between explained and unexplained variances is
relatively high, we are confident to conclude that the difference was due to the known factors and not due to
unknown factors.

The purpose of proper experimental design is to make the experimental error (unexplained variation) small
enough in relation to the treatment (explained) effect.

The test of significance (F test) involves comparing each of the explained variances with the corresponding
unexplained variance by way of a ratio (F computed or Fc to the tabular F value). The higher the Fc relative
to the tabular F, the more significant are the treatment (or factor combination) differences. The Fc may
become large if the treatment variance is considerably large compared to the experimental error; or the
experimental error is made small in relation to the mean by proper experimental design and experimental
management.

3.4.1 General procedures in analysis of variance


 State the model: symbolic representation that describes data under consideration
 State the hypothesis
 Make calculations
 Construct analysis of variance table, which is used to summarize the calculations in table form for
quick assessment of the results
 Make decision: decide either to accept or reject the hypothesis

3.4.2 General assumptions underlying the analysis of variance

I. The treatment and replication effects should be additive

21
The effect of treatment for all replications and the effect of replication for all treatments should remain
constant.
A hypothetical set of data with additive and multiplicative effect of treatments and replications.

Additive effect

Treatment Replication Replication effect


I II (I-II)
A 180 120 60
B 160 100 60
Treatment effect (A-B) 20 20

Note that the effect of treatments is constant over replications and the effect of replications is constant over
treatments.

Multiplicative effect

Replication Replication effect Log10 Treatment effect

Treatment I II (II-I) I II
A 10 20 10 1.00 1.30 0.30
B 30 60 30 1.48 1.78 0.30
Treatment effect (B- 20 40 0.48 0.48
A)

Here, the treatment effect is not constant over replications and the effect of replications is not constant over
the treatments.

The multiplicative effects are often encountered in experiments designed to evaluate the incidence of
diseases and insects. This happens because the changes in insect and disease incidence usually follow a
pattern that is in multiple of the initial incidence. When effects are multiplicative, the logarithmic
transformations of the data show the additive effect. Thus, in such cases conduct the analysis and mean
separation using the transformed data, and in tables present transformed means in parenthesis alongside
their back transformed values out of parenthesis.

II. Experimental errors (residuals) must be independent


The experimental error of one treatment should not be related to or dependent upon that of another
treatment. One example, where residuals may be dependent is in the study of pesticides: a high level of
pesticide on one plot of land may spread to adjacent plots, thereby changing the conditions on the following
plots. The assumption of independence of errors is usually assured by the use of proper randomization, i. e
treatments are applied (assigned) at random to experimental units. However, in systematic designs where the
treatments are assigned systematically instead of randomly, the assumption of independence of errors is
usually violated. The simplest way to detect independence of error is to check the experimental layout if it is
random or not.

22
III. Experimental errors (variances) must be homogeneous (homoscedasticity)
The variances of the treatments must be the same. When some treatments have errors that are exceptionally
higher or lower it is called heteroscedasticity. First type of heterogeneity of variances is usually associated
with data whose distribution is not normal. Count data, such as the number of infected plants per plot
usually follow a poisson distribution where the variance equals to the mean ( ).

Heterogeneity of variances occurs usually when some treatments have errors that are exceptionally higher or
lower than others.

IV. Experimental errors are normally distributed


It can be assumed that the observations within each ‘group’ or treatment combinations come from a
normally distributed population.

Data such as number of infested plants per plot usually follow poisson distribution and data such as percent
survival of insects or percent plants infected with a disease assume the binomial distribution.

V. Randomness
Treatments should be assigned to experimental units completely at random. Randomness is known as the
fundamental assumption of ANOVA; because, its violation also leads to violation of other assumptions.

4. Completely Randomized Design (CRD)

4.1. Uses, advantages and disadvantages

In CRD, the treatments are assigned completely at random over the whole experimental area so that each
experimental unit has the same chance of receiving any one treatment. In CRD, any difference among the
experimental units (plots) receiving the same treatment is considered as experimental error.

Uses:
1. It is useful when the experimental units (plots) are essentially homogeneous and where
environmental effects are relatively easy to control, e.g. laboratory and greenhouse experiments. For
field experiments where there is generally larger variation among experimental plots like in soil
fertility, slope, etc. the CRD is rarely used.
2. It is useful if we suspect that large fraction of the units may not respond or may be lost during the
experiment because it is easy to handle missing data in Analysis of Variance unlike in other
designs.-unequal replication is possible
3. It is useful for experiments in which the total number of experimental units is limited, because it
provides maximum degrees of freedom for error

Advantages
1. It is flexible in that the number of treatments and replications can vary, i.e. the number of
replications need not be the same from one treatment to another
2. The statistical analysis is simple even with unequal replications and it is not complicated by loss of
data or missing observations
3. Loss of information due to missing data is small as compared to other designs

23
4. The design provides the maximum degree of freedom for estimating the experimental error. This
improves the precision of the experiment and is important with small experiments where degrees of
freedom for experimental error are less than 20.

Disadvantage:
The main objection to the CRD is that it is often inefficient? as there is no way of controlling the
experimental error. Since randomization is unrestricted the experimental error includes the entire variation
over the experimental units except that due to treatment.

4.2 Randomization and layout

In this design, treatments are assigned to the experimental units completely at random.
Assume that we want to do a pot-experiment on the effect of inoculation of 6-strains of rhizobia on
nodulation of common bean using five replications.

Randomization can be done by using either lottery method or table of random numbers.

A. Lottery Method
1. Arrange 30 pots of equal size filled with the same type of soil and assign numbers from 1 to 30 in
convenient order.
2. Obtain 30 identical slips of paper, label 5 of them with treatment A, 5 of them with treatment B, with C,
with D, with E and with F (6 treatments  5 replications). Place the slips in box or hat, mix thoroughly
and pick a piece of paper at random, the treatment labeled on this paper is assigned to unit 1 (pot 1),
without returning the first slip to box, select another slip and the treatment named on this slip is assigned
to unit 2(pot 2) and continue this way until all 30 slips of paper have been drawn.

B. Use of Table of Random Numbers


1. Arrange 30 pots of equal size filled with the same type of soil and assign numbers from 1 to 30 in
convenient order.
2. Locate starting point in table of random numbers by closing your eyes and pointing a finger to any
position of random number
3. Moving up to down or right to left from the staring, record the first 30 three digit random numbers in
sequence (avoid ties). Rank the random numbers from the smallest (1) to the largest (30). The ranks will
represent the pot numbers and assign treatment A to the first five pot numbers, B to the next five pot
numbers, etc.

Example: Three treatments each replicated four times in CRD

24
Raw data of nitrogen content of common bean inoculated with 6 rhizobium strains. Treatments are
designated with letters and nitrogen content (mg) in parenthesis.

1A 2E 3C 4B 5A 6D
(19.4) (14.3) (17.0) (17.7) (32.6) (20.7)
12 B 11C 10A 9E 8D 7E
(24.8) (19.4) (27) (11.8) (21.0) (14.4)
13 F 14A 15F 16D 17B 18.C
(17.3) (32.1) (19.4) (20.5) (27.9) (9.1)
24D 23F 22B 21E 20C 19D
(18.6) (19.1) (25.2) (11.6) (11.9) (18.8)
25C 26E 27B 28F 29A 30F
(15.8) (14.2) (24.3) (16.9) (33.0) (20.8)

4.3. Analysis of variance of CRD with equal replication

1. State the model


The model for CRD expressing the relationship between the response to the treatment and the effect of other
factors unaccounted for:

where, Yij = the jth observation on the ith treatment; General mean; i = Effect of treatment i; and
Experimental error (effect due to chance)

2. Arrange the data by treatments and calculate the treatment totals (T i) and Grand total (G)

Nitrogen content of common bean inoculated with 6 rhizobium strains (mg)


___________________________________________________
Treatments (R-strains) Reps Treatment
total (Ti)
___________________________________________________
3 Dok 1 19.4 32.6 27.0 32.1 33.0 144.1
3 Dok 5 17.7 24.8 27.9 25.2 24.3 119.9
3 Dok 4 17.0 19.4 9.1 11.9 15.8 73.2
25
3 Dok 7 20.7 21.0 20.5 18.8 18.6 99.6
3 Dok 15 14.3 14.4 11.8 11.6 14.2 66.3
Composite 17.3 19.4 19.1 16.9 20.8 93.5
___________________________________________________
Grand total (G) 596.6

3. Using Yij = jth observation on the ith treatment; Ti= Treatment total; n = (r x t), the total number of
experimental unit (pots), calculate the correction factor and the various sum of squares
 C.F. = = = 11864.38

 Total Sum of Squares (TSS) = yij2 – C.F. (Sum of the square of all observations- C.F.)
= [(19.4)2 + (17.7)2 +…… + (20.8)2] - 11864.38 = 12994.36–11864.38=1129.98

 Treatment Sum of Squares (SST) =

= - 11864.38=847.05
 Error Sum of Squares (SSE) = Total SS- Treatment SS= 1129.98-847.05= 282.93
In CRD, treatment sum of squares are usually called between or among groups sum of squares while the
sum of squares among individuals treated alike is called within group or error sum of squares.

4. Calculate the mean squares (MS) for treatment and error by dividing each sum of squares by the
corresponding degree of freedom
 Treatment MS = = = 169.41

 Error MS = = = 11.79
5. Calculate F-value for testing significance of treatment effects
 F-calculated = = = 14.37
6. Obtain the tabulated F-value using treatment degree of freedom (d. f.) as numerator (n 1) and error d. f. as
denominator (n2) at 5% and 1% level of significance
F (5, 24) at 5% = 2.60; F (5, 24) at 1% = 3.90
7. Summarize all the values computed on the above steps in the ANOVA- table for quick assessment of
results.
________________________________________________________________________
Source of DF SS MS Computed F Table F
Variation 5% 1%
________________________________________________________________________
Treatment (among strains) (t-1) = 5 847.05 169.41 14.37** 2.60 3.90
Error (within strains) t(r-1) = 24 282.93 11.79
Total (rt-1) = 29 1129.98
________________________________________________________________________

26
8. Compare the calculated F- value with table F- value and decide on significance among the treatment
effects using the following rules:
a) If F-calculated > F table at 1% level of significance, the difference between treatments is highly
significant. Put two asterisks on F-calculated
b) If F-calculated > F table at 5% level of significance, but ≤ F table at 1%, the difference between
treatments is significant. Put one asterisks on F-calculated.
c) If F-calculated ≤ F table at 5% level of significance, the differences among treatments is non-
significant. Put NS on the F- calculated value in ANOVA table.
Note that a non-significant F test in the analysis of variance indicates the failure of the experiment to detect
any difference among treatments. It does not, in any way, prove that all treatments are the same. The failure
to detect treatment difference based on non- significant F-test could be the result of either a very small or nil
treatment difference or a very large experimental error or both. Thus, whenever the F-test is non-significant,
the researcher should examine the size of experimental error and the numerical difference among the
treatment means. If both values are large, the trial may be repeated and efforts should be made to reduce
experimental error so that the differences among treatments, if any can be detected. On the other hand, if
both values are small, the difference among treatments is probably too small to be of any economic value
and, thus, no additional trials are needed.

For the above example, the computed F value of 14.37 is larger than the tabulated F value at the 1% level of
significance of 3.90. Hence, the treatment difference is said to be highly significant. In other words, chances
are less than 1 in 100 that all the observed differences among the six treatment means could be due to
chance. It should also be noted that such a significant F test verifies the existence of some differences
among the treatments tested but does not specify the particular pair (or pairs) of treatments that differ
significantly. To obtain this information, procedures for comparing treatment means are used.

9. Compute the Coefficient of Variation (CV) and standard error (SE) of the treatment means
- CV= 100 =  100= 17.3%
- SE = = = 1.53mg

Coefficient of Variation indicates the degree of precision with which the treatments are compared and it is a
good index of the reliability of the experiment. The smaller the CV, the more reliable the experiment is. The
CV values greatly vary with the type of experiment, experimental material or the character measured (e.g.
data on days to flowering have smaller CV than number of nodules in common bean as within treatment
variation is usually small in the former parameter than the later). In field experiments CV up to 30% are
common and usually lesser CV for laboratory and greenhouse experiments are expected than field
experiments.

4.4 Analysis of variance of CRD with unequal replication

Example: Twenty rats (n=20) were assigned equally at random to four feed types (t= 4). Unfortunately one
of the rats died due to unknown reason. The data are rat body weight in g after being raised on these diets
for 10 days. We would like to know whether weights of rats are the same for all four diets at 5%.

27
________________________________________________________________________
Feed 1 Feed 2 Feed 3 Feed 4
________________________________________________________________________
60.8 68.7 102.6 87.9
57.0 67.7 102.1 84.2
65.0 74.0 100.2 83.1
58.6 66.3 96.5 85.7
61.7 69.8 90.3
________________________________________________________________________
Ti 303.1 346.5 401.4 431.2
ni 5 5 4 5

- C.F. = = = 115627.20

- Total Sum of Squares (TSS) = yij2 – C.F. (Sum of the square of all observations- C.F.)
= [(60.8)2 + (57.00)2 +…… + (90.3)2] – 115627.20 = 4354.698

- Treatment Sum of Squares (SST) = - C.F.

= 4226.348
- Error Sum of Squares (SSE) = Total Sum of Squares – Treatment Sum of Squares: 4354.698 -
4226.348 = 128.35
________________________________________________________________________
Source of DF SS MS Computed F F-table (1%)
Variation
________________________________________________________________________
Treatment (among groups) (t-1) = 3 4226.348 1408.788 164.6** 5.42
Error (within groups)
(total d.f. – treatment d.f.) (18-3) = 15 128.35 8.557
Total (n-1) = 18 4354.698
________________________________________________________________________
** Highly significant

CV= = = 3.75%

5. Randomized Complete Block Design (RCBD)

5.1 Uses, advantages and disadvantages

It is the most frequently used experimental design in field experiments. Completely Randomized Design is
appropriate when no sources of variation other than treatment effects are known or anticipated, i.e. the
experimental units should be similar. However, in many experiments, certain experimental units that are
treated alike will behave differently. Example: in field experiments adjacent plots are more alike in
28
response than those distant apart. Heaviest animals in a group of the same age may show a different rate of
weight gain than lighter animals.

In such situations, designs and layouts can be constructed so that the portion of variability attributable to the
known source can be measured and excluded from experimental error. Thus, difference among treatment
means will contain no contribution of the known source.

Randomized Complete Block Design (RCBD) can be used when the experimental units can be meaningfully
grouped, the number of units in a group being equal to the number of treatments. Such a group is called a
block and equals to the number of replications. Each treatment appears an equal number of times usually
once, in each block and each block contains all the treatments.

Blocking (grouping) can be done based on soil heterogeneity in a fertilizer or variety trials; initial
body weight, age, sex, and breed of animals; slope of the field, etc.

The primary purpose of blocking is to reduce experimental error by eliminating the contribution of known
sources of variation among experimental units. By blocking, variability within each block is minimized and
variability among blocks is maximized.

During the course of the experiment, all units in a block must be treated as uniformly as possible. For
example,
- if planting, weeding, fertilizer application, harvesting, data recording, etc, operations cannot be done in
one day due to some problems, then all plots in any one block should be done at the same time.
- if different individuals have to make observations of the experimental plots, then one individual should
make all the observations in a block.

This practice helps to control variation within blocks, and thus variation among blocks is mathematically
removed from experimental error.

Advantages:
1. Precision: More precision is obtained than with CRD because grouping experimental units into
blocks reduces the magnitude of experimental error.
2. Flexibility: Theoretically, there is no restriction on the number of treatments? or replications. If
extra replication is desired for certain treatments, it can be applied to two or more units per block.
3. Ease of analysis: The statistical analysis of the data is simple. If as a result of change or missing, the
data from a complete block or for certain treatments are unusable, the data may be omitted without
complicating the analysis. If data from individual units (plots) are missing, they can be estimated
easily so that simplicity of calculation is not lost.

Disadvantages:
The main disadvantage of RCBD is that when the number of treatments is large (>15), variation among
experimental units within a block becomes large, resulting in a large error term. In such situations, other
designs such as incomplete block designs should be used.

29
5.2 Randomization and layout

Step 1: Divide the experimental area (unit) into r-equal blocks, where r is the number of replications,
following the blocking technique. Blocking should be done against the gradient such as slope, soil fertility,
etc.

Step 2: Sub-divide the first block into t-equal experimental plots, where t is the number of treatments and
assign t treatments at random to t-plots using any of the randomization scheme (random numbers or lottery).

Step 3: Repeat step 2 for each of the remaining blocks.

Example: three treatments each replicated four times

ENVIRONMENTAL GRADIENT

Block 1 Block 2 Block 3 Block 4

The major difference between CRD and RCBD is that in CRD, randomization is done without any
restriction to all experimental units but in RCBD, all treatments must appear in each block and different
randomization is done for each block (randomization is done within blocks).

5.3 Analysis of variance of RCBD

Step 1. State the model


The linear model for RCBD:
Yij =  + i + j + ij

where, Yij = the observation on the j th block and the ith treatment;  = common mean effect; i = effect of
treatment i; j = effect of block j; and ij = experiment error for treatments i in block j.

Step 2. Arrange the data by treatments and blocks and calculate treatment totals (T i), Block (rep) totals (Bj)
and Grand total (G).

Example: Oil content of linseed treated at different stages of growth with N-fertilizes.
30
Oil content (g) from sample of 20 g seed

Treatment (Stages Block 1 Block 2 Block 3 Block 4 Treat.


of application) Total (Ti)
Seeding 4.4 5.9 6.0 4.1 20.4
Early blooming 3.3 1.9 4.9 7.1 17.2
Half blooming 4.4 4.0 4.5 3.1 16.0
Full blooming 6.8 6.6 7.0 6.4 26.8
Ripening 6.3 4.9 5.9 7.1 24.2
Unfertilized 6.4 7.3 7.7 6.7 28.1
(control)
Block total (Bj) 31.6 30.6 36.0 34.5
Grand total 132.7

From this table we can test:


1. Does fertilizer application have effect on oil content? Comparing treatment 6 versus 1-5. (known as
contrast analysis or group comparison)
2. Which stage of application is best in increasing oil content in linseed? Comparing the difference
between 1, 2, 3, 4 and 5

Step 3: Compute the correction factor (C.F.) and sum of squares using r as number of blocks, t as number of
treatments, Ti as total of treatment i, and Bj as total of block j.
a. C.F. =
b. Total Sum of Square (TSS) = Yij2- C.F. = (4.4)2 + (3.3)2 + …. + (6.7)2 – C.F.
= 788.23 – 733.72 = 54.51

c. Block Sum of Squares (SSB) =

=
= 736.86 – 733.72 = 3.14

d. Treatment Sum of Squares (SST) =

=
= 765.37 – 733.72 = 31.65

e. Error Sum of Squares (SSE) = Total SS – SSB – SST


= 54.51 – 3.14 – 31.65 = 19.72

Step 4: Compute the mean squares for block, treatment and error by dividing each sum of squares by its
corresponding d.f.
31
Block Mean Square (MSB) =

Treatment Mean Square (MST) =

Error Mean Square (MSE) =

Step 5: Compute the F-value for testing block and treatment differences.
F-block = for block; F-treatment = for treatments.

F-block = ; F-treatment =

Step 6: Read table F- and compare the computed F-value with tabulated F-value and make decision.
- F-table for comparing block effects, use block d. f. as numerator (n 1) and error d.f. as denominator (n2);
F (3, 15) at 5% = 3.29 and at 1% = 5.42.
- F-table for comparing treatment effects, use treatment d.f. as numerator and error d.f. as denominator F
(5, 15) at 5% = 2.90 and at 1% = 4.56. If the calculated F-value is greater than the tabulated F-value for
treatments at 1%, it means that there is a highly significant (real) difference among treatment means.

In the above example, calculated F-value for treatments (4.83) is greater than the tabulated F-value at 1%
level of significance (4.56). Thus, there is a highly significant difference among the stages of application on
nitrogen content of the linseed.

Step 7: Compute the standard error of the mean and coefficient of variability.
- Standard error (SE)

Step 8: Summarize the results of computations in analysis of variance table for quick assessment of the
result.

Analysis of variance for oil content of the linseed.

Source of Degree of Sum of Mean Computed Tabulated F


variation Freedom Squares Squares F 5% 1%
Block (r-1) = 3 3.14 1.05 0.80 3.29 5.42
Treatment (t-1) = 5 31.65 6.33 4.83** 2.90 4.56
Error (r-1) (t-1) = 15 19.72 1.31
Total (rt-1) = 23 54.51

32
5.4. Block efficiency

Blocking maximizes the difference among blocks and reduces the difference among plots of the same block
(within blocks) as small as possible. Thus, the result of every RCBD should be examined to see whether this
objective has been achieved. The procedure to measure block efficiency is:

Step 1: Determine the level of significance of block variation by computing F-value for block and test its
significance.
F-block =
By comparing it with tabulated F-value at n1 (r-1) d.f. and n2 error d.f. (r-1) (t-1) = F (3, 15) at 5% = 3.29.

If the computed F-value is greater than the tabulated F-value, blocking is said to be effective in reducing
experimental error. Also the scope of an experiment may have been increased when blocks are significantly
different since the treatments have been tested over a wider range of experimental conditions.

On the other hand, if block effects are small (calculated F for block < tabulated F value), it indicates either
that the experimenter was not successful in reducing error variance by grouping of individual units
(blocking) or that the units were essentially homogeneous to start with.

Step 2: Determine the magnitude of the reduction in experimental error due to blocking by computing the
Relative Efficiency (R.E.) as compared to CRD.

R.E. =
Where MSB = block mean square; MSE = error mean square; r = number of replications; t = number of
treatments

R.E =

R.E. (RCBD to CRD) =

If error degree of freedom of RCBD is less than 20, the R.E. should be multiplied by the adjustment factor
(k) to consider the loss in precision resulting from fewer degrees of freedom.

33
Adjusted R.E. = 0.97  K(0.98) = 0.95.

In this case, information is sacrificed in theory by using Randomized Complete Block Design, since 95
replicates in a completely randomized design give as much information as 100 blocks or replications of a
Randomized Complete Block Design. That means, RCBD was less efficient than CRD.

5.5. Missing Data in RCBD

Sometimes data for certain units may be missing or become unusable. For example,
- when an animal becomes sick or dies but not due to treatment
- when rodents destroy a plot in field
- when a flask breaks in laboratory
- when there is an obvious recording error

A method is available for estimating such data. Note that an estimate of a missing value does not supply
additional information to the experimenter; it only facilitates the analysis of the remaining data.

Case 1: When a single value is missing


When a single value is missing in RCBD, an estimate of the missing value can be calculated as:

where;

Y = estimate of the missing value


t = No. of treatments
r = No. of replications or blocks
Bo = total of observed values in block (replication) containing the missing value
To = total of observed values in treatment containing the missing value
Go = Grand total of all observed values.

The estimated value is entered in the table with the observed values and the analysis of variance is
performed as usual with one d. f. being subtracted from both total and error d.f because the estimated value
makes no contribution to the error sum of squares.

When all of the missing values are on the same block or treatment the simplest solution is to consider
as if the block or treatment had not been included in the experiment.

Example: In the table given below are yields (kg) of 4-varieties of maize (Al-composite, Rarree-1, Bukuri,
Katumani) in 4-replications planted in RCBD on a plot size of 10 m x 10 m of which one plot yield is
missing. Estimate the missing value and analyze the data

Treatment
Varieties I II III IV total (Ti)
Bukuri 18.5 15.7 16.2 14.1 64.5
Katumani 11.7 - 12.9 14.4 39(To)
Rarree-1 15.4 16.6 15.5 20.3 67.8
Al-composite 16.5 18.6 12.7 15.7 63.5
34
Block total 62.1 50.9 (Bo) 57.3 64.5
Go (Grand total) = 234.8

Solution
a. Estimate the missing value

b. Enter the estimated value and carry out the analysis following the usual procedure:
- Corrected treatment total = 39 + 13.9 = 52.9
- Corrected block total = 50.9 + 13.9 = 64.8
- Corrected grand total = 234.8 + 13.9 = 248.7
c. Analysis of variance
1. C.F. =

2. TSS=

3. Treatment SS =

4. Block SS =

5. Error SS= Total SS- Treatment SS- Block SS = 79.18 – 31.21 – 9.02 = 38.95

d. Compute the correction factor for bias (B) for treatment sum of squares as the treatment SS is biased
upwards.

B=

Bo = Total of observed values in blocks (replication) containing the missing value

Y = estimated value

B=

=
e. Subtract the computed B value from Total SS & Treatment SS
- Adjusted Treatment SS = Treatment SS – B
= 31.21 – 7.05= 24.16
35
- Adjusted Total SS = Total SS-B
= 79.18 – 7.05 = 72.13

f. Subtract 1 from error d. f. and total d. f. and complete the analysis of variance table.

Source of Degree of Sum of Mean Computed Table F


variation freedom squares squares F 5% 1%
ns
Block 3 9.02 3.01 0.61 4.07 7.59
ns
Treatment 3 24.16 8.05 1.64 4.07 7.59
Error (t-1) (r-1)-1 = 8 38.95 4.9
Total rt-1-1= 14 72.13

CV=

6. Latin Square Design

6.1. Uses, advantages and disadvantages

The major feature of the Latin Square Design is its capacity to simultaneously handle two known sources of
variation among experimental units unlike Randomized Complete Block Design (RCBD), which treats only
one known source of variation.

The two directional blocking in a Latin Square Design is commonly referred as row blocking and column
blocking. In Latin Square Design the number of treatments is equal to the number of replications that is why
it is called Latin Square.

e.g. Soil fertility gradient Soil fertility gradient

Slope gradient Soil fertility gradient

Advantages:
- Greater precision is obtained than Completely Randomized Design & Randomized Complete Block
Design (RCBD) because it is possible to estimate variation among row blocks as well as among column
blocks and remove them from the experimental error.
Disadvantages:

36
- As the number of treatments is equal to the number of replications, when the number of treatments is
large the design becomes impractical to handle. On the other hand, when the number of treatments is
small, the degree of freedom associated with the experimental error becomes too small for the error to be
reliably estimated. Thus, in practice the Latin Square Design is applicable for experiments in which the
number of treatments is not less than four and not more than eight.
- Randomization is relatively difficult.

6.2. Randomization and layout

Step 1: To randomize a five treatment Latin Square Design, select a sample of 5 x 5 Latin square plan from
appendix of statistical books. We can also create our own basic plan and the only requirement is that each
treatment must appear only once in each row and column. For our example, the basic plan can be:

A B C D E
B A E C D
C D A E B
D E B A C
E C D B A

Step 2: Randomize the row arrangement of the plan selected in step 1, following one of
the randomization schemes (either using lottery method or table of random numbers).

- Select from table of random numbers, five three digit random numbers avoiding ties if any
Random numbers: 628 846 475 902 452
Rank: (3) (4) (2) (5) (1)
- Rank the selected random numbers from the lowest (1) to the highest (5)
- Use the ranks to represent the existing row number of the selected plan and the sequence to represent the
row number of the new plan. For our example, the third row of the selected plan (rank 3) becomes the
first row (sequence) of the new plan, the fourth becomes the second row, etc.

1 2 3 4 5
3 C D A E B
4 D E B A C
2 B A E C D
5 E C D B A
1 A B C D E

Step 3: Randomize the column arrangement using the same procedure. Select five three digit random
numbers.
Random numbers: 792 032 947 293 196
Rank: (4) (1) (5) (3) (2)

The rank will be used to represent the column number of the above plan (row arranged) in step 2. For our
example, the fourth column of the plan obtained in step 2 above becomes the first column of the final plan,
the first column of the plan becomes 2, etc.

Final layout:
37
E C B A D
A D C B E
C B D E A
B E A D C
D A E C B

Note that each treatment occurs only once in each row and column

Sample layout of three treatments each replicated three times in Latin Square Design

6.3. Analysis of variance

There are four sources of variation in Latin Square Design, two more than that of CRD and one more than
that for the RCBD. The sources of variation are row, column, treatment and experimental error.

The linear model for Latin Square Design:


Yijk =  + i + j + k + ijk

where,
Yijk = the observation on the ith treatment, jth row & kth column
 = Common mean effect
i = Effect of treatment i
j = Effect of row j
k = Effect of column k
ijk = Experiment error (residual) effect

Example: Grain yield of three maize hybrids (A, B, and D) and a check variety, C, from an experiment with
Latin Square Design.

Grain yield (t/ha)

Row number C1 C2 C3 C4 Row total (R)


R1 1.640(B) 1.210(D) 1.425(C) 1.345(A) 5.620
R2 1.475(C) 1.15(A) 1.400(D) 1.290(B) 5.315
38
R3 1.670(A) 0.710(C) 1.665(B) 1.180(D) 5.225
R4 1.565(D) 1.290(B) 1.655(A) 0.660(C) 5.170
Column total(C) 6.350 4.360 6.145 4.475
Grand total (G) 21.33

Treatment Total (T) Mean


A 5.820 1.455
B 5.885 1.471
C 4.270 1.068
D 5.355 1.339

STEPS OF ANALYSIS

Step 1: Arrange the raw data according to their row and column designation, with the corresponding
treatments clearly specified for each observation and compute row total (R), column total (C), the grand
total (G) and the treatment totals (T).

Step 2: Compute the C.F. and the various Sum of Squares


C.F. =

Total SS = y2 – C.F. =

Row SS = =

Column SS= = = 0.83

Treatment SS= = = 0.43


Error SS =Total SS–Row SS–Column SS–Treatment SS=1.41–0.03–0.83–0.43 = 0.12

Step 3: Compute the mean squares for each source of variation by dividing the sum of squares by its
corresponding degrees of freedom.

Row MS = ; Column MS =

Treat. MS = ; Error MS =

Step 4: Compute the F-value for testing the treatment effect and read table F-value as:

As the computed F-value (7.15) is higher than the tabulated F-value at 5% level of significance (4.76), but
lower than the tabulated F-value at the 1% level (9.78), the treatment difference is significant at the 5% level
of significance.

39
Compute the CV as: = =
Note that although the F-test on the analysis of variance indicates significant differences among the mean
yields of the 4-maize varieties tested, it does not identify the specific pairs or groups of varieties that
differed significantly. For example, the F-test is not able to answer the question whether every one of the
three hybrids gave significantly higher yield than that of the check variety. To answer these questions, the
procedure for mean comparison should be used.

Step 5: Summarize the results of the analysis in ANOVA table

Source D.F. SS MS Computed F Table F


5% 1%
Row (t-1)=3 0.03 0.01 0.50ns 4.76 9.78
Column (t-1)=3 0.83 0.275 13.75** 4.76 9.78
Treatment (t-1)=3 0.43 0.142 7.15* 4.76 9.78
Error (t-1)(t-2)=6 0.13 0.02
Total (t2-1) =15 1.41

6.4 Relative efficiency

As in RCBD, where the efficiency of one way blocking indicates the gain in precision relative to CRD, the
efficiencies of both row and column blocking in a Latin Square Design indicate the gain in precision relative
to either the CRD or RCBD, the procedures are:
i. Compute the F-value for testing the row & column effects; and test their significance

F (row) = = = 0.50; F (column) = = 13.75**


ii. Compute the relative efficiency of LS design relative to CRD & RCBD

The relative efficiency of LS design as compared to CRD

R.E (CRD) =

= = = 3.45
This indicates that the use of Latin Square Design in the present example is estimated to increase the
experimental precision by 245% as compared to CRD. This result implies that if the CRD is used an
estimated 2.45 times more replication would have been required to detect the treatment difference of the
same magnitude as that detected with the Latin Square Design.

- The R.E. of LS design as compared to RCBD can be computed in two ways:

- When row is used as blocking factor


40
R.E. (row) = = = = 0.875
- When column used as blocking factor
R.E. (column) = = = = 4.19

When the error d. f. in the Latin Square analysis of variance is < 20, the R.E. value should be multiplied by
the adjustment factor (K) defined as:

K= = = = = 0.93

The adjusted R.E. values are computed as:

R.E. (row) = 0.875  0.93 = 0.81


R.E. (column) = 4.19  0.93 = 3.90

The results indicate that the additional column blocking made possible by the use of Latin Square Design is
estimated to have increased the experimental precision over that of RCBD by 290%, whereas the additional
row-blocking in the LS design did not increase precision over the RCBD with column as blocks. Hence, for
the above trial, a RCBD with column as blocks would have been as efficient as a Latin Square Design.

6.5. Missing data

The formula for a single missing observation


Y=
Where Ro, Co, and To are the totals of the observed values for the row, column, and treatment containing
the missing value, respectively, and Go is the grand total of the observed values, and t is the number of
treatments.

The analysis of variance is performed in the usual manner after entering the estimated value with one degree
of freedom being subtracted from total and error degrees of freedom for each missing value.

As in the case of RCBD, the treatment sum of squares is biased upward by:

Bias (B) =

Where Go, Ro, Co, To and t are as described above. Then B is subtracted from treatment SS & total SS.

Example: Yield (kg) of five rice varieties tested in Latin Square Design from plot size of 100 m 2.

41
E(12) C (-) B(11) A (10) D (8)
A (7) D (8) C(8) B(7) E (13)
C (12) B (6) D(7) E (11) A (9)
B (4) E (10) A (7) D (6) C (8)
D (5) A (8) E (15) C(9) B (5)

Estimate the missing value, complete the analysis of variance and compare the variety C with D, and A with
E at 5% level of significance using LSD test.

a. Estimate the missing value


Y=

Y=
b. Enter the estimated value and carry out the analysis following the usual procedure:
- Corrected row total = 41.0 + 11.5 = 52.5
- Corrected column total = 32.0 + 11.5 = 43.5
- Corrected treatment total = 37+ 11.5 = 48.5
- Corrected grand total = 206.0 + 11.5 = 217.50
c. Compute the C.F. and the various Sum of Squares
C.F. =

Total SS = y2 – C.F. =

Row SS = =

Column SS = = = 6.60

Treatment SS = = = 107.6
Error SS = Total SS–Row SS–Column SS – Treatment SS = 180.0–31.6–6.6–107.6= 34.2

d. Compute the correction factor for bias (B) for treatment sum of squares as the treatment SS is biased
upwards.

Bias (B) = =

e. Subtract the computed B value from Total SS & Treatment SS


Adjusted Treatment SS = Treatment SS – B = 107.6-1.56 = 106.04
Adjusted Total SS = Total SS-B = 180.0-1.56 = 178.44

f. Subtract 1 from error d. f. and total d. f. and complete the analysis of variance table.

Source D.F. SS MS Computed Table F


F 5% 1%
42
Row (5-1)=4 31.6 7.90 2.54 3.36 5.67
Column (5-1)=4 6.6 1.65 0.53 3.36 5.67
Treatment (5-1)=4 106.04 26.51 8.52** 3.36 5.67
Error (5-1)(5-2)-1 = 11 34.2 3.11
Total (t2-1)-1 =23 178.44

CV =

To compare treatments C & D, LSD5% = t 0.025(11)  where ;

= 1.226

LSD5% = 2.201 x 1.226 = 2.70 kg

Difference between the treatment means of C & D (37/4-34/5) = 9.25-6.8 = 2.45. Since the difference is less
than LSD value, there is no significant difference between treatments C & D.

To compare the treatments A & E both with equal replication (without missing value)
LSD5% = t 0.025(11)  , where = = 1.115
LSD5% = 2.201  1.115 = 2.45 kg

Difference between the treatment means of A & E (61/5-41/5) = 12.2-8.2 = 4.00. Since the difference is
greater than LSD value, there is significant difference between treatments A & E.

SAS Syntax

data Lattice;
do Column=1 to 4;
do Row=1 to 4;
input Treatment Yield @;
output;
end;
end;
datalines;
2 1.64 3 1.475 1 1.67 4 1.565
4 1.21 1 1.15 3 0.71 2 1.29
3 1.425 4 1.4 2 1.665 1 1.655
1 1.345 2 1.29 4 1.18 3 0.66
;
Proc ANOVA Data=Lattice;
class Column Row Treatment;
model Yield= Column Row Treatment;

43
Means Treatment/lsd;
run;

OR

Data Lattice;
Input Row column Treatment Yield;
cards;
1 1 2 1.64
1 2 4 1.21
1 3 3 1.425
1 4 1 1.345
2 1 3 1.475
2 2 1 1.15
2 3 4 1.4
2 4 2 1.29
3 1 1 1.67
3 2 3 0.71
3 3 2 1.665
3 4 4 1.18
4 1 4 1.565
4 2 2 1.29
4 3 1 1.655
4 4 3 0.66
;
Proc ANOVA;
Class Row Column Treatment;
Model Yield= Row Column Treatment;
Means Treatment/lsd;
run;

The ANOVA Procedure

Dependent Variable: Yield

Sum of
Source DF Squares Mean Square F Value Pr > F

Model 9 1.29244375 0.14360486 6.47 0.0170

Error 6 0.13315000 0.02219167

Corrected Total 15 1.42559375

44
R-Square Coeff Var Root MSE Yield Mean

0.906600 11.17440 0.148969 1.333125

Source DF Anova SS Mean Square F Value Pr > F

Row 3 0.03023125 0.01007708 0.45 0.7240


column 3 0.84413125 0.28137708 12.68 0.0052
Treatment 3 0.41808125 0.13936042 6.28 0.0279

7. INCOMPLETE BLOCK DESIGNS

Theoretically, the complete block designs (where each block contains all the treatments) such as Randomized
Complete Block and Latin Square are applicable to experiments with any number of treatments. However,
these complete block designs become less efficient as the number of treatments increases, mainly because block
size increases proportionally with the number of treatments which in turn increases experimental error.

An alternative set of designs for single factor experiments having a large number of treatments are the
incomplete block designs. For example, plant breeders are often interested in making comparisons among a
large number of selections in a single trial. For such trials, we use incomplete block designs. As the name
implies, the experimental units in these designs are grouped into blocks which are smaller than a complete
replication of the treatments. However, the improved precision with the use of an incomplete block designs
(where the blocks do not contain all the treatments) is achieved with some costs. The major ones are:
- inflexible number of treatments or replications or both.
- unequal degree of precision in the comparison of treatment means.
- complex data analysis.

Although there is no concrete rule as to how large the number of treatments should be before the use of an
incomplete block design, the following points may be helpful:

a. Variability in the experimental material: The advantage of an incomplete block design over complete
block designs is enhanced by an increased variability in the experimental material. In general, whenever
the block size in Randomized Complete Block Design is too large to maintain reasonable level of
uniformity among experimental units within the same block, the use of an incomplete block design should
be seriously considered.

b. Computing facilities and services: Data analysis of an incomplete block design is more complex than that
for a complete block design. Thus, the use of an incomplete block design should be considered only as the
last measure.
45
7.1. Lattice Designs

The lattice designs are the most commonly used incomplete block designs in agricultural experiments. There is
sufficient flexibility in the design to make its applications simpler than most of the other incomplete block
designs. There are two kinds of lattices: balanced lattice and partially balanced lattice designs.

Field arrangement & randomization:


- Blocks in the same replication should be made as nearly alike as possible
- Randomize the order of blocks within replication using a separate randomization in each replication.
- Randomize treatment code numbers separately in each block.

7.1 Balanced lattices


The balanced lattice design is characterized by the following basic features:
a. The number of treatments (t) must be perfect square such as 16, 25, 36, 49, 64, etc.

b. The block size (k) is equal to the square root of the number of treatments, i.e. k=
c. The number of replications (r) is one more than the block size, i.e. r = k + 1. That is, the number of
replications required is 6 for 25 treatments, 7 for 36 treatments, and so on.
As balanced lattices require large number of replications, they are not commonly used.

7.2 Partially balanced lattices

The partially balanced lattice design is more or less similar to the balanced lattice design, but it allows for a
more flexible choice of the number of replications. The partially balanced lattice design requires that the
number of treatments must be a perfect square and the block size is equal to the square root of the number of
treatments. However, any number of replications can be used in partially balanced lattice design. The partially
balanced lattice design with two replications is called simple lattice, with three replications is triple lattice and
with four replications is quadruple lattice, and so on. However, such flexibility in the number of replications
results in the loss of symmetry in the arrangement of the treatments over blocks (i.e. some treatment pairs
never appear together in the same incomplete block). Consequently, the treatment pairs that are tested in the
same incomplete block are compared with higher level of precision than for those that are not tested in the
same incomplete block. Thus, partially balanced designs are more difficult to analyze statistically, and several
different standard errors may be possible.

Example 1 (Lattice with adjustment factor): Field arrangement and broad leaved weed kill (%) in tef fields
of Debre zeit research center by 16-herbicides tested in 4  4 Triple Lattice Design (Herbicide numbers in
parenthesis).
________________________________________________________________________
Block
Replication Block % kill total (B) M Cb
_________________________________________________________________________
1 1 75(15) 57(16) 71(13) 77(14) 280 789 -51
2 78(12) 66(11) 68(10) 45(9) 257 716 -55
3 40(6) 64(5) 49(8) 42(7) 195 608 23
4 59(3) 53(1) 46(2) 54(4) 212 642 6
46
Rep total (R1) 944 -77
______________________________________________________________________
2 1 53(16) 66(4) 57(12) 47(8) 223 663 -6
2 80(14) 48(6) 73(10) 52(2) 253 700 -59
3 36(7) 63(11) 67(15) 47(3) 213 676 37
4 68(13) 60(1) 50(9) 76(5) 254 716 -46
Rep total (R2) 943 -74
_________________________________________________________________________
3 1 66(15) 46(2) 58(12) 69(5) 239 754 37
2 46(4) 40(7) 59(13) 55(10) 200 678 78
3 43(9) 55(3) 50(8) 68(14) 216 670 22
4 60(11) 58(1) 48(16) 47(6) 213 653 14
Re total (R3) 868 151

Grand Total (G) 2755 0


_________________________________________________________________________

Analysis of variance

Step 1: Calculate the block total (B), the replication total (R) and Grand Total (G) as shown above.
Step 2: Calculate the treatment totals (T) by summing the values of each treatment from the three replications.

Treat ACb Adjusted Adjusted


. Treatment Sum (0.0813Cb treatment treatment
No. total (T) Cb ) total (T') (T+Acb) mean (Ti/3)
1 171 -26 -2.11 168.89 56.29
2 144 -16 -1.30 142.70 47.57
3 161 65 5.28 166.28 55.43
4 166 78 6.34 172.34 57.45
5 209 14 1.14 210.14 70.05
6 135 -22 -1.79 133.21 44.40
7 118 138 11.22 129.22 43.07
8 146 39 3.17 149.17 49.72
9 138 -79 -6.42 131.58 43.86
10 196 -36 -2.93 193.07 64.36
11 189 -4 -0.32 188.67 62.89
12 193 -24 -1.95 191.05 63.68
13 198 -19 -1.54 196.45 65.48
14 225 -88 -7.15 217.85 72.61
15 208 23 1.87 209.87 69.96
16 158 -43 -3.49 154.50 51.50

Step 3: Using r as number of replications and k as block size, compute the total sum of squares, replication
sum of squares, treatment (unadjusted) sum of squares as:
- Correction Factor (C. F.) = = = 158125.5

47
- Total Sum of Squares = (75)2 + (57)2 + .... + (47)2 – C.F.
= 164233- 158125.5 = 6107.5
- Replication Sum of Squares = - C. F.

= -158125.5 = 158363.06- 158125.5 = 237.6

- Treatment (unadjusted) Sum of Squares = - C. F.

= -158125.5 = 163029- 158125.5 = 4903.5


Step 4: For each block, calculate block adjustment factor (Cb) as:

Cb = M -rB where M is the sum of treatment totals for all treatments appearing in that particular block, B is
the block total and r is the number of replications. For example, block 2 of replication 2 contains treatments
14, 6, 10, and 2. Hence, the M value for block 2 of replication 2 is: M = T14 + T6 + T10 + T2 = 225 + 135 +
196 + 144 = 700 and the corresponding Cb value is: Cb = 700 - (3 x 253) = -59. The Cb values for the blocks
are presented in the above table. Note that the sum of Cb values over all replications should add to zero (i. e. -
77 + -74 + 151 = 0).

Step 5: Calculate the Block (adju.) Sum of Squares as:

Block (adj.) SS = - , where Rcb is the replication total of Cb values

= -
= 888.58-356.31 = 532.27

Step 6: Calculate the intra-block error sum of squares as:


- Intra-block error SS = Total SS - Rep. SS - Treatment (unadjusted) SS – Block (adjusted) SS
= 6107.5 - 237.6 - 4903.5 - 532.27 = 434.13

Step 7: Calculate the intra-block error mean square (MS) and block (adj.) mean square (MS) as:
- Intra-block error MS = = = 20.67

- Block (adjusted) MS = = = 59.14


Note that if the adjusted block mean square is less than intra-block error mean square, no further adjustment is
done for treatment. In this case, the F-test for significance of treatment effect is made in the usual manner as
the ratio of treatment (unadjusted) mean square and intra-block error mean square and steps 8 to 13 can be
ignored. For our example, the MSB value of 59.14 is greater than the MSE value of 20.67, thus, the
adjustment factor is computed.

Step 8: Calculate adjustment factor A. For a triple lattice design, the formula is:

48
A= = = 0.0813

where Eb is the block (adju.) mean square and Ee is the intra-block error mean square, and k is block size.

Step 9: For each treatment, calculate the adjusted treatment total (T') as:
T' = T + A where the summation runs over all blocks in which that particular treatment appears. For
example, the adjusted treatment totals for treatment number 1 and 2 are computed as:
T'1 = 171 + 0.0813(6 + -46 + 14) = 168.89
T'2 = 144 + 0.0813(6 + -59 + 37) = 142.70
.
.
etc.
The adjusted treatment totals (T') and their respective means (adjusted treatment total divided by the number
of replications (3) are presented along with the unadjusted treatment totals (T) in the table above.

Step 10: Compute the adjusted Treatment Sum of Squares as:

Treatment (unadjusted) SS – [Ak(r-1)

- SSBun (unadjusted block sum of squares) = - C. F. - SSR; where B is Block total, and SSR is
Replication Sum of Squares.
= - 158125.5 - 237.6
= 160046.75-158125.5 - 237.6 = 1683.65

Thus, the adjusted treatment sum of squares


= 4903.5 – [(0.0813  4  2) = 4010.20

Step 11: Compute the treatment (adjusted) mean square as:

- Treatment (adj.) MS = = = 267.35


Step 12: Compute the F-value for testing the significance of treatment difference and compare the computed
F-value with the tabulated F-value with k2 - 1 = 15 degree of freedom as numerator and (k - 1) (rk - k - 1) = 21
as denominator.
- Computed F = = = 12.93
Step 13: Compare the computed F-value with the table F-value
F0.05 (15, 21) = 2.18
F0.01 (15, 21) = 3.03
Since the computed F-value of 12.93 is greater than the tabulated F-value at 1% level of significance (3.03),
the differences among the herbicide means are highly significant.
49
Step 14: Enter all values computed in the analysis of variance table
___________________________________________________________________________
Source of Degrees of Sum of Mean Computed Tabulated F
Variation freedom squares squares F 5%1%
___________________________________________________________________________
Replication (r-1) 2 237.6 118.8
Block (adj.) r (k-1) 9 532.27 59.14
2
Herbicide (unadj.) k -1 15 4903.5 326.9
Intra-block error (k-1) (rk-k-1)21 434.13 20.67
2
Herbicide (adju.) k -1 (15) (4010.2) 267.35 12.93 2.18 3.03
Total rk2- 1 47 6107.5
___________________________________________________________________________
Step 15: Compute the corresponding Coefficient of Variation (CV) as:

CV = =  100 = 7.9%
Note that the grand mean is same for adjusted and unadjusted treatment total.

Step 16: Compute the gain in precision of the triple lattice relative to that of Randomized Complete Block
Design as:

- % Relative precision = ; where SSB is Block (adj.) SS, SSE is intra-block

error SS, r is number of replication, and k is block size.


- Ee' (effective error mean square) = (1 + ) x Ee, where r is the number of replications, k is block
size, A is adjustment factor, Ee is intra-block error mean square.

= /(1 + )  20.67
Thus, the relative precision = (32.2/24.7)  100 = 130.4

This indicates that the precision of this experiment was increased by about 30.4% by using the triple lattice
instead of Randomized Complete Block Design.

To compare between adjusted treatment mean differences


= ; where Ee is intra-block error mean square; r is number of replications
and A is adjustment factor

LSD1% = t 0.005(21) = 2.831 3.88 = 10.97%

50
7.3. Augmented Block Design

Augmented Block Design is used:


- when there are more number of entries/genotypes
- when there is no enough seed for test entries
- when there is no sufficient fund and land resources for replicated trials
It is not powerful design, but it is used for preliminary screening of genotypes/entries. There is no
assumption of block homogeneity.

Any new material is not replicated; it appears only once in the experiment while check varieties/entries
occur as the number of the blocks.

- Error d. f. = (b-1) (c-1), where b = No. of blocks and c = Number of checks.

Thus, the minimum number of checks is 2 because error d. f. should be 12 or more for valid comparison. If
number of checks is 1, error d.f. becomes zero which is not valid.

Randomization and Layout


Divide the experimental field into blocks. The block size may not be equal, e.g., some blocks may contain
10 entries while the others may contain 12 entries.

Suppose we have 50 test lines/progenies and 4 checks, we need at least 5 blocks since the error d. f. (c - 1)
(b - 1) should be  12. The higher the number of blocks, the higher the precision.

If we have 5 equal blocks, block size will be 14 (10 test lines + 4 checks), but block size may vary.

Randomization

Two possibilities

A. Conveniency

For identification purpose, sometimes the checks are assigned at the start, end or in certain intervals in the
block.

Block 1 = P1 P2 P3 P4 …….. P10 A B C D


Block 2 = P11 P12 ...... P20 B C D A
Block 3 = P21 P22 ..... P30 C D A B
.
.
etc., where P1, P2, .......... P30 are test entries and A, B, C, D are checks.

B. Randomly allocate the checks in a block.

b1 = P1 P2 A P3 B P4 D P5 C P6 P7 ..... P10
b2 = D P1 P2 C P3 P5 A B ….....
etc.
51
One of the blocks can be 10 (test culture) + 4 (checks), while the other 12 (test cultures) + 4 (checks).
Missing test entries (Pi) do not create problem in analysis as the analysis can be done with existing
genotypes. But when checks are missing, the analysis becomes complicated.

Assume we want to test 16 new rice genotypes (test lines) = P1, P2, ..., P16 in block size of 4 with 4 checks to
screen for early maturity.
b (number of blocks) = 4
c (number of checks) = 4 (A B C D).
Block size = 4 test cultures + 4 checks = 8

Days to maturity of rice genotypes


Block 1

P1 A P2 B P3 C P4 D
120 83 100 77 90 70 85 65

Block 2

P5 B P6 C P7 A P8 D
88 76 130 71 105 84 110 64

Block 3

P9 D P10 A P11 B P12 C


102 63 140 86 135 78 138 69

Block 4

P13 A P14 D P15 C P16 B


84 82 90 63 95 68 103 75

Steps of Analysis of Variance


Preliminary steps
1. Calculate block total for each block.
b1 = 690 (120 + 83 + …. + 65); b2 = 728; b3 = 811; b4 = 660
2. Calculate total of progenies/test cultures in a particular block.
P block1 = 395; P block2= 433; P block3 = 515; P block4 = 372
3. Construct check by block two-way table and calculate check total, check mean, check effect, sum of
check totals & total of check means.

Check/ b1 b2 b3 b4 Check Check Check effect (check mean- Check total x Check
block total mean adjusted grand mean) effect
A 83 84 86 82 335 83.75 -16.67 -5584.45
B 77 76 78 75 306 76.50 -23.92 -7319.52
C 70 71 69 68 278 69.50 -30.92 -8595.76
D 65 64 63 63 255 63.75 -36.67 -9350.85
Sum 1174 -30850.58
Total of check means 293.5

52
Blocks ni Block Total of test No. of test cultures in Block T i x Be Block effect x
total culture in a a block (Ti) effect(Be) Block total
(Bj) block (Bti)
B1 8 690 395 4 0.375 1.5 258.75
B2 8 728 433 4 0.375 1.5 273.00
B3 8 811 515 4 0.625 2.5 506.87
B4 8 660 372 4 -1.375 -5.5 -907.5
 32 2889 0 0 131.12

3.1. Estimation of block effect

bi =  (Total of the ith block - total of all check means - total of all progenies/test
cultures in the block];
bi = 0

1
b1 =  (690 - 293.5 – 395) = 0.375
4
1
b2 =  (728 - 293.5 - 433) = 0.375
4
1
b3 =  (811 - 293.5 - 515) = 0.625
4
1
b4 =  (660 - 293.5 - 372) = -1.375
4
Where ni is number of entries (test culture + checks) in each block = 4 + 4 = 8, ni = N = 32.

Calculate adjusted grand mean as:

[Grand total - (b - 1) (total of all check mean)] – 


no. of test cultures in a block  corresponding block effect]

1
Adjusted grand mean = 4 16  [2889 - (4 - 1) (293.5) - 0]

= 100.42

Grand total = Bi (sum of block total) or sum of all observations = 2889.0

Estimate check effect (Ci):

Ci= Check mean - adjusted grand mean


C1= 83.75-100.42 = -16.67; C2= 76.5 - 100.42 =-23.92

53
C3 = 69.50-100.42 = -30.92; C4= 63.75-100.42 =-36.67

There are as many check effects as the number of checks (4 in this case).

Adjust progeny value per ith progeny (Pi) as:

= Observed (unadjusted) progeny value - effect of block in which the ith progeny is occurring
P1 (adjusted) = P1 - (block effect) = 120 - (+ 0.375) = 119.62, etc.
Progeny/ Observed Block Adjusted progeny Progeny effect Observed
test progeny effect value (Po- bi) (Adjusted progeny progeny
culture no. value (Po) (bi) value-Adjusted value 
(Pi) grand mean progeny
effect
1 120 0.375 119.625 19.205 2304.6
2 100 0.375 99.625 -0.795 -79.5
3 90 0.375 89.625 -10.795 -971.6
4 85 0.375 84.625 -15.795 -1343
5 88 0.375 87.625 -12.795 -1126
6 130 0.375 129.625 29.205 3796.7
7 105 0.375 104.625 4.205 441.53
8 110 0.375 109.625 9.205 1012.6
9 102 0.625 101.375 0.955 97.41
10 140 0.625 139.375 38.955 5453.7
11 135 0.625 134.375 33.955 4583.9
12 138 0.625 137.375 36.955 5099.8
13 84 -1.375 85.375 -15.045 -1264
14 90 -1.375 91.375 -9.045 -814.1
15 95 -1.375 96.375 -4.045 -384.3
16 103 -1.375 104.375 3.955 407.37
Sum 1715 17216

Step 3.5. Estimate progeny effect as: Adjusted progeny value - adjusted grand mean;
For example, progeny effect for progeny 1 = 119.62-100.42 = 19.20; etc.

Analysis of variance
1. Correction Factor (C.F.) = = = 260822.53
2. Total SS = Y2 – C.F. = Sum of the squares of all observations – C.F. = (120)2 + (83)2 + ... + (75)2 -
260822.53 = 276621-260822.53
= 15798.47

3. Crude block SS =  = + + +

54
= 262425.62
4. True block SS: Crude block SS – C.F. = 262425.62 - 260822 .53
= 1603.09
5. Adjusted SS Due to entries (C + P) = (Adjusted grand mean  Observed grand total) + [(
)+( )+(
(Crude block sum squares)] = (100.42 

2889) + [(131.12) + (-30850.58) + 17216– 262425.62] = 14184.3

16. Unadjusted SS due to entries


=

= +

+ (120)2 + (100)2 + …. + (103)2 – 260822.53 = 87042.5 + 189557 -


260822.53 = 15776.97

Partition SS due to entries (C + P) to components:

SS due to checks =

= - = 87042.5 – 86142.25 = 900.25

SS due to test cultures/progenies


=
= 189557 – 183826.56 = 5730.44

SS due to checks  test cultures/progenies


= Unadju sted SS due to entries (C + P) - SS due to checks - SS due to test culture/progenies
= 15776.97 – 900.25 – 5730.44 = 9146.28
SS due to error: Total SS – True blocks SS - SS due to entries (adjusted)
= 15798.47 – 1603.09 – 14184.3 = 11.08

55
Summarize the results of analysis in ANOVA table
________________________________________________________________________
Source D.F. SS MS F-cal. F-Table
5% 1%
__________________________________________________________________
Block (b-1) 3 1603.09 534.36
Adjusted entries (C+P)-1= 19 14184.3 746.54
Unadjusted entries(C+P)-1= 19 15776.97 830.37
. Checks (C-1) 3 900.25 300.08 243.97** 3.86 6.99
. Test culture (P-1) 15 5730.44 382.03 310.59** 3.01 4.96
. Test culture vs. check 1 9146.28 9146.28 7436.00** 5.12 10.56
Error (b-1) (C-1) 9 11.08 1.23
Total (N – 1) 31 15798.47

Mean Comparison

i. To compare any two check means at 5% level of significance

LSD 5% = t 0.025 (9) x = 2.262  = 2.262 


= 1.77 days
ii. To compare two progenies/test materials occurring in the same block at 5% level of significance
t0.025(9) x = 2.262  = LSD5% = 2.262 
= 3.55 days
iii. To compare two test cultures (progenies) occurring in different blocks at 5% level of significance
LSD 5% = t 0.025 (9)  ; c = No. of checks

= 2.262 ) = 3.97 days


iv. To compare a progeny/test culture with any check at 5% level of significance
LSD 5% = t 0.025 (9)  ;
= 2.262  ) = 3.14 days

56
8. FACTORIAL EXPERIMENTS

8.1 Simple Effects, Main Effects and Interaction

Factorial experiments are experiments in which two or more factors are studied together. Factor is a kind of
treatment and in a factorial experiment any factor will supply several treatments. In factorial experiment, the
treatments consist of combinations of two or more factors each at two or more levels.

Factorial experiment can be done in CRD, RCBD and Latin Square Design as long as the treatments allow.
Thus, the term factorial describes specific way in which the treatments are formed and it does not refer to
the experimental design used, e.g. nitrogen & phosphorus rates:
N = 0, 50, 100, 150 kg/ha
P = 0, 50, 100, 150 kg/ha

Kinds (noug cake, groundnut cake) and levels of protein supplement (25%, 50%, and 75%).

The term level refers to the several treatments within any factor, e.g. if 5-varieties of sorghum are tested
using 3-different row spacing, the experiment is called 5 x 3 factorial experiment with 5 levels of variety
factor (A) and three levels of spacing factor (B). An experiment involving 3 factors (variety, N-rate,
weeding method) each at 2 levels is referred as 2  2  2 or 23 factors; 3 refers to the number of factors and
2 refers to levels. Here we have 8 treatment combinations variety (x, y): N-rate (0, 50 kg/ha), weeding (with
or without weeding). The 23  3 is a four factor experiment in which three factors each at 2-levels and the
4th factor at 3 levels.

If the above 23 factorial experiment is done in RCBD, the correct description of the experiment will be 2 3
factorial experiment in RCBD.

Interaction

Sometimes the factors act independent of each other. By this we mean that changing the level of one factor
produces the same effect at all levels of another factor. Often, however, the effects of two or more factors
are not independent. Interaction occurs when the effect of one factor changes as the level of the other factor
changes, e.g. if the effect of 50kg N on variety X is 10 Q/ha and its effect on a variety Y is 15 Q/ha, then
there is interaction. When factors interact, the factors are not independent and a single factor experiment
will lead to disconnected or misleading information. However, if there is no interaction it is concluded that
the factors under consideration act independently of each other. Thus, results from separate single factor
experiments are equivalent to those from a factorial experiment.

Example: A tall maize variety might out yield a short variety in high fertilizer rates due to high dry matter
production.

Interaction is the failure of the differences in response to changes in levels of one factor to be the same at all
levels of another factor or when the effect of one factor changes as the level of the other factor changes.

57
2 x 2 Factorial data of wheat yield (t/ha)
_________________________________________________________________
N-rate (kg/ha) (Factor B)
_____________________________ Simple effect of
Variety (Factor A) 0 (b0) 50 (b1) nitrogen on variety
__________________________________________________________________
X (a0) 1.0 1.0 (a0b1-a0b0) = 0
Y (a1) 2.0 4.0 (a1b1-a1b0) = 2
Simple effect of
Variety (a1b0-a0b0) = 1 (a1b1-a0b1) = 3

Simple effects
- Simple effect of variety at N0: 2-1 = 1
- Simple effect of variety at N1: 4-1 = 3
- Simple effect of N on variety X: 1-1 = 0
- Simple effect of N on variety Y: 4-2 = 2

Main effects are the averages of the simple effects


- Main effect of variety = ½ (Simple effect of A at bo + Simple effect of A at b1)
= ½ [(a1b0-a0b0) + (a1b1-a0b1)] = ½ [(2-1) + (4-1)] = 2
- Main effect of nitrogen = ½ (Simple effect of factor B at ao + Simple effect of factor A at a1) = ½ [(a0b1-
a0b0) + (a1b1-a1b0)] = ½ [(1-1) + (4-2)] = 1

Interaction
It is calculated as the average of difference between simple effects of A at the two levels of B or the
difference between the simple effects of B at the two levels of A.
= ½ (Simple effect of A at b1 – simple effect of A at b0)
= ½ [(a1b1-a0b1) - (a1b0-a0b0)] = ½ [(4-1) - (2-1)] = 1
or
= ½ (Simple effect of B at a1 – simple effect of B at a0)
= ½ (a1b1-a1b0) - (a0b1-a0b0)] = ½ [(4-2) - (1-1)] = 1

In factorial experiments, the following points should be noted:


a. Factorial experiments are useful to study relationships among several factors and to determine the
presence and magnitude of interaction
b. An interaction effect between two factors can be measured only if the two factors are tested together
in the same experiment.
c. When interaction is absent, the simple effect of a factor is the same for all levels of the other factors
and equals to the main effect.
d. When interaction is present, the simple effect of a factor changes as the level of the other factor
changes.

8.2 Two Factor Factorial in Randomized Complete Block Design


Example: An agronomist wanted to study the effect of different rates of phosphorus fertilizer on two
varieties of common bean (Phaseolus vulgaris). He thought that the varieties might respond differently to
58
fertilizer and he decided to use a factorial experiment with 2- factors: variety at two levels (T 1 =
Indeterminate T2 = Determinate) and phosphorus rate at 3 levels (P1 = none, P2 = 25 kg/ha, P3 = 50 kg/ha).
Using the full factorial set of combinations, he had six treatments:
T1P1; T1P2; T1P3; T2P1; T2P2; T2P3,

He conducted this experiment using Randomized Complete Block Design with four blocks of six plots each.

Field layout and yield of common bean (Q/ha)

Block-I
T2P2 T2P1 T1P1 T2P3 T1P3 T1P2
8.3 11.0 11.5 15.7 18.2 17.1

Block-II
T2P1 T2P2 T2P3 T1P2 T1P1 T1P3
11.2 10.5 16.7 17.6 13.6 17.6

Block-III
T1P2 T1P1 T2P1 T1P3 T2P3 T2P2
17.6 14.3 12.1 18.2 16.6 9.1

Block-IV
T1P3 T2P2 T2P3 T2P1 T1P2 T1P1
18.9 12.8 17.5 12.6 18.1 14.5

The linear model for Two Factor Randomized Block Design:


Yijk =  + i + j + k + ik + ijk
where, Yijk = the value of the response variable;  = Common mean effect; i = Effect of factor A; j =
Effect of block; k = Effect of factor B; ik = Interaction effect of factor A & factor B; and ijk =
Experiment error (residual) effect

Steps of Analysis of Variance

1. Construct two way table for factors and calculate factor A total, Factor B total and grand total
________________________________________________________
Phosphorus (Factor B)
_________________________________________
Variety (Factor A) P1 P2 P3 Factor A total (A)
_______________________________________________________
T1 (indeterminate) 53.9 70.4 72.9 197.2
T2 (determinate) 46.9 40.7 66.5 154.1
Factor B total (B) 100.8 111.1 139.4 351.3(G)
________________________________________________________
Block total
Block I II III IV
Total 81.8 87.2 87.9 94.4

59
2. Using r as number of blocks, a as level of factor A, b level of factor B, compute C.F., total SS, block
SS, treatment SS and Error SS

- C.F. = = = 5142.15; where r is number of replications, a is level of factor A and b is


level of factor B

- Total SS =(8.3)2+(11)2+... + (14.5)2 – 5142.15 = 243.38


- Block SS = - 5142.15 = 13.32

- Treatment SS = - 5142.15= 221.38


- Error SS = Total SS– Block SS – Treatment SS= 243.38 -13.32 - 221.38= 8.68

3. Compute the three factorial components of treatment SS [partition treatments SS in to factor A SS,
factor B SS, and A x B (interaction) SS]

- Factor A (variety) SS = - C.F = - 5142.15


= 77.40
- Factor B (P-rate ) SS = - C.F = – 5142.15 = 99.87
- A  B SS = Treatment SS – Factor A SS – Factor B SS =221.38 –77.40 – 99.87 = 44.11

ANOVA TABLE
__________________________________________________________
Source DF SS MS F-calcul. F-table
5% 1%
__________________________________________________________
Block r-1(4-1) = 3 13.32 4.44 7.65** 3.29 5.42
Variety (V) a-1 (2-1) = 1 77.40 77.40 133.45** 4.54 8.68
Phosphorus (P) b-1 (3-1) = 2 99.87 49.93 86.09** 3.68 6.36
VP (a-1) (b-1) = 2 44.11 22.05 38.03** 3.68 6.36
Error (r-1) (ab-1) = 15 8.68 0.58
Total rab -1= 23 243.38
_________________________________________________________

CV =  100; CV =  100 = 5.2%

Interpretation of a factorial experiment

The interpretation of the results of factorial experiment depends on the outcome of the significance tests. If
factor A  factor B interaction is significant, the main effects have no real meaning whether significant or
not. In our case, since A  B interaction is highly significant, the results of experiment are best summarized

60
in a two way table means of various A  B combinations. If interaction is not significant, then all of the
information in the trial is contained in the significant main effects. In this case the results may be
summarized in tables of mean for factors with significant main effects.

Mean Comparisons

There are three types of means in a two factor factorial experiment.


- Factor A means
- Factor B means
- Factor combinations (AB) or treatment means

Variety  Phosphorus rate means


_______________________________________________
Phosphorus (B)
____________________________________
Variety (A) P1 P2 P3 Variety mean
_______________________________________________
T1 (indeterminate) 13.47 17.60 18.22 16.43
T2 (determinate) 11.72 10.17 16.62 12.84

Phosphorus mean 12.59 13.88 17.42=G


_________________________________________________

Standard error of mean differences ( )

- to compare any two factor A means: A= =


= 0.31 Q
- to compare any two factor B means: B= =
= 0.38 Q
- to compare any two factor combination (treatment) means: AB = = =
0.54 Q

Remark

Three or more factor experimental designs: Read Gomez & Gomez, Chapter 4, starting from page
130

8.3. Split-Plot Design

8.3.1 Uses, advantages and disadvantages

Split-plot design is frequently used for factorial experiments where the nature of experimental material
makes it difficult to handle all factor combination. The principle underlying is that the levels of one factor
61
are assigned at random to large experimental units. The large units are then divided into smaller units
and then the levels of the second factor are assigned at random to small units within large units.

The large units are called the whole units or main-plots whereas the small units are called the split-plots or
sub-plots (units). Thus, each main plot becomes a block for the sub-plot treatments. In split-plot design, the
main plot factor effects are estimated from larger units, while the sub-plot factor effects and the interactions
of the main-plot and sub-plot factors are estimated from small units.

As there are two sizes of experimental units, there are two types of experimental error, one for the main
plot factor and the other for the sub-plot factor. Generally, the error associated with the sub-plots is smaller
than that for the whole plots due to the fact that error degrees of freedom for the main plot are usually less
than those for the sub-plots.

In split-plot design, the precision for the measurement of the effect of main plot factor is sacrificed to
improve the precision of the measurement of the sub-plot factors.

Situations when to use split-plot design?


a. When the level of one or more of the factors require larger amounts of experimental units than
another. For instance, in field experiments, one of the factors could be method of land preparation
(tractor, oxen, hand) and method of fertilizer application (broad cast, drill). These factors usually
require larger experimental plots (units). The other factor could be varieties which can be compared
using smaller units (plots). In this case methods of land preparation and fertilizer application can be
assigned to main-plots and the varieties to the sub-plots.

b. When an additional factor is to be incorporated in an experiment to increase its scope. For example,
if the major purpose of an experiment is to compare the effect of several vaccines as a protectant
against infection from certain disease of animals, to increase the scope of the experiment, several
breeds of animals can be included which are known to differ in their resistance to disease. Here, the
breeds of animals could be arranged in main units and the vaccines to the subunits.

c. When greater precision is desired for comparison of certain factors than others.

Since in a split-plot design, plot size and precision of measurement of the effects are not the same for both
factors, the assignment of a particular factor to either the main-plot or to the sub-plot is extremely important.

Guidelines to apply factors either to main-plots or sub-plots:


a. Degree of precision required: Factors which require greater degree of precision should be assigned to
the sub-plot. For example, animal breeder testing three breeds of dairy cows under different types of
feed stuff, will assign the breeds of animals to sub-units and the feed stuffs to the main unit. On the
other hand, animal nutritionist may assign the feeds to the sub-units and the breeds of animals to the
main-units as he is more interested on feed stuffs than breeds.

b. Relative size of the main effect: If the main effect of one factor (factor A) is expected to be much
larger and easier to detect than factor B, then factor A can be assigned to the main unit and factor B
to the sub-unit. For instance, in fertilizer and variety experiments, the researcher may assign variety

62
to the sub-unit and fertilizer rate to the main-unit, because he expects fertilizer effect to be much
large and easier to detect than the varietal effect.

c. Management practice: The factors, which require smaller amounts of experimental material, should
be assigned to sub-plots. For example, in an experiment to evaluate the frequency of irrigation (5,
10, 15 days), on performance of different tree seedlings on nursery, the irrigation frequency factor
could be assigned to the main plot and the different tree species to the sub-plots to minimize water
movement to adjacent plots.

Advantages
a. It permits the efficient use of some factors, which require large experimental units in combination
with other factors, which require small experimental units.
b. It provides increased precision in comparison of some of the factors (sub-plot factors).
c. It promotes the introduction of new treatments into an experiment, which is already in progress.

Disadvantages:
a. Statistical analysis is complicated because different factors have different error mean squares.
b. Low precision for the main plot factor can result in large differences being non-significant, while
small differences on the sub-plot factor may be statically significant even though they are of no
practical significance.

8.3.2 Randomization and layout

There are two separate randomization process in split-plot design, one for the main plot factor and another
for the sub-plot factor.

In each block, the main plot factors are first randomly applied to the main plots followed by random
assignment of the sub-plot factors. Each of the randomization is done by any of the randomization schemes.

Example: An experiment was designed to test the effect of feeding four forage crops (Rhodes grass, Vetch,
Alfalfa and Oat) on weight gain (kg/month) of the two breeds of cows (Zebu, Holstein). At the start of the
experiment, it was assumed that breeds of cows would respond differently to the feed stuffs. Therefore, it
was decided to use factorial experiment. The objective of the experiment was to compare the effect of
forage crops as precisely as possible. Therefore, the experimenter assigned the breeds of animals to the
main-plot and the four forage crops to the sub-plots. The experiment was replicated in three blocks (barns)
based on initial body weight of animals as a blocking factor.

Procedures of randomization
Step 1: Divide the experimental area into r = 3 blocks, and divide each block into two main plots. Then
randomly assign the two breeds of animals (H, Z) in each of the blocks.

Note that the arrangement of the main-plot factor can follow any of the designs: CRD, RCBD and LATIN
square.

Step 2: Divide each of the main plot (unit) into 4-sub plots (units) and randomly assign the four feed stuffs
(A, V, O, R) to each of the six-main plots (units).
Note:
63
Each main-plot factor is tested r-times where r is the number of blocks while each sub-plot factor is tested a
 r times where a is level of factor A and r is the number of blocks. This is the primary reason for more
precision for the sub-plot factors as compared to the main-plot factors.

The layout and the weight gain (kg/month) of the animals for feeding are given below:
Block I Block II Block III
H Z H Z Z H
A R O V O V
25.9 15.5 18.0 22.7 13.2 28.4

V A A O A A
25.3 18.9 26.7 13.5 19.6 27.6

O O V R V R
19.3 13.8 24.8 15.0 22.3 25.4

R V R A R O
22.2 21.0 24.2 18.3 15.2 20.5

- Main-plot factor is breed of animals: Holstein (H), Zebu (Z).


- Split-plot (units) factor is feed stuffs: Alfalfa (A), Vetch (V), Rhodes grass (R), Oat (O)

8.3.3 Analysis of variance

The linear model for Split Plot Design:


Yijk =  + i + j + k + ()ij + ()ik + ijk
where, Yijk = the value of the response variable;  = Common mean effect; i = Effect of factor A (main plot
factor); j = Effect of block; k = Effect of factor B (sub-plot factor); ()ij = Interaction effect of factor A
& Block (error a); ik = Interaction effect of factor A & B and ijk = Experiment error (residual) effect
(error b)

Steps of Analysis
Step 1: Arrange data by treatments (main-plot, sub-plot) and blocks and calculate main-plot total, and sub-
plot total.

Treatments Blocks
Breeds Feeds I II III
Holstein Alfalfa 25.9 26.7 27.6
Vetch 25.3 24.8 28.4
Oat 19.3 18.0 20.5
Rhodes grass 22.2 24.2 25.4
Main-plot totals 92.7 93.7 101.9
Zebu Alfalfa 18.9 18.3 19.6
Vetch 21.0 22.7 22.3
Oat 13.8 13.5 13.2

64
Rhodes grass 15.5 15.0 15.2
Main-plot totals 69.2 69.5 70.3

Step 2: Construct two way tables of totals.


2.1. Block by factor A two-way table and compute block total, factor A total and grand total.

Weight gain totals of block by factor A.

Breeds (A) Block 1 Block 2 Block 3 Factor A


(breeds) total
Holstein 92.7 93.7 101.9 288.3
Zebu 69.2 69.5 70.3 209.0
Block total 161.9 163.2 172.2
Grand total 497.3

2.2. Factor A by factor B total two-way table and calculate factor B totals
Feeds (B)
Breeds (A) Alfalfa (b1) Vetch (b2) Oat (b3) R. Grass (b4)

Holstein (a1) 80.2 78.5 57.8 71.8


Zebu (a2) 56.8 66.0 40.5 45.7
Factor B total 137.0 144.5 98.3 117.5

Step 3: Compute the correction factor and sum of squares for the main-plot analysis.
C.F.=
Total SS = = 10820.59 – 10304.47
= 516.12
Block SS =
= 10312.34 –10304.47 = 7.87
Factor A (breeds) (main plot factor) SS=
10566.49 – 10304.47 = 262.02
Error (a) SS = Block SS-factor A SS

Error (a) SS= = 5.03

Step 4: Compute the sum of squares for the sub-plot analysis.


Factor B (feed stuff) SS =

65
= = 215.26

Factor A  Factor B SS = Factor A SS - Factor B SS

= 10800.45 – 10304.47 – 262.02 – 215.26 = 18.7

Error (b) SS = Total SS – total of all other sum of square


= Total SS – Block SS - Factor A SS- Error (a) SS - Factor B SS –
AB SS = 516.12 – 7.87 – 262.02 – 5.03 – 215.26 – 18.7 = 7.24

Step 5: For each source- of variation compute the mean squares by dividing the SS by its corresponding
degrees of freedom.
Block MS =

Factor A MS =

Error (a) MS =

Factor B MS =

A  B MS =

Error (b) MS =

Step 6: Compute the F-value for each effect that needs to be tested.
F(block) =

F(A) =

F(B) =

F(AB) =
Step 7: Construct the ANOVA Table, obtain the corresponding tabulated F-value and compare it with the
calculated F-value at prescribed level of significance.

Source of D.F. SS MS Calculated F Tabulated F


variation 5% 1%
Block (r-1) = 2 7.87 3.93 1.56 19.0 99.0
Breeds (A) (a-1) = 1 262.02 262.02 104.39** 18.51 98.5
Error (a) (r-1) (a-1) = 2 5.03 2.51
Feeds (B) (b-1) = 3 215.26 71.75 119.58** 3.49 5.95
66
AB (a-1) (b-1) = 3 18.70 6.23 10.38** 3.49 5.95
Error (b) a(r-1)(b-1) = 12 7.24 0.60
Total (rab-1) = 23 516.12

Step 8. Compute the two coefficients of variation, one corresponding to the main-plot analysis and another
to the sub-plot analysis.
CV (a) =

CV (b) =
Note that CV (a) is greater than CV (b), this is because factors assigned to the main-plot are expected to be
measured with less precision than that assigned to the sub-plot.

Mean comparisons:
Mean weight (kg/month) of Holstein and Zebu breeds fed with four types of forage species

Breeds (A) Alfalfa (b1) Vetch (b2) Oat (b3) R. Grass (b4) Factor A mean
(A total/rb)
Holstein (a1) 26.7 26.2 19.3 23.9 24.0 (A1)
Zebu (a2) 18.9 22.0 13.5 15.2 17.4 (A2)
Factor B mean
(B total/ra) 22.8 (B1) 24.1 (B2) 16.4 (B3) 19.6 (B4)

Standard Errors of the Mean Differences

a. to compare two main plot factor (A) means (A1 & A2): A=

A= = 0.65 kg

LSD1%= t 0.005 [error (a) d. f.]  =LSD1% =t 0.005 (2)0.65 kg


= 9.9250.65 kg = 6.45 kg

Since the mean difference (d) between A 1 and A2 (24.0-17.4 = 6.6) > LSD value at 1% (6.45), the difference
between the two means is highly significant.
b. to compare two sub-plot factor (B) means: B= =
= 0.45 kg
Example: To compare means of B3 and B2, LSD1% = t 0.005 [error (b) d. f.]  =
LSD1% = t 0.005 (12)  0.45 kg = 3.055  0.45 = 1.37 kg.

Since the mean difference (d) between B 2 and B3 (24.1 – 16.4 = 7.7) is greater than LSD value at 1% (1.37),
the difference between the two means is highly significant.

67
c. to compare two sub-plot treatment means at the same level of main plot:

= = = 0.63 kg

To compare (a1b1= 26.7) with (a1b3 = 19.3); LSD1% = t0.005 [error (b) d. f.] 
= LSD1% = t 0.005 (12)  0.63 kg = 3.055  0.63 = 1.92 kg.

 Since the mean difference (d) between (a1b1 & a1b3 = 26.7-19.3 = 7.4) is greater than LSD value at
1% (1.92 kg), the difference between the two means is highly significant.

d. to compare two sub-plot treatment means at different main-plot factor level:

= = 0.85 kg

that contains both MSE(a) and MSE(b) has no exact value for the d.f. associated with it. To obtain an
approximation:

d.f. =

To compare (a1b1= 26.7) with (a2b2 = 22.0); LSD1% = t0.005 [error d.f.] 
= LSD1% = t 0.005(5)  0.85 kg = 4.032  0.85 kg = 3.43 kg.

 Since the mean difference (d) between (a1b1 & a1b3 = 26.7-22.0 = 4.7) is greater than LSD value at 1%
(3.43 kg), the difference between the two means is highly significant.

Presentation
If interaction of the factors is significant, results are summarized in two way table of means. However, if
interaction is non-significant, the results are summarized in one way table of means for the significant
factor.

9. COMPARISON OF TREATMENT MEANS

The F-test (ANOVA) shows whether there is significant difference among treatments or not. But, it does not
show us which means are different from each other. There are many ways to compare the means of
treatments tested in an experiment. One of these is pair comparison, the simplest and most commonly used
comparisons in agricultural research.

There are two types of pair comparisons:


68
A. Planned pair comparison: In which the specific pair of treatments to be compared are identified before
the start of the experiment, e.g. comparing the control treatment with each of the other treatments- A priori
test

B. Unplanned pair comparison: In which no specific comparison is chosen in advance. Instead, every
possible pair of treatment means are compared to identify pairs of treatments that are significantly different,
e.g. variety trials. A posteriori test or Post hoc test

The most commonly used test procedures for pair comparison in agricultural research are the Least
Significant Difference and Tukey’s test which are suitable for planned pair comparison and Duncan’s
Multiple Range Test (DMRT) which is applicable to an unplanned pair comparison.

9.1 Least Significant Difference (LSD) Test

LSD is the simplest and the most commonly used procedure for making pair comparisons. The procedure
provides a single value at a prescribed level of significance, which serves as the boundary between
significant and non-significant differences between any pair of treatment means. That is, two treatments are
declared significantly different at a prescribed level of significance if their mean difference exceed the
computed LSD value, otherwise they are not significantly different.

The LSD test is not valid for comparing all possible pair of means, especially when the number of
treatments is large. This is so because the number of possible pairs of treatment means increase rapidly as
the number of treatments increase. In experiments where no real difference exists among all treatments, the
numerical difference between the largest and smallest treatment means is expected to exceed the LSD value
when the number of treatments is large.

To avoid this problem, the LSD test is used only when the F-test for treatment effect is significant and the
number of treatments is not too large (less than six).

The procedure for applying the LSD test to compare any two treatments means

1. Rank the treatment means from the largest to the smallest in the column and from the smallest to largest
in rows.
2. Compute all possible differences between the two treatment means to be compared.
3. Compute the LSD value at α level of significance
LSDα = tα/2 (n)  s
where s = standard error of the treatment mean difference; t α/2 (n) is the table t-value at α/2 level of
significance and with n error degree of freedom

Example: Oil content (g) of linseed treated at different six stages of growth with N-fertilizes tested in
RCBD in four replications with error mean square of 1.31.
LSD5% = t0.025(15)  , where MSE is error mean square; r is the number of replications = 2.131 

1.72 g
69
LSD1% = t 0.005(15)  = 2.947 

4. Compare the mean difference (d) in step 2 with LSD value computed in step (3) using the following rule:
- if /d/ >LSD value at 1% level of significance, there is highly significant difference between the two
treatment means compared (put two asterisks on differences).
- if /d/>LSD value at 5% level of significance but < LSD value at 1% level of significance, there is
significant difference between the two treatment means compared (put one asterisks on differences)
- if /d/  LSD value at 5% level of significance, the two treatment means compared are not significantly
different (put n.s.)

No Treatments (stage of application Treatment mean (g)


of N)
1 Seedling 5.10
2 Early blooming 4.30
3 Half blooming 4.00
4 Full Blooming 6.70
5 Ripening 6.05
6 Unfertilized (control) 7.03

Treatments 4.00 4.30 5.10 6.05 6.70 7.03


(T3) (T2) (T1) (T5) (T4) (T6)
7.03 (T6) 3.03** 2.73** 1.93* 0.98ns 0.33ns -
6.70 (T4) 2.70** 2.40** 1.60ns 0.65ns -
6.05 (T5) 2.05* 1.75* 0.95ns -
5.10 (T1) 1.10ns 0.80ns -
4.30 (T2) 0.30ns -
4.00 (T3) -

Thus, the differences between T6 & T3, T6 & T2, T4 & T3, T4 & T2 are highly significant; while the
differences between T3 & T5, T1 & T6, T2 & T5 are significant.

Note that there are possible (unplanned) pair comparisons and (t-1) planned pair comparisons where
t is the number of treatments. In the above example, 15 unplanned pair comparisons and five planned pair
comparisons are possible, since we have one control.

Presentation of data using LSD

Table __. Mean oil content of linseed treated with nitrogen fertilizer at different stages
70
___________________________________________
Stage of application Oil content (g)
___________________________________________
Seedling 5.10
Early blooming 4.30
Half-blooming 4.00
Full- blooming 6.70
Ripening 6.05
Unfertilized 7.03
___________________________________________
LSD(0.05) 1.72 g
CV (%) 20.7

9.2 Duncan’s Multiple Range Test (DMRT)

It is most widely used to make all possible pair comparisons. The procedure for applying the DMRT is
similar to LSD test but it requires progressively larger values for significance between the treatment means
as they are more widely separated in the array.

The test is more appropriate when the total number of treatments is large. It involves the calculation of the
shortest significant difference (SSD).

The SSD is calculated for all possible relative positions (P) between the treatment means when the means
are arranged in order of magnitude (in decreasing or increasing order).

Procedure
Step 1: Arrange all the treatment means in increasing or decreasing order.
 Data such as crop yield are usually arranged from the highest to the lowest.

Example: Yields (kg/plot) of wheat varieties grown in 4 x 4 Latin Square Design with error mean square
of 0.45:

B (12.3), A (12.00), C (10.8), D (6.7)

Step 2: Calculate (the standard error of the treatment mean difference) as:
=

Step 3: Calculate the shortest significant difference (SSD) for relative positions (P) in the array of means.
 Since we have four treatment means, they can be 2, 3 and 4 distance apart.
 B and A are 2 distance apart (P = 2); B and C are 3 distance apart (P = 3); B and D are 4 distance
apart (P = 4); A and D are 3 distance apart (P = 3); etc.

For the above example, the R values with error d. f. of 6 at 1% level of significance are found from R-table
(see Appendix F in Gomez & Gomez)
P= 2 3 4
71
R0.01 = 5.24 5.51 5.65
SSD = 1.74 1.83 1.88
P = the distance in ranks between the pairs of treatment means to be compared.
R= significant studentized range at error d. f. (6).

SSD (Shortest Significant Difference) =

SSD at P = 2 =

SSD at P = 3 =

SSD at P = 4 =
 Note that SSD values increase as the distance between treatments (P) to be compared increases.

Step 4:Test the difference between treatment means in the following order.
 Largest – Smallest = 12.3 – 6.7 = 5.6; compare with SSD value at (P = 4) = 1.88; d (5.6) > SSD at P
= 4 (1.88); thus the difference is significant at 1% level of significance.
 Largest – 2nd smallest = 12.3 – 10.8 = 1.5; compare with SSD at (P=3) = 1.84; d (1.5) < SSD at P= 3
(1.84); thus, the difference is non-significant at 1% level of significance.
 Largest – 2nd largest = 12.3 – 12.0 = 0.3 < SSD at P = 2 (1.75); thus, the difference is non-significant
at 1% level of significance.
 2nd largest – smallest = 12.0 - 6.7 = 5.3; compare with SSD at P = 3 (1.84); d (5.3) > SSD (1.84) at P
= 3; significant at 1% level of significance
 2nd smallest – smallest = 10.8 – 6.7 = 4.1 compared with SSD at P = 2 (1.75); d (4.1) > SSD (1.75) at
P = 2; significant
 etc

 Note that SSD value at (P=2) is equals to LSD value

________________________________________________________
B (12.3) A (12.0) C (10.8) D (6.7)
_____________________________________________
D (6.7) 5.6** (P=4) 5.3** (P=3) P = 4.1** (P=2)-
ns ns
C (10.8) 1.5 (P=3) 1.2 (P=2) -
A (12.0) 0.3ns (P=2) -
B (12.3) -
________________________________________________________

SSD (P = 4) = 1.88; SSD (P =3) = 1.83; SSD (P =2) = 1.74

Treatments B & D, A & D, C & D are significantly different at 1%, while treatments B & C, B & A, and A
& C are not significantly different at 1% level of significance.

Step 5: Present the test result in one of the following two ways

72
A. Use a line notation if the sequence of results can be arranged according to their ranks.

 Any two means underscored by the same line are not significantly different at 1% level of significance
according to DMRT.
B(12.3) A(12.0) C(10.8) D(6.7)

B. Use the alphabet notation if the desired sequence of the results is not based on their rank which is
commonly used.

 The alphabet notation can be derived from line notation simply by assigning the same alphabet to all
treatment means connected by the same horizontal line.
 It is usual practice to assign letter a for the first line, b for the second line, c for third line and so on.
 Note that letter a can be used for the largest or smallest treatment mean depending on the rank of
arrangement.

Presentation of data using DMRT


Table __ Mean yields of wheat varieties planted at Debrezeit Agricultural Research Center.

Variety Yield (kg)


A 12.0 a
B 12.3 a
C 10.8 a
D 6.7 b

 Note that we have to put a footnote below the table stating that any two means in the same column
followed by the same letter are not significantly different at 1% level of significance according to
DMRT. Note also that both LSD and DMRT are not used in the same table. Use either of them
depending on the appropriateness of the test.

9.3 Tukey’s Test

 It is more conservative than LSD test because it requires the largest treatment mean differences for
significance.

 It is computed in a manner similar to the LSD test except that standard error of the mean is used
instead of standard error of the mean difference ( ), and

 Studentized range (q-table) is used in place of t-table

The procedure:
1. Select a value from q table, which depends on the number of means (n) and error degree of freedom
(v).
2. Compute the Critical Difference (CD) as = q(n, v)  where MSE is error mean square; n is
number of means to be compared; v is error degrees of freedom and r is number of
replications.

73
3. For any pair of means, if the absolute value of the difference /d/ > critical value, the difference is
judged to be significant at a prescribed level of significance.

Example: The following analysis of variance table is from CRD with six varieties replicated four times in
glass house (mean rust incidence)
Source d. f. MS F-cal. F-table (5%)
____________
Variety (t-1) = 5 2976.44 24.80** 2.77
Error t(r-1) = 18 120.00
__________________________________________________________
Variety: 1 2 3 4 5 6
Mean stem rust incidence (%): 50.3 69.0 24.0 94.0 75.0 95.3

n= 6; V= 18; q0.05 (6, 18) = 4.495


CD = q  = 4.495  = 24.62%

Difference between means


________________________________________________________________________
24.0(3) 50.3(1) 69.0(2) 75(5) 94(4) 95.3(6)
_________________________________________________________
95.3(6) 71.3* 45.0* 26.3* 20.3ns 1.3ns -
94(4) 70.0* 43.7* 25.0* 19.0ns -
ns
75(5) 51.0* 24.7* 6.0 -
69(2) 45.0* 18.7ns -
50.3(1) 26.3* -
24.0(3) -
________________________________________________________________________

Thus, differences between varieties 6&3, 4&3, 5&3, 2&3, etc. are significant while differences between
varieties 2&1, 5&2, etc. are non-significant.

9.4. Pair Comparisons with Missing Data

In applying the LSD test and DMRT, it is important that the appropriate standard error of the mean
difference (s ) should be used. s is affected by the experimental design used, the number of replications of
the two treatments being compared, and the specific type of means to be compared.

A). In CRD, RCBD and Latin Square Design where the number of replications for all treatments is equal,
the s for any pair of treatment means is computed as:
74
Where, MSE is mean square for error; r = number of replications that is common to all treatments.
Thus, lsd = where n is error degree freedom.

B). When the two treatments do not have the same number of replications in CRD. is computed as:

where MSE is mean square error; r i and rj are the number of replications of the two treatment means (i & j)
to be compared.

Thus, LSD=

Example: CRD with an unequal replications, effect of 4 – types of feedstuff on weight gain of chicks.
Treatment A = given to 5- chicks (5-replications) = 43.8 g
Treatment B = given to 4 chicks (4 replications) = 73.0 g
Treatment C = given to 3 chicks (3 replications) = 73.33 g
Treatment D = given to 5 chicks (5 replications) = 142.8 g

To compare treatment B with treatment D:

Given error mean square of 843.1 and error degree of freedom of 13, test if there is significant difference
between treatments B & D.

Difference between treatment mean: (D – B) = = 142.8-73.0 = 69.8g


LSD5% = t0.025 (13)  = 2.160  = 42.07g

LSD1% = t0.005 (13)  = 3.012  = 58.67g


Compare treatment mean difference with the calculated lsd value. Since 69.8 > LSD value at 1%
(58.67), there is a highly significant difference between treatments B and D or treatment D significantly
increased the weight of chicks as compared to treatment B.

C). for the treatments with a single missing value and that of any other treatment without missing values.

a) For RCBD: . Thus, LSD=t/2(error d.f.)  where MSE is mean

square of error; t = no. of treatments; r = no. of replications.

75
b) Latin Square Design: . Thus, LSD= t/2(error d.f.)  where MSE

is mean square error of the analysis of variance of Latin Square Design with a single missing value; r
= number of replications.

10. ANALYSIS OF COVARIANCE

The analysis of covariance simultaneously examines the variance and covariance of selected variables so
that the character of primary interest is more accurately characterized than by the use of analysis of variance
only. Analysis of covariance requires measurement of the character of interest and the measurement of one
or more variable(s) known as covariate(s). It also requires that the functional relationship of the co-variates
(x) with the character of primary interest (y) is known before hand.

Examples: Consider wheat variety trial in which weed infestation is used as a co-variate with a known
functional relationship between weed incidence and grain yield (the character of primary interest), the
covariance analysis can adjust grain yield in each plot to a common level of weed incidence. With the
covariance analysis, the variation in yield due to weed incidence is quantified and effectively separated from
that due to varieties.

Similarly, age or initial body weight of experimental animals can be used as a covariate and weight gain due
to rations as character of interest.

Covariance analysis can be applied to any number of covariates and to any type of functional relationships
between variables. In this section, however, we will deal with the case of a single covariate whose
relationship to character of primary interest is linear.

10.1 Uses of Covariance Analysis

1. To Control Experimental Error


One way to reduce experimental error is by using proper type of blocking. However, blocking can not cope
with certain types of variability such as spotty soil heterogeneity and unpredictable insect/disease incidence.
In such cases, heterogeneity between experimental plots does not follow a definite pattern. Thus, use of
76
covariance analysis should be considered in experiments in which blocking can not adequately reduce the
experimental error. By measuring an additional variable (i.e covariate) that is known to be linearly related to
the primary variable (y), the source of variation associated with the covariate can be deducted from
experimental error.

The experimental error is reduced and the precision for comparing treatment increased, e.g. in a cattle
feeding experiment to compare the effects of several rations on weight gain, animals assigned to any one
block may vary in initial weight. Now if the initial weight is correlated with gain in weight, a portion of
experimental error for gain can be the result of differences in initial weight. By covariance analysis, a
contribution, which can be attributed to differences in initial weight, can be computed and eliminated from
experimental error.

2. Adjustment of Treatment Mean


With the covariance analysis, the primary variable (y) can be adjusted linearly up wards or down wards
depending on the relative size of its respective covariate (x). The treatment means of dependent variable is
adjusted to a value that it would have had there been no differences in the value of co-variate, e.g. age of
cows, initial body weight of the cow, number of plants harvested per plot, etc. In situations where real
differences among treatments for the independent variable do occur but are not the direct effect of the
treatments, adjustment is warranted, e.g. in variety trial, if seeds may differ widely in germination, not
because of inherent differences of the varieties, adjustment can be done, but if the density is treatment by
itself there is no need to adjust the means.

3. Interpretation of Experimental Results


Covariance analysis aids the experimenter in understanding the principles underlying the results of an
investigation. By examining the primary character of interest (y) together with other characters (x) whose
functional relationship to y are known, the biological process governing the treatment effects on y can be
characterized more clearly.

4. Estimation of Missing Data


The missing data formula technique biases the treatment sum of squares upwards. The use of covariance
analysis to estimate the missing value(s) results in a minimum residual sum of squares and unbiased
treatment sum of squares.

10.2. Computation Procedure

Covariance analysis is an extension of analysis of variance. It is a combination of analysis of variance and


linear regression. Covariance analysis can be used for CRD; RCBD and split-plot designs, but the
computation procedures vary some what.

Computation procedure for RCBD

The following data show ascorbic acid content (y) of ten varieties of common bean. From the previous
experience, it was known that increase in maturity resulted in decrease in vitamin C content (linear r/ship).
Since all varieties were not of the same level of maturity on the same day, it was not possible to harvest all

77
plots at the same stage of maturity. Hence, the percentage of dry matter based on 100 g of freshly harvested
beans was observed as an index of maturity and used as a covariate.

Ascorbic acid content (ASAC, mg/100 g of seed) and percentage of dry matter
(% DM) for common bean varieties
___________________________________________________________
Block I Block 2 Block 3
____________ _____________ ______________
Variety %DM ASAC %DM ASAC %DM ASAC
________(X)___(Y)____(X)__(Y)______(X)___(Y)__
1 34 93 33 95 35 92
2 40 47 40 51 51 33
3 32 81 30 100 34 72
4 38 67 38 74 40 65
5 25 119 24 128 25 125
6 30 106 29 111 32 99
7 33 106 34 107 35 97
8 34 61 31 83 31 94
9 31 80 30 106 35 77
10 21 149 25 151 23 170
_______________________________________________________________
Conduct the analysis of covariance & calculate standard error of mean difference.

Variety Block I Block II Block III Variety total


% DM ASAC % DM ASAC % DM ASAC
X Y X Y X Y X Y
1 34 93 33 95 35 92 102 280
2 40 47 40 51 51 33 131 131
3 32 81 30 100 34 72 96 253
4 38 67 38 74 40 65 116 206
5 25 119 24 128 25 125 74 372
6 30 106 29 111 32 99 91 316
7 33 106 34 107 35 97 102 310
8 34 61 31 83 31 94 96 238
9 31 80 30 106 35 77 96 263
10 21 149 25 151 23 170 69 470
Block 318 909 314 1006 341 924 973 2839
total
78
Steps of Analysis

1. Conduct analysis of variance for each of the variables, covariance and sum
of square of treatment and error sum of squares
_______________________________________________________________
Source D. F. SS of ASAC (Y) SS of % DM (X) SS of XY
________________________________________________________________
Block 2 545.3 42.47 -75.23
Treatment 9 25689.0 972.70 -4633.23
Error 18 1608.7 86.20 -251.77
Treatment + Error 27 27297.7 1058.90 -4885.00
_________________________________________________________________
2. Analyse covariance
C.F. = = = 92078.23

Total Sum of Products = = (34  93) + (40  47) + ... + (23  170) - 92078.23 =
87118-92078.23 = -4960.23
Sum of Products due to Blocks = - C.F. -
92078.23 = 92003-92078.23 = -75.23
Sum of Products for Treatments: - C.F. = -
92078.23 = -4633.23
Error Sum of Squares of Products = Total Sum of Products – Block SS of Products – Treatment SS of
Products = -4960.23-(-75.23)-(-4633.23)
= -251.77
3. Compute the adjusted error SS of Y as = Error SS due to Y –

= 1608.7 – = 873.34
4. Compute (treatment + error) adjusted SS of Y as:

From the = (Treatment + Error SS of Y) -


above
ANCOVA = 27297.7- = 4761.84
table

5. Compute the treatment adjusted SS of Y = Treatment + Error adjusted SS of Y - Error adjusted SS of Y


= 4761.84-873.34 = 3888.5
79
_____________________________________________________________________
Source D.F SS MS F-cal F table (1%)
_____________________________________________________________________
Treatment (adjust) (t-1) 10-1 = 9 3888.50 432.0 8.41** 3.68
Error (adjusted) (r-1) (t-1)-1= 17 873.34 51.4
____________________________________________________________________

6. Compute regression coefficient or slope of data


= = = -2.92
7. Compute adjusted treatment mean as:
adjusted = – b( – X grand mean).
 For example, for treatment 1= adjusted = 93.33-(-2.921.57)
= 97.92
_________________________________________________________________
No Average ASAC Average Deviation of Adjusted
(Unadjusted) % DM ) % DM (xmean– xgrand mean) ASAC
__________________________________________________________________
1 93.33 34.00 1.57 97.92
2 43.67 43.67 11.24 76.48
3 84.33 32.00 -0.43 83.08
4 68.67 38.67 6.24 86.88
5 124.00 24.67 -7.76 101.33
6 105.33 30.33 -2.10 99.21
7 103.33 34.00 1.57 107.92
8 79.33 32.00 -0.43 78.08
9 87.67 32.00 -0.43 86.41
10 156.67 23.00 -9.43 129.13
______________________________________________________________________

Compute the relative efficiency (R.E.) of covariance analysis compared to standard analysis of variance

( Error unadjusted MS of Y ) (1608.7/18)


100  100
R.E. = ( Error adjusted MS of Y )(1  Treat MS of X ) 972.7/9 =
(51.4) (1  )
Error SS of X 86.2
77.15%

Thus, the result indicates that the use of % dry matter as the covariate has not increased precision in ascorbic
acid content which would have been obtained had the ANOVA is done without covariance.

Adjusted Error MS of Y
CV = x 100 = = 7.6%
Grand Mean of Y

80
Mean comparison

S to compare two adjusted treatment means:

, where xi & xj are the covariate means of ith and


jth treatment; r is the number of replications common to both treatments.
 For instance, to compare means of T1 & T2:
S = where 34 & 43.67 are the covariate means of 1 st and 2nd treatments; 3 is the
number of replications common to both treatments.
S = = 9.49

11. Combine Analysis of Data

In field experiments, it is necessary to repeat the experiments over a number of locations, seasons, or both,
e.g. varietal trials, plant spacing, fertilizer trials. The purpose of repeating the experiments is to find
recommendation that can be applied over space (location), time (season) or both.
In such repeated experiments, appropriate statistical procedures for a combined analysis of data have to be
used. The main purposes of combined analysis of data are:
 to estimate the average response to a given experiment
 to test the consistency of the response from place to place or year to year, i.e. to determine if there is
interaction effect of the treatments, e.g. stability analysis of varieties.
If the response is consistent from place to place and year to year, it shows the absence of interaction.

Steps of combined analysis

1. Construct an outline of combined analysis over years or locations or both on the basis of experimental
design used.
 For example, the outline of ANOVA for the experiment conducted at six environments with five
treatments and six replications in RCBD is given as:

___________________________________________________________________
Source Degrees of
freedom Mean squares Computed F
________________________________________________________________________
Environment (E) (e-1) = 5 EMS EMS/RMS
Blocks/within environment e(r-1) = 30 RMS
Treatments (T) (t-1) = 4 TMS TMS/MSE
TxE (e-1) (t-1) = 2 ExTMS E xTMS/MSE
Pooled error e(r-1) (t-1) = 120 MSE
________________________________________________________________________
81
Where e = no. of environments; r = no. of replications; t = no. of treatments; EMS = environment mean
square; RMS is blocks/replications mean square; TMS = treatment mean squares; and MSE = mean square
error.

2. Compute the usual ANOVA for each environment according to RCBD and obtain the mean squares of
error and degrees of freedom from individual environment analysis.
________________________________________________________________________
Environment: E1 E2 E3 E4 E5 E6
Error d.f. (r-1)(t-1) 20 20 20 20 20 20
Error MS 5776 4028 4516 9526 7056 5535

3. Test the homogeneity of the variances

Case 1: When there are only two environments, use F-test as:

F calculated = and compare with table F-value at d.f. for larger error mean
square as numerator and d.f. for smaller error mean square as denominator.
 If F-calculated is > F-table, the null hypothesis of homogeneity of variances is rejected, i.e. the error
variances are heterogeneous. Thus, combined analysis cannot be conducted directly. We can analyze the
data separately or transform the data to homogenize the error variances to use combined analysis.
 On the other hand, if the F-calculate value is ≤F-table, the error variances are homogeneous, thus we
can directly proceed to combined analysis.

Case 2: When environment is more than two, use Bartlett’s chi-square test

2c =

Where
Kj = Error d.f. for each environment

= Pooled error mean square = = = 6072.83


MSEi = Error mean square for each environment
e = number of environment

1/kj
Environment Kj MSE log MSE Kj x MSE Kj x log MSEi
E1 20 5776 3.761627 75.6 75.23254 0.05
E2 20 4028 3.605089 75.6 72.10179 0.05
E3 20 4516 3.654754 75.6 73.09508 0.05
E4 20 9526 3.978911 75.6 79.57821 0.05
82
E5 20 7056 3.848559 75.6 76.97117 0.05
E6 20 5535 3.743118 75.6 74.86235 0.05
453.6 451.8411 0.3
= Pooled error mean square = =
= 6072.83
Log (6072.83) = 3.78
2c = ; 2c = 3.97

2 (e-1) 5% = 2c (5) 5% = 11.07


Since 2c (3.97) is less than 2 table at 5% (11.07), accept the null hypothesis. Thus, this implies the errors
are homogeneous

12. REGRESSION AND CORRELATION ANALYSIS

12.1 Types of Regression & Correlation

Regression analysis describes the effect of one or more variables (designated as independent variables) on a
single variable (designated as the dependent variable). It expresses the dependent variable as a function of
independent variable(s).

For regression analysis, it is important to clearly distinguish between the dependent and independent
variables.

Examples:
- Weight gain in animals depends on feed
- Number of growth rings in a tree depends on age of the tree
- Grain yield of maize depends on a fertilizer rate

In the above cases, weight gain, number of growth rings and grain yield are dependent variables, while feed,
age and fertilizer rates are independent variables.

The independent variable is designated by x and the dependent variable by y.

Correlation analysis, on the other hand, provides a measure of the degree of association between the
variables, e.g. the association between height and weight of students; body weight of cows and milk
production; grain yield of maize and thousand kernel weight.

Regression and correlation analysis can be classified:


83
a. Based on the number of independent variables as:
- Simple: one independent variable and one dependent variable.
- Multiple: if more than one independent variables and a dependent variable is involved
b. Based on the form of functional relationship as:
- Linear: if the form of underlying relationship is linear
- Non-linear: if the form of the relationship is non-linear

Thus, regression and correlation analysis can be classified into 4:


 Simple linear regression and correlation analysis
 Multiple linear regression and correlation analysis
 Simple non-linear regression and correlation analysis
 Multiple non-linear regression and correlation analysis

Linear Relationships
The relationship between any two variables (independent and dependent) is linear if the change in y is
constant as x changes through out the range of x under consideration.

The functional form of linear relationship between a dependent variable y and an independent variable x is
represented by the equation.
y = a + bx

y = is the dependent variable


a = the intercept of the line on the y-axis (the value of y when x is 0)
b = linear regression coefficient, is the slope of the line or the amount of change in y
for each unit change in x.

When there are more than one independent variables as say k-independent variables (x 1, x2, ………, xk), the
simple linear regression equation y = a + βx can be extended to the multiple linear functional form of:
y = α + β1x1 + β2x2 +……. + βkxk

where α is the y intercept (the value of y when all x’s are 0); β1, β 2, ….. βk are partial regression coefficients
associated with the independent variables.

12.2 Simple Linear Regression and Correlation Analysis

12.2.1 Simple linear regression analysis

The simple linear regression analysis deals with the estimation and test of significance concerning the two
parameter α and β in the equation:
Y = α + βx
The data required for the application of the simple linear regression are the n-pairs (with n >2) of y and x
values.

Steps to estimate α and β

84
Step 1: Compute the means ( and ), deviation from means [ , ], square of the deviates (x2, y2)
and product of deviates (xy).

Example: Determine the regression equation for dependence of wing length of 13 sparrows of various ages.

Age (days), Wing Deviations from Square of Product of


x length the mean deviates deviates
(cm), Y
x2 y2 xy
3 1.4 -7 -2.015 49 4.060225 14.105
4 1.5 -6 -1.915 36 3.667225 11.49
5 2.2 -5 -1.215 25 1.476225 6.075
6 2.4 -4 -1.015 16 1.030225 4.06
8 3.1 -2 -0.315 4 0.099225 0.63
9 3.2 -1 -0.215 1 0.046225 0.215
10 3.2 0 -0.215 0 0.046225 0
11 3.9 1 0.485 1 0.235225 0.485
12 4.1 2 0.685 4 0.469225 1.37
14 4.7 4 1.285 16 1.651225 5.14
15 4.5 5 1.085 25 1.177225 5.425
16 5.2 6 1.785 36 3.186225 10.71
17 5 7 1.585 49 2.512225 11.095
= 130 44.4 262 19.6569 70.8

Mean = 10.0 3.415

Step 2: Compute the estimates of the regression parameter α and β as:

where a is the estimate of α (the y intercept) and b is the estimate of β (linear regression
coefficient, slope).

b= ;a=

a = 3.415 – 0.27  10 = 0.715 cm

Thus, the estimated linear regression equation is:


for 3 x 17 (avoid extrapolating the regression line beyond the range of
observations)

This is the estimated linear functional relationship between age (days) and wing length (cm). Thus, wing
length increases by 0.27 cm every day.

Step 3: Test the significance of β (linear regression coefficient)

85
 To test β, compute the residual mean square as:

=0.05

The residual mean square denotes the variance of y after taking into account the dependence of y on x.

Compute the test statistic (tb) value as:

tb =

Step 4: Compare the calculated tb value with tabulated t-value at α/2 level of significance, at n-2 (13-2) = 11
d.f.; where n is pair of observations.

At 5% level, t 0.025 (11) = 2.201, and at 1% = t 0.005 (11) = 3.106

Since calculated /tb/ (19.5) is > the tabulated t-value at the 1% level of significance, the linear response of
wing length to changes in the days within the range of 3 to 17 days is highly significant.

12.2.2 Simple linear correlation analysis

The simple linear correlation analysis deals with the estimation and test of significance of the simple linear
correlation coefficient (r), which is a measure of the degree of linear association between two variables x
and y (there is no need to have a dependent and independent variable).

The value of r lies within the range of –1 to +1, with extreme values indicating the perfect linear association
and the mid-value of 0 indicates no-linear association between the two variables. The value of r is negative
when a positive change in one variable is associated with a negative change in another and positive when
the values of two variables change in the same direction (increase or decrease).

Even though the zero r value indicates the absence of linear association between two variables, it does not
indicate the absence of association between them. It is possible for the two variables to have a non-linear
association such as quadratic form. The procedure for the estimation and test of significance of a simple
linear correlation coefficient between two variables x and y are:

Step 1: Compute the means ( ), the sum of square of the deviates ( and ), and the sum of the
cross product of deviates of the two variables.

Step 2: Compute the simple linear correlation coefficient for the above example as:

r=

86
Step 3: Test the significance of the simple linear correlation coefficient (r) by comparing the computed r-
value with the tabulated r-value at n-2 d.f.
 The simple linear correlation coefficient (r) is declared significant at α level of significance if the
absolute value of the computed r-value > the corresponding tabulated r-value.

Computed r = 0.98 with d.f. of n-2 = 13-2 =11


r – table at 5% (11) = 0.55, r at 1% (11) = 0.68

Thus, the simple linear correlation coefficient is significant at 1% level of significance which indicates the
presence of a highly significant and positive linear association between ages and wing length of sparrows.

12.3 Multiple Linear Regression

The simple linear regression and correlation analysis is applicable only in cases with one independent
variable. However, in many situations Y may be dependent on more than one independent variables. Linear
regression analysis involving more than one independent variables is called multiple linear regression. The
relationship of the dependent variable Y to the K independent variables X1, X2, ... Xk can be expressed as:

Y =  + β1X1 + 2X2 +... + βkXk


The data required for the application of multiple linear regression analysis involving K independent
variables are (n) (k + 1) observations, where n is number of pairs (n> 2).

Linear regression involving two independent variables can be expressed as: Y =  + 1X1 + 2X2 where 1 &
2 are partial regression coefficients. 1 measures a change in Y for unit change in X 1, if X2 is held constant.
Similarly, 2 measures the rate of change in Y for a unit change in X 2 where X1 is held constant. 
(sometimes designated as 0) is the value of Y when both X1 & X2 are zero.

For applying multiple linear regression analysis:


- the effect of each of K independent variables on Y must be linear
- the effect of each Xi on Y is independent of the other X, i.e. no interaction.

Example: The following data show the weight gain, initial body weight & age of five chicks fed with
certain type of rations for a month.

Initial age (days) (X1) 5 5 5 4 6


Initial body weight (g) (X2) 10 15 12 15 20
Weight gain (g) (Y) 5 6 7 8 6

Fit the multiple linear regression model Y =  + 1X1 + 2 X2 to the data

Step 1. Calculate the means and square of the deviates

X1 X2 Y y2 x12 x22 x1y x2y x1x2


5 10 5 1.96 0 19.36 0 6.16 0
87
5 15 6 0.16 0 0.36 0 -0.24 0
5 12 7 0.36 0 5.76 0 -1.44 0
4 15 8 2.56 1 0.36 -1.6 0.96 -0.6
6 20 6 0.16 1 31.36 -0.4 -2.24 5.6
Sum 25 72 32 5.2 2 57.2 -2 3.2 5
Mean 5 14.4 6.4
Step 2. Solve for b1 & b2

b1= ; = =-

1.46

b2 = ; = = 0.18

Step 3. Compute the estimate of the intercept as:

a= – b1 -b2

a = 6.4 – (-1.46  5) – 0.18  14.4= 6.4 +7.3-2.59 = 11.11

Thus, the estimated multiple linear regression equation for initial age (days) and initial body weight (g) with
weight gain (g) is: Ŷ= 11.1 - 1.46 X1 + 0.18 X2 for 4  X1  6; and 10  X2 20.

Step 4: Compute:
The sum of squares due to regression (SSR) =

SSR = b1 = (-1.46  -2) + (0.18  3.2) = 3.496

Residual (error) SS =

Coefficient of determination (R2) = = = 0.67

R2 measures the amount of explained variation of Y due to the independent variables. Thus, in the above
example 67% of the total variation in weight gain (g) of chicks can be accounted for a linear function
involving initial age (days) and initial body weight (g).

Step 5: Test the significance of R2


Compute F value as: = = = 2.05

88
where k is number of independent variables (2) and n is number of data pairs (5)

Read table F-value as F (k, n-k-1); F (2, 2) 5% = 19.00; F (2, 2) 1% = 99.00

Since the computed F-value (2.05) is less than the table F value at 5% (19.00) the estimated multiple linear
regression Ŷ= 11.1 - 1.46 X1 + 0.18 X2 is not significant at the 5% level.
 Thus, the combined linear effect of initial age (days) and initial body weight (g) on weight gain (g)
of chicks is not significant.

Remark: The larger the R2 value, the more important the regression equation in characterizing Y.
 On the other hand, if the value of R2 is low, even if the F-test is significant, the estimated linear
regression equation may not be useful.
 For example an R2 value of 0.26, even if significant indicates that only 26% of the total variation in
the dependent variable (Y) is explained by the linear function of the independent variables
considered.

13. DATA TRANSFORMATION

For valid applications of parametric analysis like ANOVA, t-test, etc certain basic assumptions must be met.
If the data violate such assumptions transformation of data can be used.

Data transformation can also be used to reduce the CV

The appropriate type of data transformation to be used depends on the specific type of relationship between
the variances and the means. Conduct the analysis using the transformed data, and in tables present
transformed means in parenthesis alongside their back transformed values out of parenthesis.

The three commonly used data transformation methods are:

13.1 Logarithmic Transformation

More appropriate when the treatment effects are multiplicative rather than additive, then the logarithmic
transformation of the data will exhibit additivity. Such conditions are generally found on count data such as
number of insects per plot, number of eggs per plant, etc.

- X’=log (x + 1); where x is original data and x + 1 is preferred especially when some of the observed
values are small. Logarithmic of base 10 are generally are used but any base would be satisfactory.
- Number of eggs/plant = 9; log (9+1); log (10) = 1

13.2. Square Root Transformation

The square root transformation is applicable when the data consists of counts of rare events such as the
number of infested plants in a plot, the number of insects caught in traps. For such data the variance tends to
be proportional to the mean. Square root transformation is also appropriate for percentage data where the
range is between 0-30% or between 70-100%.

89
If most of the values in the data set are small especially with zeros present.

X’ = ; x = 0; 0.707

Statistical computation can be done on the transformed data. The mean can be expressed in terms of the
original data by squaring the transformed value and subtractions of 0.5

13.3 Arcsine Transformation


An Arcsine or angular transformation is appropriate for data on proportions and data expressed in decimal
fractions or percentages. The arcsine transformation abbreviated “arcsine” frequently referred as “angular
transformation" or “inverse sine” or “sin -1”. However, not all percentage data need to be transformed and
Arcsine is not the only transformation possible.
- Rule 1: For percentage data lying within the range of 30 to 70%, no transformation is needed.
- Rule 2: For percentage data lying within the range of 0-30 or 70 – 100%, but not both, use the
square root transformation.
- Rule 3: For percentage data that do not follow the ranges specified in either rule 1 or rule 2, the
arcsine transformation should be used.

P’ = Sin -1 , P = 100 %= Sin -1 = 90. An angle whose sin is 1 = 90.

For proportion of 0 to 1 (0-100%), the transformed values will range between 0 and 90 degrees, percentage
44%: Sin-1 = 41.55

Transformed values can be transformed back to proportion as:


P = (Sin P’)2 P = (Sin 90)2 = 1 = 100%, P = (Sin 41.55)2 = 0.44 or 44%

90
APPENDIX

Appendix Table I. Area under the Standard Normal Probability Distribution (Z Distribution) -Top Tail
Probabilities (Area to the right of the given Z-values)

Z .00 .01 .02 .03 .04 .05 0.06 0.07 0.08


0.09

0.0 .5000 .4960 .4920 .4880 .4840 .4801 .4761 .4721 .468
0.1 1 .4641
0.2 .4602 .4562 .4522 .4483 .4443 .4404 .4364 .4325 .428
0.3 6 .4247
0.4 .4207 .4168 .4129 .4090 .4052 .4013 .3974 .3936 .389
0.5 7 .3859
0.6 .3821 .3783 .3745 .3707 .3669 .3632 .3594 .3557 .352
0.7 0 .3483
0.8 .3446 .3409 .3372 .3336 .3300 .3264 .3228 .3192 .315
0.9 6 .3121
1.0 .3085 .3050 .3015 .2981 .2946 .2912 .2877 .2843 .281
1.1 0 .2776
1.2 .2743 .2709 .2676 .2643 .2611 .2578 .2546 .2514 .248
1.3 3 .2451
1.4 .2420 .2389 .2358 .2327 .2296 .2266 .2236 .2206 .217
1.5 7 .2148
1.6 .2119 .2090 .2061 .2033 .2005 .1977 .1949 .1922 .189
1.7 4 .1867
1.8 .1841 .1814 .1788 .1762 .1736 .1711 .1685 .1660 .163
1.9 5 .1611
2.0 .1587 .1562 .1539 .1515 .1492 .1469 .1446 .1423 .140
2.1 1 .1379
2.2 .1357 .1335 .1314 .1292 .1271 .1251 .1230 .1210 .119
2.3 0 .1170
2.4 .1151 .1131 .1112 .1093 .1075 .1056 .1038 .1020 .100
2.5 3 .0985
2.6 .0968 .0951 .0934 .0918 .0901 .0885 .0869 .0853 .083
2.7 8 .0823
2.8 .0808 .0793 .0778 .0764 .0749 .0735 .0721 .0708 .069
2.9 4 .0681
3.0 .0668 .0655 .0643 .0630 .0618 .0606 .0594 .0582 .057
1 .0559
.0548 .0537 .0526 .0516 .0505 .0495 .0485 .0475 .046
5 .0455
.0446 .0436 .0427 .0418 .0413 .0406 .0392 .0384 .037
5 .0367
91
Z .00 .01 .02 .03 .04 .05 0.06 0.07 0.08
0.09
.0359 .0351 .0344 .0336 .0329 .0322 .0314 .0307 .030
1 .0294
.0287 .0281 .0274 .0268 .0262 .0256 .0250 .0244 .023
9 .0233
.0228 .0222 .0217 .0212 .0207 .0202 .0197 .0192 .018
8 .0183
.0179 .0174 .0170 .0166 .0162 .0158 .0154 .0150 .014
6 .0143
.0139 .0136 .0132 .0129 .0125 .0122 .0119 .0116 .011
3 .0110
.0107 .0104 .0102 .0099 .0096 .0094 .0091 .0089 .008
7 .0084
.0082 .0080 .0078 .0075 .0073 .0071 .0069 .0068 .006
6 .0064
.0062 .0060 .0059 .0057 .0055 .0054 .0052 .0051 .004
9 .0048
.0047 .0045 .0044 .0043 .0041 .0040 .0039 .0038 .003
7 .0036
.0035 .0034 .0033 .0032 .0031 .0030 .0029 .0028 .002
7 .0026
.0026 .0025 .0024 .0023 .0023 .0022 .0021 .0021 .002
0 .0019
.0019 .0018 .0018 .0017 .0016 .0016 .0015 .0015 .001
4 .0014
.0013 .0013 .0013 .0012 .0012 .0011 .0011 .0011 .001
0 .0010

Remark
The values in the table show area under the Standard Normal Curve to the right of the Z-values.
The Z-value is the combination of the numbers in the left most column and in the heading row at whose
intersection the value in the table appears.
Example: For 95% Confidence, =5% (total of both tails), it is 2.5%= 0.025 in each tail, which is equal to
the Area to the left of –Zα/2=the Area to the right of Zα/2. Therefore, we simply look for this are (0.025) in
the table, and the Z-value is read from the corresponding values in the left most column and the heading
row. For this example (i.e., for 95%, the Z-value is 1.9 from the left most row plus 0.06 from the heading
row, which is 1.96).

92
Appendix Table II. Percentage Points of the t distribution

_______________________________________________________________________________
α (one-tailed)
Degree of _________________________________________________________________
freedom (n-1) 0.25 0.10 0.05 0.025 0.01 0.005 0.0025 0.001 0.0005
________________________________________________________________________________
1 1.000 3.078 6.314 12.706 31.821 63.657 127.32 318.31 636.62
2 0.816 1.886 2.920 4.303 6.965 9.925 14.089 23.328 31.598
3 0.765 1.638 2.353 3.182 4.541 5.841 7.453 10.213 12.924
4 0.741 1.533 2.132 2.776 3.747 4.604 5.598 7.173 8.610
5 0.727 1.475 2.015 2.571 3.365 4.032 4.773 5.893 6.869
6 0.727 1.440 1.943 2.447 3.143 3.707 4.317 5.208 5.959
7 0.711 1.415 1.895 2.365 2.998 3.499 4.019 4.785 5.408
8 0.706 1.397 1.860 2.306 2.896 3.355 3.833 4.501 5.041
9 0.703 1.383 1.833 2.262 2.821 3.250 3.690 4.297 4.780
10 0.700 1.372 1.812 2.228 2.764 3.169 3.581 4.144 4.587
11 0.697 1.363 1.796 2.201 2.718 3.106 3.497 4.025 4.437
12 0.695 1.356 1.782 2.179 2.681 3.055 3.428 3.930 4.318
13 0.694 1.350 1.771 2.160 2.650 3.012 3.372 3.852 4.221
14 0.692 1.345 1.761 2.145 2.624 2.977 3.326 3.787 4.140
15 0.691 1.341 1.753 2.131 2.620 2.947 3.286 3.733 4.073
16 0.690 1.337 1.746 2.120 2.583 2.921 3.252 3.686 4.015
17 0.689 1.333 1.740 2.110 2.567 2.898 3.222 3.646 3.965
18 0.688 1.330 1.734 2.101 2.552 2.878 3.197 3.610 3.922
19 0.688 1.328 1.729 2.093 2.539 2.861 3.174 3.579 3.883
20 0.687 1.325 1.725 2.086 2.528 2.845 3.153 3.552 3.850
21 0.687 1.323 1.721 2.080 2.518 2.831 3.135 3.527 3.819
22 0.686 1.321 1.717 2.074 2.508 2.819 3.119 3.505 3.792
23 0.685 1.319 1.714 2.069 2.500 2.807 3.104 3.485 3.767
24 0.685 1.318 1.711 2.064 2.492 2.797 3.091 3.467 3.745
25 0.684 1.316 1.708 2.060 2.485 2.787 3.078 3.450 3.725
26 0.684 1.315 1.706 2.056 2.479 2.779 3.067 3.435 3.707
27 0.684 1.314 1.703 2.052 2.473 2.771 3.057 3.421 3.690
28 0.683 1.313 1.701 2.048 2.467 2.763 3.047 3.408 3.674
29 0.683 1.311 1.699 2.045 2.462 2.756 3.038 3.396 3.659
30 0.683 1.310 1.697 2.042 2.457 2.750 3.030 3.385 3.646
40 0.681 1.303 1.684 2.021 2.423 2.704 2.971 3.307 3.551
60 0.679 1.296 1.671 2.000 2.390 2.660 2.915 3.232 3.460
120 0.677 1.289 1.658 1.980 2.358 2.617 2.860 3.160 3.373
 0.674 1.282 1.645 1.960 2.326 2.576 2.807 3.090 3.291
______________________________________________________________________________

93
Appendix Table III. Chi-square (2) Table

D.F. 0.99 0.95 0.90 0.75 0.50 0.25 0.20 0.10 0.05 0.02 0.01
1 0.000 0.000 0.016 0.102 0.455 1.32 1.642 2.706 3.841 5.412 6.635
2 0.020 0.103 0.211 0.575 1.386 2.77 3.219 4.605 5.991 7.824 9.210
3 0.115 0.352 0.584 1.213 2.366 4.11 4.642 6.251 7.815 9.837 11.345
4 0.297 0.711 1.064 1.923 3.357 5.38 5.989 7.779 9.488 11.668 13.277
5 0.554 1.145 1.610 2.675 4.351 6.63 7.289 9.236 11.070 13.388 15.086
6 0.872 1.635 2.204 3.455 5.348 7.84 8.558 10.645 12.592 15.033 16.812
7 1.239 2.167 2.833 4.255 6.346 9.04 9.803 12.017 14.067 16.622 18.475
8 1.646 2.733 3.490 5.017 7.344 10.22 11.030 13.362 15.507 18.168 20.090
9 2.088 3.325 4.168 5.899 8.343 11.39 12.242 14.684 16.919 19.679 21.666
10 2.568 3.940 4.865 6.737 9.342 12.55 13.442 15.987 18.307 21.161 23.209
11 3.053 4.575 5.578 7.584 10.341 13.70 14.631 17.275 19.675 22.618 24.725
12 3.572 5.226 6.304 8.438 11.340 14.84 15.812 18.549 21.026 24.054 26.217
13 4.107 5.892 7.042 9.299 12.340 15.98 16.985 19.812 22.362 25.472 27.688
14 4.660 6.571 7.790 10.165 13.339 17.12 18.151 21.064 23.685 26.873 29.141
15 5.229 7.261 8.547 11.036 14.339 18.25 19.311 22.307 24.996 28.259 30.578
16 5.812 7.962 9.312 11.912 15.338 19.37 20.465 23.542 26.296 29.633 32.000
17 6.408 8.672 10.085 12.792 16.338 20.49 21.615 24.769 27.587 30.995 33.409
18 7.015 9.390 10.865 13.675 17.338 21.60 22.760 25.989 28.869 32.346 34.805
19 7.633 10.117 11.651 14.562 18.338 22.72 23.900 27.204 30.144 33.687 36.191
20 8.260 10.851 12.443 15.452 19.337 23.84 25.038 28.412 31.410 35.020 37.566
21 26.171 29.615 32.671 36.343 38.932
22 9.542 12.338 14.041 17.240 21.337 26.04 27.301 30.813 33.924 37.659 40.289
23 28.429 32.007 35.172 38.968 41.638
24 10.856 13.848 15.659 19.037 23.337 28.24 29.553 33.196 36.415 40.270 42.980
25 30.675 34.382 37.652 41.566 44.314
26 12.198 15.379 17.292 20.843 25.336 30.43 31.795 35.563 38.885 42.856 45.642
27 32.912 36.741 40.113 44.140 46.963
28 13.565 16.928 18.939 22.657 27.336 32.62 34.027 37.916 41.337 45.419 48.278
29 35.139 39.087 42.557 46.693 49.588
30 14.953 18.493 20.599 24.478 29.336 34.80 36.250 40.256 43.773 47.962 50.892

94
Appendix Table IV. The 5% and 1% Point for the F-distribution

Degree of freedom for numerator (n1)


DF for P
Denomin
ator (n2) 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16
1 .05 161 200 216 225 230 234 237 239 241 242 243 244 245 245 246 246
.01 4052 4999 5403 5625 5764 5859 5928 5981 6022 6056 6083 6106 6126 6143 6157 6170
2 .05 18.51 19.00 19.16 19.25 19.30 19.33 19.35 19.37 19.38 19.40 19.40 19.41 19.42 19.42 19.43 19.43
.01 98.50 99.00 99.17 99.25 99.30 99.33 99.36 99.37 99.39 99.40 99.41 99.42 99.42 99.43 99.43 99.44
3 .05 10.13 9.55 9.28 9.12 9.01 8.94 8.89 8.85 8.81 8.79 8.76 8.74 8.73 8.71 8.70 8.69
.01 34.12 30.82 29.46 28.71 28.24 27.91 27.67 27.49 27.35 27.23 27.13 27.05 26.98 26.92 26.87 26.83
4 .05 7.71 6.94 6.59 6.39 6.26 6.16 6.09 6.04 6.00 5.96 5.94 5.91 5.89 5.87 5.86 5.84
.01 21.20 18.00 16.69 15.98 15.52 15.21 14.98 14.80 14.66 14.55 14.45 14.37 14.31 14.25 14.20 14.15
5 .05 6.61 5.79 5.41 5.19 5.05 4.95 4.88 4.82 4.77 4.74 4.70 4.68 4.66 4.64 4.62 4.60
.01 16.26 13.27 12.06 11.39 10.97 10.67 10.46 10.29 10.16 10.05 9.96 9.89 9.82 9.77 9.72 9.68
6 .05 5.99 5.14 4.76 4.53 4.39 4.28 4.21 4.15 4.10 4.06 4.03 4.00 3.98 3.96 3.94 3.92
.01 13.75 10.92 9.78 9.15 8.75 8.47 8.26 8.10 7.98 7.87 7.79 7.72 7.66 7.60 7.56 7.52
7 .05 5.59 4.74 4.35 4.12 3.97 3.87 3.79 3.73 3.68 3.64 3.60 3.57 3.55 3.53 3.51 3.49
.01 12.25 9.55 8.45 7.85 7.46 7.19 6.99 6.84 6.72 6.62 6.54 6.47 6.41 6.36 6.31 6.28
8 .05 5.32 4.46 4.07 3.84 3.69 3.58 3.50 3.44 3.39 3.35 3.31 3.28 3.26 3.24 3.22 3.20
.01 11.26 8.65 7.59 7.01 6.63 6.37 6.18 6.03 5.91 5.81 5.73 5.67 5.61 5.56 5.52 5.48
9 .05 5.12 4.26 3.86 3.63 3.48 3.37 3.29 3.23 3.18 3.14 3.10 3.07 3.05 3.03 3.01 2.99
.01 10.56 8.02 6.99 6.42 6.06 5.80 5.61 5.47 5.35 5.26 5.18 5.11 5.05 5.01 4.96 4.92
10 .05 4.96 4.10 3.71 3.48 3.33 3.22 3.14 3.07 3.02 2.98 2.94 2.91 2.89 2.86 2.85 2.83
.01 10.04 7.56 6.55 5.99 5.64 5.39 5.20 5.06 4.94 4.85 4.77 4.71 4.65 4.60 4.56 4.52
11 .05 4.84 3.98 3.59 3.36 3.20 3.09 3.01 2.95 2.90 2.85 2.82 2.79 2.76 2.74 2.72 2.70
.01 9.65 7.21 6.22 5.67 5.32 5.07 4.89 4.74 4.63 4.54 4.46 4.40 4.34 4.29 4.25 4.21
12 .05 4.75 3.89 3.49 3.26 3.11 3.00 2.91 2.85 2.80 2.75 2.72 2.69 2.66 2.64 2.62 2.60
.01 9.33 6.93 5.95 5.41 5.06 4.82 4.64 4.50 4.39 4.30 4.22 4.16 4.10 4.05 4.01 3.97
13 .05 4.67 3.81 3.41 3.18 3.03 2.92 2.83 2.77 2.71 2.67 2.63 2.60 2.58 2.55 2.53 2.51
.01 9.07 6.70 5.74 5.21 4.86 4.62 4.44 4.30 4.19 4.10 4.02 3.96 3.91 3.86 3.82 3.78
14 .05 4.60 3.74 3.34 3.11 2.96 2.85 2.76 2.70 2.65 2.60 2.57 2.53 2.51 2.48 2.46 2.44
.01 8.86 6.51 5.56 5.04 4.69 4.46 4.28 4.14 4.03 3.94 3.86 3.80 3.75 3.70 3.66 3.62
15 .05 4.54 3.68 3.29 3.06 2.90 2.79 2.71 2.64 2.59 2.54 2.51 2.48 2.45 2.42 2.40 2.38
.01 8.68 6.36 5.42 4.89 4.56 4.32 4.14 4.00 3.89 3.80 3.73 3.67 3.61 3.56 3.57 3.49
16 .05 4.49 3.63 3.24 3.01 2.85 2.74 2.66 2.59 2.54 2.49 2.46 2.42 2.40 2.37 2.35 2.33
.01 8.53 6.23 5.29 4.77 4.44 4.20 4.03 3.89 3.78 3.69 3.62 3.55 3.50 3.45 3.41 3.37
17 .05 4.45 3.59 3.20 2.96 2.81 2.70 2.61 2.55 2.49 2.45 2.41 2.38 2.35 2.33 2.31 2.29
.01 8.40 6.11 5.18 4.67 4.34 4.10 3.93 3.79 3.68 3.59 3.52 3.46 3.40 3.35 3.31 3.27
18 .05 4.41 3.55 3.16 2.93 2.77 2.66 2.58 2.51 2.46 2.41 2.37 2.34 2.31 2.29 2.27 2.25
.01 8.29 6.01 5.09 4.58 4.25 4.01 3.84 3.71 3.60 3.51 3.43 3.37 3.32 3.27 3.23 3.19
19 .05 4.38 3.52 3.13 2.90 2.74 2.63 2.54 2.48 2.42 2.38 2.34 2.31 2.28 2.26 2.23 2.21
.01 8.18 5.93 5.01 4.50 4.17 3.94 3.77 3.63 3.52 3.43 3.36 3.30 3.24 3.19 3.15 3.12
20 .05 4.35 3.49 3.10 2.87 2.71 2.60 2.51 2.45 2.39 2.35 2.31 2.28 2.25 2.22 2.20 2.18
.01 8.10 5.85 4.94 4.43 4.10 3.87 3.70 3.56 3.46 3.37 3.29 3.23 3.18 3.13 3.09 3.05
21 .05 4.32 3.47 3.07 2.84 2.68 2.57 2.49 2.42 2.37 2.32 2.28 2.25 2.22 2.20 2.18 2.16
.01 8.02 5.78 4.87 4.37 4.04 3.81 3.64 3.51 3.40 3.31 3.24 3.17 3.12 3.07 3.03 2.99
22 .05 4.30 3.41 3.05 2.82 2.66 2.55 2.46 2.40 2.34 2.30 2.26 2.23 2.20 2.17 2.15 2.13
.01 7.95 5.72 4.82 4.31 3.99 3.76 3.59 3.45 3.35 3.26 3.18 3.12 3.07 3.02 2.98 2.94
23 .05 4.28 3.42 3.03 2.80 2.64 2.53 2.44 2.37 2.32 2.77 2.24 2.20 2.18 2.15 2.13 2.11
.01 7.88 5.66 4.76 4.26 3.94 3.71 3.54 3.41 3.30 3.21 3.14 3.07 3.02 2.97 2.93 2.89
24 .05 4.26 3.40 3.01 2.78 2.62 2.51 2.42 2.36 2.30 2.25 2.22 2.18 2.15 2.13 2.11 2.09
.01 7.82 5.61 4.72 4.22 3.90 3.67 3.50 3.36 3.26 3.17 3.09 3.03 2.98 2.93 2.89 2.85
25 .05 4.24 3.39 2.99 2.76 2.60 2.49 2.40 2.34 2.28 2.24 2.20 2.16 2.14 2.11 2.09 2.07
.01 7.77 5.57 4.68 4.18 3.85 3.63 3.46 3.32 3.22 3.13 3.06 2.99 2.94 2.89 2.85 2.81
26 .05 4.23 3.37 2.98 2.74 2.59 2.47 2.39 2.32 2.27 2.22 2.18 2.15 2.12 2.09 2.07 2.05
.01 7.72 5.53 4.64 4.14 3.82 3.59 3.42 3.29 3.18 3.09 3.02 2.96 2.90 2.86 2.81 2.78
27 .05 4.21 3.35 2.96 2.73 2.57 2.46 2.37 2.31 2.25 2.20 2.17 2.13 2.10 2.08 2.06 2.04
.01 7.68 5.49 4.60 4.11 3.78 3.56 3.39 3.26 3.15 3.06 2.99 2.93 2.87 2.82 2.78 2.75
28 .05 4.20 3.34 2.95 2.71 2.56 2.45 2.36 2.29 2.24 2.19 2.15 2.12 2.09 2.06 2.04 2.02
.01 7.64 5.45 4.57 4.07 3.75 3.53 3.36 3.23 3.12 3.03 2.96 2.90 2.84 2.79 2.75 2.72
29 .05 4.18 3.33 2.93 2.70 2.55 2.43 2.35 2.28 2.22 2.18 2.14 2.10 2.08 2.05 2.03 2.01
.01 7.60 5.42 4.54 4.04 3.73 3.50 3.33 3.20 3.09 3.00 2.93 2.87 2.81 2.77 2.73 2.69
30 .05 4.17 3.32 2.92 2.69 2.53 2.42 2.33 2.27 2.21 2.16 2.13 2.09 2.06 2.04 2.01 1.99
.01 7.56 5.39 4.51 4.02 3.70 3.47 3.30 3.17 3.07 2.98 2.91 2.84 2.79 2.74 2.70 2.66
32 .05 4.15 3.29 2.90 2.67 2.51 2.40 2.31 2.24 2.19 2.14 2.10 2.07 2.04 2.01 1.99 1.97
.01 7.50 5.34 4.46 3.97 3.65 3.43 3.26 3.13 3.02 2.93 2.86 2.80 2.74 2.70 2.65 2.62
34 .05 4.13 3.28 2.88 2.65 2.49 2.38 2.29 2.23 2.17 2.12 2.08 2.05 2.02 1.99 1.97 1.95
.01 7.44 5.29 4.42 3.93 3.61 3.39 3.22 3.09 2.98 2.89 2.82 2.76 2.70 2.66 2.61 2.58
36 .05 4.11 3.26 2.87 2.63 2.48 2.36 2.28 2.21 2.15 2.11 2.07 2.03 2.00 1.98 1.95 1.93
.01 7.40 5.25 4.38 3.89 3.57 3.35 3.18 3.05 2.95 2.86 2.79 2.72 2.67 2.62 2.58 2.54
38 .05 4.10 3.24 2.85 2.62 2.46 2.35 2.26 2.19 2.14 2.09 2.05 2.02 1.99 1.96 1.94 1.92
.01 7.35 5.21 4.34 3.86 3.54 3.32 3.15 3.02 2.92 2.83 2.75 2.69 2.64 2.59 2.55 2.51
40 .05 4.08 3.23 2.84 2.61 2.45 2.34 2.25 2.18 2.12 2.08 2.04 2.00 1.97 1.95 1.92 1.90
.01 7.31 5.18 4.31 3.83 3.51 3.29 3.12 2.99 2.89 2.80 2.73 2.66 2.61 2.56 2.52 2.48
42 .05 4.07 3.22 2.83 2.59 2.44 2.32 2.24 2.17 2.11 2.06 2.03 1.99 1.96 1.94 1.91 1.89
.01 7.28 5.15 4.29 3.80 3.49 3.27 3.10 2.97 2.86 2.78 2.70 2.64 2.59 2.54 2.50 2.46
44 .05 4.06 3.21 2.82 2.58 2.43 2.31 2.23 2.16 2.10 2.05 2.01 1.98 1.95 1.92 1.90 1.88
.01 7.25 5.12 4.26 3.78 3.47 3.24 3.08 2.95 2.84 2.75 2.68 2.62 2.56 2.52 2.47 2.44
46 .05 4.05 3.20 2.81 2.57 2.42 2.30 2.22 2.15 2.09 2.04 2.00 1.97 1.94 1.91 1.89 1.87
.01 7.22 5.10 4.42 3.76 3.44 3.22 3.06 2.93 2.82 2.73 2..66 2.60 2.54 2.50 2.45 2.42
48 .05 4.04 3.19 2.80 2.57 2.41 2.29 2.21 2.14 2.08 2.03 1.99 1.96 1.93 1.90 1.88 1.86
.01 7.19 5.08 4.22 3.74 3.43 3.20 3.04 2.91 2.80 2.71 2.64 2.58 2.53 2.48 2.44 2.40
50 .05 4.03 3.18 2.79 2.56 2.40 2.29 2.20 2.13 2.07 2.03 1.99 1.95 1.92 1.89 1.87 1.85
.01 7.17 5.06 4.20 3.72 3.41 3.19 3.02 2.89 2.78 2.70 2.63 2.56 2.51 2.46 2.42 2.38
55 .05 4.02 3.16 2.77 2.54 2.38 2.27 2.18 2.11 2.06 2.01 1.97 1.93 1.90 1.88 1.85 1.83
.01 7.12 5.01 4.16 3.68 3.37 3.15 2.98 2.85 2.75 2.66 2.59 2.53 2.47 2.42 2.38 2.34
60 .05 4.00 3.15 2.76 2.53 2.37 2.25 2.17 2.10 2.04 1.99 1.95 1.92 1.89 1.86 1.84 1.82
.01 7.08 4.98 4.13 3.65 3.34 3.12 2.95 2.82 2.72 2.63 2.56 2.50 2.44 2.39 2.35 2.31
65 .05 3.99 3.14 2.75 2.51 2.36 2.24 2.15 2.08 2.03 1.98 1.94 1.90 1.87 1.85 1.82 1.80
.01 7.04 4.95 4.10 3.62 3.31 3.09 2.93 2.80 2.69 2.61 2.53 2.47 2.42 2.37 2.33 2.29
70 .05 3.98 3.13 2.74 2.50 2.35 2.23 2.14 2.07 2.02 1.97 1.93 1.89 1.86 1.84 1.81 1.79
.01 7.01 4.92 4.07 3.60 3.29 3.07 2.91 2.78 2.67 2.59 2.51 2.45 2.40 2.35 2.31 2.27
80 .05 3.96 3.11 2.72 2.49 2.33 2.21 2.13 2.06 2.00 1.95 1.91 1.88 1.84 1.82 1.79 1.77
.01 6.96 4.88 4.04 3.56 3.26 3.04 2.87 2.74 2.64 2.55 2.48 2.42 2.36 2.31 2.27 2.23
100 .05 3.94 3.09 2.70 2.46 2.31 2.19 2.10 2.03 1.97 1.93 1.89 1.85 1.82 1.79 1.77 1.75
.01 6.90 4.82 3.98 3.51 3.21 2.99 2.82 2.69 2.59 2.50 2.43 2.37 2.31 2.27 2.22 2.19
120 .05 3.92 3.07 2.68 2.45 2.29 2.18 2.09 2.02 1.96 1.91 1.87 1.83 1.80 1.78 1.75 1.73
.01 6.85 4.79 3.95 3.48 3.17 2.96 2.79 2.66 2.56 2.47 2.40 2.34 2.28 2.23 2.19 2.15
150 .05 3.90 3.06 2.66 2.43 2.27 2.16 2.07 2.00 1.94 1.89 1.85 1.82 1.79 1.76 1.73 1.71
.01 6.81 4.75 3.91 3.45 3.14 2.92 2.76 2.63 2.53 2.44 2.37 2.31 2.25 2.20 2.16 2.12
200 .05 3.89 3.01 2.65 2.42 2.26 2.14 2.06 1.98 1.93 1.88 1.84 1.80 1.77 1.74 1.72 1.69
.01 6.76 4.71 3.88 3.41 3.11 2.89 2.73 2.60 2.50 2.41 2.34 2.27 2.22 2.17 2.13 2.09
400 .05 3.86 3.02 2.63 2.39 2.24 2.12 2.03 1.96 1.90 1.85 1.81 1.78 1.74 1.72 1.69 1.67
.01 6.70 4.66 3.83 3.37 3.06 2.85 2.68 2.56 2.45 2.37 2.29 2.23 2.17 2.13 2.08 2.05
1000 .05 3.85 3.00 2.61 2.38 2.22 2.11 2.02 1.95 1.89 1.84 1.80 1.76 1.73 1.70 1.68 1.65
.01 6.66 4.63 3.80 3.34 3.04 2.82 2.66 2.53 2.43 2.34 2.27 2.20 2.15 2.10 2.06 2.02
 .05 3.84 3.00 2.60 2.37 2.21 2.10 2.01 1.94 1.88 1.83 1.79 1.75 1.72 1.69 1.67 1.64
.01 6.63 4.61 3.78 3.32 3.02 2.80 2.64 2.51 2.41 2.32 2.25 2.18 2.13 2.08 2.04 1.99
Appendix Table V. Significant Studentized Ranges for 5% and 1% Level for Duncan’s Multiple Range Test

Significance P = Number of means for range being tested


Error Level 2 3 4 5 6 7 8 9 10 12 14 16 18 20
d.f.
.05 18.0 18.0 18.0 18.0 18.0 18.0 18.0 18.0 18.0 18.0 18.0 18.0 18.0 18.0
1 .01 90.0 90.0 90.0 90.0 90.0 90.0 90.0 90.0 90.0 90.0 90.0 90.0 90.0 90.0
.05 6.09 6.09 6.09 6.09 6.09 6.09 6.09 6.09 6.09 6.09 6.09 6.09 6.09 6.09
2 .01 14.0 14.0 14.0 14.0 14.0 14.0 14.0 14.0 14.0 14.0 14.0 14.0 14.0 14.0
.05 4.50 4.50 4.50 4.50 4.50 4.50 4.50 4.50 4.50 4.50 4.50 4.50 4.50 4.50
3 .01 8.26 8.5 8.6 8.7 8.8 8.9 8.9 9.0 9.0 9.0 9.1 9.2 9.3 9.3
.05 3.93 4.01 4.02 4.02 4.02 4.02 4.02 4.02 4.02 4.02 4.02 4.02 4.02 4.02
4 .01 6.51 6.8 6.9 7.0 7.1 7.1 7.2 7.2 7.3 7.3 7.4 7.4 7.5 7.5
.05 3.64 3.74 3.79 3.83 3.83 3.83 3.83 3.83 3.83 3.83 3.83 3.83 3.83 3.83
5 .01 5.70 5.96 6.11 6.18 6.26 6.33 6.40 6.44 6.5 6.6 6.6 6.7 6.7 6.8
.05 3.46 3.58 3.64 3.68 3.68 3.68 3.68 3.68 3.68 3.68 3.68 3.68 3.68 3.68
6 .01 5.24 5.51 5.65 5.73 5.81 5.88 5.95 6.00 6.0 6.1 6.2 6.2 6.3 6.3
.05 3.35 3.47 3.54 3.58 3.60 3.61 3.61 3.61 3.61 3.61 3.61 3.61 3.61 3.61
7 .01 4.95 5.22 5.37 5.45 5.53 5.61 5.69 5.73 5.8 5.8 5.9 5.9 6.0 6.0
.05 3.26 3.39 3.47 3.52 3.55 3.56 3.56 3.56 3.56 3.56 3.56 3.56 3.56 3.56
8 .01 4.74 5.00 5.14 5.23 5.32 5.40 5.47 5.51 5.5 5.6 5.7 5.7 5.8 5.8
.05 3.20 3.34 3.41 3.47 3.50 3.52 3.52 3.52 3.52 3.52 3.52 3.52 3.52 3.52
9 .01 4.60 4.86 4.99 5.08 5.17 5.25 5.32 5.36 5.4 5.5 5.5 5.6 5.7 5.7
.05 3.15 3.30 3.37 3.43 3.46 3.47 3.47 3.47 3.47 3.47 3.47 3.47 3.47 3.48
10 .01 4.48 4.73 4.88 4.96 5.06 5.13 5.20 5.24 5.28 5.36 5.42 5.48 5.54 5.55
.05 3.11 3.27 3.35 3.39 3.43 3.44 3.45 3.46 4.46 3.46 3.46 3.46 3.47 3.48
11 .01 4.39 4.63 4.77 4.86 4.94 5.01 5.06 5.12 5.15 5.24 5.28 5.34 5.38 5.39
.05 3.08 3.23 3.33 3.36 3.40 3.42 3.44 3.44 3.46 3.46 3.46 3.46 3.47 3.48
12 .01 4.32 4.55 4.68 4.76 4.81 4.92 4.96 5.02 5.07 5.13 5.17 5.22 5.24 5.26
.05 3.06 3.21 3.30 3.35 3.38 3.41 3.42 3.44 3.45 3.45 3.46 3.46 3.47 3.47
13 .01 4.26 4.48 4.62 4.69 4.74 4.84 4.88 4.94 4.98 5.04 5.08 5.13 5.14 5.15
.05 3.03 3.18 3.27 3.33 3.37 3.39 3.41 3.42 3.44 3.45 3.46 3.46 3.47 3.47
14 .01 4.21 4.42 4.55 4.63 4.70 4.78 4.83 3.87 4.91 4.96 5.00 5.04 5.06 5.07
.05 3.01 3.16 3.25 3.31 3.36 3.38 3.40 3.42 3.43 3.44 3.45 3.46 3.47 3.47
15 .01 4.17 4.37 4.50 4.58 4.64 4.72 4.77 4.81 4.84 4.90 4.94 4.97 4.99 5.00
.05 3.00 3.15 3.23 3.30 3.34 3.37 3.39 3.41 3.43 3.44 3.45 3.46 3.47 3.47
16 .01 4.13 4.34 4.45 4.54 4.60 4.67 4.72 4.76 4.79 4.84 4.88 4.91 4.93 4.94
.05 2.98 3.13 3.22 3.28 3.33 3.36 3.38 3.40 3.42 3.44 3.45 3.46 3.47 3.47
17 .01 4.10 4.30 4.41 4.50 4.56 4.63 4.68 4.72 4.75 4.80 4.83 4.86 4.88 4.89
.05 2.97 3.12 3.21 3.27 3.32 3.35 3.37 3.39 3.41 3.43 3.45 3.46 3.47 3.47
18 .01 4.07 4.27 4.38 4.46 4.53 4.59 4.64 4.68 4.71 4.76 4.79 4.82 4.84 4.85
.05 2.96 3.11 3.19 3.26 3.31 3.35 3.37 3.39 3.41 3.43 3.44 3.46 3.47 3.47
19 .01 4.05 4.24 4.35 4.43 4.50 4.56 4.61 4.64 4.67 4.72 4.76 4.79 4.81 4.82
.05 2.95 3.10 3.18 3.25 3.30 3.34 3.36 3.38 3.40 3.43 3.44 3.46 3.46 3.47
20 .01 4.02 4.22 4.33 4.40 4.47 4.53 4.58 4.61 4.65 4.69 4.73 3.76 4.78 4.79
.05 2.93 3.08 3.17 3.24 3.29 3.32 3.35 3.37 3.39 3.42 3.44 3.45 3.46 3.47
22 .01 3.99 4.17 4.28 4.36 4.42 4.48 4.53 4.57 4.60 4.65 4.68 4.71 4.74 4.75
.05 2.92 3.07 3.15 3.22 3.28 3.31 3.34 3.37 3.38 3.41 3.44 3.45 3.46 3.47
24 .01 3.96 4.14 4.24 4.33 4.39 4.44 4.49 4.53 4.57 4.62 4.64 4.67 4.70 4.72
.05 2.91 3.06 3.14 3.21 3.27 3.30 3.34 3.36 3.38 3.41 3.43 3.45 3.46 3.47
26 .01 3.93 4.11 4.21 4.30 4.36 4.41 4.46 4.50 4.53 4.58 4.62 4.65 4.67 4.69
.05 2.90 3.04 3.13 3.20 3.26 3.30 3.33 3.35 3.37 3.40 3.43 3.45 3.46 3.47
28 .01 3.91 4.08 4.18 4.28 4.34 4.39 4.43 4.47 4.51 4.56 4.60 4.62 4.65 4.67
.05 2.89 3.04 3.12 3.20 3.25 3.29 3.32 3.35 3.37 3.40 4.43 3.44 3.46 3.47
30 .01 3.89 4.06 4.16 4.22 4.32 4.36 4.41 4.45 4.48 4.54 4.58 4.61 4.63 4.65
Appendix Table VI. Critical Values of the q Distribution ( = 0.05); k or p = number of mean to be compared; v = error degree of
freedom.

v k(or p):2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20
1 17.97 26.98 32.82 37.08 40.41 43.12 45.4 47.36 49.07 50.59 51.96 53.20 54.33 55.36 56.32 57.22 58.04 58.83 59.56
2 6.085 8.331 9.798 10.88 11.74 12.44 13.03 13.54 13.99 14.39 14.75 15.08 15.38 15.65 15.91 16.14 16.37 16.57 16.77
3 4.501 5.910 6.825 7.502 8.037 8.478 8.853 9.177 9.462 9.717 9.946 10.15 10.35 10.53 10.69 10.84 10.98 11.11 11.24
4 3.927 5.040 5.757 6.287 6.707 7.053 7.347 7.602 7.826 8.027 8.208 8.373 8.525 8.664 8.794 8.914 9.028 9.134 9.233
5 3.635 4.602 5.218 5.673 6.033 6.330 6.582 6.802 6.995 7.168 7.324 7.466 7.596 7.717 7.828 7.932 8.030 8.122 8.208
6 3.461 4.339 4.896 5.305 5.628 5.895 6.122 6.319 6.493 6.649 6.789 6.917 7.034 7.143 7.244 7.338 7.426 7.508 7.587
7 3.344 4.165 4.681 5.060 5.359 5.606 5.815 5.998 6.158 6.302 6.431 6.550 6.658 6.759 6.852 6.939 7.020 7.097 7.17
8 3.261 4.041 4.529 4.886 5.167 5.399 5.597 5.767 5.918 6.054 6.175 6.287 6.389 6.483 6.571 6.653 6.729 6.802 6.87
9 3.199 3.949 4.415 4.756 5.024 5.244 5.432 5.595 5.739 5.867 5.983 6.089 6.186 6.276 6.359 6.437 6.510 6.579 6.644
10 3.151 3.877 4.327 4.654 4.912 5.124 5.305 5.461 5.599 5.722 5.833 5.935 6.028 6.114 6.194 6.269 6.339 6.405 6.467
11 3.113 3.82 4.256 4.574 4.823 5.028 5.202 5.353 5.487 5.605 5.713 5.811 5.901 5.984 6.062 6.134 6.202 6.265 6.326
12 3.082 3.773 4.199 4.508 4.751 4.950 5.119 5.265 5.395 5.511 5.615 5.710 5.798 5.878 5.953 6.023 6.089 6.151 6.209
13 3.055 3.735 4.151 4.453 4.690 4.885 5.049 5.192 5.318 5.431 5.533 5.625 5.711 5.789 5.862 5.931 5.995 6.055 6.112
14 3.033 3.702 4.111 4.407 4.639 4.829 4.990 5.131 5.254 5.364 5.463 5.554 5.637 5.714 5.786 5.852 5.915 5.974 6.029
15 3.014 3.674 4.076 4.367 4.595 4.782 4.94 5.077 5.198 5.306 5.404 5.493 5.574 5.649 5.72 5.785 5.846 5.904 5.958
16 2.998 3.649 4.046 4.333 4.557 4.741 4.897 5.031 5.150 5.256 5.352 5.439 5.520 5.593 5.662 5.727 5.786 5.843 5.897
17 2.984 3.628 4.020 4.303 4.524 4.705 4.858 4.991 5.108 5.212 5.307 5.392 5.471 5.544 5.612 5.675 5.734 5.790 5.842
18 2.971 3.609 3.997 4.277 4.495 4.673 4.824 4.956 5.071 5.174 5.267 5.352 5.429 5.501 5.568 5.63 5.688 5.743 5.794
19 2.960 3.593 3.977 4.253 4.469 4.645 4.794 4.924 5.038 5.14 5.231 5.315 5.391 5.462 5.528 5.589 5.647 5.701 5.752
20 2.950 3.578 3.958 4.232 4.445 4.620 4.768 4.896 5.008 5.108 5.199 5.282 5.357 5.427 5.493 5.553 5.610 5.663 5.714
24 2.919 3.532 3.901 4.166 4.373 4.541 4.684 4.807 4.915 5.012 5.099 5.179 5.251 5.319 5.381 5.439 5.494 5.545 5.594
30 2.888 3.486 3.845 4.102 4.302 4.464 4.602 4.720 4.824 4.917 5.001 5.077 5.147 5.211 5.271 5.327 5.379 5.429 5.475
40 2.858 3.442 3.791 4.039 4.232 4.389 4.521 4.635 4.735 4.824 4.904 4.977 5.044 5.106 5.163 5.216 5.266 5.313 5.358
60 2.829 3.399 3.737 3.977 4.163 4.314 4.441 4.55 4.646 4.732 4.808 4.878 4.942 5.001 5.056 5.107 5.154 5.199 5.241
120 2.800 3.356 3.685 3.917 4.096 4.241 4.363 4.468 4.560 4.641 4.714 4.781 4.842 4.898 4.950 4.998 5.044 5.086 5.126
 2.772 3.314 3.633 3.858 4.030 4.170 4.286 4.387 4.474 4.552 4.622 4.685 4.743 4.796 4.845 4.891 4.934 4.974 5.012
Critical Values of the q Distribution ( = 0.01); k or p = number of mean to be compared; v = error degree of freedom.

V k(or p):2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20
1 90.03 135.0 164.3 185.6 202.2 215.8 227.2 237.0 245.6 253.2 260.0 266.2 271.8 277.0 281.8 286.3 290.4 294.3 298.0
2 14.04 19.02 22.29 24.72 26.63 28.20 29.53 30.68 31.69 32.59 33.4 34.13 34.81 35.43 36.00 36.53 37.03 37.5 37.95
3 8.261 10.62 12.17 13.33 14.24 15.00 15.64 16.20 16.69 17.13 17.53 17.89 18.22 18.52 18.81 19.07 19.32 19.55 19.77
4 6.512 8.120 9.173 9.958 10.58 11.10 11.55 11.93 12.27 12.57 12.84 13.09 13.32 13.53 13.73 13.91 14.08 14.24 14.4
5 5.702 6.976 7.804 8.421 8.913 9.321 9.669 9.972 10.24 10.48 10.70 10.89 11.08 11.24 11.40 11.55 11.68 11.81 11.93
6 5.243 6.331 7.033 7.556 7.973 8.318 8.613 8.869 9.097 9.301 9.485 9.653 9.808 9.951 10.08 10.21 10.32 10.43 10.54
7 4.949 5.919 6.543 7.005 7.373 7.679 7.939 8.166 8.368 8.548 8.711 8.86 8.997 9.124 9.242 9.353 9.456 9.554 9.646
8 4.746 5.635 6.204 6.625 6.960 7.237 7.474 7.681 7.863 8.027 8.176 8.312 8.436 8.552 8.659 8.760 8.854 8.943 9.027
9 4.596 5.428 5.957 6.348 6.658 6.915 7.134 7.325 7.495 7.647 7.784 7.910 8.025 8.132 8.232 8.325 8.412 8.495 8.573
10 4.482 5.270 5.769 6.136 6.428 6.669 6.875 7.055 7.213 7.356 7.485 7.603 7.712 7.812 7.906 7.993 8.076 8.153 8.226
11 4.392 5.146 5.621 5.97 6.247 6.476 6.672 6.842 6.992 7.128 7.250 7.362 7.465 7.560 7.649 7.732 7.809 7.883 7.952
12 4.320 5.046 5.502 5.836 6.101 6.321 6.507 6.670 6.814 6.943 7.060 7.167 7.265 7.356 7.441 7.520 7.594 7.665 7.731
13 4.260 4.964 5.404 5.727 5.981 6.192 6.372 6.528 6.667 6.791 6.903 7.006 7.101 7.188 7.269 7.345 7.417 7.485 7.548
14 4.210 4.895 5.322 5.634 5.881 6.085 6.258 6.409 6 .543 6.664 6.772 6.871 6.962 7.047 7.126 7.199 7.268 7.333 7.395
15 4.168 4.836 5.252 5.556 5.796 5.994 6.162 6.309 6.439 6.555 6.660 6.757 6.845 6.927 7.003 7.074 7.142 7.204 7.264
16 4.131 4.786 5.192 5.489 5.722 5.915 6.079 6.222 6.349 6.462 6.564 6.658 6.744 6.823 6.898 6.967 7.032 7.093 7.152
17 4.099 4.742 5.140 5.430 5.659 5.847 6.007 6.147 6.270 6.381 6.480 6.572 6.656 6.734 6.806 6.873 6.937 6.997 7.053
18 4.071 4.703 5.094 5.379 5.603 5.788 5.944 6.081 6.201 6.310 6.407 6.497 6.579 6.655 6.725 6.792 6.854 6.912 6.968
19 4.046 4.67 5.054 5.334 5.554 5.735 5.889 6.022 6.141 6.247 6.342 6.430 6.510 6.585 6.654 6.719 6.780 6.837 6.891
20 4.024 4.639 5.018 5.294 5.510 5.688 5.839 5.970 6.087 6.191 6.285 6.371 6.450 6.523 6.591 6.654 6.714 6.771 6.823
24 3.956 4.546 4.907 5.168 5.374 5.542 5.685 5.809 5.919 6.017 6.106 6.186 6.261 6.330 6.394 6.453 6.510 6.563 6.612
30 3.889 4.455 4.799 5.048 5.242 5.401 5.536 5.653 5.756 5.849 5.932 6.008 6.078 6.143 6.203 6.259 6.311 6.361 6.407
40 3.825 4.367 4.696 4.931 5.114 5.265 5.392 5.502 5.559 5.686 5.764 5.835 5.900 5.961 6.017 6.069 6.119 6.165 6.209
60 3.762 4.282 4.595 4.818 4.991 5.133 5.253 5.356 5.447 5.528 5.601 5.667 5.728 5.785 5.837 5.886 5.931 5.974 6.015
120 3.702 4.2 4.497 4.709 4.872 5.005 5.118 5.214 5.299 5.375 5.443 5.505 5.562 5.614 5.662 5.708 5.750 5.790 5.827
 3.643 4.12 4.403 4.603 4.757 4.882 4.987 5.078 5.157 5.227 5.290 5.348 5.400 5.448 5.493 5.535 5.574 5.611 5.645

Common questions

Powered by AI

Logarithmic transformations help in cases where effects are multiplicative by converting them into additive effects, which simplifies analysis and mean separation . This is important for properly interpreting data and ensuring valid comparisons. Checking for homogeneity of experimental errors (homoscedasticity) is crucial as unequal variances, known as heteroscedasticity, can bias results and invalidate statistical tests . Homogeneity ensures that all treatment variances are equivalent, which is fundamental for the assumptions underlying analysis of variance (ANOVA). Inconsistent variance, such as in data with non-normal distribution, can skew significance tests, necessitating transformations for accurate results .

A Completely Randomized Design (CRD) is appropriate when experimental units are homogenous and environmental effects are easily controlled, such as in laboratory or greenhouse experiments . Its limitations include increased variation among plots in field experiments due to heterogeneity in factors like soil fertility and slope, making CRD rarely used in such scenarios . Additionally, while CRD simplifies data handling in cases of missing data, it treats the variation among plots as experimental error, which can lead to biased results in the presence of substantial environmental variability .

The Triple Lattice Design offers advantages in terms of increased precision compared to traditional designs like the Completely Randomized Design or Randomized Complete Block Design. It is particularly beneficial in large experiments with numerous treatments where including all treatments within a block isn’t feasible . The design improves precision by balancing out variability within blocks and using an adjustment factor to correct treatment totals, minimizing bias in the treatment sum of squares, enhancing the reliability of results . Most beneficial in contexts with a large number of treatments and potential heterogeneity, the Triple Lattice Design facilitates more precise comparisons by accommodating multiple sources of variation in a systematic and efficient manner .

In experimental research, factors such as loss of data due to unforeseen circumstances or experimental errors necessitate the estimation of missing data to maintain the dataset's completeness and reliability . Estimation processes help maintain data integrity by allowing the recovery of incomplete datasets, enabling analysis of variance and other statistical tests to proceed without bias. This ensures that missing data doesn't disproportionately affect results, as incomplete datasets can skew error estimates and lead to misinterpretation of treatment effects . Techniques like imputing missing values using row, column, and treatment totals are employed to adjust the sum of squares for a balanced analysis .

A Latin Square Design is used when there are two known sources of variation that need to be controlled simultaneously, unlike a Randomized Complete Block Design (RCBD), which addresses only one. It is particularly beneficial when the number of treatments equals the number of replications . The advantage of a Latin Square Design is its greater precision in controlling variation across two directions, often referred to as row and column blocking, which results in better isolation of treatment effects by accounting for variability . This leads to increased precision over both CRD and RCBD when appropriate and is useful in experimental conditions where controlling two directional gradients is necessary, such as soil fertility and irrigation conditions .

Blocking in a Randomized Complete Block Design (RCBD) involves grouping experimental units into blocks based on a known source of variation such as soil heterogeneity or animal characteristics, with each block containing all treatments . This technique reduces experimental error by minimizing variability within blocks and maximizing it among blocks, thereby isolating and removing the contribution of the known source of variability from the experimental error . RCBD offers greater precision than a Completely Randomized Design (CRD) by controlling within-block variation . Its flexibility allows the inclusion of extra replications for certain treatments, and it simplifies analysis even when data from some blocks or treatments are missing . However, its effectiveness can diminish when the number of treatments is large, which may lead to large within-block variation .

Replication in experiments improves precision by increasing the number of times a treatment is applied, which narrows the estimates of treatment means closer to true values, and provides multiple observations to estimate experimental error, thus enhancing the reliability of results . The number of replications required is determined by the desired precision level, variability of experimental units, the number of treatments, and the type of experimental design . High precision requires more replications, while more uniform experimental units require fewer replications compared to variable units . More treatments generally need fewer replications than experiments with fewer treatments .

Systematic designs can violate the assumption of independence of errors when treatments are assigned non-randomly, potentially leading to dependencies where the error of one treatment affects another, such as through environmental spillover effects like pesticide drift . This violates the fundamental assumption needed for valid ANOVA results. Randomization plays a crucial role in mitigating this by ensuring treatments are independently assigned to experimental units, thereby minimizing the risk of correlated errors and enhancing the robustness of error independence . Proper randomization ensures that variations due to systematic errors are averaged out across treatments, thus maintaining the validity of statistical inferences.

Homoscedasticity refers to the condition where experimental errors (variances) among treatments are uniform, meaning they have the same variance across all levels . This assumption is critical for validly applying ANOVA as it ensures that the treatment variance is adequately estimated and that any observed differences are due to actual treatment effects rather than unequal error variance . In cases where variances are not equal (heteroscedasticity), it may indicate that certain treatments have unusually high or low variances, potentially affecting the reliability of conclusions drawn from statistical tests . Ensuring homoscedasticity is crucial to maintain the integrity of statistical comparisons across treatments.

Strategies to reduce experimental error include increasing the size of the experiment, refining experimental techniques, utilizing blocking, and applying replication. Increasing the size of the experiment through more replicates or additional treatments enhances the precision of mean estimates and provides multiple observations to estimate experimental error . Refining techniques ensures uniform application of treatments and controls external influences . Blocking helps to measure contributions of extraneous factors to total variability by dividing the field into homogenous parts . Replication enhances precision, provides error estimates, and increases the inference scope due to its repetition over time and locations, allowing more generalized conclusions . Collectively, these strategies ensure comparability and reliability of the experiment's results by minimizing variation not due to the treatment effects.

You might also like