History and Applications of Statistics
History and Applications of Statistics
INTRODUCTION
In biology, including agricultural sciences, “the laws of nature” are not that simple. Biological phenomena
often show variation that obscures the law we want to establish. For instance, if we treat two fields in
identical way, the yields obtained from the same variety of a crop will not be the same. Similarly, two cows
of the same age, breed, and body weight fed with the same type and amount of fodder will produce different
amounts of milk. Thus, variation is a typical feature of biological data. Such variation is systematically
studied, quantified, and interpreted using statistical sciences.
Statistics is defined as study of numerical data based on variation in nature. It is a science, which deals with
collection, classification, tabulation, summarization and analysis of quantitative data or numerical facts.
Statistics was developed to deal with problems in which, for the individual observations, laws of cause and
effect are not apparent to the observer and where an objective approach is needed. In such problems, there
must always be some uncertainty about any inference based on a limited number of observations.
The word statistics also refers to numerical and quantitative data such as statistics of births, deaths,
marriage, production, yield, etc. The application of statistical methods to the solution of biological problems
is biometry, biological statistics or bio-statistics. The word biometry comes from two Greek roots: 'bios'
mean life and metron mean 'to measure'. Thus, biometry literally means the measurement of life.
The term statistics is an old one. Statistics must have started as a state arithmetic technique to assist a ruler
who needed to know the wealth and number of his subjects in order to levy a tax or wage a war. We know
that Caesar Augustus sent out a decree that the entire world should be taxed. Consequently, he required that
all persons report to the nearest statistician-in that day the tax collector.
William, the conqueror ordered a survey of the lands of England for purposes of taxation and military
service. This was called the Domesday Book, which is a manuscript record of the "Great Survey" of much
of England and parts of Wales completed in 1086 by order of King William the Conqueror.
Several centuries after the Domesday Book, we find an application of empirical probability in ship
insurance, which seems to have been available to Flemish shipping in the fourteenth century. This can be
little more than speculation or gambling, but it developed in to the very respectable form of statistics called
insurance.
Gambling in the form of games of chance led to the theory of probability originated by Pascal and Fermat
about the middle of the seventeenth century because of their interest in the gambling experiences of the
Chevalier de Mere. To the statistician and the experimental scientist, the theory contains much of practical
use for the processing of data.
The normal curve or normal curve of error has been very important in the development of statistics. The
equation of this curve was first published in 1733 by de Moivre. De Moivre had no idea of applying his
result to experimental observations and his paper remained unknown until Karl Pearson found it in a library
1
in 1924. However, the same result was later developed by two mathematical astronomers, Laplace (1749-
1827) and Gauss (1777-1855) independently of one another.
Charles Darwin (1809-1882), a biologist, received the second volume of Lyell’s book while on the Beagle.
Darwin formed his theories later and he may have been stimulated by his reading of this book. Darwin’s
work was largely biometrical or statistical in nature and he certainly renewed enthusiasm in biology. Gregor
Mendel (1822-1884) too, with his studies of plant hybrids published in 1866, had a biometrical or statistical
problem.
In the nineteenth century, the need for a sounder basis for statistics became apparent. Karl Pearson (1857-
1936) initially a mathematical physicist applied his mathematics to evolution as a result of the enthusiasm in
biology created by Darwin. Pearson spent nearly half a century in serious statistical research. In addition, he
founded the journal Biometrica and a school of statistics; as a result the study of statistics gained impetus.
While Pearson was concerned with large samples, large-sample theory was proved to be somewhat
inadequate for experimenters with necessarily small samples. Among these was W. S. Gosset (1876-1937),
a student of Karl Pearson and a scientist of the Guinness firm of brewers. Gosset’s mathematics appears to
have been insufficient to the task of finding exact distributions of the sample standard deviation, of the ratio
of the sample mean to the sample standard deviation, and of the correlation coefficient, statistics with which
he was particularly concerned. Consequently, he resorted to drawing shuffled cards, computing and
compiling empirical frequency distributions. Papers on the results appeared in Biometrika in 1908 under the
name of student, Gosset’s pseudonym. Today Student’s t is a basic tool of statisticians and experimenters.
Now that the use of Student’s t distribution is so widespread, it is interesting to note that the German
astronomer, Helmert, had obtained it mathematically as early as 1875.
R.A. Fisher (1890-1962) was influenced by Karl Pearson and Student and made numerous and important
contributions to statistics. He and his students gave considerable impetus to the use of statistical procedures
in many fields, particularly in agriculture, biology and genetics.
Abrahm Wald (1902-1950) has contributed two books on Sequential Analysis and Statistical Decision
Functions. Thus, it is in the 20th century that most of the statistical methods presently used have been
developed.
Statistics is important in the field of social science, agriculture, medical, engineering, etc because it provides
tools to analyze collected data. Scientists frequently use statistics to analyze research data. Statistics
provides scientific methods for appropriate data collection, analysis and summarization of data and
inferential statistical methods for drawing conclusions in the face of uncertainty. Statistical methodologies
have wide applicability to almost any branch of science dealing with the study of uncertain phenomena.
Statistics has already become a very important and useful subject and the various techniques are being used
to analyze and solve the problems in different discipline. Thus, currently statistics is used as an analytical
tool in many fields of research.
2
However, there are some of the limitations, i.e. it is not suited to study the qualitative phenomenon; statistics
does not study individuals as it deals with an aggregate of objects and does not give any special importance
to the individual of a series; statistical laws are not exact like as physical or natural law of sciences,
statistical analysis is only in terms of probability and chance not an exact; and statistics is liable to be
misused as statistical methods are more dangerous tools in the hand of the inexpert.
Variable: A property with respect to which individuals in a sample differ in some way, e.g. length, weight,
height, color, etc. Characteristics, which show variations, are called random variables.
3. Attribute or nominal variables: Variables that cannot be measured but are expressed qualitatively, e.g.
color of common bean seed: white, black, red; sex of animals could be male or female, blood group of
humans could be A, AB, O, etc.
When such attribute data are combined with frequency (number of occurrences), they can be treated
statistically. Such data are called enumeration data. Suppose we have 18 mixed bean seeds, they can be
grouped as: black (3), white (5) and red (10).
Variate (datum): A single reading, score or observation of a given variable. If we measure height of 5
plants from a plot, each of the 5 readings of height will be a variate: 10 cm, 5 cm, 20 cm, 25 cm, 15 cm.
The basic method of collecting the observations in a sample is called simple random sampling. This is
where any observation has the same probability of being collected, e.g. giving each student in a class equal
3
chance in measuring height. The aim is always to sample in a manner that does not create a bias in favour of
any observation being selected. Nearly all applied statistical procedures that are concerned with using
samples to make inferences (i.e. draw conclusions) about populations assume some form of random
sampling. If the sampling is not random, then we are never sure as to what population is represented by our
sample. When random sampling from clearly defined populations is not possible, then interpretation of
standard methods of estimation becomes more difficult. [Read the other different types of probability and
non-probability sampling methods and their applications from statistical books].
Populations must be defined at the start of any study and this definition should include the
spatial and temporal limits to the population. Our formal statistical inference is restricted to these limits. For
example, if we sample from a population of animals at a certain location in December 2010, then our
inference is restricted to that location in December 2010. We cannot infer what the population might be like
at any other time or in any other place, although we can speculate or make predictions.
Parameter: A population value, which we generally do not know, but would like to infer (estimate) about.
For example, national average yields of maize in 2010 in Ethiopia. Parameters are designated using Greek
letters such as µ, , , etc.
Statistics: Are sample estimates of population value (parameters) and designated using Latin letters such as
, s, p, etc.
The population parameters cannot be measured directly because the populations are usually too large, i.e.
they contain too many observations for practical measurement. It is important to remember that population
parameters are usually considered to be fixed, but unknown, values so they are not random variables and
do not have probability distributions. Sample statistics are random variables, because their values
depend on the outcome of the sampling experiment, and therefore they do have probability distributions,
called sampling distributions.
1. Descriptive (deductive) statistics: Are methods, which are used to describe a set of data without
involving generalization. Deal with the presentation of research data or any numerical information. Help
in summarizing and organizing data so as to make them readable for users, e.g. mean, median, mode,
standard deviation, etc.
2. Inferential (inductive) statistics: It is a statistics, which helps, in drawing conclusion about the whole
(population) based on data from some of its parts (samples).
2. STATISTICAL INFERENCE
2.1 Estimation
Most of the time, we make decisions about population on the basis of sample information because it is
generally difficult and sometimes impossible to consider all the individuals in a population for economic,
4
time and other reasons. For example, we take random sample and calculate the sample mean ( ) as an
estimate of population mean (µ) and sample variance (s2) as an estimate of population variance (2).
Estimation is a technique, which enables us to estimate the value of population parameter based on sample
values. The formula that is used to make estimation is known as estimator whereas the resulting sample
value is called an estimate.
There are two kinds of estimation: point estimation and interval estimation.
Point estimation is a technique by which a single value is obtained as an estimate of a population parameter
like = 10; s2 = 0.6 are point estimates of µ and 2, respectively.
Interval estimation is an estimation technique in which two limits within which a parameter is expected
to be found are determined. It provides a range of values that might include the parameter with a known
probability, e.g. confidence intervals. For example, population mean (µ) lies between two points such that
a<µ<b where a and b are lower and higher limits, respectively, and are obtained from sample
observations, e.g. the average weight of students in a class lies between 50 kg and 60 kg.
Unbiasedness: the expected value of the sample statistic (the mean of its probability distribution)
should be equal to the parameter. Repeated samples should produce estimates which do not consistently
under or over estimate the population parameter.
Consistency: as the sample size increases, then the estimate will get closer to the population
parameter. Once the sample includes the whole population, the sample statistic will obviously equals the
population parameter.
Efficiency: it has the lowest variance among all competing estimators. For example, the sample
mean is a more efficient estimator of the population mean of a variable with a normal probability
distribution than the sample median, despite the two statistics being numerically equivalent.
Sufficiency: the estimates should have all the required information about the parameter, which is
being estimated. For example, the sample mean is a more sufficient estimator than the range as it includes
all the observations.
The sample mean ( ) and variance (s2) are point estimates of population mean (µ) and population variance
(2), respectively. These sample estimates may or may not be equal to the population mean. Thus, it is better
to give our estimation in interval and state that the population mean (µ) is included in the interval with some
measure of confidence.
Confidence interval is a probability statement concerning the limit within which a given parameter lies,
e.g. we can say the probability that µ lies within the limit a and b is 0.95 where a< b; P (a<µ< b) = 0.95
5
Case 1: When population standard deviation () is known (given), the sample size could be large (n
30) or small (n<30)
A (100- ) % Confidence Interval for population mean (µ) = Z/2 where
- = allowable error rate
- - Z/2 = L1 (lower confidence limit)
Example: A certain population has standard deviation of 10 and a sample of 100 observations were taken
from a population and the mean of the sample was 4.00
Find the point estimate and construct the 95% confidence interval for the population mean (µ).
Given: =10; = 4.00; n = 100. Thus,
- Point estimate = 4.00
- A 95% confidence interval for population mean
= Z/2 = 4.00 Z0.025 = 4.00 1.961.00
- L1 = 4.00-(1.96 x 1.00) = 2.04
- L2= 4.00 + (1.96 x 1.00) = 5.96
Thus, 95% C. I. for population mean is 2.04 < µ < 5.96. This means, the probability that the population
mean lies in between 2.04 and 5.96 is 0.95 or we are 95% confident that the population mean is included in
this range.
If we increase the confidence coefficient from 0.95 to 0.99, the width of the interval increases and the more
certain we can be that the true mean is included in this estimated range. However, as we increase the
confidence level, the estimate will loose some precision (will be more vague).
6
- A 99% C. I. = Z/2 = 60 Z0.005 = 60 2.58 0.83
Thus, 99% C.I. for population mean = 57.86 g<µ<62.14 g. The probability that the mean weight gain per
month is between 57.86 g and 62.14 g is 0.99, or we are 99% confident that the population mean is found in
this interval.
The central limit theorem states that as sample size increases, the sampling distribution of means approaches
a normal distribution. However, samples of small size do not follow the normal distribution curve, but t-
distribution.
Thus, A (100- )% C.I. for population mean (µ) = t/2 (n-1) where n-1 is degree of freedom:
Properties of t-distribution:
- symmetrical in shape and has a mean equal to zero like normal (Z) distribution.
- the shape of t-distribution is flatter than the Z distribution, but as sample size approaches 30, the
flatness associated with the t-distribution disappears and the curve approximates the shape of Z-
distribution.
- there are different distributions for each degree of freedom
t for n = 30
t for n = 20
t for n =10
μ =0
Example: A sample of 25 horses has average age of 24 years with standard deviation of 4 years. Calculate
95% C.I. for population mean.
A 95% C.I. = t/2 (n-1) = 24 t0.025 (25-1) = 24 2.064 0.8. Thus, a 95% C.I. = 22.35
years < < 25.65 years
Case 1: When 12 & 22 (population variances) are known (n1 and n2) could be large or small)
7
- Point estimate = - if > or - if >
- Z/2 if >
where and are sample mean of 1st and 2nd group; n1 and n2 are sample sizes of 1st and 2nd group; 12 & 22
are population variances of 1st and 2nd group.
Case 2: 12 & 22 are unknown, but with large sample sizes (n1 and n2 30)
For large sample size, s12 is an estimate of 12 and s22 is an estimate of 22
- Point estimate = - if > or - if >
Example: The performance of two breeds of dairy cattle (Holstein & Newjersy) was tested by taking a
sample of 100 cattle each. The daily milk yield was as follows. Find point estimate and construct a 90% C.
I. for population mean difference.
- Holstein (x): = 15 liters; s1 = 2 liters.
- Newjersy(y): = 10 liters; s2 = 1.5 liters
Case 3: 12 & 22 are unknown and n1 and n2 are small (< 30)
- Point estimate = - if > or - if >
- A (100-)% C.I. for population mean difference
= - t/2 (n1+n2-2) sp
8
Example: Two groups of nine framers were selected, one group using the new type of plough and another
group using old type of plough. Assume the soil, weather, etc. conditions are the same for both groups
(variation is only due to plough types). The yields and standard deviation were:
Estimate the difference between two plough types and construct a 99% confidence interval for the
difference of the population means:
- Point estimate = - = 35.22-31.55 = 3.67 qt
Thus, 99% C.I. for pop mean difference = -2.81qt <1–2< 10.15 qt
In attempting to reach at decisions, we make assumptions or guesses about the population. Such assumption,
which may or may not be true is called statistical hypothesis.
Examples:
- The ratio of male to female in Ethiopia is 1:1
- The cross of 2 heterozygous varieties for color of tomato will produce plants with red and white
flowers in the ratio of 3:1. Rr Rr RR: 2Rr: rr where red is dominant over white.
A procedure that enables us to decide whether to accept or reject the null hypothesis or procedure used to
determine whether observed samples differ significantly from expected results are called test of hypothesis
or rules of decision.
9
Rejection and Acceptance of Hypothesis
A hypothesis is rejected if the probability that it is true is less than some predetermined probability. The
predetermined probability is selected by the investigator before he collects his data based on: his research
experience, the consequences of an incorrect decision and the kind of risk the researcher prepares to
take.
When the null hypothesis is rejected, the finding is said to be statistically significant at that specified level
of significance. When the available evidence does not support the rejection of the null hypothesis, the
finding is said to be non-significant.
The two kinds of correct decisions are accepting a true null hypothesis and rejecting a false null
hypothesis.
The probability of type I error is the significance level, i.e. in a given hypothesis, the maximum
probability which we will be willing to risk a type I error is called the level of significance of the test. It is
denoted by and is generally specified before samples are taken. In practice 5% or 1% levels of
significance are in common use. A 5% level of significance means that there are about 5 chances in 100 that
we would reject null hypothesis (H O) when it should be accepted, i.e. we are 95% sure that we make a
correct decision.
The probability of type II error (accepting a false null hypothesis) is represented by . Type II error usually
occurs when the sample size is small. The more an investigator protects his experiment against type I error,
the greater the likelihood of committing type II error.
Thus, the value of must be chosen to minimize type I as well as type II error.
Power of Test: This is defined as the probability of rejecting the hypothesis when it is false and
symbolically given as (1-). We have to design our experiment in such away that the power of test is
maximized.
In most cases, the HO (null hypothesis) is one of no effect (e.g. no difference between two means) and the
HA (the alternative hypothesis) can be in either direction; the H O is rejected if one mean is bigger than the
other mean or vice versa. This is termed a two-tailed test because large values of the test statistic at either
end of the sampling distribution will result in rejection of H O. In two-tailed test, to do a test with α = 0.05,
then we use critical values of the test statistic at α/2 = 0.025 at each end of the sampling distribution.
10
Sometimes, our HO is more specific than just no difference. We might only be interested in whether one
mean is bigger or smaller than the other mean but not the other way. In one-tailed test, to do a test with α =
0.05, then we use critical values of the test statistic at α = 0.05 at one end of the sampling distribution.
Thus, in one tailed test, the area corresponding to the level of significance is located at one end of the
sampling distribution while in a two tailed test the region of rejection is divided into two equal parts, one at
each end of the distribution.
Most statistical tables either provide critical values for both one- and two-tailed tests but some just have
either one- or two-tailed critical values depending on the statistic, so make sure that you look-up the correct
α values when you use tables. Statistical software usually produces two-tailed α values so you should
compare the α value to 0.10 for a one-tailed test at 0.05.
Limit
Limit
Procedure:
i. State the null and alternate hypothesis
ii. Determine the significance level (). The level of significance establishes a criterion for rejection or
acceptance of the null hypothesis (10%, 5%, 1%)
iii. Compute the test statistic: The test statistic is the value used to determine whether the null
hypothesis should be rejected or accepted. The test statistic depends on sample size or whether the
population standard deviation is known or not:
a. when is known n could be large or small, the test statistic
Zc = (normal distribution)
11
Zc (test statistic) = (normal distribution)
iv. Determine the critical regions (value). The critical value is the point of demarcation between the
acceptance or rejection regions. To get the critical value, read Z or t-table depending on the test statistic
used.
v. Compare calculated value (test statistic) with table value of Z or t and make decision.
If the absolute value of calculated value (test statistic) is greater than table Z or t-value, reject the null
hypothesis and accept the alternate hypothesis.
If /Zc/ or /tc/ > Z or t table value at the specified level of significance, reject the H O (null hypothesis)
If /Zc/ or/tc/ Z or t-table value at the specified level of significance, accept H O
Example 1: A standard examination has been given for several years with mean (µ) score of 80 and
variance (σ2) of 49. A group of 25 students were taught with special emphasis on reading skills. If the 25
students obtained a mean grade of 83 on the examination, is there reason to believe that the special emphasis
changed the result on the test at 5% level of significance assuming that the grades are normally distributed?
µ = 80; σ2 = 49; = 83; n =25; α = 0.05.
Zc = = = 2.14
Example 2: It is known that under good management, Zebu cows give an average milk yield of 6 liters/day.
A cross was made between Zebu and Holstein and the result from a sample of 49 cross breeds gave average
daily milk yield of 6.6 liters with standard deviation of 2 liters. Do the cross breeds perform better than the
pure Zebu cows at 5% level of significance?
Given: =6 liters; n = 49; =6.2; s = 2
Zc = = = 2.1
Procedure:
1. State the null and alternate hypothesis
2. Compute the test statistic
Case I: When 12 & 22 (population variances) are known (n1 and n2) could be large or small
Zc = ( - )/
Case II: When 12 & 22 are unknown, n1 & n2 are large (30)
Zc = ( - )/
Case III: When 12 & 22 unknown, n1 & n2 small ( 30);
tc = ( - )/sp where sp (pooled standard
deviation) =
Independent t-test
In this case, the allocation of treatments on experimental units is done completely at random, e.g. varieties A
(filled) & B (open) are allocated to 12 fields, each on 6 fields randomly.
13
Example: A researcher wanted to compare the potentials of 2 different types of fertilizer. He conducted an
experiment at 10 different locations and the two fertilizer types were randomly applied in five of the fields
each on maize. The following yields were obtained in tons/ha.
Is there any difference in potential yield effect of the fertilizers at 5% level of significance?
Solution
1. Calculate the means and standard deviations: = 10.24; s1 = 1.32; = 9.76; s2 = 1.33; both n1 & n2 = 5
= sp
sp = = 1.32
= 1.32 = 0.58
4. Read t- table value for t(/2) at 5 + 5 – 2 d.f., t0.025 (8) = 2.306.
5. Make decision. Since t-calculated (0.58) < t-table (2.306), we accept the null hypothesis, i.e. the data do
not give sufficient evidence to indicate the difference in potential yield of two fertilizers. Thus, the
difference between the two fertilizer types is not statistically significant.
If five plots were divided in to two half; one fertilized (Nitrogen - N) and other was control (H)
14
Example: An experiment was done to compare the yields (qt/ha) of two varieties of maize, data were
collected from seven farms. At each farm, variety A was planted on one plot and variety B on the
neighboring plot.
______________________________________________________________
Farm: 1 2 3 4 5 6 7
______________________________________________________________
Variety A: 82 68 109 95 112 76 81
Variety B: 88 66 121 106 116 79 89
______________________________________________________________
Test whether there is a significant difference between two varieties at 5% level of significance.
Solution
1. Calculate the differences, mean of the differences, and variance of the differences as shown below
______________________________________________________________
Farm: 1 2 3 4 5 6 7
______________________________________________________________
Variety A: 82 68 109 95 112 76 81
Variety B: 88 66 121 106 116 79 89
Deterrence (d) -6 2 -12 -11 -4 -3 -8 -42
(d- ) 2 0 64 36 25 4 9 4 142
______________________________________________________________
Note that in paired data some farms are poor (2, 6), thus low yields for both varieties are obtained while
others are good (3, 5) and thus high yield for both varieties.
15
tc = = = = -3.26
5. Make decision: Since /tc/ (3.26) > t-table (2.447), there is a significant difference between the varieties of
maize at 5% level of significance.
3.1 Introduction
In research, a scientist identifies solution to problems through experimentation. Research can be broadly
defined as systematic investigation in to a subject to discover new facts or principles or to confirm or
deny the results of previous finding. Such investigation will help in decision making such as
recommending a new procedure, a new fertilizer rate, a new pesticide, etc.
The procedure for research is generally known as the scientific method, which, although difficult to define
precisely, it usually involves, the following steps:
a. Formulation of hypothesis: a tentative explanation or solution
b. Planning an experiment to objectively test the hypothesis
c. Careful observation and collection of the data
d. Analysis and interpretation of the experimental results.
Remark: Not all researches involve experimentation! Example: observational researches, surveys, etc. in
which no variable is manipulated by the researcher. These are not experiments!!
Experiment is a collection of research designs which use manipulation and controlled testing to understand
causal processes. Generally, one or more variables are manipulated to determine their effect on a dependent
variable.
The term refers to five interrelated activities required in the investigation. These are:
a. Formulating statistical hypothesis and making plans for laying out, collection and analysis of data
b. Stating the decision rules to be followed in testing statistical hypothesis (e.g., F-test)
c. Collecting data according to plan
d. Analyzing data according to plan
e. Making decisions based on decision rules
Treatment: It is an amount of material or a method that is to be tested in the experiment such as crop
varieties, insecticides, feedstuffs, fertilizer rates, method of land preparation, irrigation frequency, etc.
Experimental unit: It is an object on which the treatment is applied to observe an effect, e.g. cows, plot of
land, petri-dishes, pots, etc
In the study of the effect of different rations on milk production, ration is a treatment and animal is the
experimental unit; while in the study of different fertilizer rates on yield of maize, the fertilizer rates are
treatment and plot of land is experimental unit.
Experimental error: It is a measure of the variation, which exists among observations on experimental
units treated alike. Variation generally comes from two main sources:
17
1. Inherent variability that exists in the experimental material to which treatments are applied.
2. Lack of uniformity in the physical conduct of an experiment or failure to standardize the
experimental techniques such as lack of accuracy in measurement, recording data on different days,
etc.
Therefore, every possible effort should be made to reduce the experimental error.
Replication: A situation where a treatment appears more than once in an experiment, it is said to be
replicated. The functions of replication are:
a) It provides an estimate of experimental error because it provides several observations on experimental
units receiving the same treatment. For an experiment on which each treatment appears only once, no
estimate of experimental error is possible and when there is no method of estimating the experimental
error, there is no way to determine whether observed differences indicate the real differences or due to
inherent variability.
b) It improves the precision of an experiment: As the number of replicates increase, the estimates of
population means as observed treatment means become closer to the true value.
c) It increases the scope of inference and conclusion of the experiments: Field experiments are normally
repeated over years and locations because conditions vary from year to year and location to location.
The purpose of replication in space and time is to increase the scope of inference. The results of an
experiment are applicable only to conditions that are similar to that condition.
Usually three replications are taken as the minimum number for standard experiments
Two procedures for calculating the number of replications are described below:
Procedure 1: This method takes into consideration the variability of experimental material and field, which
is measured in terms of coefficient of variation (CV) and standard error of means (SEM). To calculate the
number of replications we can use the following formula:
N = (CV/SEM)2
Procedure 2: If CV & SEM are not known, then the number of replications can be arrived at using the
principle that the precision of treatment comparisons increases if the experimental error is kept to minimum.
The experimental error can be kept to the minimum by providing more degrees of freedom for the
experimental error. In other words, a lower number of degrees of freedom for experimental error results in
enlarged experimental error. Based on this principle, the number of degrees of freedom for error should not
be less than 15 (not less than 10 in any case).
When ‘t’ treatments are replicated ‘r’ times, the error is based on (t-1) (r-1) degrees of freedom in
Randomized Complete Block Design (RCBD) and t(r-1) in Completely Randomized Design (CRD), which
should not be less than 15.
No. of treatments: 2 3 4 5 6
Minimum no. of replications (CRD): 9 6 5 4 4
Minimum no. of replications (RCBD): 16 9 6 5 4
Increasing either number of replications or plot size can improve precision, but the improvement achieved
by doubling plot size is almost always less than the improvement achieved by doubling replications.
Randomization
Assigning the treatments to the experimental units in such away that any unit has equal chance to receive
any treatment, i.e. every treatment should have an equal chance of being assigned to any experimental units.
Thus, a particular treatment should not be consistently favored or disfavored.
Purposes of randomization
a) To eliminate bias: randomization ensures that no treatment is favored or discriminated against the
systematic assignment to units in a design
b) To ensure independence among the observations. This is necessary to provide valid significance tests.
Randomization is usually done by using tables of random numbers or by drawing cards or lots.
19
Confounding
It occurs when the differences due to experimental treatments, i.e. the contrast specified in your hypothesis,
cannot be separated from other factors that might be causing the observed differences. Example, if you
wished to test the effect of a particular hormone on some behavioral response of sheep. You create two
groups of sheep, males and females, and inject the hormone into the male sheep and leave the females as the
control group. Even if other aspects of the design are ok, differences between the means of the two groups
cannot be definitely attributed to effects of the hormone alone. The two groups are also different in sex and
this may also be, at least partly, determining the behavioral responses of the sheep. In this example, the
effects of hormone are confounded with the effects of sex.
Controls
A control is a part of the experiment that is not affected by the factor or factors studied, but otherwise
encounter exactly the same circumstances as the experimental units treated with the investigated factor(s).
For example, when investigating the effect of spraying micronutrients, the crop being in the control should
also be sprayed with the same amount of water except the micro nutrients. The reason for this way to work
is to make sure that it is only the effect of the substance of interest that is investigated. Otherwise, it may
not be possible to draw conclusions of the reason(s) to the outcome of the experiment.
Local Control
The principle of local control is another important principle of experimental designs. In local control, the
extraneous factor, i.e., the known source of variability, is made to vary deliberately over as wide a range as
necessary; and this needs to be done in such a way that the variability it causes can be measured and hence
eliminated from the experimental error. In other words, according to the principle of local control, we first
divide the field into several homogeneous parts, known as blocks, and then each such block is divided into
parts (units) equal to the number of treatments. Then the treatments are randomly assigned to these parts of
a block. Dividing the field into several homogenous parts is known as ‘blocking’. In general, blocks are the
levels at which we hold an extraneous factor fixed, so that we can measure its contribution to the total
variability of the data by means of analysis of variance. In brief, through the principle of local control we
can eliminate the variability due to extraneous factor(s) from the experimental error.
Remark: Blocking direction should be against the variability gradient!
Analysis of variance was first introduced by R.A Fisher in 1930s. Analysis of variance is used in all fields
of research where data are quantitatively measured and it is used: to estimate and test about population
variance and to estimate and test about population means.
20
It is defined as an arithmetic technique whereby the total variation in a set of data is divided (or partitioned)
into meaningful components or parts to have a better idea how a particular object (person, plant, animal,
etc.) reacts to a change in the conditions applied to it and how the environment plays a role in the expression
of this reaction.
The test of significance deals with computing the ratio between the explained and unexplained variances
and comparing with a probability value. If the ratio between explained and unexplained variances is
relatively high, we are confident to conclude that the difference was due to the known factors and not due to
unknown factors.
The purpose of proper experimental design is to make the experimental error (unexplained variation) small
enough in relation to the treatment (explained) effect.
The test of significance (F test) involves comparing each of the explained variances with the corresponding
unexplained variance by way of a ratio (F computed or Fc to the tabular F value). The higher the Fc relative
to the tabular F, the more significant are the treatment (or factor combination) differences. The Fc may
become large if the treatment variance is considerably large compared to the experimental error; or the
experimental error is made small in relation to the mean by proper experimental design and experimental
management.
21
The effect of treatment for all replications and the effect of replication for all treatments should remain
constant.
A hypothetical set of data with additive and multiplicative effect of treatments and replications.
Additive effect
Note that the effect of treatments is constant over replications and the effect of replications is constant over
treatments.
Multiplicative effect
Treatment I II (II-I) I II
A 10 20 10 1.00 1.30 0.30
B 30 60 30 1.48 1.78 0.30
Treatment effect (B- 20 40 0.48 0.48
A)
Here, the treatment effect is not constant over replications and the effect of replications is not constant over
the treatments.
The multiplicative effects are often encountered in experiments designed to evaluate the incidence of
diseases and insects. This happens because the changes in insect and disease incidence usually follow a
pattern that is in multiple of the initial incidence. When effects are multiplicative, the logarithmic
transformations of the data show the additive effect. Thus, in such cases conduct the analysis and mean
separation using the transformed data, and in tables present transformed means in parenthesis alongside
their back transformed values out of parenthesis.
22
III. Experimental errors (variances) must be homogeneous (homoscedasticity)
The variances of the treatments must be the same. When some treatments have errors that are exceptionally
higher or lower it is called heteroscedasticity. First type of heterogeneity of variances is usually associated
with data whose distribution is not normal. Count data, such as the number of infected plants per plot
usually follow a poisson distribution where the variance equals to the mean ( ).
Heterogeneity of variances occurs usually when some treatments have errors that are exceptionally higher or
lower than others.
Data such as number of infested plants per plot usually follow poisson distribution and data such as percent
survival of insects or percent plants infected with a disease assume the binomial distribution.
V. Randomness
Treatments should be assigned to experimental units completely at random. Randomness is known as the
fundamental assumption of ANOVA; because, its violation also leads to violation of other assumptions.
In CRD, the treatments are assigned completely at random over the whole experimental area so that each
experimental unit has the same chance of receiving any one treatment. In CRD, any difference among the
experimental units (plots) receiving the same treatment is considered as experimental error.
Uses:
1. It is useful when the experimental units (plots) are essentially homogeneous and where
environmental effects are relatively easy to control, e.g. laboratory and greenhouse experiments. For
field experiments where there is generally larger variation among experimental plots like in soil
fertility, slope, etc. the CRD is rarely used.
2. It is useful if we suspect that large fraction of the units may not respond or may be lost during the
experiment because it is easy to handle missing data in Analysis of Variance unlike in other
designs.-unequal replication is possible
3. It is useful for experiments in which the total number of experimental units is limited, because it
provides maximum degrees of freedom for error
Advantages
1. It is flexible in that the number of treatments and replications can vary, i.e. the number of
replications need not be the same from one treatment to another
2. The statistical analysis is simple even with unequal replications and it is not complicated by loss of
data or missing observations
3. Loss of information due to missing data is small as compared to other designs
23
4. The design provides the maximum degree of freedom for estimating the experimental error. This
improves the precision of the experiment and is important with small experiments where degrees of
freedom for experimental error are less than 20.
Disadvantage:
The main objection to the CRD is that it is often inefficient? as there is no way of controlling the
experimental error. Since randomization is unrestricted the experimental error includes the entire variation
over the experimental units except that due to treatment.
In this design, treatments are assigned to the experimental units completely at random.
Assume that we want to do a pot-experiment on the effect of inoculation of 6-strains of rhizobia on
nodulation of common bean using five replications.
Randomization can be done by using either lottery method or table of random numbers.
A. Lottery Method
1. Arrange 30 pots of equal size filled with the same type of soil and assign numbers from 1 to 30 in
convenient order.
2. Obtain 30 identical slips of paper, label 5 of them with treatment A, 5 of them with treatment B, with C,
with D, with E and with F (6 treatments 5 replications). Place the slips in box or hat, mix thoroughly
and pick a piece of paper at random, the treatment labeled on this paper is assigned to unit 1 (pot 1),
without returning the first slip to box, select another slip and the treatment named on this slip is assigned
to unit 2(pot 2) and continue this way until all 30 slips of paper have been drawn.
24
Raw data of nitrogen content of common bean inoculated with 6 rhizobium strains. Treatments are
designated with letters and nitrogen content (mg) in parenthesis.
1A 2E 3C 4B 5A 6D
(19.4) (14.3) (17.0) (17.7) (32.6) (20.7)
12 B 11C 10A 9E 8D 7E
(24.8) (19.4) (27) (11.8) (21.0) (14.4)
13 F 14A 15F 16D 17B 18.C
(17.3) (32.1) (19.4) (20.5) (27.9) (9.1)
24D 23F 22B 21E 20C 19D
(18.6) (19.1) (25.2) (11.6) (11.9) (18.8)
25C 26E 27B 28F 29A 30F
(15.8) (14.2) (24.3) (16.9) (33.0) (20.8)
where, Yij = the jth observation on the ith treatment; General mean; i = Effect of treatment i; and
Experimental error (effect due to chance)
2. Arrange the data by treatments and calculate the treatment totals (T i) and Grand total (G)
3. Using Yij = jth observation on the ith treatment; Ti= Treatment total; n = (r x t), the total number of
experimental unit (pots), calculate the correction factor and the various sum of squares
C.F. = = = 11864.38
Total Sum of Squares (TSS) = yij2 – C.F. (Sum of the square of all observations- C.F.)
= [(19.4)2 + (17.7)2 +…… + (20.8)2] - 11864.38 = 12994.36–11864.38=1129.98
= - 11864.38=847.05
Error Sum of Squares (SSE) = Total SS- Treatment SS= 1129.98-847.05= 282.93
In CRD, treatment sum of squares are usually called between or among groups sum of squares while the
sum of squares among individuals treated alike is called within group or error sum of squares.
4. Calculate the mean squares (MS) for treatment and error by dividing each sum of squares by the
corresponding degree of freedom
Treatment MS = = = 169.41
Error MS = = = 11.79
5. Calculate F-value for testing significance of treatment effects
F-calculated = = = 14.37
6. Obtain the tabulated F-value using treatment degree of freedom (d. f.) as numerator (n 1) and error d. f. as
denominator (n2) at 5% and 1% level of significance
F (5, 24) at 5% = 2.60; F (5, 24) at 1% = 3.90
7. Summarize all the values computed on the above steps in the ANOVA- table for quick assessment of
results.
________________________________________________________________________
Source of DF SS MS Computed F Table F
Variation 5% 1%
________________________________________________________________________
Treatment (among strains) (t-1) = 5 847.05 169.41 14.37** 2.60 3.90
Error (within strains) t(r-1) = 24 282.93 11.79
Total (rt-1) = 29 1129.98
________________________________________________________________________
26
8. Compare the calculated F- value with table F- value and decide on significance among the treatment
effects using the following rules:
a) If F-calculated > F table at 1% level of significance, the difference between treatments is highly
significant. Put two asterisks on F-calculated
b) If F-calculated > F table at 5% level of significance, but ≤ F table at 1%, the difference between
treatments is significant. Put one asterisks on F-calculated.
c) If F-calculated ≤ F table at 5% level of significance, the differences among treatments is non-
significant. Put NS on the F- calculated value in ANOVA table.
Note that a non-significant F test in the analysis of variance indicates the failure of the experiment to detect
any difference among treatments. It does not, in any way, prove that all treatments are the same. The failure
to detect treatment difference based on non- significant F-test could be the result of either a very small or nil
treatment difference or a very large experimental error or both. Thus, whenever the F-test is non-significant,
the researcher should examine the size of experimental error and the numerical difference among the
treatment means. If both values are large, the trial may be repeated and efforts should be made to reduce
experimental error so that the differences among treatments, if any can be detected. On the other hand, if
both values are small, the difference among treatments is probably too small to be of any economic value
and, thus, no additional trials are needed.
For the above example, the computed F value of 14.37 is larger than the tabulated F value at the 1% level of
significance of 3.90. Hence, the treatment difference is said to be highly significant. In other words, chances
are less than 1 in 100 that all the observed differences among the six treatment means could be due to
chance. It should also be noted that such a significant F test verifies the existence of some differences
among the treatments tested but does not specify the particular pair (or pairs) of treatments that differ
significantly. To obtain this information, procedures for comparing treatment means are used.
9. Compute the Coefficient of Variation (CV) and standard error (SE) of the treatment means
- CV= 100 = 100= 17.3%
- SE = = = 1.53mg
Coefficient of Variation indicates the degree of precision with which the treatments are compared and it is a
good index of the reliability of the experiment. The smaller the CV, the more reliable the experiment is. The
CV values greatly vary with the type of experiment, experimental material or the character measured (e.g.
data on days to flowering have smaller CV than number of nodules in common bean as within treatment
variation is usually small in the former parameter than the later). In field experiments CV up to 30% are
common and usually lesser CV for laboratory and greenhouse experiments are expected than field
experiments.
Example: Twenty rats (n=20) were assigned equally at random to four feed types (t= 4). Unfortunately one
of the rats died due to unknown reason. The data are rat body weight in g after being raised on these diets
for 10 days. We would like to know whether weights of rats are the same for all four diets at 5%.
27
________________________________________________________________________
Feed 1 Feed 2 Feed 3 Feed 4
________________________________________________________________________
60.8 68.7 102.6 87.9
57.0 67.7 102.1 84.2
65.0 74.0 100.2 83.1
58.6 66.3 96.5 85.7
61.7 69.8 90.3
________________________________________________________________________
Ti 303.1 346.5 401.4 431.2
ni 5 5 4 5
- C.F. = = = 115627.20
- Total Sum of Squares (TSS) = yij2 – C.F. (Sum of the square of all observations- C.F.)
= [(60.8)2 + (57.00)2 +…… + (90.3)2] – 115627.20 = 4354.698
= 4226.348
- Error Sum of Squares (SSE) = Total Sum of Squares – Treatment Sum of Squares: 4354.698 -
4226.348 = 128.35
________________________________________________________________________
Source of DF SS MS Computed F F-table (1%)
Variation
________________________________________________________________________
Treatment (among groups) (t-1) = 3 4226.348 1408.788 164.6** 5.42
Error (within groups)
(total d.f. – treatment d.f.) (18-3) = 15 128.35 8.557
Total (n-1) = 18 4354.698
________________________________________________________________________
** Highly significant
CV= = = 3.75%
It is the most frequently used experimental design in field experiments. Completely Randomized Design is
appropriate when no sources of variation other than treatment effects are known or anticipated, i.e. the
experimental units should be similar. However, in many experiments, certain experimental units that are
treated alike will behave differently. Example: in field experiments adjacent plots are more alike in
28
response than those distant apart. Heaviest animals in a group of the same age may show a different rate of
weight gain than lighter animals.
In such situations, designs and layouts can be constructed so that the portion of variability attributable to the
known source can be measured and excluded from experimental error. Thus, difference among treatment
means will contain no contribution of the known source.
Randomized Complete Block Design (RCBD) can be used when the experimental units can be meaningfully
grouped, the number of units in a group being equal to the number of treatments. Such a group is called a
block and equals to the number of replications. Each treatment appears an equal number of times usually
once, in each block and each block contains all the treatments.
Blocking (grouping) can be done based on soil heterogeneity in a fertilizer or variety trials; initial
body weight, age, sex, and breed of animals; slope of the field, etc.
The primary purpose of blocking is to reduce experimental error by eliminating the contribution of known
sources of variation among experimental units. By blocking, variability within each block is minimized and
variability among blocks is maximized.
During the course of the experiment, all units in a block must be treated as uniformly as possible. For
example,
- if planting, weeding, fertilizer application, harvesting, data recording, etc, operations cannot be done in
one day due to some problems, then all plots in any one block should be done at the same time.
- if different individuals have to make observations of the experimental plots, then one individual should
make all the observations in a block.
This practice helps to control variation within blocks, and thus variation among blocks is mathematically
removed from experimental error.
Advantages:
1. Precision: More precision is obtained than with CRD because grouping experimental units into
blocks reduces the magnitude of experimental error.
2. Flexibility: Theoretically, there is no restriction on the number of treatments? or replications. If
extra replication is desired for certain treatments, it can be applied to two or more units per block.
3. Ease of analysis: The statistical analysis of the data is simple. If as a result of change or missing, the
data from a complete block or for certain treatments are unusable, the data may be omitted without
complicating the analysis. If data from individual units (plots) are missing, they can be estimated
easily so that simplicity of calculation is not lost.
Disadvantages:
The main disadvantage of RCBD is that when the number of treatments is large (>15), variation among
experimental units within a block becomes large, resulting in a large error term. In such situations, other
designs such as incomplete block designs should be used.
29
5.2 Randomization and layout
Step 1: Divide the experimental area (unit) into r-equal blocks, where r is the number of replications,
following the blocking technique. Blocking should be done against the gradient such as slope, soil fertility,
etc.
Step 2: Sub-divide the first block into t-equal experimental plots, where t is the number of treatments and
assign t treatments at random to t-plots using any of the randomization scheme (random numbers or lottery).
ENVIRONMENTAL GRADIENT
The major difference between CRD and RCBD is that in CRD, randomization is done without any
restriction to all experimental units but in RCBD, all treatments must appear in each block and different
randomization is done for each block (randomization is done within blocks).
where, Yij = the observation on the j th block and the ith treatment; = common mean effect; i = effect of
treatment i; j = effect of block j; and ij = experiment error for treatments i in block j.
Step 2. Arrange the data by treatments and blocks and calculate treatment totals (T i), Block (rep) totals (Bj)
and Grand total (G).
Example: Oil content of linseed treated at different stages of growth with N-fertilizes.
30
Oil content (g) from sample of 20 g seed
Step 3: Compute the correction factor (C.F.) and sum of squares using r as number of blocks, t as number of
treatments, Ti as total of treatment i, and Bj as total of block j.
a. C.F. =
b. Total Sum of Square (TSS) = Yij2- C.F. = (4.4)2 + (3.3)2 + …. + (6.7)2 – C.F.
= 788.23 – 733.72 = 54.51
=
= 736.86 – 733.72 = 3.14
=
= 765.37 – 733.72 = 31.65
Step 4: Compute the mean squares for block, treatment and error by dividing each sum of squares by its
corresponding d.f.
31
Block Mean Square (MSB) =
Step 5: Compute the F-value for testing block and treatment differences.
F-block = for block; F-treatment = for treatments.
F-block = ; F-treatment =
Step 6: Read table F- and compare the computed F-value with tabulated F-value and make decision.
- F-table for comparing block effects, use block d. f. as numerator (n 1) and error d.f. as denominator (n2);
F (3, 15) at 5% = 3.29 and at 1% = 5.42.
- F-table for comparing treatment effects, use treatment d.f. as numerator and error d.f. as denominator F
(5, 15) at 5% = 2.90 and at 1% = 4.56. If the calculated F-value is greater than the tabulated F-value for
treatments at 1%, it means that there is a highly significant (real) difference among treatment means.
In the above example, calculated F-value for treatments (4.83) is greater than the tabulated F-value at 1%
level of significance (4.56). Thus, there is a highly significant difference among the stages of application on
nitrogen content of the linseed.
Step 7: Compute the standard error of the mean and coefficient of variability.
- Standard error (SE)
Step 8: Summarize the results of computations in analysis of variance table for quick assessment of the
result.
32
5.4. Block efficiency
Blocking maximizes the difference among blocks and reduces the difference among plots of the same block
(within blocks) as small as possible. Thus, the result of every RCBD should be examined to see whether this
objective has been achieved. The procedure to measure block efficiency is:
Step 1: Determine the level of significance of block variation by computing F-value for block and test its
significance.
F-block =
By comparing it with tabulated F-value at n1 (r-1) d.f. and n2 error d.f. (r-1) (t-1) = F (3, 15) at 5% = 3.29.
If the computed F-value is greater than the tabulated F-value, blocking is said to be effective in reducing
experimental error. Also the scope of an experiment may have been increased when blocks are significantly
different since the treatments have been tested over a wider range of experimental conditions.
On the other hand, if block effects are small (calculated F for block < tabulated F value), it indicates either
that the experimenter was not successful in reducing error variance by grouping of individual units
(blocking) or that the units were essentially homogeneous to start with.
Step 2: Determine the magnitude of the reduction in experimental error due to blocking by computing the
Relative Efficiency (R.E.) as compared to CRD.
R.E. =
Where MSB = block mean square; MSE = error mean square; r = number of replications; t = number of
treatments
R.E =
If error degree of freedom of RCBD is less than 20, the R.E. should be multiplied by the adjustment factor
(k) to consider the loss in precision resulting from fewer degrees of freedom.
33
Adjusted R.E. = 0.97 K(0.98) = 0.95.
In this case, information is sacrificed in theory by using Randomized Complete Block Design, since 95
replicates in a completely randomized design give as much information as 100 blocks or replications of a
Randomized Complete Block Design. That means, RCBD was less efficient than CRD.
Sometimes data for certain units may be missing or become unusable. For example,
- when an animal becomes sick or dies but not due to treatment
- when rodents destroy a plot in field
- when a flask breaks in laboratory
- when there is an obvious recording error
A method is available for estimating such data. Note that an estimate of a missing value does not supply
additional information to the experimenter; it only facilitates the analysis of the remaining data.
where;
The estimated value is entered in the table with the observed values and the analysis of variance is
performed as usual with one d. f. being subtracted from both total and error d.f because the estimated value
makes no contribution to the error sum of squares.
When all of the missing values are on the same block or treatment the simplest solution is to consider
as if the block or treatment had not been included in the experiment.
Example: In the table given below are yields (kg) of 4-varieties of maize (Al-composite, Rarree-1, Bukuri,
Katumani) in 4-replications planted in RCBD on a plot size of 10 m x 10 m of which one plot yield is
missing. Estimate the missing value and analyze the data
Treatment
Varieties I II III IV total (Ti)
Bukuri 18.5 15.7 16.2 14.1 64.5
Katumani 11.7 - 12.9 14.4 39(To)
Rarree-1 15.4 16.6 15.5 20.3 67.8
Al-composite 16.5 18.6 12.7 15.7 63.5
34
Block total 62.1 50.9 (Bo) 57.3 64.5
Go (Grand total) = 234.8
Solution
a. Estimate the missing value
b. Enter the estimated value and carry out the analysis following the usual procedure:
- Corrected treatment total = 39 + 13.9 = 52.9
- Corrected block total = 50.9 + 13.9 = 64.8
- Corrected grand total = 234.8 + 13.9 = 248.7
c. Analysis of variance
1. C.F. =
2. TSS=
3. Treatment SS =
4. Block SS =
5. Error SS= Total SS- Treatment SS- Block SS = 79.18 – 31.21 – 9.02 = 38.95
d. Compute the correction factor for bias (B) for treatment sum of squares as the treatment SS is biased
upwards.
B=
Y = estimated value
B=
=
e. Subtract the computed B value from Total SS & Treatment SS
- Adjusted Treatment SS = Treatment SS – B
= 31.21 – 7.05= 24.16
35
- Adjusted Total SS = Total SS-B
= 79.18 – 7.05 = 72.13
f. Subtract 1 from error d. f. and total d. f. and complete the analysis of variance table.
CV=
The major feature of the Latin Square Design is its capacity to simultaneously handle two known sources of
variation among experimental units unlike Randomized Complete Block Design (RCBD), which treats only
one known source of variation.
The two directional blocking in a Latin Square Design is commonly referred as row blocking and column
blocking. In Latin Square Design the number of treatments is equal to the number of replications that is why
it is called Latin Square.
Advantages:
- Greater precision is obtained than Completely Randomized Design & Randomized Complete Block
Design (RCBD) because it is possible to estimate variation among row blocks as well as among column
blocks and remove them from the experimental error.
Disadvantages:
36
- As the number of treatments is equal to the number of replications, when the number of treatments is
large the design becomes impractical to handle. On the other hand, when the number of treatments is
small, the degree of freedom associated with the experimental error becomes too small for the error to be
reliably estimated. Thus, in practice the Latin Square Design is applicable for experiments in which the
number of treatments is not less than four and not more than eight.
- Randomization is relatively difficult.
Step 1: To randomize a five treatment Latin Square Design, select a sample of 5 x 5 Latin square plan from
appendix of statistical books. We can also create our own basic plan and the only requirement is that each
treatment must appear only once in each row and column. For our example, the basic plan can be:
A B C D E
B A E C D
C D A E B
D E B A C
E C D B A
Step 2: Randomize the row arrangement of the plan selected in step 1, following one of
the randomization schemes (either using lottery method or table of random numbers).
- Select from table of random numbers, five three digit random numbers avoiding ties if any
Random numbers: 628 846 475 902 452
Rank: (3) (4) (2) (5) (1)
- Rank the selected random numbers from the lowest (1) to the highest (5)
- Use the ranks to represent the existing row number of the selected plan and the sequence to represent the
row number of the new plan. For our example, the third row of the selected plan (rank 3) becomes the
first row (sequence) of the new plan, the fourth becomes the second row, etc.
1 2 3 4 5
3 C D A E B
4 D E B A C
2 B A E C D
5 E C D B A
1 A B C D E
Step 3: Randomize the column arrangement using the same procedure. Select five three digit random
numbers.
Random numbers: 792 032 947 293 196
Rank: (4) (1) (5) (3) (2)
The rank will be used to represent the column number of the above plan (row arranged) in step 2. For our
example, the fourth column of the plan obtained in step 2 above becomes the first column of the final plan,
the first column of the plan becomes 2, etc.
Final layout:
37
E C B A D
A D C B E
C B D E A
B E A D C
D A E C B
Note that each treatment occurs only once in each row and column
Sample layout of three treatments each replicated three times in Latin Square Design
There are four sources of variation in Latin Square Design, two more than that of CRD and one more than
that for the RCBD. The sources of variation are row, column, treatment and experimental error.
where,
Yijk = the observation on the ith treatment, jth row & kth column
= Common mean effect
i = Effect of treatment i
j = Effect of row j
k = Effect of column k
ijk = Experiment error (residual) effect
Example: Grain yield of three maize hybrids (A, B, and D) and a check variety, C, from an experiment with
Latin Square Design.
STEPS OF ANALYSIS
Step 1: Arrange the raw data according to their row and column designation, with the corresponding
treatments clearly specified for each observation and compute row total (R), column total (C), the grand
total (G) and the treatment totals (T).
Row SS = =
Step 3: Compute the mean squares for each source of variation by dividing the sum of squares by its
corresponding degrees of freedom.
Row MS = ; Column MS =
Treat. MS = ; Error MS =
Step 4: Compute the F-value for testing the treatment effect and read table F-value as:
As the computed F-value (7.15) is higher than the tabulated F-value at 5% level of significance (4.76), but
lower than the tabulated F-value at the 1% level (9.78), the treatment difference is significant at the 5% level
of significance.
39
Compute the CV as: = =
Note that although the F-test on the analysis of variance indicates significant differences among the mean
yields of the 4-maize varieties tested, it does not identify the specific pairs or groups of varieties that
differed significantly. For example, the F-test is not able to answer the question whether every one of the
three hybrids gave significantly higher yield than that of the check variety. To answer these questions, the
procedure for mean comparison should be used.
As in RCBD, where the efficiency of one way blocking indicates the gain in precision relative to CRD, the
efficiencies of both row and column blocking in a Latin Square Design indicate the gain in precision relative
to either the CRD or RCBD, the procedures are:
i. Compute the F-value for testing the row & column effects; and test their significance
R.E (CRD) =
= = = 3.45
This indicates that the use of Latin Square Design in the present example is estimated to increase the
experimental precision by 245% as compared to CRD. This result implies that if the CRD is used an
estimated 2.45 times more replication would have been required to detect the treatment difference of the
same magnitude as that detected with the Latin Square Design.
When the error d. f. in the Latin Square analysis of variance is < 20, the R.E. value should be multiplied by
the adjustment factor (K) defined as:
K= = = = = 0.93
The results indicate that the additional column blocking made possible by the use of Latin Square Design is
estimated to have increased the experimental precision over that of RCBD by 290%, whereas the additional
row-blocking in the LS design did not increase precision over the RCBD with column as blocks. Hence, for
the above trial, a RCBD with column as blocks would have been as efficient as a Latin Square Design.
The analysis of variance is performed in the usual manner after entering the estimated value with one degree
of freedom being subtracted from total and error degrees of freedom for each missing value.
As in the case of RCBD, the treatment sum of squares is biased upward by:
Bias (B) =
Where Go, Ro, Co, To and t are as described above. Then B is subtracted from treatment SS & total SS.
Example: Yield (kg) of five rice varieties tested in Latin Square Design from plot size of 100 m 2.
41
E(12) C (-) B(11) A (10) D (8)
A (7) D (8) C(8) B(7) E (13)
C (12) B (6) D(7) E (11) A (9)
B (4) E (10) A (7) D (6) C (8)
D (5) A (8) E (15) C(9) B (5)
Estimate the missing value, complete the analysis of variance and compare the variety C with D, and A with
E at 5% level of significance using LSD test.
Y=
b. Enter the estimated value and carry out the analysis following the usual procedure:
- Corrected row total = 41.0 + 11.5 = 52.5
- Corrected column total = 32.0 + 11.5 = 43.5
- Corrected treatment total = 37+ 11.5 = 48.5
- Corrected grand total = 206.0 + 11.5 = 217.50
c. Compute the C.F. and the various Sum of Squares
C.F. =
Row SS = =
Column SS = = = 6.60
Treatment SS = = = 107.6
Error SS = Total SS–Row SS–Column SS – Treatment SS = 180.0–31.6–6.6–107.6= 34.2
d. Compute the correction factor for bias (B) for treatment sum of squares as the treatment SS is biased
upwards.
Bias (B) = =
f. Subtract 1 from error d. f. and total d. f. and complete the analysis of variance table.
CV =
= 1.226
Difference between the treatment means of C & D (37/4-34/5) = 9.25-6.8 = 2.45. Since the difference is less
than LSD value, there is no significant difference between treatments C & D.
To compare the treatments A & E both with equal replication (without missing value)
LSD5% = t 0.025(11) , where = = 1.115
LSD5% = 2.201 1.115 = 2.45 kg
Difference between the treatment means of A & E (61/5-41/5) = 12.2-8.2 = 4.00. Since the difference is
greater than LSD value, there is significant difference between treatments A & E.
SAS Syntax
data Lattice;
do Column=1 to 4;
do Row=1 to 4;
input Treatment Yield @;
output;
end;
end;
datalines;
2 1.64 3 1.475 1 1.67 4 1.565
4 1.21 1 1.15 3 0.71 2 1.29
3 1.425 4 1.4 2 1.665 1 1.655
1 1.345 2 1.29 4 1.18 3 0.66
;
Proc ANOVA Data=Lattice;
class Column Row Treatment;
model Yield= Column Row Treatment;
43
Means Treatment/lsd;
run;
OR
Data Lattice;
Input Row column Treatment Yield;
cards;
1 1 2 1.64
1 2 4 1.21
1 3 3 1.425
1 4 1 1.345
2 1 3 1.475
2 2 1 1.15
2 3 4 1.4
2 4 2 1.29
3 1 1 1.67
3 2 3 0.71
3 3 2 1.665
3 4 4 1.18
4 1 4 1.565
4 2 2 1.29
4 3 1 1.655
4 4 3 0.66
;
Proc ANOVA;
Class Row Column Treatment;
Model Yield= Row Column Treatment;
Means Treatment/lsd;
run;
Sum of
Source DF Squares Mean Square F Value Pr > F
44
R-Square Coeff Var Root MSE Yield Mean
Theoretically, the complete block designs (where each block contains all the treatments) such as Randomized
Complete Block and Latin Square are applicable to experiments with any number of treatments. However,
these complete block designs become less efficient as the number of treatments increases, mainly because block
size increases proportionally with the number of treatments which in turn increases experimental error.
An alternative set of designs for single factor experiments having a large number of treatments are the
incomplete block designs. For example, plant breeders are often interested in making comparisons among a
large number of selections in a single trial. For such trials, we use incomplete block designs. As the name
implies, the experimental units in these designs are grouped into blocks which are smaller than a complete
replication of the treatments. However, the improved precision with the use of an incomplete block designs
(where the blocks do not contain all the treatments) is achieved with some costs. The major ones are:
- inflexible number of treatments or replications or both.
- unequal degree of precision in the comparison of treatment means.
- complex data analysis.
Although there is no concrete rule as to how large the number of treatments should be before the use of an
incomplete block design, the following points may be helpful:
a. Variability in the experimental material: The advantage of an incomplete block design over complete
block designs is enhanced by an increased variability in the experimental material. In general, whenever
the block size in Randomized Complete Block Design is too large to maintain reasonable level of
uniformity among experimental units within the same block, the use of an incomplete block design should
be seriously considered.
b. Computing facilities and services: Data analysis of an incomplete block design is more complex than that
for a complete block design. Thus, the use of an incomplete block design should be considered only as the
last measure.
45
7.1. Lattice Designs
The lattice designs are the most commonly used incomplete block designs in agricultural experiments. There is
sufficient flexibility in the design to make its applications simpler than most of the other incomplete block
designs. There are two kinds of lattices: balanced lattice and partially balanced lattice designs.
b. The block size (k) is equal to the square root of the number of treatments, i.e. k=
c. The number of replications (r) is one more than the block size, i.e. r = k + 1. That is, the number of
replications required is 6 for 25 treatments, 7 for 36 treatments, and so on.
As balanced lattices require large number of replications, they are not commonly used.
The partially balanced lattice design is more or less similar to the balanced lattice design, but it allows for a
more flexible choice of the number of replications. The partially balanced lattice design requires that the
number of treatments must be a perfect square and the block size is equal to the square root of the number of
treatments. However, any number of replications can be used in partially balanced lattice design. The partially
balanced lattice design with two replications is called simple lattice, with three replications is triple lattice and
with four replications is quadruple lattice, and so on. However, such flexibility in the number of replications
results in the loss of symmetry in the arrangement of the treatments over blocks (i.e. some treatment pairs
never appear together in the same incomplete block). Consequently, the treatment pairs that are tested in the
same incomplete block are compared with higher level of precision than for those that are not tested in the
same incomplete block. Thus, partially balanced designs are more difficult to analyze statistically, and several
different standard errors may be possible.
Example 1 (Lattice with adjustment factor): Field arrangement and broad leaved weed kill (%) in tef fields
of Debre zeit research center by 16-herbicides tested in 4 4 Triple Lattice Design (Herbicide numbers in
parenthesis).
________________________________________________________________________
Block
Replication Block % kill total (B) M Cb
_________________________________________________________________________
1 1 75(15) 57(16) 71(13) 77(14) 280 789 -51
2 78(12) 66(11) 68(10) 45(9) 257 716 -55
3 40(6) 64(5) 49(8) 42(7) 195 608 23
4 59(3) 53(1) 46(2) 54(4) 212 642 6
46
Rep total (R1) 944 -77
______________________________________________________________________
2 1 53(16) 66(4) 57(12) 47(8) 223 663 -6
2 80(14) 48(6) 73(10) 52(2) 253 700 -59
3 36(7) 63(11) 67(15) 47(3) 213 676 37
4 68(13) 60(1) 50(9) 76(5) 254 716 -46
Rep total (R2) 943 -74
_________________________________________________________________________
3 1 66(15) 46(2) 58(12) 69(5) 239 754 37
2 46(4) 40(7) 59(13) 55(10) 200 678 78
3 43(9) 55(3) 50(8) 68(14) 216 670 22
4 60(11) 58(1) 48(16) 47(6) 213 653 14
Re total (R3) 868 151
Analysis of variance
Step 1: Calculate the block total (B), the replication total (R) and Grand Total (G) as shown above.
Step 2: Calculate the treatment totals (T) by summing the values of each treatment from the three replications.
Step 3: Using r as number of replications and k as block size, compute the total sum of squares, replication
sum of squares, treatment (unadjusted) sum of squares as:
- Correction Factor (C. F.) = = = 158125.5
47
- Total Sum of Squares = (75)2 + (57)2 + .... + (47)2 – C.F.
= 164233- 158125.5 = 6107.5
- Replication Sum of Squares = - C. F.
Cb = M -rB where M is the sum of treatment totals for all treatments appearing in that particular block, B is
the block total and r is the number of replications. For example, block 2 of replication 2 contains treatments
14, 6, 10, and 2. Hence, the M value for block 2 of replication 2 is: M = T14 + T6 + T10 + T2 = 225 + 135 +
196 + 144 = 700 and the corresponding Cb value is: Cb = 700 - (3 x 253) = -59. The Cb values for the blocks
are presented in the above table. Note that the sum of Cb values over all replications should add to zero (i. e. -
77 + -74 + 151 = 0).
= -
= 888.58-356.31 = 532.27
Step 7: Calculate the intra-block error mean square (MS) and block (adj.) mean square (MS) as:
- Intra-block error MS = = = 20.67
Step 8: Calculate adjustment factor A. For a triple lattice design, the formula is:
48
A= = = 0.0813
where Eb is the block (adju.) mean square and Ee is the intra-block error mean square, and k is block size.
Step 9: For each treatment, calculate the adjusted treatment total (T') as:
T' = T + A where the summation runs over all blocks in which that particular treatment appears. For
example, the adjusted treatment totals for treatment number 1 and 2 are computed as:
T'1 = 171 + 0.0813(6 + -46 + 14) = 168.89
T'2 = 144 + 0.0813(6 + -59 + 37) = 142.70
.
.
etc.
The adjusted treatment totals (T') and their respective means (adjusted treatment total divided by the number
of replications (3) are presented along with the unadjusted treatment totals (T) in the table above.
- SSBun (unadjusted block sum of squares) = - C. F. - SSR; where B is Block total, and SSR is
Replication Sum of Squares.
= - 158125.5 - 237.6
= 160046.75-158125.5 - 237.6 = 1683.65
CV = = 100 = 7.9%
Note that the grand mean is same for adjusted and unadjusted treatment total.
Step 16: Compute the gain in precision of the triple lattice relative to that of Randomized Complete Block
Design as:
= /(1 + ) 20.67
Thus, the relative precision = (32.2/24.7) 100 = 130.4
This indicates that the precision of this experiment was increased by about 30.4% by using the triple lattice
instead of Randomized Complete Block Design.
50
7.3. Augmented Block Design
Any new material is not replicated; it appears only once in the experiment while check varieties/entries
occur as the number of the blocks.
Thus, the minimum number of checks is 2 because error d. f. should be 12 or more for valid comparison. If
number of checks is 1, error d.f. becomes zero which is not valid.
Suppose we have 50 test lines/progenies and 4 checks, we need at least 5 blocks since the error d. f. (c - 1)
(b - 1) should be 12. The higher the number of blocks, the higher the precision.
If we have 5 equal blocks, block size will be 14 (10 test lines + 4 checks), but block size may vary.
Randomization
Two possibilities
A. Conveniency
For identification purpose, sometimes the checks are assigned at the start, end or in certain intervals in the
block.
b1 = P1 P2 A P3 B P4 D P5 C P6 P7 ..... P10
b2 = D P1 P2 C P3 P5 A B ….....
etc.
51
One of the blocks can be 10 (test culture) + 4 (checks), while the other 12 (test cultures) + 4 (checks).
Missing test entries (Pi) do not create problem in analysis as the analysis can be done with existing
genotypes. But when checks are missing, the analysis becomes complicated.
Assume we want to test 16 new rice genotypes (test lines) = P1, P2, ..., P16 in block size of 4 with 4 checks to
screen for early maturity.
b (number of blocks) = 4
c (number of checks) = 4 (A B C D).
Block size = 4 test cultures + 4 checks = 8
P1 A P2 B P3 C P4 D
120 83 100 77 90 70 85 65
Block 2
P5 B P6 C P7 A P8 D
88 76 130 71 105 84 110 64
Block 3
Block 4
Check/ b1 b2 b3 b4 Check Check Check effect (check mean- Check total x Check
block total mean adjusted grand mean) effect
A 83 84 86 82 335 83.75 -16.67 -5584.45
B 77 76 78 75 306 76.50 -23.92 -7319.52
C 70 71 69 68 278 69.50 -30.92 -8595.76
D 65 64 63 63 255 63.75 -36.67 -9350.85
Sum 1174 -30850.58
Total of check means 293.5
52
Blocks ni Block Total of test No. of test cultures in Block T i x Be Block effect x
total culture in a a block (Ti) effect(Be) Block total
(Bj) block (Bti)
B1 8 690 395 4 0.375 1.5 258.75
B2 8 728 433 4 0.375 1.5 273.00
B3 8 811 515 4 0.625 2.5 506.87
B4 8 660 372 4 -1.375 -5.5 -907.5
32 2889 0 0 131.12
bi = (Total of the ith block - total of all check means - total of all progenies/test
cultures in the block];
bi = 0
1
b1 = (690 - 293.5 – 395) = 0.375
4
1
b2 = (728 - 293.5 - 433) = 0.375
4
1
b3 = (811 - 293.5 - 515) = 0.625
4
1
b4 = (660 - 293.5 - 372) = -1.375
4
Where ni is number of entries (test culture + checks) in each block = 4 + 4 = 8, ni = N = 32.
1
Adjusted grand mean = 4 16 [2889 - (4 - 1) (293.5) - 0]
= 100.42
Grand total = Bi (sum of block total) or sum of all observations = 2889.0
53
C3 = 69.50-100.42 = -30.92; C4= 63.75-100.42 =-36.67
There are as many check effects as the number of checks (4 in this case).
= Observed (unadjusted) progeny value - effect of block in which the ith progeny is occurring
P1 (adjusted) = P1 - (block effect) = 120 - (+ 0.375) = 119.62, etc.
Progeny/ Observed Block Adjusted progeny Progeny effect Observed
test progeny effect value (Po- bi) (Adjusted progeny progeny
culture no. value (Po) (bi) value-Adjusted value
(Pi) grand mean progeny
effect
1 120 0.375 119.625 19.205 2304.6
2 100 0.375 99.625 -0.795 -79.5
3 90 0.375 89.625 -10.795 -971.6
4 85 0.375 84.625 -15.795 -1343
5 88 0.375 87.625 -12.795 -1126
6 130 0.375 129.625 29.205 3796.7
7 105 0.375 104.625 4.205 441.53
8 110 0.375 109.625 9.205 1012.6
9 102 0.625 101.375 0.955 97.41
10 140 0.625 139.375 38.955 5453.7
11 135 0.625 134.375 33.955 4583.9
12 138 0.625 137.375 36.955 5099.8
13 84 -1.375 85.375 -15.045 -1264
14 90 -1.375 91.375 -9.045 -814.1
15 95 -1.375 96.375 -4.045 -384.3
16 103 -1.375 104.375 3.955 407.37
Sum 1715 17216
Step 3.5. Estimate progeny effect as: Adjusted progeny value - adjusted grand mean;
For example, progeny effect for progeny 1 = 119.62-100.42 = 19.20; etc.
Analysis of variance
1. Correction Factor (C.F.) = = = 260822.53
2. Total SS = Y2 – C.F. = Sum of the squares of all observations – C.F. = (120)2 + (83)2 + ... + (75)2 -
260822.53 = 276621-260822.53
= 15798.47
3. Crude block SS = = + + +
54
= 262425.62
4. True block SS: Crude block SS – C.F. = 262425.62 - 260822 .53
= 1603.09
5. Adjusted SS Due to entries (C + P) = (Adjusted grand mean Observed grand total) + [(
)+( )+(
(Crude block sum squares)] = (100.42
= +
SS due to checks =
55
Summarize the results of analysis in ANOVA table
________________________________________________________________________
Source D.F. SS MS F-cal. F-Table
5% 1%
__________________________________________________________________
Block (b-1) 3 1603.09 534.36
Adjusted entries (C+P)-1= 19 14184.3 746.54
Unadjusted entries(C+P)-1= 19 15776.97 830.37
. Checks (C-1) 3 900.25 300.08 243.97** 3.86 6.99
. Test culture (P-1) 15 5730.44 382.03 310.59** 3.01 4.96
. Test culture vs. check 1 9146.28 9146.28 7436.00** 5.12 10.56
Error (b-1) (C-1) 9 11.08 1.23
Total (N – 1) 31 15798.47
Mean Comparison
56
8. FACTORIAL EXPERIMENTS
Factorial experiments are experiments in which two or more factors are studied together. Factor is a kind of
treatment and in a factorial experiment any factor will supply several treatments. In factorial experiment, the
treatments consist of combinations of two or more factors each at two or more levels.
Factorial experiment can be done in CRD, RCBD and Latin Square Design as long as the treatments allow.
Thus, the term factorial describes specific way in which the treatments are formed and it does not refer to
the experimental design used, e.g. nitrogen & phosphorus rates:
N = 0, 50, 100, 150 kg/ha
P = 0, 50, 100, 150 kg/ha
Kinds (noug cake, groundnut cake) and levels of protein supplement (25%, 50%, and 75%).
The term level refers to the several treatments within any factor, e.g. if 5-varieties of sorghum are tested
using 3-different row spacing, the experiment is called 5 x 3 factorial experiment with 5 levels of variety
factor (A) and three levels of spacing factor (B). An experiment involving 3 factors (variety, N-rate,
weeding method) each at 2 levels is referred as 2 2 2 or 23 factors; 3 refers to the number of factors and
2 refers to levels. Here we have 8 treatment combinations variety (x, y): N-rate (0, 50 kg/ha), weeding (with
or without weeding). The 23 3 is a four factor experiment in which three factors each at 2-levels and the
4th factor at 3 levels.
If the above 23 factorial experiment is done in RCBD, the correct description of the experiment will be 2 3
factorial experiment in RCBD.
Interaction
Sometimes the factors act independent of each other. By this we mean that changing the level of one factor
produces the same effect at all levels of another factor. Often, however, the effects of two or more factors
are not independent. Interaction occurs when the effect of one factor changes as the level of the other factor
changes, e.g. if the effect of 50kg N on variety X is 10 Q/ha and its effect on a variety Y is 15 Q/ha, then
there is interaction. When factors interact, the factors are not independent and a single factor experiment
will lead to disconnected or misleading information. However, if there is no interaction it is concluded that
the factors under consideration act independently of each other. Thus, results from separate single factor
experiments are equivalent to those from a factorial experiment.
Example: A tall maize variety might out yield a short variety in high fertilizer rates due to high dry matter
production.
Interaction is the failure of the differences in response to changes in levels of one factor to be the same at all
levels of another factor or when the effect of one factor changes as the level of the other factor changes.
57
2 x 2 Factorial data of wheat yield (t/ha)
_________________________________________________________________
N-rate (kg/ha) (Factor B)
_____________________________ Simple effect of
Variety (Factor A) 0 (b0) 50 (b1) nitrogen on variety
__________________________________________________________________
X (a0) 1.0 1.0 (a0b1-a0b0) = 0
Y (a1) 2.0 4.0 (a1b1-a1b0) = 2
Simple effect of
Variety (a1b0-a0b0) = 1 (a1b1-a0b1) = 3
Simple effects
- Simple effect of variety at N0: 2-1 = 1
- Simple effect of variety at N1: 4-1 = 3
- Simple effect of N on variety X: 1-1 = 0
- Simple effect of N on variety Y: 4-2 = 2
Interaction
It is calculated as the average of difference between simple effects of A at the two levels of B or the
difference between the simple effects of B at the two levels of A.
= ½ (Simple effect of A at b1 – simple effect of A at b0)
= ½ [(a1b1-a0b1) - (a1b0-a0b0)] = ½ [(4-1) - (2-1)] = 1
or
= ½ (Simple effect of B at a1 – simple effect of B at a0)
= ½ (a1b1-a1b0) - (a0b1-a0b0)] = ½ [(4-2) - (1-1)] = 1
He conducted this experiment using Randomized Complete Block Design with four blocks of six plots each.
Block-I
T2P2 T2P1 T1P1 T2P3 T1P3 T1P2
8.3 11.0 11.5 15.7 18.2 17.1
Block-II
T2P1 T2P2 T2P3 T1P2 T1P1 T1P3
11.2 10.5 16.7 17.6 13.6 17.6
Block-III
T1P2 T1P1 T2P1 T1P3 T2P3 T2P2
17.6 14.3 12.1 18.2 16.6 9.1
Block-IV
T1P3 T2P2 T2P3 T2P1 T1P2 T1P1
18.9 12.8 17.5 12.6 18.1 14.5
1. Construct two way table for factors and calculate factor A total, Factor B total and grand total
________________________________________________________
Phosphorus (Factor B)
_________________________________________
Variety (Factor A) P1 P2 P3 Factor A total (A)
_______________________________________________________
T1 (indeterminate) 53.9 70.4 72.9 197.2
T2 (determinate) 46.9 40.7 66.5 154.1
Factor B total (B) 100.8 111.1 139.4 351.3(G)
________________________________________________________
Block total
Block I II III IV
Total 81.8 87.2 87.9 94.4
59
2. Using r as number of blocks, a as level of factor A, b level of factor B, compute C.F., total SS, block
SS, treatment SS and Error SS
3. Compute the three factorial components of treatment SS [partition treatments SS in to factor A SS,
factor B SS, and A x B (interaction) SS]
ANOVA TABLE
__________________________________________________________
Source DF SS MS F-calcul. F-table
5% 1%
__________________________________________________________
Block r-1(4-1) = 3 13.32 4.44 7.65** 3.29 5.42
Variety (V) a-1 (2-1) = 1 77.40 77.40 133.45** 4.54 8.68
Phosphorus (P) b-1 (3-1) = 2 99.87 49.93 86.09** 3.68 6.36
VP (a-1) (b-1) = 2 44.11 22.05 38.03** 3.68 6.36
Error (r-1) (ab-1) = 15 8.68 0.58
Total rab -1= 23 243.38
_________________________________________________________
The interpretation of the results of factorial experiment depends on the outcome of the significance tests. If
factor A factor B interaction is significant, the main effects have no real meaning whether significant or
not. In our case, since A B interaction is highly significant, the results of experiment are best summarized
60
in a two way table means of various A B combinations. If interaction is not significant, then all of the
information in the trial is contained in the significant main effects. In this case the results may be
summarized in tables of mean for factors with significant main effects.
Mean Comparisons
Remark
Three or more factor experimental designs: Read Gomez & Gomez, Chapter 4, starting from page
130
Split-plot design is frequently used for factorial experiments where the nature of experimental material
makes it difficult to handle all factor combination. The principle underlying is that the levels of one factor
61
are assigned at random to large experimental units. The large units are then divided into smaller units
and then the levels of the second factor are assigned at random to small units within large units.
The large units are called the whole units or main-plots whereas the small units are called the split-plots or
sub-plots (units). Thus, each main plot becomes a block for the sub-plot treatments. In split-plot design, the
main plot factor effects are estimated from larger units, while the sub-plot factor effects and the interactions
of the main-plot and sub-plot factors are estimated from small units.
As there are two sizes of experimental units, there are two types of experimental error, one for the main
plot factor and the other for the sub-plot factor. Generally, the error associated with the sub-plots is smaller
than that for the whole plots due to the fact that error degrees of freedom for the main plot are usually less
than those for the sub-plots.
In split-plot design, the precision for the measurement of the effect of main plot factor is sacrificed to
improve the precision of the measurement of the sub-plot factors.
b. When an additional factor is to be incorporated in an experiment to increase its scope. For example,
if the major purpose of an experiment is to compare the effect of several vaccines as a protectant
against infection from certain disease of animals, to increase the scope of the experiment, several
breeds of animals can be included which are known to differ in their resistance to disease. Here, the
breeds of animals could be arranged in main units and the vaccines to the subunits.
c. When greater precision is desired for comparison of certain factors than others.
Since in a split-plot design, plot size and precision of measurement of the effects are not the same for both
factors, the assignment of a particular factor to either the main-plot or to the sub-plot is extremely important.
b. Relative size of the main effect: If the main effect of one factor (factor A) is expected to be much
larger and easier to detect than factor B, then factor A can be assigned to the main unit and factor B
to the sub-unit. For instance, in fertilizer and variety experiments, the researcher may assign variety
62
to the sub-unit and fertilizer rate to the main-unit, because he expects fertilizer effect to be much
large and easier to detect than the varietal effect.
c. Management practice: The factors, which require smaller amounts of experimental material, should
be assigned to sub-plots. For example, in an experiment to evaluate the frequency of irrigation (5,
10, 15 days), on performance of different tree seedlings on nursery, the irrigation frequency factor
could be assigned to the main plot and the different tree species to the sub-plots to minimize water
movement to adjacent plots.
Advantages
a. It permits the efficient use of some factors, which require large experimental units in combination
with other factors, which require small experimental units.
b. It provides increased precision in comparison of some of the factors (sub-plot factors).
c. It promotes the introduction of new treatments into an experiment, which is already in progress.
Disadvantages:
a. Statistical analysis is complicated because different factors have different error mean squares.
b. Low precision for the main plot factor can result in large differences being non-significant, while
small differences on the sub-plot factor may be statically significant even though they are of no
practical significance.
There are two separate randomization process in split-plot design, one for the main plot factor and another
for the sub-plot factor.
In each block, the main plot factors are first randomly applied to the main plots followed by random
assignment of the sub-plot factors. Each of the randomization is done by any of the randomization schemes.
Example: An experiment was designed to test the effect of feeding four forage crops (Rhodes grass, Vetch,
Alfalfa and Oat) on weight gain (kg/month) of the two breeds of cows (Zebu, Holstein). At the start of the
experiment, it was assumed that breeds of cows would respond differently to the feed stuffs. Therefore, it
was decided to use factorial experiment. The objective of the experiment was to compare the effect of
forage crops as precisely as possible. Therefore, the experimenter assigned the breeds of animals to the
main-plot and the four forage crops to the sub-plots. The experiment was replicated in three blocks (barns)
based on initial body weight of animals as a blocking factor.
Procedures of randomization
Step 1: Divide the experimental area into r = 3 blocks, and divide each block into two main plots. Then
randomly assign the two breeds of animals (H, Z) in each of the blocks.
Note that the arrangement of the main-plot factor can follow any of the designs: CRD, RCBD and LATIN
square.
Step 2: Divide each of the main plot (unit) into 4-sub plots (units) and randomly assign the four feed stuffs
(A, V, O, R) to each of the six-main plots (units).
Note:
63
Each main-plot factor is tested r-times where r is the number of blocks while each sub-plot factor is tested a
r times where a is level of factor A and r is the number of blocks. This is the primary reason for more
precision for the sub-plot factors as compared to the main-plot factors.
The layout and the weight gain (kg/month) of the animals for feeding are given below:
Block I Block II Block III
H Z H Z Z H
A R O V O V
25.9 15.5 18.0 22.7 13.2 28.4
V A A O A A
25.3 18.9 26.7 13.5 19.6 27.6
O O V R V R
19.3 13.8 24.8 15.0 22.3 25.4
R V R A R O
22.2 21.0 24.2 18.3 15.2 20.5
Steps of Analysis
Step 1: Arrange data by treatments (main-plot, sub-plot) and blocks and calculate main-plot total, and sub-
plot total.
Treatments Blocks
Breeds Feeds I II III
Holstein Alfalfa 25.9 26.7 27.6
Vetch 25.3 24.8 28.4
Oat 19.3 18.0 20.5
Rhodes grass 22.2 24.2 25.4
Main-plot totals 92.7 93.7 101.9
Zebu Alfalfa 18.9 18.3 19.6
Vetch 21.0 22.7 22.3
Oat 13.8 13.5 13.2
64
Rhodes grass 15.5 15.0 15.2
Main-plot totals 69.2 69.5 70.3
2.2. Factor A by factor B total two-way table and calculate factor B totals
Feeds (B)
Breeds (A) Alfalfa (b1) Vetch (b2) Oat (b3) R. Grass (b4)
Step 3: Compute the correction factor and sum of squares for the main-plot analysis.
C.F.=
Total SS = = 10820.59 – 10304.47
= 516.12
Block SS =
= 10312.34 –10304.47 = 7.87
Factor A (breeds) (main plot factor) SS=
10566.49 – 10304.47 = 262.02
Error (a) SS = Block SS-factor A SS
65
= = 215.26
Step 5: For each source- of variation compute the mean squares by dividing the SS by its corresponding
degrees of freedom.
Block MS =
Factor A MS =
Error (a) MS =
Factor B MS =
A B MS =
Error (b) MS =
Step 6: Compute the F-value for each effect that needs to be tested.
F(block) =
F(A) =
F(B) =
F(AB) =
Step 7: Construct the ANOVA Table, obtain the corresponding tabulated F-value and compare it with the
calculated F-value at prescribed level of significance.
Step 8. Compute the two coefficients of variation, one corresponding to the main-plot analysis and another
to the sub-plot analysis.
CV (a) =
CV (b) =
Note that CV (a) is greater than CV (b), this is because factors assigned to the main-plot are expected to be
measured with less precision than that assigned to the sub-plot.
Mean comparisons:
Mean weight (kg/month) of Holstein and Zebu breeds fed with four types of forage species
Breeds (A) Alfalfa (b1) Vetch (b2) Oat (b3) R. Grass (b4) Factor A mean
(A total/rb)
Holstein (a1) 26.7 26.2 19.3 23.9 24.0 (A1)
Zebu (a2) 18.9 22.0 13.5 15.2 17.4 (A2)
Factor B mean
(B total/ra) 22.8 (B1) 24.1 (B2) 16.4 (B3) 19.6 (B4)
a. to compare two main plot factor (A) means (A1 & A2): A=
A= = 0.65 kg
Since the mean difference (d) between A 1 and A2 (24.0-17.4 = 6.6) > LSD value at 1% (6.45), the difference
between the two means is highly significant.
b. to compare two sub-plot factor (B) means: B= =
= 0.45 kg
Example: To compare means of B3 and B2, LSD1% = t 0.005 [error (b) d. f.] =
LSD1% = t 0.005 (12) 0.45 kg = 3.055 0.45 = 1.37 kg.
Since the mean difference (d) between B 2 and B3 (24.1 – 16.4 = 7.7) is greater than LSD value at 1% (1.37),
the difference between the two means is highly significant.
67
c. to compare two sub-plot treatment means at the same level of main plot:
= = = 0.63 kg
To compare (a1b1= 26.7) with (a1b3 = 19.3); LSD1% = t0.005 [error (b) d. f.]
= LSD1% = t 0.005 (12) 0.63 kg = 3.055 0.63 = 1.92 kg.
Since the mean difference (d) between (a1b1 & a1b3 = 26.7-19.3 = 7.4) is greater than LSD value at
1% (1.92 kg), the difference between the two means is highly significant.
= = 0.85 kg
that contains both MSE(a) and MSE(b) has no exact value for the d.f. associated with it. To obtain an
approximation:
d.f. =
To compare (a1b1= 26.7) with (a2b2 = 22.0); LSD1% = t0.005 [error d.f.]
= LSD1% = t 0.005(5) 0.85 kg = 4.032 0.85 kg = 3.43 kg.
Since the mean difference (d) between (a1b1 & a1b3 = 26.7-22.0 = 4.7) is greater than LSD value at 1%
(3.43 kg), the difference between the two means is highly significant.
Presentation
If interaction of the factors is significant, results are summarized in two way table of means. However, if
interaction is non-significant, the results are summarized in one way table of means for the significant
factor.
The F-test (ANOVA) shows whether there is significant difference among treatments or not. But, it does not
show us which means are different from each other. There are many ways to compare the means of
treatments tested in an experiment. One of these is pair comparison, the simplest and most commonly used
comparisons in agricultural research.
B. Unplanned pair comparison: In which no specific comparison is chosen in advance. Instead, every
possible pair of treatment means are compared to identify pairs of treatments that are significantly different,
e.g. variety trials. A posteriori test or Post hoc test
The most commonly used test procedures for pair comparison in agricultural research are the Least
Significant Difference and Tukey’s test which are suitable for planned pair comparison and Duncan’s
Multiple Range Test (DMRT) which is applicable to an unplanned pair comparison.
LSD is the simplest and the most commonly used procedure for making pair comparisons. The procedure
provides a single value at a prescribed level of significance, which serves as the boundary between
significant and non-significant differences between any pair of treatment means. That is, two treatments are
declared significantly different at a prescribed level of significance if their mean difference exceed the
computed LSD value, otherwise they are not significantly different.
The LSD test is not valid for comparing all possible pair of means, especially when the number of
treatments is large. This is so because the number of possible pairs of treatment means increase rapidly as
the number of treatments increase. In experiments where no real difference exists among all treatments, the
numerical difference between the largest and smallest treatment means is expected to exceed the LSD value
when the number of treatments is large.
To avoid this problem, the LSD test is used only when the F-test for treatment effect is significant and the
number of treatments is not too large (less than six).
The procedure for applying the LSD test to compare any two treatments means
1. Rank the treatment means from the largest to the smallest in the column and from the smallest to largest
in rows.
2. Compute all possible differences between the two treatment means to be compared.
3. Compute the LSD value at α level of significance
LSDα = tα/2 (n) s
where s = standard error of the treatment mean difference; t α/2 (n) is the table t-value at α/2 level of
significance and with n error degree of freedom
Example: Oil content (g) of linseed treated at different six stages of growth with N-fertilizes tested in
RCBD in four replications with error mean square of 1.31.
LSD5% = t0.025(15) , where MSE is error mean square; r is the number of replications = 2.131
1.72 g
69
LSD1% = t 0.005(15) = 2.947
4. Compare the mean difference (d) in step 2 with LSD value computed in step (3) using the following rule:
- if /d/ >LSD value at 1% level of significance, there is highly significant difference between the two
treatment means compared (put two asterisks on differences).
- if /d/>LSD value at 5% level of significance but < LSD value at 1% level of significance, there is
significant difference between the two treatment means compared (put one asterisks on differences)
- if /d/ LSD value at 5% level of significance, the two treatment means compared are not significantly
different (put n.s.)
Thus, the differences between T6 & T3, T6 & T2, T4 & T3, T4 & T2 are highly significant; while the
differences between T3 & T5, T1 & T6, T2 & T5 are significant.
Note that there are possible (unplanned) pair comparisons and (t-1) planned pair comparisons where
t is the number of treatments. In the above example, 15 unplanned pair comparisons and five planned pair
comparisons are possible, since we have one control.
Table __. Mean oil content of linseed treated with nitrogen fertilizer at different stages
70
___________________________________________
Stage of application Oil content (g)
___________________________________________
Seedling 5.10
Early blooming 4.30
Half-blooming 4.00
Full- blooming 6.70
Ripening 6.05
Unfertilized 7.03
___________________________________________
LSD(0.05) 1.72 g
CV (%) 20.7
It is most widely used to make all possible pair comparisons. The procedure for applying the DMRT is
similar to LSD test but it requires progressively larger values for significance between the treatment means
as they are more widely separated in the array.
The test is more appropriate when the total number of treatments is large. It involves the calculation of the
shortest significant difference (SSD).
The SSD is calculated for all possible relative positions (P) between the treatment means when the means
are arranged in order of magnitude (in decreasing or increasing order).
Procedure
Step 1: Arrange all the treatment means in increasing or decreasing order.
Data such as crop yield are usually arranged from the highest to the lowest.
Example: Yields (kg/plot) of wheat varieties grown in 4 x 4 Latin Square Design with error mean square
of 0.45:
Step 2: Calculate (the standard error of the treatment mean difference) as:
=
Step 3: Calculate the shortest significant difference (SSD) for relative positions (P) in the array of means.
Since we have four treatment means, they can be 2, 3 and 4 distance apart.
B and A are 2 distance apart (P = 2); B and C are 3 distance apart (P = 3); B and D are 4 distance
apart (P = 4); A and D are 3 distance apart (P = 3); etc.
For the above example, the R values with error d. f. of 6 at 1% level of significance are found from R-table
(see Appendix F in Gomez & Gomez)
P= 2 3 4
71
R0.01 = 5.24 5.51 5.65
SSD = 1.74 1.83 1.88
P = the distance in ranks between the pairs of treatment means to be compared.
R= significant studentized range at error d. f. (6).
SSD at P = 2 =
SSD at P = 3 =
SSD at P = 4 =
Note that SSD values increase as the distance between treatments (P) to be compared increases.
Step 4:Test the difference between treatment means in the following order.
Largest – Smallest = 12.3 – 6.7 = 5.6; compare with SSD value at (P = 4) = 1.88; d (5.6) > SSD at P
= 4 (1.88); thus the difference is significant at 1% level of significance.
Largest – 2nd smallest = 12.3 – 10.8 = 1.5; compare with SSD at (P=3) = 1.84; d (1.5) < SSD at P= 3
(1.84); thus, the difference is non-significant at 1% level of significance.
Largest – 2nd largest = 12.3 – 12.0 = 0.3 < SSD at P = 2 (1.75); thus, the difference is non-significant
at 1% level of significance.
2nd largest – smallest = 12.0 - 6.7 = 5.3; compare with SSD at P = 3 (1.84); d (5.3) > SSD (1.84) at P
= 3; significant at 1% level of significance
2nd smallest – smallest = 10.8 – 6.7 = 4.1 compared with SSD at P = 2 (1.75); d (4.1) > SSD (1.75) at
P = 2; significant
etc
________________________________________________________
B (12.3) A (12.0) C (10.8) D (6.7)
_____________________________________________
D (6.7) 5.6** (P=4) 5.3** (P=3) P = 4.1** (P=2)-
ns ns
C (10.8) 1.5 (P=3) 1.2 (P=2) -
A (12.0) 0.3ns (P=2) -
B (12.3) -
________________________________________________________
Treatments B & D, A & D, C & D are significantly different at 1%, while treatments B & C, B & A, and A
& C are not significantly different at 1% level of significance.
Step 5: Present the test result in one of the following two ways
72
A. Use a line notation if the sequence of results can be arranged according to their ranks.
Any two means underscored by the same line are not significantly different at 1% level of significance
according to DMRT.
B(12.3) A(12.0) C(10.8) D(6.7)
B. Use the alphabet notation if the desired sequence of the results is not based on their rank which is
commonly used.
The alphabet notation can be derived from line notation simply by assigning the same alphabet to all
treatment means connected by the same horizontal line.
It is usual practice to assign letter a for the first line, b for the second line, c for third line and so on.
Note that letter a can be used for the largest or smallest treatment mean depending on the rank of
arrangement.
Note that we have to put a footnote below the table stating that any two means in the same column
followed by the same letter are not significantly different at 1% level of significance according to
DMRT. Note also that both LSD and DMRT are not used in the same table. Use either of them
depending on the appropriateness of the test.
It is more conservative than LSD test because it requires the largest treatment mean differences for
significance.
It is computed in a manner similar to the LSD test except that standard error of the mean is used
instead of standard error of the mean difference ( ), and
The procedure:
1. Select a value from q table, which depends on the number of means (n) and error degree of freedom
(v).
2. Compute the Critical Difference (CD) as = q(n, v) where MSE is error mean square; n is
number of means to be compared; v is error degrees of freedom and r is number of
replications.
73
3. For any pair of means, if the absolute value of the difference /d/ > critical value, the difference is
judged to be significant at a prescribed level of significance.
Example: The following analysis of variance table is from CRD with six varieties replicated four times in
glass house (mean rust incidence)
Source d. f. MS F-cal. F-table (5%)
____________
Variety (t-1) = 5 2976.44 24.80** 2.77
Error t(r-1) = 18 120.00
__________________________________________________________
Variety: 1 2 3 4 5 6
Mean stem rust incidence (%): 50.3 69.0 24.0 94.0 75.0 95.3
Thus, differences between varieties 6&3, 4&3, 5&3, 2&3, etc. are significant while differences between
varieties 2&1, 5&2, etc. are non-significant.
In applying the LSD test and DMRT, it is important that the appropriate standard error of the mean
difference (s ) should be used. s is affected by the experimental design used, the number of replications of
the two treatments being compared, and the specific type of means to be compared.
A). In CRD, RCBD and Latin Square Design where the number of replications for all treatments is equal,
the s for any pair of treatment means is computed as:
74
Where, MSE is mean square for error; r = number of replications that is common to all treatments.
Thus, lsd = where n is error degree freedom.
B). When the two treatments do not have the same number of replications in CRD. is computed as:
where MSE is mean square error; r i and rj are the number of replications of the two treatment means (i & j)
to be compared.
Thus, LSD=
Example: CRD with an unequal replications, effect of 4 – types of feedstuff on weight gain of chicks.
Treatment A = given to 5- chicks (5-replications) = 43.8 g
Treatment B = given to 4 chicks (4 replications) = 73.0 g
Treatment C = given to 3 chicks (3 replications) = 73.33 g
Treatment D = given to 5 chicks (5 replications) = 142.8 g
Given error mean square of 843.1 and error degree of freedom of 13, test if there is significant difference
between treatments B & D.
C). for the treatments with a single missing value and that of any other treatment without missing values.
75
b) Latin Square Design: . Thus, LSD= t/2(error d.f.) where MSE
is mean square error of the analysis of variance of Latin Square Design with a single missing value; r
= number of replications.
The analysis of covariance simultaneously examines the variance and covariance of selected variables so
that the character of primary interest is more accurately characterized than by the use of analysis of variance
only. Analysis of covariance requires measurement of the character of interest and the measurement of one
or more variable(s) known as covariate(s). It also requires that the functional relationship of the co-variates
(x) with the character of primary interest (y) is known before hand.
Examples: Consider wheat variety trial in which weed infestation is used as a co-variate with a known
functional relationship between weed incidence and grain yield (the character of primary interest), the
covariance analysis can adjust grain yield in each plot to a common level of weed incidence. With the
covariance analysis, the variation in yield due to weed incidence is quantified and effectively separated from
that due to varieties.
Similarly, age or initial body weight of experimental animals can be used as a covariate and weight gain due
to rations as character of interest.
Covariance analysis can be applied to any number of covariates and to any type of functional relationships
between variables. In this section, however, we will deal with the case of a single covariate whose
relationship to character of primary interest is linear.
The experimental error is reduced and the precision for comparing treatment increased, e.g. in a cattle
feeding experiment to compare the effects of several rations on weight gain, animals assigned to any one
block may vary in initial weight. Now if the initial weight is correlated with gain in weight, a portion of
experimental error for gain can be the result of differences in initial weight. By covariance analysis, a
contribution, which can be attributed to differences in initial weight, can be computed and eliminated from
experimental error.
The following data show ascorbic acid content (y) of ten varieties of common bean. From the previous
experience, it was known that increase in maturity resulted in decrease in vitamin C content (linear r/ship).
Since all varieties were not of the same level of maturity on the same day, it was not possible to harvest all
77
plots at the same stage of maturity. Hence, the percentage of dry matter based on 100 g of freshly harvested
beans was observed as an index of maturity and used as a covariate.
Ascorbic acid content (ASAC, mg/100 g of seed) and percentage of dry matter
(% DM) for common bean varieties
___________________________________________________________
Block I Block 2 Block 3
____________ _____________ ______________
Variety %DM ASAC %DM ASAC %DM ASAC
________(X)___(Y)____(X)__(Y)______(X)___(Y)__
1 34 93 33 95 35 92
2 40 47 40 51 51 33
3 32 81 30 100 34 72
4 38 67 38 74 40 65
5 25 119 24 128 25 125
6 30 106 29 111 32 99
7 33 106 34 107 35 97
8 34 61 31 83 31 94
9 31 80 30 106 35 77
10 21 149 25 151 23 170
_______________________________________________________________
Conduct the analysis of covariance & calculate standard error of mean difference.
1. Conduct analysis of variance for each of the variables, covariance and sum
of square of treatment and error sum of squares
_______________________________________________________________
Source D. F. SS of ASAC (Y) SS of % DM (X) SS of XY
________________________________________________________________
Block 2 545.3 42.47 -75.23
Treatment 9 25689.0 972.70 -4633.23
Error 18 1608.7 86.20 -251.77
Treatment + Error 27 27297.7 1058.90 -4885.00
_________________________________________________________________
2. Analyse covariance
C.F. = = = 92078.23
Total Sum of Products = = (34 93) + (40 47) + ... + (23 170) - 92078.23 =
87118-92078.23 = -4960.23
Sum of Products due to Blocks = - C.F. -
92078.23 = 92003-92078.23 = -75.23
Sum of Products for Treatments: - C.F. = -
92078.23 = -4633.23
Error Sum of Squares of Products = Total Sum of Products – Block SS of Products – Treatment SS of
Products = -4960.23-(-75.23)-(-4633.23)
= -251.77
3. Compute the adjusted error SS of Y as = Error SS due to Y –
= 1608.7 – = 873.34
4. Compute (treatment + error) adjusted SS of Y as:
Compute the relative efficiency (R.E.) of covariance analysis compared to standard analysis of variance
Thus, the result indicates that the use of % dry matter as the covariate has not increased precision in ascorbic
acid content which would have been obtained had the ANOVA is done without covariance.
Adjusted Error MS of Y
CV = x 100 = = 7.6%
Grand Mean of Y
80
Mean comparison
In field experiments, it is necessary to repeat the experiments over a number of locations, seasons, or both,
e.g. varietal trials, plant spacing, fertilizer trials. The purpose of repeating the experiments is to find
recommendation that can be applied over space (location), time (season) or both.
In such repeated experiments, appropriate statistical procedures for a combined analysis of data have to be
used. The main purposes of combined analysis of data are:
to estimate the average response to a given experiment
to test the consistency of the response from place to place or year to year, i.e. to determine if there is
interaction effect of the treatments, e.g. stability analysis of varieties.
If the response is consistent from place to place and year to year, it shows the absence of interaction.
1. Construct an outline of combined analysis over years or locations or both on the basis of experimental
design used.
For example, the outline of ANOVA for the experiment conducted at six environments with five
treatments and six replications in RCBD is given as:
___________________________________________________________________
Source Degrees of
freedom Mean squares Computed F
________________________________________________________________________
Environment (E) (e-1) = 5 EMS EMS/RMS
Blocks/within environment e(r-1) = 30 RMS
Treatments (T) (t-1) = 4 TMS TMS/MSE
TxE (e-1) (t-1) = 2 ExTMS E xTMS/MSE
Pooled error e(r-1) (t-1) = 120 MSE
________________________________________________________________________
81
Where e = no. of environments; r = no. of replications; t = no. of treatments; EMS = environment mean
square; RMS is blocks/replications mean square; TMS = treatment mean squares; and MSE = mean square
error.
2. Compute the usual ANOVA for each environment according to RCBD and obtain the mean squares of
error and degrees of freedom from individual environment analysis.
________________________________________________________________________
Environment: E1 E2 E3 E4 E5 E6
Error d.f. (r-1)(t-1) 20 20 20 20 20 20
Error MS 5776 4028 4516 9526 7056 5535
Case 1: When there are only two environments, use F-test as:
F calculated = and compare with table F-value at d.f. for larger error mean
square as numerator and d.f. for smaller error mean square as denominator.
If F-calculated is > F-table, the null hypothesis of homogeneity of variances is rejected, i.e. the error
variances are heterogeneous. Thus, combined analysis cannot be conducted directly. We can analyze the
data separately or transform the data to homogenize the error variances to use combined analysis.
On the other hand, if the F-calculate value is ≤F-table, the error variances are homogeneous, thus we
can directly proceed to combined analysis.
Case 2: When environment is more than two, use Bartlett’s chi-square test
2c =
Where
Kj = Error d.f. for each environment
1/kj
Environment Kj MSE log MSE Kj x MSE Kj x log MSEi
E1 20 5776 3.761627 75.6 75.23254 0.05
E2 20 4028 3.605089 75.6 72.10179 0.05
E3 20 4516 3.654754 75.6 73.09508 0.05
E4 20 9526 3.978911 75.6 79.57821 0.05
82
E5 20 7056 3.848559 75.6 76.97117 0.05
E6 20 5535 3.743118 75.6 74.86235 0.05
453.6 451.8411 0.3
= Pooled error mean square = =
= 6072.83
Log (6072.83) = 3.78
2c = ; 2c = 3.97
Regression analysis describes the effect of one or more variables (designated as independent variables) on a
single variable (designated as the dependent variable). It expresses the dependent variable as a function of
independent variable(s).
For regression analysis, it is important to clearly distinguish between the dependent and independent
variables.
Examples:
- Weight gain in animals depends on feed
- Number of growth rings in a tree depends on age of the tree
- Grain yield of maize depends on a fertilizer rate
In the above cases, weight gain, number of growth rings and grain yield are dependent variables, while feed,
age and fertilizer rates are independent variables.
Correlation analysis, on the other hand, provides a measure of the degree of association between the
variables, e.g. the association between height and weight of students; body weight of cows and milk
production; grain yield of maize and thousand kernel weight.
Linear Relationships
The relationship between any two variables (independent and dependent) is linear if the change in y is
constant as x changes through out the range of x under consideration.
The functional form of linear relationship between a dependent variable y and an independent variable x is
represented by the equation.
y = a + bx
When there are more than one independent variables as say k-independent variables (x 1, x2, ………, xk), the
simple linear regression equation y = a + βx can be extended to the multiple linear functional form of:
y = α + β1x1 + β2x2 +……. + βkxk
where α is the y intercept (the value of y when all x’s are 0); β1, β 2, ….. βk are partial regression coefficients
associated with the independent variables.
The simple linear regression analysis deals with the estimation and test of significance concerning the two
parameter α and β in the equation:
Y = α + βx
The data required for the application of the simple linear regression are the n-pairs (with n >2) of y and x
values.
84
Step 1: Compute the means ( and ), deviation from means [ , ], square of the deviates (x2, y2)
and product of deviates (xy).
Example: Determine the regression equation for dependence of wing length of 13 sparrows of various ages.
where a is the estimate of α (the y intercept) and b is the estimate of β (linear regression
coefficient, slope).
b= ;a=
This is the estimated linear functional relationship between age (days) and wing length (cm). Thus, wing
length increases by 0.27 cm every day.
85
To test β, compute the residual mean square as:
=0.05
The residual mean square denotes the variance of y after taking into account the dependence of y on x.
tb =
Step 4: Compare the calculated tb value with tabulated t-value at α/2 level of significance, at n-2 (13-2) = 11
d.f.; where n is pair of observations.
Since calculated /tb/ (19.5) is > the tabulated t-value at the 1% level of significance, the linear response of
wing length to changes in the days within the range of 3 to 17 days is highly significant.
The simple linear correlation analysis deals with the estimation and test of significance of the simple linear
correlation coefficient (r), which is a measure of the degree of linear association between two variables x
and y (there is no need to have a dependent and independent variable).
The value of r lies within the range of –1 to +1, with extreme values indicating the perfect linear association
and the mid-value of 0 indicates no-linear association between the two variables. The value of r is negative
when a positive change in one variable is associated with a negative change in another and positive when
the values of two variables change in the same direction (increase or decrease).
Even though the zero r value indicates the absence of linear association between two variables, it does not
indicate the absence of association between them. It is possible for the two variables to have a non-linear
association such as quadratic form. The procedure for the estimation and test of significance of a simple
linear correlation coefficient between two variables x and y are:
Step 1: Compute the means ( ), the sum of square of the deviates ( and ), and the sum of the
cross product of deviates of the two variables.
Step 2: Compute the simple linear correlation coefficient for the above example as:
r=
86
Step 3: Test the significance of the simple linear correlation coefficient (r) by comparing the computed r-
value with the tabulated r-value at n-2 d.f.
The simple linear correlation coefficient (r) is declared significant at α level of significance if the
absolute value of the computed r-value > the corresponding tabulated r-value.
Thus, the simple linear correlation coefficient is significant at 1% level of significance which indicates the
presence of a highly significant and positive linear association between ages and wing length of sparrows.
The simple linear regression and correlation analysis is applicable only in cases with one independent
variable. However, in many situations Y may be dependent on more than one independent variables. Linear
regression analysis involving more than one independent variables is called multiple linear regression. The
relationship of the dependent variable Y to the K independent variables X1, X2, ... Xk can be expressed as:
Linear regression involving two independent variables can be expressed as: Y = + 1X1 + 2X2 where 1 &
2 are partial regression coefficients. 1 measures a change in Y for unit change in X 1, if X2 is held constant.
Similarly, 2 measures the rate of change in Y for a unit change in X 2 where X1 is held constant.
(sometimes designated as 0) is the value of Y when both X1 & X2 are zero.
Example: The following data show the weight gain, initial body weight & age of five chicks fed with
certain type of rations for a month.
b1= ; = =-
1.46
b2 = ; = = 0.18
a= – b1 -b2
Thus, the estimated multiple linear regression equation for initial age (days) and initial body weight (g) with
weight gain (g) is: Ŷ= 11.1 - 1.46 X1 + 0.18 X2 for 4 X1 6; and 10 X2 20.
Step 4: Compute:
The sum of squares due to regression (SSR) =
Residual (error) SS =
R2 measures the amount of explained variation of Y due to the independent variables. Thus, in the above
example 67% of the total variation in weight gain (g) of chicks can be accounted for a linear function
involving initial age (days) and initial body weight (g).
88
where k is number of independent variables (2) and n is number of data pairs (5)
Since the computed F-value (2.05) is less than the table F value at 5% (19.00) the estimated multiple linear
regression Ŷ= 11.1 - 1.46 X1 + 0.18 X2 is not significant at the 5% level.
Thus, the combined linear effect of initial age (days) and initial body weight (g) on weight gain (g)
of chicks is not significant.
Remark: The larger the R2 value, the more important the regression equation in characterizing Y.
On the other hand, if the value of R2 is low, even if the F-test is significant, the estimated linear
regression equation may not be useful.
For example an R2 value of 0.26, even if significant indicates that only 26% of the total variation in
the dependent variable (Y) is explained by the linear function of the independent variables
considered.
For valid applications of parametric analysis like ANOVA, t-test, etc certain basic assumptions must be met.
If the data violate such assumptions transformation of data can be used.
The appropriate type of data transformation to be used depends on the specific type of relationship between
the variances and the means. Conduct the analysis using the transformed data, and in tables present
transformed means in parenthesis alongside their back transformed values out of parenthesis.
More appropriate when the treatment effects are multiplicative rather than additive, then the logarithmic
transformation of the data will exhibit additivity. Such conditions are generally found on count data such as
number of insects per plot, number of eggs per plant, etc.
- X’=log (x + 1); where x is original data and x + 1 is preferred especially when some of the observed
values are small. Logarithmic of base 10 are generally are used but any base would be satisfactory.
- Number of eggs/plant = 9; log (9+1); log (10) = 1
The square root transformation is applicable when the data consists of counts of rare events such as the
number of infested plants in a plot, the number of insects caught in traps. For such data the variance tends to
be proportional to the mean. Square root transformation is also appropriate for percentage data where the
range is between 0-30% or between 70-100%.
89
If most of the values in the data set are small especially with zeros present.
X’ = ; x = 0; 0.707
Statistical computation can be done on the transformed data. The mean can be expressed in terms of the
original data by squaring the transformed value and subtractions of 0.5
For proportion of 0 to 1 (0-100%), the transformed values will range between 0 and 90 degrees, percentage
44%: Sin-1 = 41.55
90
APPENDIX
Appendix Table I. Area under the Standard Normal Probability Distribution (Z Distribution) -Top Tail
Probabilities (Area to the right of the given Z-values)
0.0 .5000 .4960 .4920 .4880 .4840 .4801 .4761 .4721 .468
0.1 1 .4641
0.2 .4602 .4562 .4522 .4483 .4443 .4404 .4364 .4325 .428
0.3 6 .4247
0.4 .4207 .4168 .4129 .4090 .4052 .4013 .3974 .3936 .389
0.5 7 .3859
0.6 .3821 .3783 .3745 .3707 .3669 .3632 .3594 .3557 .352
0.7 0 .3483
0.8 .3446 .3409 .3372 .3336 .3300 .3264 .3228 .3192 .315
0.9 6 .3121
1.0 .3085 .3050 .3015 .2981 .2946 .2912 .2877 .2843 .281
1.1 0 .2776
1.2 .2743 .2709 .2676 .2643 .2611 .2578 .2546 .2514 .248
1.3 3 .2451
1.4 .2420 .2389 .2358 .2327 .2296 .2266 .2236 .2206 .217
1.5 7 .2148
1.6 .2119 .2090 .2061 .2033 .2005 .1977 .1949 .1922 .189
1.7 4 .1867
1.8 .1841 .1814 .1788 .1762 .1736 .1711 .1685 .1660 .163
1.9 5 .1611
2.0 .1587 .1562 .1539 .1515 .1492 .1469 .1446 .1423 .140
2.1 1 .1379
2.2 .1357 .1335 .1314 .1292 .1271 .1251 .1230 .1210 .119
2.3 0 .1170
2.4 .1151 .1131 .1112 .1093 .1075 .1056 .1038 .1020 .100
2.5 3 .0985
2.6 .0968 .0951 .0934 .0918 .0901 .0885 .0869 .0853 .083
2.7 8 .0823
2.8 .0808 .0793 .0778 .0764 .0749 .0735 .0721 .0708 .069
2.9 4 .0681
3.0 .0668 .0655 .0643 .0630 .0618 .0606 .0594 .0582 .057
1 .0559
.0548 .0537 .0526 .0516 .0505 .0495 .0485 .0475 .046
5 .0455
.0446 .0436 .0427 .0418 .0413 .0406 .0392 .0384 .037
5 .0367
91
Z .00 .01 .02 .03 .04 .05 0.06 0.07 0.08
0.09
.0359 .0351 .0344 .0336 .0329 .0322 .0314 .0307 .030
1 .0294
.0287 .0281 .0274 .0268 .0262 .0256 .0250 .0244 .023
9 .0233
.0228 .0222 .0217 .0212 .0207 .0202 .0197 .0192 .018
8 .0183
.0179 .0174 .0170 .0166 .0162 .0158 .0154 .0150 .014
6 .0143
.0139 .0136 .0132 .0129 .0125 .0122 .0119 .0116 .011
3 .0110
.0107 .0104 .0102 .0099 .0096 .0094 .0091 .0089 .008
7 .0084
.0082 .0080 .0078 .0075 .0073 .0071 .0069 .0068 .006
6 .0064
.0062 .0060 .0059 .0057 .0055 .0054 .0052 .0051 .004
9 .0048
.0047 .0045 .0044 .0043 .0041 .0040 .0039 .0038 .003
7 .0036
.0035 .0034 .0033 .0032 .0031 .0030 .0029 .0028 .002
7 .0026
.0026 .0025 .0024 .0023 .0023 .0022 .0021 .0021 .002
0 .0019
.0019 .0018 .0018 .0017 .0016 .0016 .0015 .0015 .001
4 .0014
.0013 .0013 .0013 .0012 .0012 .0011 .0011 .0011 .001
0 .0010
Remark
The values in the table show area under the Standard Normal Curve to the right of the Z-values.
The Z-value is the combination of the numbers in the left most column and in the heading row at whose
intersection the value in the table appears.
Example: For 95% Confidence, =5% (total of both tails), it is 2.5%= 0.025 in each tail, which is equal to
the Area to the left of –Zα/2=the Area to the right of Zα/2. Therefore, we simply look for this are (0.025) in
the table, and the Z-value is read from the corresponding values in the left most column and the heading
row. For this example (i.e., for 95%, the Z-value is 1.9 from the left most row plus 0.06 from the heading
row, which is 1.96).
92
Appendix Table II. Percentage Points of the t distribution
_______________________________________________________________________________
α (one-tailed)
Degree of _________________________________________________________________
freedom (n-1) 0.25 0.10 0.05 0.025 0.01 0.005 0.0025 0.001 0.0005
________________________________________________________________________________
1 1.000 3.078 6.314 12.706 31.821 63.657 127.32 318.31 636.62
2 0.816 1.886 2.920 4.303 6.965 9.925 14.089 23.328 31.598
3 0.765 1.638 2.353 3.182 4.541 5.841 7.453 10.213 12.924
4 0.741 1.533 2.132 2.776 3.747 4.604 5.598 7.173 8.610
5 0.727 1.475 2.015 2.571 3.365 4.032 4.773 5.893 6.869
6 0.727 1.440 1.943 2.447 3.143 3.707 4.317 5.208 5.959
7 0.711 1.415 1.895 2.365 2.998 3.499 4.019 4.785 5.408
8 0.706 1.397 1.860 2.306 2.896 3.355 3.833 4.501 5.041
9 0.703 1.383 1.833 2.262 2.821 3.250 3.690 4.297 4.780
10 0.700 1.372 1.812 2.228 2.764 3.169 3.581 4.144 4.587
11 0.697 1.363 1.796 2.201 2.718 3.106 3.497 4.025 4.437
12 0.695 1.356 1.782 2.179 2.681 3.055 3.428 3.930 4.318
13 0.694 1.350 1.771 2.160 2.650 3.012 3.372 3.852 4.221
14 0.692 1.345 1.761 2.145 2.624 2.977 3.326 3.787 4.140
15 0.691 1.341 1.753 2.131 2.620 2.947 3.286 3.733 4.073
16 0.690 1.337 1.746 2.120 2.583 2.921 3.252 3.686 4.015
17 0.689 1.333 1.740 2.110 2.567 2.898 3.222 3.646 3.965
18 0.688 1.330 1.734 2.101 2.552 2.878 3.197 3.610 3.922
19 0.688 1.328 1.729 2.093 2.539 2.861 3.174 3.579 3.883
20 0.687 1.325 1.725 2.086 2.528 2.845 3.153 3.552 3.850
21 0.687 1.323 1.721 2.080 2.518 2.831 3.135 3.527 3.819
22 0.686 1.321 1.717 2.074 2.508 2.819 3.119 3.505 3.792
23 0.685 1.319 1.714 2.069 2.500 2.807 3.104 3.485 3.767
24 0.685 1.318 1.711 2.064 2.492 2.797 3.091 3.467 3.745
25 0.684 1.316 1.708 2.060 2.485 2.787 3.078 3.450 3.725
26 0.684 1.315 1.706 2.056 2.479 2.779 3.067 3.435 3.707
27 0.684 1.314 1.703 2.052 2.473 2.771 3.057 3.421 3.690
28 0.683 1.313 1.701 2.048 2.467 2.763 3.047 3.408 3.674
29 0.683 1.311 1.699 2.045 2.462 2.756 3.038 3.396 3.659
30 0.683 1.310 1.697 2.042 2.457 2.750 3.030 3.385 3.646
40 0.681 1.303 1.684 2.021 2.423 2.704 2.971 3.307 3.551
60 0.679 1.296 1.671 2.000 2.390 2.660 2.915 3.232 3.460
120 0.677 1.289 1.658 1.980 2.358 2.617 2.860 3.160 3.373
0.674 1.282 1.645 1.960 2.326 2.576 2.807 3.090 3.291
______________________________________________________________________________
93
Appendix Table III. Chi-square (2) Table
D.F. 0.99 0.95 0.90 0.75 0.50 0.25 0.20 0.10 0.05 0.02 0.01
1 0.000 0.000 0.016 0.102 0.455 1.32 1.642 2.706 3.841 5.412 6.635
2 0.020 0.103 0.211 0.575 1.386 2.77 3.219 4.605 5.991 7.824 9.210
3 0.115 0.352 0.584 1.213 2.366 4.11 4.642 6.251 7.815 9.837 11.345
4 0.297 0.711 1.064 1.923 3.357 5.38 5.989 7.779 9.488 11.668 13.277
5 0.554 1.145 1.610 2.675 4.351 6.63 7.289 9.236 11.070 13.388 15.086
6 0.872 1.635 2.204 3.455 5.348 7.84 8.558 10.645 12.592 15.033 16.812
7 1.239 2.167 2.833 4.255 6.346 9.04 9.803 12.017 14.067 16.622 18.475
8 1.646 2.733 3.490 5.017 7.344 10.22 11.030 13.362 15.507 18.168 20.090
9 2.088 3.325 4.168 5.899 8.343 11.39 12.242 14.684 16.919 19.679 21.666
10 2.568 3.940 4.865 6.737 9.342 12.55 13.442 15.987 18.307 21.161 23.209
11 3.053 4.575 5.578 7.584 10.341 13.70 14.631 17.275 19.675 22.618 24.725
12 3.572 5.226 6.304 8.438 11.340 14.84 15.812 18.549 21.026 24.054 26.217
13 4.107 5.892 7.042 9.299 12.340 15.98 16.985 19.812 22.362 25.472 27.688
14 4.660 6.571 7.790 10.165 13.339 17.12 18.151 21.064 23.685 26.873 29.141
15 5.229 7.261 8.547 11.036 14.339 18.25 19.311 22.307 24.996 28.259 30.578
16 5.812 7.962 9.312 11.912 15.338 19.37 20.465 23.542 26.296 29.633 32.000
17 6.408 8.672 10.085 12.792 16.338 20.49 21.615 24.769 27.587 30.995 33.409
18 7.015 9.390 10.865 13.675 17.338 21.60 22.760 25.989 28.869 32.346 34.805
19 7.633 10.117 11.651 14.562 18.338 22.72 23.900 27.204 30.144 33.687 36.191
20 8.260 10.851 12.443 15.452 19.337 23.84 25.038 28.412 31.410 35.020 37.566
21 26.171 29.615 32.671 36.343 38.932
22 9.542 12.338 14.041 17.240 21.337 26.04 27.301 30.813 33.924 37.659 40.289
23 28.429 32.007 35.172 38.968 41.638
24 10.856 13.848 15.659 19.037 23.337 28.24 29.553 33.196 36.415 40.270 42.980
25 30.675 34.382 37.652 41.566 44.314
26 12.198 15.379 17.292 20.843 25.336 30.43 31.795 35.563 38.885 42.856 45.642
27 32.912 36.741 40.113 44.140 46.963
28 13.565 16.928 18.939 22.657 27.336 32.62 34.027 37.916 41.337 45.419 48.278
29 35.139 39.087 42.557 46.693 49.588
30 14.953 18.493 20.599 24.478 29.336 34.80 36.250 40.256 43.773 47.962 50.892
94
Appendix Table IV. The 5% and 1% Point for the F-distribution
v k(or p):2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20
1 17.97 26.98 32.82 37.08 40.41 43.12 45.4 47.36 49.07 50.59 51.96 53.20 54.33 55.36 56.32 57.22 58.04 58.83 59.56
2 6.085 8.331 9.798 10.88 11.74 12.44 13.03 13.54 13.99 14.39 14.75 15.08 15.38 15.65 15.91 16.14 16.37 16.57 16.77
3 4.501 5.910 6.825 7.502 8.037 8.478 8.853 9.177 9.462 9.717 9.946 10.15 10.35 10.53 10.69 10.84 10.98 11.11 11.24
4 3.927 5.040 5.757 6.287 6.707 7.053 7.347 7.602 7.826 8.027 8.208 8.373 8.525 8.664 8.794 8.914 9.028 9.134 9.233
5 3.635 4.602 5.218 5.673 6.033 6.330 6.582 6.802 6.995 7.168 7.324 7.466 7.596 7.717 7.828 7.932 8.030 8.122 8.208
6 3.461 4.339 4.896 5.305 5.628 5.895 6.122 6.319 6.493 6.649 6.789 6.917 7.034 7.143 7.244 7.338 7.426 7.508 7.587
7 3.344 4.165 4.681 5.060 5.359 5.606 5.815 5.998 6.158 6.302 6.431 6.550 6.658 6.759 6.852 6.939 7.020 7.097 7.17
8 3.261 4.041 4.529 4.886 5.167 5.399 5.597 5.767 5.918 6.054 6.175 6.287 6.389 6.483 6.571 6.653 6.729 6.802 6.87
9 3.199 3.949 4.415 4.756 5.024 5.244 5.432 5.595 5.739 5.867 5.983 6.089 6.186 6.276 6.359 6.437 6.510 6.579 6.644
10 3.151 3.877 4.327 4.654 4.912 5.124 5.305 5.461 5.599 5.722 5.833 5.935 6.028 6.114 6.194 6.269 6.339 6.405 6.467
11 3.113 3.82 4.256 4.574 4.823 5.028 5.202 5.353 5.487 5.605 5.713 5.811 5.901 5.984 6.062 6.134 6.202 6.265 6.326
12 3.082 3.773 4.199 4.508 4.751 4.950 5.119 5.265 5.395 5.511 5.615 5.710 5.798 5.878 5.953 6.023 6.089 6.151 6.209
13 3.055 3.735 4.151 4.453 4.690 4.885 5.049 5.192 5.318 5.431 5.533 5.625 5.711 5.789 5.862 5.931 5.995 6.055 6.112
14 3.033 3.702 4.111 4.407 4.639 4.829 4.990 5.131 5.254 5.364 5.463 5.554 5.637 5.714 5.786 5.852 5.915 5.974 6.029
15 3.014 3.674 4.076 4.367 4.595 4.782 4.94 5.077 5.198 5.306 5.404 5.493 5.574 5.649 5.72 5.785 5.846 5.904 5.958
16 2.998 3.649 4.046 4.333 4.557 4.741 4.897 5.031 5.150 5.256 5.352 5.439 5.520 5.593 5.662 5.727 5.786 5.843 5.897
17 2.984 3.628 4.020 4.303 4.524 4.705 4.858 4.991 5.108 5.212 5.307 5.392 5.471 5.544 5.612 5.675 5.734 5.790 5.842
18 2.971 3.609 3.997 4.277 4.495 4.673 4.824 4.956 5.071 5.174 5.267 5.352 5.429 5.501 5.568 5.63 5.688 5.743 5.794
19 2.960 3.593 3.977 4.253 4.469 4.645 4.794 4.924 5.038 5.14 5.231 5.315 5.391 5.462 5.528 5.589 5.647 5.701 5.752
20 2.950 3.578 3.958 4.232 4.445 4.620 4.768 4.896 5.008 5.108 5.199 5.282 5.357 5.427 5.493 5.553 5.610 5.663 5.714
24 2.919 3.532 3.901 4.166 4.373 4.541 4.684 4.807 4.915 5.012 5.099 5.179 5.251 5.319 5.381 5.439 5.494 5.545 5.594
30 2.888 3.486 3.845 4.102 4.302 4.464 4.602 4.720 4.824 4.917 5.001 5.077 5.147 5.211 5.271 5.327 5.379 5.429 5.475
40 2.858 3.442 3.791 4.039 4.232 4.389 4.521 4.635 4.735 4.824 4.904 4.977 5.044 5.106 5.163 5.216 5.266 5.313 5.358
60 2.829 3.399 3.737 3.977 4.163 4.314 4.441 4.55 4.646 4.732 4.808 4.878 4.942 5.001 5.056 5.107 5.154 5.199 5.241
120 2.800 3.356 3.685 3.917 4.096 4.241 4.363 4.468 4.560 4.641 4.714 4.781 4.842 4.898 4.950 4.998 5.044 5.086 5.126
2.772 3.314 3.633 3.858 4.030 4.170 4.286 4.387 4.474 4.552 4.622 4.685 4.743 4.796 4.845 4.891 4.934 4.974 5.012
Critical Values of the q Distribution ( = 0.01); k or p = number of mean to be compared; v = error degree of freedom.
V k(or p):2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20
1 90.03 135.0 164.3 185.6 202.2 215.8 227.2 237.0 245.6 253.2 260.0 266.2 271.8 277.0 281.8 286.3 290.4 294.3 298.0
2 14.04 19.02 22.29 24.72 26.63 28.20 29.53 30.68 31.69 32.59 33.4 34.13 34.81 35.43 36.00 36.53 37.03 37.5 37.95
3 8.261 10.62 12.17 13.33 14.24 15.00 15.64 16.20 16.69 17.13 17.53 17.89 18.22 18.52 18.81 19.07 19.32 19.55 19.77
4 6.512 8.120 9.173 9.958 10.58 11.10 11.55 11.93 12.27 12.57 12.84 13.09 13.32 13.53 13.73 13.91 14.08 14.24 14.4
5 5.702 6.976 7.804 8.421 8.913 9.321 9.669 9.972 10.24 10.48 10.70 10.89 11.08 11.24 11.40 11.55 11.68 11.81 11.93
6 5.243 6.331 7.033 7.556 7.973 8.318 8.613 8.869 9.097 9.301 9.485 9.653 9.808 9.951 10.08 10.21 10.32 10.43 10.54
7 4.949 5.919 6.543 7.005 7.373 7.679 7.939 8.166 8.368 8.548 8.711 8.86 8.997 9.124 9.242 9.353 9.456 9.554 9.646
8 4.746 5.635 6.204 6.625 6.960 7.237 7.474 7.681 7.863 8.027 8.176 8.312 8.436 8.552 8.659 8.760 8.854 8.943 9.027
9 4.596 5.428 5.957 6.348 6.658 6.915 7.134 7.325 7.495 7.647 7.784 7.910 8.025 8.132 8.232 8.325 8.412 8.495 8.573
10 4.482 5.270 5.769 6.136 6.428 6.669 6.875 7.055 7.213 7.356 7.485 7.603 7.712 7.812 7.906 7.993 8.076 8.153 8.226
11 4.392 5.146 5.621 5.97 6.247 6.476 6.672 6.842 6.992 7.128 7.250 7.362 7.465 7.560 7.649 7.732 7.809 7.883 7.952
12 4.320 5.046 5.502 5.836 6.101 6.321 6.507 6.670 6.814 6.943 7.060 7.167 7.265 7.356 7.441 7.520 7.594 7.665 7.731
13 4.260 4.964 5.404 5.727 5.981 6.192 6.372 6.528 6.667 6.791 6.903 7.006 7.101 7.188 7.269 7.345 7.417 7.485 7.548
14 4.210 4.895 5.322 5.634 5.881 6.085 6.258 6.409 6 .543 6.664 6.772 6.871 6.962 7.047 7.126 7.199 7.268 7.333 7.395
15 4.168 4.836 5.252 5.556 5.796 5.994 6.162 6.309 6.439 6.555 6.660 6.757 6.845 6.927 7.003 7.074 7.142 7.204 7.264
16 4.131 4.786 5.192 5.489 5.722 5.915 6.079 6.222 6.349 6.462 6.564 6.658 6.744 6.823 6.898 6.967 7.032 7.093 7.152
17 4.099 4.742 5.140 5.430 5.659 5.847 6.007 6.147 6.270 6.381 6.480 6.572 6.656 6.734 6.806 6.873 6.937 6.997 7.053
18 4.071 4.703 5.094 5.379 5.603 5.788 5.944 6.081 6.201 6.310 6.407 6.497 6.579 6.655 6.725 6.792 6.854 6.912 6.968
19 4.046 4.67 5.054 5.334 5.554 5.735 5.889 6.022 6.141 6.247 6.342 6.430 6.510 6.585 6.654 6.719 6.780 6.837 6.891
20 4.024 4.639 5.018 5.294 5.510 5.688 5.839 5.970 6.087 6.191 6.285 6.371 6.450 6.523 6.591 6.654 6.714 6.771 6.823
24 3.956 4.546 4.907 5.168 5.374 5.542 5.685 5.809 5.919 6.017 6.106 6.186 6.261 6.330 6.394 6.453 6.510 6.563 6.612
30 3.889 4.455 4.799 5.048 5.242 5.401 5.536 5.653 5.756 5.849 5.932 6.008 6.078 6.143 6.203 6.259 6.311 6.361 6.407
40 3.825 4.367 4.696 4.931 5.114 5.265 5.392 5.502 5.559 5.686 5.764 5.835 5.900 5.961 6.017 6.069 6.119 6.165 6.209
60 3.762 4.282 4.595 4.818 4.991 5.133 5.253 5.356 5.447 5.528 5.601 5.667 5.728 5.785 5.837 5.886 5.931 5.974 6.015
120 3.702 4.2 4.497 4.709 4.872 5.005 5.118 5.214 5.299 5.375 5.443 5.505 5.562 5.614 5.662 5.708 5.750 5.790 5.827
3.643 4.12 4.403 4.603 4.757 4.882 4.987 5.078 5.157 5.227 5.290 5.348 5.400 5.448 5.493 5.535 5.574 5.611 5.645
Logarithmic transformations help in cases where effects are multiplicative by converting them into additive effects, which simplifies analysis and mean separation . This is important for properly interpreting data and ensuring valid comparisons. Checking for homogeneity of experimental errors (homoscedasticity) is crucial as unequal variances, known as heteroscedasticity, can bias results and invalidate statistical tests . Homogeneity ensures that all treatment variances are equivalent, which is fundamental for the assumptions underlying analysis of variance (ANOVA). Inconsistent variance, such as in data with non-normal distribution, can skew significance tests, necessitating transformations for accurate results .
A Completely Randomized Design (CRD) is appropriate when experimental units are homogenous and environmental effects are easily controlled, such as in laboratory or greenhouse experiments . Its limitations include increased variation among plots in field experiments due to heterogeneity in factors like soil fertility and slope, making CRD rarely used in such scenarios . Additionally, while CRD simplifies data handling in cases of missing data, it treats the variation among plots as experimental error, which can lead to biased results in the presence of substantial environmental variability .
The Triple Lattice Design offers advantages in terms of increased precision compared to traditional designs like the Completely Randomized Design or Randomized Complete Block Design. It is particularly beneficial in large experiments with numerous treatments where including all treatments within a block isn’t feasible . The design improves precision by balancing out variability within blocks and using an adjustment factor to correct treatment totals, minimizing bias in the treatment sum of squares, enhancing the reliability of results . Most beneficial in contexts with a large number of treatments and potential heterogeneity, the Triple Lattice Design facilitates more precise comparisons by accommodating multiple sources of variation in a systematic and efficient manner .
In experimental research, factors such as loss of data due to unforeseen circumstances or experimental errors necessitate the estimation of missing data to maintain the dataset's completeness and reliability . Estimation processes help maintain data integrity by allowing the recovery of incomplete datasets, enabling analysis of variance and other statistical tests to proceed without bias. This ensures that missing data doesn't disproportionately affect results, as incomplete datasets can skew error estimates and lead to misinterpretation of treatment effects . Techniques like imputing missing values using row, column, and treatment totals are employed to adjust the sum of squares for a balanced analysis .
A Latin Square Design is used when there are two known sources of variation that need to be controlled simultaneously, unlike a Randomized Complete Block Design (RCBD), which addresses only one. It is particularly beneficial when the number of treatments equals the number of replications . The advantage of a Latin Square Design is its greater precision in controlling variation across two directions, often referred to as row and column blocking, which results in better isolation of treatment effects by accounting for variability . This leads to increased precision over both CRD and RCBD when appropriate and is useful in experimental conditions where controlling two directional gradients is necessary, such as soil fertility and irrigation conditions .
Blocking in a Randomized Complete Block Design (RCBD) involves grouping experimental units into blocks based on a known source of variation such as soil heterogeneity or animal characteristics, with each block containing all treatments . This technique reduces experimental error by minimizing variability within blocks and maximizing it among blocks, thereby isolating and removing the contribution of the known source of variability from the experimental error . RCBD offers greater precision than a Completely Randomized Design (CRD) by controlling within-block variation . Its flexibility allows the inclusion of extra replications for certain treatments, and it simplifies analysis even when data from some blocks or treatments are missing . However, its effectiveness can diminish when the number of treatments is large, which may lead to large within-block variation .
Replication in experiments improves precision by increasing the number of times a treatment is applied, which narrows the estimates of treatment means closer to true values, and provides multiple observations to estimate experimental error, thus enhancing the reliability of results . The number of replications required is determined by the desired precision level, variability of experimental units, the number of treatments, and the type of experimental design . High precision requires more replications, while more uniform experimental units require fewer replications compared to variable units . More treatments generally need fewer replications than experiments with fewer treatments .
Systematic designs can violate the assumption of independence of errors when treatments are assigned non-randomly, potentially leading to dependencies where the error of one treatment affects another, such as through environmental spillover effects like pesticide drift . This violates the fundamental assumption needed for valid ANOVA results. Randomization plays a crucial role in mitigating this by ensuring treatments are independently assigned to experimental units, thereby minimizing the risk of correlated errors and enhancing the robustness of error independence . Proper randomization ensures that variations due to systematic errors are averaged out across treatments, thus maintaining the validity of statistical inferences.
Homoscedasticity refers to the condition where experimental errors (variances) among treatments are uniform, meaning they have the same variance across all levels . This assumption is critical for validly applying ANOVA as it ensures that the treatment variance is adequately estimated and that any observed differences are due to actual treatment effects rather than unequal error variance . In cases where variances are not equal (heteroscedasticity), it may indicate that certain treatments have unusually high or low variances, potentially affecting the reliability of conclusions drawn from statistical tests . Ensuring homoscedasticity is crucial to maintain the integrity of statistical comparisons across treatments.
Strategies to reduce experimental error include increasing the size of the experiment, refining experimental techniques, utilizing blocking, and applying replication. Increasing the size of the experiment through more replicates or additional treatments enhances the precision of mean estimates and provides multiple observations to estimate experimental error . Refining techniques ensures uniform application of treatments and controls external influences . Blocking helps to measure contributions of extraneous factors to total variability by dividing the field into homogenous parts . Replication enhances precision, provides error estimates, and increases the inference scope due to its repetition over time and locations, allowing more generalized conclusions . Collectively, these strategies ensure comparability and reliability of the experiment's results by minimizing variation not due to the treatment effects.