1(a) Bioassay: meaning, components, scope and examples
Meaning of bioassay
Bioassay, or biological assay, is an experiment used to estimate the potency, strength, or effectiveness
of a substance by observing its effect on living organisms, tissues, cells, or biological systems.
Components of bioassay
The main components of a bioassay are:
1. Stimulus: The substance or treatment applied to the biological subject.
Example: drug, vitamin, hormone, insecticide, or antibiotic.
2. Subject:The living organism or biological material on which the stimulus is applied.
Example: rat, mouse, chick, insect, bacteria, tissue, or plant.
3. Dose:The amount or concentration of the stimulus given to the subject.
4. Response:The reaction produced by the subject after receiving the stimulus.
Responses may be:
Quantitative response: measured numerically, such as body weight gain, milk production, blood
pressure, growth rate.
Qualitative response: expressed as presence or absence of an effect, such as death/survival,
cured/not cured, germinated/not germinated.
Scope of bioassay
Bioassay is widely used in biological, agricultural, medical, pharmaceutical, and veterinary research.
Important scopes are:
1. Estimation of potency of drugs, hormones and vitamins.
2. Testing effectiveness of antibiotics.
3. Determination of toxicity of insecticides and pesticides.
4. Comparison of two or more biological preparations.
5. Quality control of pharmaceutical and veterinary products.
6. Assessment of response of animals or organisms to chemicals.
7. Evaluation of new chemicals where chemical analysis is difficult or insufficient.
Examples of bioassay
1. Assay of vitamin D by measuring bone healing or growth response in rats or chicks.
2. Assay of insecticides by observing mortality of houseflies, mosquitoes, or other insects after
treatment.
Other examples include insulin assay, antibiotic assay, hormone assay, and toxicity testing of pesticides.
1(b) Standard and test preparations, tolerance, and Direct Assays
Standard preparation
A standard preparation is a biological or chemical preparation whose potency or strength is already
known. It is used as a reference material in bioassay.
The response produced by the standard preparation is compared with the response produced by the
test preparation.
Example: A standard vitamin D solution with known potency.
Test preparation
A test preparation is the preparation whose potency is unknown and is to be estimated by bioassay.
It should contain the same active principle as the standard preparation and should produce the same
type of biological response.
Example: An unknown vitamin D sample whose potency is to be estimated.
Tolerance
Tolerance refers to the ability of a biological subject to withstand a dose of a substance without showing
a specified response or harmful effect.
In bioassay, tolerance is important because different individuals may respond differently to the same
dose. Some subjects may respond at a low dose, while others may need a higher dose. This biological
variation affects the precision of bioassay results.
For example, in an insecticide assay, some insects may die at a low dose, while others may tolerate the
same dose and survive.
Direct Assays
A direct assay is a type of bioassay in which the dose of standard and test preparations is administered
directly to randomly selected, identical subjects until a specified response occurs.
In this method, the dose required to produce the response is measured directly.
The basic idea is:
• Give the standard preparation to one group of subjects.
• Give the test preparation to another similar group.
• Record the amount of dose needed to produce the same biological response.
• Compare the mean doses of standard and test preparations.
• Estimate the relative potency of the test preparation.
If:
𝑥ˉ𝑆 = mean dose of standard preparation
𝑥ˉ𝑇 = mean dose of test preparation
then relative potency may be estimated as:
𝑥ˉ𝑆
𝑅=
𝑥ˉ𝑇
Features of direct assays
1. Subjects are selected randomly.
2. Standard and test preparations are applied to similar subjects.
3. The dose is increased until the required response occurs.
4. The response is measured directly.
5. The method is useful only when the response can be clearly observed.
6. It is difficult to use when the exact tolerance limit of each subject cannot be determined.
Limitations of direct assays
1. It is possible only for certain stimuli and subjects.
2. It is difficult when individual tolerance varies greatly.
3. It may be unsuitable when the response is death, because the subject cannot be reused.
4. Time lag between administration and response may create error.
5. It usually requires carefully matched or identical subjects.
2. Parallel Line Assays
Meaning
A parallel line assay is an indirect bioassay used when the relationship between the response and the
logarithm of dose is linear.
In this assay, both the standard and test preparations are given at different dose levels. The responses
are recorded, and two dose-response lines are fitted:
• one line for the standard preparation,
• one line for the test preparation.
If the two lines are straight and parallel, the relative potency of the test preparation can be estimated.
Basic principle
The response is related to log dose as:
𝑌 = 𝑎 + 𝑏log 𝐷
For the standard preparation:
𝑌𝑆 = 𝑎𝑆 + 𝑏log 𝐷𝑆
For the test preparation:
𝑌𝑇 = 𝑎𝑇 + 𝑏log 𝐷𝑇
Here, both lines have the same slope 𝑏, so they are parallel.
The difference between the two lines represents the difference in potency between the standard and
test preparations.
Assumptions of parallel line assay
The main assumptions are:
1. The response is linearly related to log dose.
2. The dose-response lines of standard and test preparations are parallel.
3. The two preparations produce the same type of biological response.
4. Experimental errors are normally and independently distributed.
5. The variance of responses is homogeneous.
6. Subjects are randomly assigned to treatments.
Design of parallel line assay
Suppose there are:
• 𝑘doses of standard preparation,
• 𝑘doses of test preparation,
• 𝑛subjects per dose.
Then the total number of observations is:
2𝑘𝑛
If the number of doses and replications are equal for standard and test preparations, it is called a
symmetrical parallel line assay.
Common forms are:
• 4-point assay: 2 doses of standard and 2 doses of test.
• 6-point assay: 3 doses of standard and 3 doses of test.
• 8-point assay: 4 doses of standard and 4 doses of test.
Statistical analysis
The analysis is usually done by ANOVA. The total variation is divided into:
1. Preparation variation
Difference between standard and test preparations.
2. Regression variation
Variation due to linear dose-response relationship.
3. Parallelism variation
Tests whether the two dose-response lines are parallel.
4. Linearity variation
Tests whether the dose-response relationship is linear.
5. Error variation
Random experimental error.
ANOVA table for parallel line assay
Source of variation Purpose
Preparation Tests difference between standard and test
Regression Tests linear relationship between response and log dose
Parallelism Tests whether the two lines are parallel
Linearity Tests departure from linearity
Error Measures random variation
Total Total variation
Validity tests
Before estimating relative potency, the assay must satisfy validity tests.
1. Test of regression
The regression should be significant.
This means response changes significantly with dose.
2. Test of parallelism
Parallelism should not be significant.
This means the standard and test dose-response lines are parallel.
3. Test of linearity
Linearity deviation should not be significant.
This means the relationship between response and log dose is linear.
4. Test of preparation difference
A very large preparation difference may indicate unsatisfactory assay precision.
Estimation of relative potency
The relative potency is estimated from the horizontal distance between the two parallel dose-response
lines.
If the test preparation requires a smaller dose than the standard to produce the same response, the test
preparation is more potent.
If it requires a larger dose, it is less potent.
In general:
𝑎𝑇 − 𝑎𝑆
log 𝑅 =
𝑏
where:
𝑅 = relative potency
𝑎𝑇 = intercept of test preparation
𝑎𝑆 = intercept of standard preparation
𝑏 = common slope
Advantages of parallel line assay
1. It is more precise than direct assay.
2. It uses several dose levels.
3. It allows testing of linearity and parallelism.
4. It is useful for estimating relative potency.
5. It can be applied to vitamins, hormones, drugs, antibiotics, and insecticides.
Conclusion
Parallel line assay is an important bioassay method in which the responses of standard and test
preparations are compared over different dose levels. It is valid only when the dose-response
relationship is linear and the two lines are parallel. When these conditions are satisfied, the relative
potency of the test preparation can be estimated accurately.
3(a) Different growth models
Growth models are used to study the trend or growth rate of a variable over time. In
biostatistics/econometrics, the following growth models are commonly used.
1. Linear growth model
The linear growth model assumes that the dependent variable increases or decreases by a constant
absolute amount per unit of time.
𝑌𝑡 = 𝛽1 + 𝛽2 𝑡 + 𝑢𝑡
where,
𝑌𝑡 = value of the variable at time 𝑡
𝑡= time
𝛽1 = intercept
𝛽2 = absolute change per unit time
𝑢𝑡 = error term
The growth rate is:
𝛽2
𝜌= × 100
𝑌ˉ
where 𝑌ˉis the average value of 𝑌.
This model is useful when the variable changes by nearly the same amount every year.
2. Compound growth model
The compound growth model assumes that the variable grows at a constant percentage rate over time.
𝑌𝑡 = 𝑌0 (1 + 𝜌)𝑡
where,
𝑌0 = initial value
𝜌= compound growth rate
𝑡= time
Taking natural logarithm:
ln 𝑌𝑡 = ln 𝑌0 + 𝑡ln(1 + 𝜌)
or,
ln 𝑌𝑡 = 𝛽0 + 𝛽1 𝑡 + 𝑢𝑡
The compound growth rate is:
𝜌 = 𝑒 𝛽1 − 1
or in percentage,
𝜌 = (𝑒 𝛽1 − 1) × 100
This model is widely used in population, production, finance and agricultural growth studies.
3. Exponential growth model
The exponential growth model is also a log-linear model. It assumes that growth takes place
continuously at a proportional rate.
ln 𝑌𝑡 = 𝛽0 + 𝛽1 𝑡 + 𝑢𝑡
Here, 𝛽1 represents the exponential growth rate.
The approximate growth rate is:
𝛽1 × 100
If 𝛽1 = 0.027, then growth rate is approximately:
0.027 × 100 = 2.7%
This model is suitable when the variable grows rapidly over time.
4. Kinked exponential growth model
The kinked exponential growth model is used when there is a structural break in the growth pattern at a
particular year 𝑘.
For example, growth before and after a policy change, technological change, or economic reform may
be different.
The model is:
ln 𝑄𝑡 = 𝜇1 + 𝛽1 (𝐷1 𝑡 + 𝐷2 𝑘) + 𝛽2 (𝐷2 𝑡 − 𝐷2 𝑘) + 𝜖𝑡
where,
𝐷1 = 1before the break year and 0otherwise
𝐷2 = 1after the break year and 0otherwise
𝑘= break point
𝛽1 = growth rate before the break
𝛽2 = growth rate after the break
The hypothesis tested is:
𝐻0 : 𝛽1 = 𝛽2
If 𝐻0 is rejected, it means that a structural change has occurred and the growth rates of the two periods
are significantly different.
Thus, the kinked exponential model is useful for comparing growth rates before and after a structural
change.
3(b) Comparing two regressions using dummy variable approach
To compare two regressions, we pool the observations of two groups or periods and use a dummy
variable.
Let:
𝑫𝒊 = 𝟏
for group or period I, and
𝑫𝒊 = 𝟎
for group or period II.
The pooled regression model is:
𝒀𝒊 = 𝜶𝟏 + 𝜶𝟐 𝑫𝒊 + 𝜷𝟏 𝑿𝒊 + 𝜷𝟐 (𝑫𝒊 𝑿𝒊 ) + 𝒖𝒊
For 𝑫𝒊 = 𝟎:
𝑬(𝒀𝒊 ∣ 𝑫𝒊 = 𝟎, 𝑿𝒊 ) = 𝜶𝟏 + 𝜷𝟏 𝑿𝒊
This is the regression for group II.
For 𝑫𝒊 = 𝟏:
𝑬(𝒀𝒊 ∣ 𝑫𝒊 = 𝟏, 𝑿𝒊 ) = (𝜶𝟏 + 𝜶𝟐 ) + (𝜷𝟏 + 𝜷𝟐 )𝑿𝒊
This is the regression for group I.
Interpretation
• 𝜶𝟐 measures the difference between intercepts.
• 𝜷𝟐 measures the difference between slopes.
• If 𝜶𝟐 is significant, the intercepts are different.
• If 𝜷𝟐 is significant, the slopes are different.
• If both are not significant, the two regressions are not significantly different.
Test
The null hypothesis is:
𝑯𝟎 : 𝜶𝟐 = 𝜷𝟐 = 𝟎
If 𝑯𝟎 is rejected, the two regressions are significantly different. This dummy variable method gives the
same conclusion as the Chow test but is simpler.
Conclusion
Thus, by using a dummy variable and an interaction term, two regression lines can be compared in a
single pooled regression model. The dummy variable tests the difference in intercepts, while the
interaction term tests the difference in slopes. This method is easy and gives conclusions similar to the
Chow test.
4. Design of experiments, Latin square design and analysis
Design of experiments
Design of experiments is the statistical plan of conducting an experiment so that treatment effects can
be measured accurately. It involves proper randomization, replication and local control to reduce
experimental error and obtain valid conclusions.
Latin square design
Latin square design (LSD) is an experimental design in which experimental units are arranged in rows
and columns, and each treatment occurs once in each row and once in each column. It is used when
two sources of variation are to be controlled.
In this experiment:
• Treatments = 4 feeds
• Row factor = age group
• Column factor = breed
• Observations = 4 × 4 = 16
So, the design used is a 4 × 4 Latin Square Design.
Linear model
𝑌𝑖𝑗(𝑘) = 𝜇 + 𝜌𝑖 + 𝛾𝑗 + 𝜏𝑘 + 𝑒𝑖𝑗𝑘
where,
𝑌𝑖𝑗(𝑘)
= milk yield receiving 𝑘th feed in 𝑖th row and 𝑗th column
= general mean
𝜌𝑖
= effect of 𝑖th age group
𝛾𝑗
= effect of 𝑗th breed
𝜏𝑘
= effect of 𝑘th feed
𝑒𝑖𝑗𝑘
= random error
ANOVA table
Source of variation d.f. S.S. M.S. F Sig.
Age group 3 2.000 0.667 0.131 0.938
Breed 3 10.500 3.500 0.689 0.591
Feed 3 186.000 62.000 12.197 0.006
Error 6 30.500 5.083 — —
Corrected Total 15 229.000 — — —
Testing of significance
For feed effect
Null hypothesis:
𝐻0 : 𝑓1 = 𝑓2 = 𝑓3 = 𝑓4
Since,
𝑝 = 0.006 < 0.01
the feed effect is highly significant. Therefore, the four feeds differ significantly in milk production.
For breed effect
Null hypothesis:
𝐻0 : 𝑏1 = 𝑏2 = 𝑏3 = 𝑏4
Since,
𝑝 = 0.591 > 0.05
the breed effect is not significant. Therefore, breeds do not differ significantly in milk production.
Identification of the best feed
Estimated feed means:
Feed Mean milk yield
Feed 1 8.75
Feed 2 10.25
Feed 3 12.25
Feed 4 17.75
The highest mean is for Feed 4 = 17.75 litres.
Therefore, Feed 4 is the best feed.
Estimated breed means:
Breed Mean milk yield
Breed 1 12.25
Breed 2 11.00
Breed 3 12.50
Breed 4 13.25
Breed 4 has the highest mean, but breed effect is not significant, so no breed can be declared
statistically best.
Conclusion
The experiment was conducted in a 4 × 4 Latin square design. Feed effect was significant, but breed
effect was not significant. Among the feeds, Feed 4 produced the highest milk yield and may be
considered the best feed.
5(a) Factorial experiment
A factorial experiment is an experiment in which two or more factors, each at two or more levels, are
studied simultaneously to observe their main effects and interaction effects. It is not a design by itself;
it can be conducted in CRD, RBD or LSD.
For example, if two factors 𝐴and 𝐵are each at two levels, the experiment is called a:
22 factorial experiment
The four treatment combinations are:
𝑎0 𝑏0 , 𝑎1 𝑏0 , 𝑎0 𝑏1 , 𝑎1 𝑏1
or commonly written as:
1, 𝑎, 𝑏, 𝑎𝑏
Advantages of factorial experiment over single factor experiment
1. It studies more than one factor in the same experiment.
2. It estimates main effects of all factors.
3. It estimates interaction effects between factors.
4. It is more efficient, because several factors are tested together.
5. It saves time, labour and cost compared with separate single factor experiments.
6. It gives more complete information about the combined effect of factors.
7. The hypotheses for main effects and interactions can be tested separately.
Linear model for 𝟐𝟐 factorial experiment in Randomized Complete Block Design
Let factor 𝐴have two levels and factor 𝐵have two levels, with 𝑟blocks.
The linear model is:
𝑌𝑖𝑗𝑘 = 𝜇 + 𝛼𝑖 + 𝛽𝑗 + (𝛼𝛽)𝑖𝑗 + 𝜌𝑘 + 𝜖𝑖𝑗𝑘
where,
𝑖, 𝑗 = 0,1; 𝑘 = 1,2, … , 𝑟
𝑌𝑖𝑗𝑘 = observation from 𝑖th level of 𝐴, 𝑗th level of 𝐵, in 𝑘th block
𝜇= general mean
𝛼𝑖 = effect of factor 𝐴
𝛽𝑗 = effect of factor 𝐵
(𝛼𝛽)𝑖𝑗 = interaction effect of 𝐴and 𝐵
𝜌𝑘 = effect of 𝑘th block
𝜖𝑖𝑗𝑘 = random error
ANOVA table for 𝟐𝟐 factorial experiment in RBD
Source of variation d.f. Sum of squares Mean square F-ratio
Block 𝑟−1 SS(Block) MS(Block) —
Treatment 3 SS(Treatment) MS(Treatment) 𝑀𝑆𝑇 /𝑀𝑆𝐸
Factor A 1 SS(A) MS(A) 𝑀𝑆𝐴 /𝑀𝑆𝐸
Factor B 1 SS(B) MS(B) 𝑀𝑆𝐵 /𝑀𝑆𝐸
A × B interaction 1 SS(AB) MS(AB) 𝑀𝑆𝐴𝐵 /𝑀𝑆𝐸
Error 3(𝑟 − 1) SSE MSE —
Total 4𝑟 − 1 SS(Total) — —
Partition of total variation
𝑆𝑆(𝑇𝑜𝑡𝑎𝑙) = 𝑆𝑆(𝐵𝑙𝑜𝑐𝑘) + 𝑆𝑆(𝐴) + 𝑆𝑆(𝐵) + 𝑆𝑆(𝐴𝐵) + 𝑆𝑆(𝐸𝑟𝑟𝑜𝑟)
Here,
𝑆𝑆(𝐴) + 𝑆𝑆(𝐵) + 𝑆𝑆(𝐴𝐵) = 𝑆𝑆(𝑇𝑟𝑒𝑎𝑡𝑚𝑒𝑛𝑡)
Tests of significance
For factor 𝐴:
𝑀𝑆𝐴
𝐹𝐴 =
𝑀𝑆𝐸
For factor 𝐵:
𝑀𝑆𝐵
𝐹𝐵 =
𝑀𝑆𝐸
For interaction 𝐴𝐵:
𝑀𝑆𝐴𝐵
𝐹𝐴𝐵 =
𝑀𝑆𝐸
If the calculated 𝐹-value is greater than the tabulated 𝐹-value, the corresponding factor or interaction is
significant.
Conclusion
Thus, a 22 factorial experiment studies two factors at two levels each. In RBD, the total variation is
divided into block, factor 𝐴, factor 𝐵, interaction 𝐴𝐵, and error. Its main advantage over a single factor
experiment is that it can estimate both main effects and interaction effects in one experiment.
5(b) Multiple comparison test and LSD test
Multiple comparison test
After ANOVA, if the F-ratio is significant, we conclude that all treatment means are not equal. But
ANOVA does not tell which treatment means differ. So, we compare the treatment means pairwise to
identify the best treatment.
This process is called multiple comparison, and the tests used for this purpose are called multiple
comparison tests. Examples are LSD test, DMRT and Tukey’s test.
Least Significant Difference test / LSD test
The Least Significant Difference test is the oldest and simplest multiple comparison test. It is also called
the multiple t-test because it is based on Student’s 𝑡-test.
It is used to compare two treatment means after the ANOVA F-test is significant.
Formula
For equal replications,
2𝑀𝑆𝐸
𝐿𝑆𝐷 = 𝑡𝛼,𝑒𝑟𝑟𝑜𝑟 𝑑𝑓 × √
𝑟
where,
𝑀𝑆𝐸 = error mean square
𝑟 = number of replications
𝑡𝛼,𝑒𝑟𝑟𝑜𝑟 𝑑𝑓 = tabulated t-value at desired level
At 5% level,
2𝑀𝑆𝐸
𝐿𝑆𝐷5% = 𝑡0.05 × √
𝑟
At 1% level,
2𝑀𝑆𝐸
𝐿𝑆𝐷1% = 𝑡0.01 × √
𝑟
The notes also call 𝐿𝑆𝐷1% as MSD, or most significant difference.
Procedure of LSD test
1. First perform ANOVA.
2. If the treatment F-ratio is significant, calculate LSD.
3. Arrange treatment means in descending or ascending order.
4. Find the difference between every pair of treatment means.
5. Compare each mean difference with LSD.
Decision rule:
∣ 𝑌ˉ𝑖 − 𝑌ˉ𝑗 ∣> 𝐿𝑆𝐷
then the two treatment means are significantly different.
If,
∣ 𝑌ˉ𝑖 − 𝑌ˉ𝑗 ∣≤ 𝐿𝑆𝐷
then the two treatment means are not significantly different.
Interpretation
If two means differ by more than LSD, they are significantly different. If the difference is less than LSD,
they are statistically identical.
The treatment with the highest mean, which is significantly superior to others, is considered the best
treatment.
Conclusion
Thus, LSD is a simple multiple comparison test used after significant ANOVA to identify which treatment
means are significantly different and to select the best treatment.
6(a) Split-plot design
A split-plot design is a two-factor experimental design in which one factor is applied to large
experimental units, called whole plots, and another factor is applied to smaller units within each whole
plot, called subplots or split plots. It is used when one factor requires larger units or when more
precision is desired for the second factor and its interaction.
Example from animal husbandry
In a dairy experiment, types of milking machines may be assigned to whole plots because they require a
large amount of milk or animals. Then methods of milk cooling or pasteurization may be assigned to
subplots because they require smaller quantities of milk.
So,
• Whole plot factor: type of milking machine
• Subplot factor: cooling or pasteurization method
• Response: milk quality or milk yield
Advantages
1. Larger experimental units can be used for factors that are difficult to apply on small units.
2. Greater precision is obtained for the subplot factor.
3. The interaction between whole plot and subplot factors can be studied with good precision.
4. It is useful when one factor is more important than the other.
Disadvantages
1. Whole plot factor is measured with less precision.
2. Analysis is more complicated than RBD or CRD.
3. There are two error terms: Error I and Error II.
4. Missing data create more difficulty in analysis.
Linear model
Suppose factor 𝐴has 𝑝levels and is assigned to whole plots, factor 𝐵has 𝑞levels and is assigned to
subplots, and there are 𝑟blocks.
𝑌𝑖𝑗𝑘 = 𝜇 + 𝜌𝑖 + 𝛼𝑗 + 𝜖𝑖𝑗 + 𝛽𝑘 + (𝛼𝛽)𝑗𝑘 + 𝜖𝑖𝑗𝑘
where,
𝑖 = 1,2, … , 𝑟
𝑗 = 1,2, … , 𝑝
𝑘 = 1,2, … , 𝑞
𝑌𝑖𝑗𝑘 = observation from 𝑗th level of 𝐴and 𝑘th level of 𝐵in 𝑖th block
𝜇= general mean
𝜌𝑖 = effect of 𝑖th block
𝛼𝑗 = effect of whole plot factor 𝐴
𝛽𝑘 = effect of subplot factor 𝐵
(𝛼𝛽)𝑗𝑘 = interaction effect of 𝐴and 𝐵
𝜖𝑖𝑗 = whole plot error or Error I
𝜖𝑖𝑗𝑘 = subplot error or Error II
Partition of total sum of squares
𝑆𝑆(𝑇𝑜𝑡𝑎𝑙) = 𝑆𝑆(𝐵𝑙𝑜𝑐𝑘) + 𝑆𝑆(𝐴) + 𝑆𝑆(𝐸𝑟𝑟𝑜𝑟 𝐼) + 𝑆𝑆(𝐵) + 𝑆𝑆(𝐴𝐵) + 𝑆𝑆(𝐸𝑟𝑟𝑜𝑟 𝐼𝐼)
Here, Error I is used to test whole plot factor 𝐴, and Error II is used to test subplot factor 𝐵and
interaction 𝐴𝐵.
ANOVA table for split-plot design
Source of variation d.f. Mean square F-ratio
Block 𝑟−1 MS(Block) —
Factor A 𝑝−1 MS(A) 𝑀𝑆(𝐴)/𝑀𝑆(𝐸𝑟𝑟𝑜𝑟 𝐼)
Error I (𝑝 − 1)(𝑟 − 1) MS(Error I) —
Factor B 𝑞−1 MS(B) 𝑀𝑆(𝐵)/𝑀𝑆(𝐸𝑟𝑟𝑜𝑟 𝐼𝐼)
A×B (𝑝 − 1)(𝑞 − 1) MS(AB) 𝑀𝑆(𝐴𝐵)/𝑀𝑆(𝐸𝑟𝑟𝑜𝑟 𝐼𝐼)
Error II 𝑝(𝑞 − 1)(𝑟 − 1) MS(Error II) —
Total 𝑝𝑞𝑟 − 1 — —
6(b) Analysis of covariance / ANCOVA
Analysis of covariance, or ANCOVA, is a statistical technique that combines the features of analysis of
variance and linear regression. It is used to compare treatment effects after adjusting the response
variable for the effect of one or more related variables called covariates or concomitant variables.
For example, in a feeding trial on goats, final weight gain may depend on the initial body weight of
goats. Here,
• Response variable 𝑌: weight gain
• Covariate 𝑋: initial body weight
• Treatment: different feeds
ANCOVA adjusts the weight gain for initial body weight before comparing feed effects.
When ANCOVA can be applied
ANCOVA can be applied when:
1. There is a primary response variable 𝑌.
2. There is one or more covariates 𝑋related to 𝑌.
3. The relationship between 𝑌and 𝑋is linear.
4. The covariate is not affected by the treatment.
5. The covariate is measured before or along with the experiment.
6. Errors are independent, normally distributed and have common variance.
7. The regression of 𝑌on 𝑋is independent of treatment, meaning the regression slope is common
for all treatments.
Linear model of ANCOVA
For CRD, the model is:
𝑌𝑖𝑗 = 𝜇 + 𝜏𝑖 + 𝛽𝑋𝑖𝑗 + 𝜖𝑖𝑗
or,
𝑌𝑖𝑗 = 𝜇∗ + 𝜏𝑖 + 𝛽(𝑋𝑖𝑗 − 𝑋ˉ.. ) + 𝜖𝑖𝑗
where,
𝑌𝑖𝑗 = response variable
𝑋𝑖𝑗 = covariate
𝜇= general mean
𝜏𝑖 = treatment effect
𝛽= common regression coefficient
𝜖𝑖𝑗 = random error
Difference between ANOVA and ANCOVA
ANOVA ANCOVA
Compares treatment means without covariate Compares treatment means after adjusting for
adjustment. covariates.
Uses only the response variable. Uses response variable and one or more covariates.
Partitions variation into treatment, block and Partitions variation and also removes variation due to
error components. covariate.
Does not use regression. Combines ANOVA with regression.
Compares raw treatment means. Compares adjusted treatment means.
Error variation may be larger. Error variation is usually reduced.
Suitable when experimental units are Suitable when experimental units differ in some
homogeneous. measurable initial character.
Conclusion
Thus, ANCOVA is used when an additional variable affects the response. By adjusting for this covariate,
ANCOVA gives a more precise comparison of treatment effects than ordinary ANOVA.
Questions from 2024-25
5. Discuss vital statistics. What are the uses of it? How can you measure a population? Define crude
death rate, age-specific death rate, infant mortality rate, standardized death rate and general fertility
rate.
Vital statistics
Vital statistics is the branch of statistics which deals with data relating to births, deaths, marriages,
migration, morbidity and mortality of human population. It helps to study the size, structure, growth
and health condition of a population.
Uses of vital statistics
1. To study population trend, growth and composition.
2. To calculate birth rate, death rate, fertility rate and mortality rate.
3. To help in public health planning and disease control.
4. To help government in administration and policy making.
5. To compare mortality and fertility between different regions or countries.
6. To prepare life tables and estimate expectation of life.
7. To help insurance companies in calculating life risk and premium.
Measurement of population
Population at any time may be measured by census or estimated from vital events.
If,
𝑃0 = population at last census
𝐵 = births, 𝐷 = deaths
𝐼 = immigration, 𝐸 = emigration
then population at time 𝑡is:
𝑃𝑡 = 𝑃0 + (𝐵 − 𝐷) + (𝐼 − 𝐸)
Here,
𝐵 − 𝐷 = natural increase
𝐼 − 𝐸 = net migration
For rate calculation, usually mid-year population or average population is used.
Crude Death Rate (CDR)
Crude death rate is the number of deaths per 1000 population in a year.
𝐷
𝐶𝐷𝑅 = × 1000
𝑃
where,
𝐷= total deaths during the year
𝑃= mid-year population
It gives the general level of mortality in a population.
Age-Specific Death Rate (ASDR)
Age-specific death rate is the death rate for a particular age group.
𝐷𝑥
𝐴𝑆𝐷𝑅 = × 1000
𝑃𝑥
where,
𝐷𝑥 = deaths in age group 𝑥
𝑃𝑥 = population in age group 𝑥
It is better than CDR for comparing mortality because mortality differs by age.
Infant Mortality Rate (IMR)
Infant mortality rate is the number of deaths of infants below one year of age per 1000 live births in a
year.
𝐷0
𝐼𝑀𝑅 = × 1000
𝐵
where,
𝐷0= deaths of children below 1 year
𝐵= total live births during the year
It is an important indicator of health and socio-economic condition.
Standardized Death Rate (STDR)
Standardized death rate is an adjusted death rate used to compare mortality between two populations
having different age structures.
By direct method:
∑𝑚𝑥 𝑃𝑥𝑠
𝑆𝑇𝐷𝑅 =
∑𝑃𝑥𝑠
where,
𝑚𝑥 = age-specific death rate
𝑃𝑥𝑠 = standard population in age group 𝑥
It removes the effect of different age compositions and gives a fair comparison of mortality.
General Fertility Rate (GFR)
General fertility rate is the number of live births per 1000 women of reproductive age, usually 15–49
years.
𝐵
𝐺𝐹𝑅 = × 1000
𝑊15−49
where,
𝐵= total live births during the year
𝑊15−49 = number of women aged 15–49 years
It is more refined than crude birth rate because it considers only women in the reproductive age group.
Conclusion
Vital statistics are essential for studying population growth, mortality, fertility and public health. Rates
such as CDR, ASDR, IMR, STDR and GFR help to compare population conditions and guide planning and
policy decisions.
4. What is indirect assay based on quantitative response?
Indirect assay based on quantitative response
An indirect assay based on quantitative response is a type of bioassay in which specified doses of
standard and test preparations are given to subjects, and the measurable responses are recorded.
Here, the response is quantitative, meaning it can be measured numerically, such as:
• change in body weight,
• growth response,
• degree of healing,
• change in analytical value,
• time of survival.
In this assay, the relationship between dose and response is first established. Then the dose required to
produce a particular response is estimated from the dose-response relationship.
The relative potency of the test preparation is obtained by comparing it with the standard preparation.
Main features
1. Doses are fixed before the experiment.
2. Responses are measured quantitatively.
3. Dose-response relationship is established.
4. Potency is estimated indirectly from the response curve.
5. Both standard and test preparations are compared.
Types
Indirect assays based on quantitative response are mainly of two types:
1. Parallel line assay
2. Slope-ratio assay
Thus, this assay is called indirect because potency is not measured directly, but estimated from the
dose-response relationship.