Statistics notes for Botany Optionals
The variation or deviation of the different values of variable form the average is known as
dispersion.
For example, consider the yield of two crops A & B.
Crop:A(yield per plant in gms): 5,6,7
Crop: B (yield per plant in gms) 2,6,10
If mean alone is taken to explain the data, then it will be concluded that the two crops are equal
with respect to yield per plant (because the mean value of crop A and Crop B are equal to 6).
Actually the two crops are different when the variation of observations is concerned. Therefore,
measure of central tendency alone is not sufficient to explain a data. The variation should also be
considered simultaneously.
Similarly, measure of variation alone will not explain a data completely. For example, consider the
two sets of data which relate to the yield per plant (in gms) of two crops A & B.
Crop A: 100,102,103
Crop B: 1,2,3
The variation is equal to 1 gm in both the crops A and B. This means that the yield of the two crops
are equal (with respect to variation). But it is evident that means are different. Therefore, measure
of central tendency and measure of dispersion should go hand in hand to explain a data.
Absolute measure of dispersion
A measure of dispersion is expressed in the same unit in which the original data are given called
absolute measure of dispersion. By using this measure of dispersion the variability of two or more
distributions with same units can be compared.
Types of absolute measure of dispersion.
1. Range
2. Quartile deviation
3. Mean deviation
4. Standard deviation
Relative measure of dispersion
An absolute measure of dispersion is expressed as a percentage of measure of central tendency
(Mean, Median, Mode etc.,) gives relative measure of dispersion. It is unit free measurement. By
using this measure of dispersion the variability of two or more distribution with different units of
measurements can be compared.
[Link] 1 of 8
Type of relative measure of dispersion
1. Co – efficient of variation
2. Co – efficient of Quartile deviation
3. Co – efficient of Mean deviation
4. Co – efficient of Range.
Standard deviation:
It is defined as the positive square root of the mean of the squared deviations from mean. The
square of S.D is called variance. It posses almost all the characteristics of good measure of
dispersion.
Relative measure of dispersion
Case 1 : Raw data
2
S.D = Sqrt ( (∑(x – x ) ) / n - 1 )
n-1 = degree of freedom
2 2
S.D = Sqrt ( (∑x –( ( ∑x ) )/n ) / n - 1 )
n = total number of observation
This is known as variable square method.
Case 2 : Discrete data
2 2
S.D = Sqrt ( ∑fx –( ∑fx ) ) / n - 1 )
n = total frequency
Case 3 : Continuous data
2 2
S.D = Sqrt ( (∑fd –( ( ∑fd ) )/n ) / n - 1 ) x i
d= ( m – A ) / i
i= common class interval
m=midpoint
A= Assumed value.
[Link] of variation:
It is the most important relative measure of dispersion. It is the ratio of S.D to mean expressed as a
percentage. It is used to compare the variability of different sets of data having different units of
measurement. It is used to know the reliability of data. For example, in an experiment the C.V. % of
yield data should be within 5% and 15% .If it is not in this range the data will not be reliable to
make conclusion of the experiment.
Note: Various formulas to calculate the above measure of dispersion can be referred with the
practical record note.
[Link] 2 of 8
Co-efficient of Variation:
CV (%) = ((SD) / Mean) X 100
CORRELATION AND REGRESSION
Correlation
If the change of one variable affects the change of another variable then the variables are said to
be correlated.
The systematic interrelationship between the variables is termed as correlation
Eg. i) Price and demand
ii) Population and unemployment
iii) Yield and pest incidence
Positive correlation
An increase in one variable may cause an increase in other variable, or a decrease in one variable
may cause a decrease in the other variable. In other words if the movement of the variables are in
the same direction then they are said to be positively correlated. It is also called direct correlation.
E.g. Yield and fertilizer application, Incidence of pests and weather factors.
Negative correlation
If the movements of the variable are in opposite direction then the variables are said to be
negatively correlated. It is also called indirect correlation.
E.g. Yield and pest incidence
Scatter Diagram
In correlation problems, first we have to investigate whether there is any relation between the
variables say, X and Y. For this purpose we use scatter diagram.
Let (X1, Y1), (X2, Y2), (X3, Y3) . . . . . . . . . . . . (Xn, Yn) be n pairs of observations. If the values of
variables X and Y are plotted along the X axis and Y axis respectively in the XY plane of graph
sheet, the resultant diagram of dots is known as scatter diagram. From the scatter diagram we can
say whether there is any correlation between X and Y, correlation is positive or negative and the
correlation is linear or curve linear.
Advantages:
1. Scatter diagram method is easy to understand and simple to follow.
2. It does not involve too much of mathematics.
3. It helps to get a preliminary idea of correlation in a data.
Disadvantages:
1. It gives only the direction but not magnitude of correlation between the two variables.
[Link] 3 of 8
Correlation coefficient
The scatter diagram will give only a vague idea about the presence or absence of correlation and
the nature of the correlation. It will not indicate about the strength or degree or relationship
between two variables. The index of the degree of relationship between two variables is known as
correlation coefficient. It can be determined by the following formula
where , SPxy = [ Σ xy – (Σx Σy) / n ]
2 2
SSx = [ Σ x – ( Σ x ) / n) ]
2 2
SSy = [ Σ y – ( Σ y ) / n) ]
The two variables X & Y are said to be positively, negatively correlated according as r is positive or
[Link] r is equal to zero then the two variables are said to be uncorrelated.
Properties of correlation co-efficient
• ‘r’ will always lie between – 1 & +1
• the value or ‘r’ is unaltered if X & Y are interchanged. This property is called symmetric
property.
• Suppose the values of data on two variables are altered uniformly then the value of ‘r’ will
not be altered. This property is called invariance property of ‘r’.
Regression
Of two variables under study one may represent the cause and the other may represent the effect.
The variable representing the cause is known as independent variable and it is denoted by X. The
variable X is sometimes called predictor variable or regressor. The variable representing the effect
is known as dependent variable and is denoted by Y. The variable Y is sometimes called predicted
variable. The relationship between the independent and dependent variables may be expressed as
a function. Such functional relationship between two variables is termed as regression.
When only two variables are involved the functional relationship is known as simple regression. If
the relationship between the two variables a straight line, it is known as simple linear regression.
Otherwise it is called as simple non linear regression. When there are more than two variables
and one of them is assumed to be dependent upon the others, the functional relationship between
the variables is known as multiple regression.
[Link] 4 of 8
Curve fitting
In general when two variables are studied simultaneously the interest will be to know the
functional relationship between them. The scatter diagram is used to identify the type of
relationship between the two variables for example, when n pairs of observations (X1, Y1), (X12 Y2),.
. . . . . . . . . . . . . . . (Xn, Yn) are plotted on a graph sheet.
The scatter diagram for a given data may result in non – linear forms. The functional relationships
may be given by algebraic expressions like
Y = a + bX
The above algebraic expression is known as curve. The method of finding such relationship is
known as curve fitting. A best fitting line is one for which the sum of the squares of the residuals,
(or errors) is minimum. For this purpose the principle ,known as method of least squares is used.
Method of least square
A best fitting line is one for which sum of the squares of the residuals (errors) is minimum. The
principle which minimise the error in known as method of least squares.
Regression coefficient
In the equation Y = a +bX, ‘b’ is the slope of the line, also called Regression coefficient and ‘a’ is
the intercept of the line with the Y-axis.
bYX = regression coefficient Y on X
bXY = regression coefficient X on Y
Properties of regression coefficients
• Correlation coefficient is the geometric mean between the regression coefficient
• If one of the regression Coefficients is greater than unity, the other must be less than unity.
• Arithmetic mean of regression coefficients is greater than the correlation coefficient.
• Regression coefficient are independent of the origin but not the scale
PROBABILITY AND ITS DISTRIBUTION
The theory of probability is a branch of applied mathematics dealing with the effects of chance.
If there are ‘n’ mutually exclusive, equally likely and exhaustive cases, among which ‘m’ are
favourable to the event ‘E’ then the probability of that event is given by the ratio m/n
P(E) = favourable number of cases / Total number of cases
Binomial distribution
⁃ The simplest probability distribution is binomial distribution. It is used to represent the
probability distribution of discrete random variables.
[Link] 5 of 8
⁃ An experiment which has two possible outcomes is called a Bernoulli trail. The two
outcomes are usually called success or failure.
⁃ An experiment consisting of a repeated number of Bernoulli trails is called a binomial
experiment.
Example1: Consider the population of seeds of a crop.
If the seeds are allowed to germinate, then the observation will be such that each seed has
only any one of the two results viz., ‘germination’ or ‘no germination’. The individuals or
objects in the population are seeds. Therefore distribution of population of seeds is a
binomial distribution as far as the variable ‘germination’ is concerned.
Example 2: Suppose a pathologist wants to know the proportion of sesame plants which will be
affected by phyllody disease.
The individual or object will be a sesames plant in the study. The possible outcomes of
individual(or plant) are any one of the two namely ‘diseased’ or ‘not diseased’.Therefore the
distribution of sesame plants will be a binomial distribution as far as the variable ‘incidence
of disease’ is concerned.
Mean
The means of the binomial distribution with parameters ‘n’ and ‘p’
Mean = np
Variance = npq
Poisson Distribution
The Poison distribution is also used to represent the probability distribution of a discrete random
variable. It is employed in describing random events that occur rarely.
The Poison distribution is given by the probability mass function,
In the formula,
λ = np = mean number of times an event occurs,
x = the number of times the event occurs.
[Link] 6 of 8
Poisson distribution is a limiting case of Binomial Distribution
• When the number of trails is very large(n→∝)
• When the Probability of success is very small(p→0)
• The product np always remains a constant(np→λ)
Then binomial distribution can be approximate to a poison distribution.
Example of Poisson distribution
• The number of typographical errors per page in typed material.
• The number of deaths per day due to a specific disease in a certain town.
• The number of born blind per year in a town.
Normal distribution
⁃ The most important and widely used probability distribution is normal distribution. It is also
known as Gaussian distribution.
⁃ The normal distribution is defined as to represent the probability distribution of a continuous
random variable. Its probability density function is expressed by the relation,
In the above formula,
e = Naperian base equality 2.7183
x = a given value of the random variable in the range --∞ ≤ x ≤ ∞
Properties of Normal distribution
1. The normal curve is perfectly symmetrical about the mean. Also the curve is bell shaped
2. Mean = Median = Mode
3. It has only one mode. It is uni modal.
4. The ordinate at the mean of the distribution divides the total area under the normal curve in to
two equal parts.
5. Skewness is zero.
[Link] 7 of 8
6. Half of the values that is 50% are less than the mean and half of the values are greater than
the mean.
By Ramsundar IFS.
[Link] 8 of 8