Statistical Methods and Applications Overview
Statistical Methods and Applications Overview
For
STAT-211 (Statistical Methods)
Year-II, Sem-I, Academic Year: 2020-21
1|Page
Lecture No. 1: Introduction to Statistics and Its Application
Statistics
Statistics is the science of application of mathematics on collected data.
OR Meaning of statistics Statistics is concerned with scientific methods for collecting,
organising, summarising, presenting and analysing data as well as deriving valid
conclusions and making reasonable decisions on the basis of this analysis.
Father of Statistics: R. A. Fisher
Father of Indian Statistics: P.C. Mahalonobis
The word statistics is generally used in two different ways.
(1) When it is used in plural, it means the quantitative data affected to a marked extent
by a multiplicity of causes. When we say `collect statistics’ means collect the
numerical data which are to be analyzed and interpreted. e.g.
(i) Wheat production affected by various causes.
(ii) Collect the data of height of the students of second semester.
(2) When it is used in singular, it means "the science of collecting, classifying and using
the data for further statistical treatments". It involves the methods of analysis used in
the analysis and interpretation of data and they are known as statistical methods.
Definitions:
(1):It is a study of population, variation and the methods for reduction of the data.
Biometry: When the principles of statistics are applied on living thing or organisms, the
science is called biometry.
Statistical methods: The methods by which statistical data are analyzed are called
statistical methods.
Aims of studying statistics
(1) To study the population: The study of population of any kind agricultural activities on
the basis of sample data.
(2) To understand the nature of variability e.g. Height of plants. The biological
phenomena observed under one set of conditions are never duplicated exactly under
another set of similar conditions. Therefore, repetition of experiment is necessary to
account all the factors causing variation. In biological phenomena were variation is a
rule rather than exception it is this function that has wide application.
2|Page
(3) To express the facts in summary form (the facts that are based on large number of
observations) e. g. It is not possible for one to form a precise idea about the income
position of the population of India from the records of individuals.
(4) To provide correct method(s) for taking sample (sampling).
(5) To provide proper method for comparison of two or more things.
(6) It helps in prediction/forecasting the yield of a particular crop for a particular year on
the basis of the past records.
Limitations of Statistics
Statistics with its wide applications in almost every sphere of human activity is
not without limitations. The following are the important limitations.
1) It does not deal with individual.
2) It deals only with quantitative characters.
3) Statistical results are true only on an average.
4) Statistics can be misused.
5) It does not reveal the entire story.
6) Expert knowledge is must to handle the statistical data.
(1) It does not study individuals
Statistics deals with an aggregate of objects and does not give any specific
recognition to the individual items of a series. e. g. the individual figures of
agricultural production of any country for a particular year are meaningless unless,
to facilitate comparison, similar figures of other countries or of the same country of
different years are given. Height of Mr. X is 5' 8" does not constitute statistical
statement. The average height of an Indian is 5' 8".
(2) It deals only with quantitative characters.
Efficiency, honesty, intelligence factors can be measure indirectly e. g. efficiency of
selling agent can be judged by studying the no. of articles sold by him.
(3) Statistical results are true only on an average.
Average consumption of milk per head in a certain locality is 1/2 liter but it does
not give any idea of the shortage of milk faced by the poor. The conclusions
obtained statistically are not universally true; they are true only under certain
conditions. This is because statistics as a science less exact as compared to natural
sciences.
3|Page
(4) Statistics can be misused.
Because if conclusions are based on incomplete information, statistics can prove
anything. There are three types lies: lies, dammed lies and statistics. Statistics are
like clay of which one can make a God or Devil as he please.
Importance/Application in agricultural research
(1) It helps to understand nature of variability or differences.
(2) To arrive at the meaningful conclusion on the basis of sample study in the field.
(3) Express the data/result of the field experiment in summary form.
(4) Sampling
a) In state agricultural survey, for estimation of areas and yield of crops.
b) In price fixation policy of various agricultural commodities.
c) In agricultural extension survey, to study the impact of programs.
d) In agricultural economics survey, to study the demand-supply policy, the
growth rate of population and cost of production of various crops.
(5) In agricultural meteorology for weather forecasting and to correlate weather
parameters with crop production.
Variable : The characteristics which show variation or variability are called variables or
variats e.g. cabbage yield, wheat yield per hectares of the growers. Variable
can of two types,
(i) Qualitative : The characteristics which can not be measured numerically or in terms
of magnitude e.g. flower color, nature of surface.
(ii) Quantitative: The characteristics which can be measured in terms of magnitude e.g.
yield of crop, height, weight. The quantitative characteristics are of
two types.
(a) Discrete : Character which takes only integer values/or whole value. There is a
definite gap between two values. e.g. No. of students in a class, No. of
bacteria in given area.
(b) Continuous: The quantity which can take any numerical value within a certain range.
Height, weight (They are in fraction and there is no definite gap between values).
4|Page
Lecture 2. Graphical Representation:
Graph: A display of point and line. In a graph of pair value (x, y), so called the co-
ordinate of o point, are plotted on a graph pepper suitably choosing the scales along X-
axis (abscissa) and Y-axis (ordinate). The plotted points are joined by the straight line in
their sequence of occurrence. The figure so obtain are called graph.
Advantage of Graphical or Diagramatics representation:
1. Diagram give as bird’s-eye view of complex data.
2. They have long lasting impression
3. Easy to understand even by common man
4. They save time and Labour.
5. They facilitate comparison
One dimensional diagram or graph
Bar diagram and Line Diagram
Different Types of Bar diagram are.
1. Simple bar chart
2. Multiple bar chart
3. Sub-divided bar chart
Two dimensional diagram or graph
Circle, Rectangle, Pie Chart
Three Dimensional Diagram
1. Cubes, Cylinder and spheres
Frequency Distribution
Objectives:
1. To condense the mass of data in such a manner that similarities and dissimilarities can
be easily understand.
2. To enable statistical treatment to the data collected.
Frequency: The no. or individual of items occurring in each class is termed as frequency.
Frequency distribution: The manner in which the frequencies are distributed over the
different class is called frequency distribution of the character under study and the table
indicating frequency distribution is called frequency table.
5|Page
Class limit: It is the lowest and highest values of the distribution that can be included in
the class e.g 10-20, 20-30 etc. Two boundaries of a class are known as the lower limit and
upper limit of a class.
Class interval: The width of a class that is the difference of upper and lower limit of the
class is known as class interval.
Class mid point: It is the value lying half way between the lower limit (LL) and upper
limit (UL) of a class interval i.e. (LL + UL)/2.
Points while deciding class interval/classes:
1. It should be of uniform width which facilitates the statistical computation.
2. Range of the class should cover the data and should be continuous.
3. It should be convenient to make the mid-point of a class.
4. It should not be over lapping.
Types of frequency distribution:
(1) Discrete frequency distribution
(2) Continuous frequency distribution
Methods of classifying the data according to class interval:
Exclusive method: When the class intervals are so fixed that the upper limit of one class
is the lower limit of the next class. This method is known as exclusive method e.g. 0-10,
10-20. Usually this method is preferred for continuous type of data. The data observed
up to 9.99 would be included in 0-10 class while 10 or greater than 10 will be included in
10-20 class.
Inclusive method: In this method of classification, the upper limit of one class is
included in that class itself e.g. 100-199, 200-299. The value of 100 and 199 will be
included in the class of 100-199. This method is preferred for discrete type of data.
**Procedure to form frequency distribution:
Step 1: Find range of the data. Range = Highest value – Lowest value.
Step 2: Fix the number of classes. Number of classes should preferably between 5 to 15
and should not be less than 5 and more than 30.
Approximate no. of classes = K =1 + 3.322 log N (Sturge’s rule) where N = no. of
observations under study.
Step 3: Fix the class interval = CI = Range/No. of classes or (L-S)/K where L = largest
value and S = smallest value
Step 4: Arrange different classes in ascending order of magnitude
6|Page
Step 5: Pick up the values of observation and make tally mark against respective classes.
Step 6: Find total tally mark of each class which will give the no. of frequencies in the
respective classes.
Graphical Representation
Graphical representation is used when we have to represent the data of a
frequency distribution and of a time series. It is represented by points which are plotted
on a graph paper.
Advantages of graphical representation
1. Easy to understand and interpret data at a glance.
2. It facilitates comparisons.
3. It gives eye view of complex data.
4. It has long lasting impression.
5. It gives an attractive and interesting view.
Limitations of graphical representation
1. It cannot show all those facts which are there in the tables.
2. It shows tendency and fluctuations, actual values are not known.
3. The charts take more time to be drawn than the tables.
Graphs of Frequency Distribution
Histogram: It is a bar diagram which is suitable for frequency distributions with
continuous classes. The width of all bars is equal to class interval and heights of the bars
are in proportion to the frequencies of the respective classes. In this diagram bars touch
each other but one bar never overlaps the other.
Frequency polygon: When the mid points of the tops of the adjacent bars of a histogram
are joined in order by a straight line, then the graph of lines so obtained is called a
frequency polygon.
7|Page
Frequency curve: A frequency curve is a graphical representation of frequencies
corresponding to their variates values by a smooth curve. A smoothened frequency
polygon represents a frequency curve.
Ogive or Cumulative frequency curve: it is a graph plotted for the variates values and
their corresponding cumulative frequencies and joined by a free hand smooth curve. The
curve is ‘S’ shaped. There are two methods of constructing ogive viz. (i) “less than”
method and (ii) “more than” method. In the “less than” method, we start with the upper
limit of classes and go on adding the frequencies however, in the “more than” methods,
we start with the lower limit of classes. The first method gives a rising curve whereas
second method shows a declining curve.
Example1:
Following are the make obtain by 60 students in an examination of 100 mark
85 38 55 49 52 60 65 42 31 30
29 18 60 32 13 80 9 15 51 40
54 59 90 69 36 95 5 59 66 70
69 28 50 59 61 41 85 76 30 45
58 47 20 75 56 52 60 23 18 48
45 72 86 66 42 28 35 37 60 80
8|Page
Prepare frequency table of the data and draw Histogram, Frequency polygon and ogive.
[no. of
Soln:
Step 1. Range= Highest –lowest value = 95-5 = 90
Here, Highest value = 95, Lowest value= 5;
Step 2. Number of class, K=1+3.322* log60 = 6.90 ~7
Range 90
Step 3. Class interval= 12.86 13
Numberof class 7
For convenience, the class interval is taken as 10 instead of 9. The lowest class may be
started from 0 instead of 5 and similarly to include the highest value (95) the highest class
is taken with a upper limit 100.
For For Polygon For Ogive
Histogram
Class Tally mark Frequency Mid-value of Cumulative
(X) X Frequency
5-18 1111 4 11.5 4
18-31 1111 1111 1 9 24.5 13
31-44 1111 1111 11 10 37.5 23
44-57 1111 1111 1111 12 50.5 35
57-70 11111111111111 14 63.5 49
70-83 1111 11 6 80.5 55
83-96 1111 11 5 93.5 60
TOTAL N=60
Histogram
16
Number of student (Frequency)
14
12
10
8
6
4
2
0
5 18 31 44 57 70 83 96
Classe
s
9|Page
Frequency Polygon
16
ogive
70
60
Cummulative Frequency
50
40
30
Y-Values
20
10
0
0 20 40 60 80 100
Class mid Values
10 | P a g e
Lecture 3: Measures of Central Tendency
Different groups of data or statistical series or frequency distributions differ in
four characteristics.
1) Central tendency or location 3) Skewness or symmetry
2) Dispersion or variation 4) Kurtosis or peaked ness
Central tendency: A values of the variable tend to cluster around a central value or
centrally located observation of the distribution. This characteristic is known as central
tendency.
This centrally located value which represents the group of values is termed as the
measure of central tendency e.g. an average is called measure of central tendency.
Objectives
1) To get one single value that describes the characteristics of the entire
series/group.
2) To compare two or more distributions.
11 | P a g e
(1) Arithmetic mean or Mean
It is the most common and ideal measure of central tendency. It is defined as the
sum of the observed values of the character (or variable) divided by the number of
observations considered in obtained sum (total).
_
Symbolically X: Sample mean
: Population mean
X i
X i1
n
Raw data or ungrouped data
X i
X i1
n
Grouped data
fX i i
X i 1
k
f i 1
i n
(X-5)2 4 1 0 1 4 0 1 11
x 3 4 5 4 3 5 4 Product=14400
log x 0.477 0.602 0.698 0.602 0.4771 0.698 0.602 Sum=4.158
1/x 0.333 0.25 0.2 0.25 0.333 0.2 0.25 Sum=1.8166
(3 4 ..... 4) 28
Mean ( X ) = 4
7 7
GM= (3x4……x4)1/7 = (14400)1/7 = 3.926 or
12 | P a g e
n
log xi
Anti log( i 1 = Antilog( log 3 log 4...... log 4
GM=
n n
0.477 0.602...... 0.602
Anti log Anti log( 0.594) 3.926
7
AM˃GM˃HM
13 | P a g e
Merits and demerits of Arithmetic mean
Merits
1) It is rigidly defined.
2) It is based on all the observations.
3) It is readily comprehensible.
4) It is easy to calculate.
5) Its algebraic (Mathematical) treatment is especially easy and definitely possible.
6) It is also least affected by the fluctuation of sampling.
Demerits
1) It is affected by the extreme values.
2) If there is large variation in the data then A.M. becomes some times meaning less.
3) It is not used to measure rate of growth or rate of speed directly.
Uses: It is most popular and simple estimate and used widely in almost all the fields
of studies such as social science, economics, business, agriculture, medical sciences,
engineering and such other sciences.
Weighted Mean
When different observations are to be given different weights, arithmetic mean
does not prove to be a good measure of central tendency. In such cases weighted mean is
to be calculated.
If X1, X2, X3,... Xn are different observation and W1, W2, W3,... Wn are their
respective weights then,
W.M.
W1X1 W2 X 2 ...... Wn Xn
=
W X i i
W1 W2 ...... Wn W i
Uses
1) Used when the number of individuals in different classes of grouped widely
varying.
2) Used when the importance of all the items in a series is not same.
14 | P a g e
3) Used when the ratios, percentages or rates e.g. rupees per kilogram, rupees per
meter etc. are to be averaged.
4) Weighted mean is particularly used in calculating birth rates, death rates, index
numbers, average yield etc.
Geometric Mean
AM gives equal weightage to all the items and has got a tendency towards the
higher values. Sometimes it is necessary to get average having a tendency towards the
lower values. In such case, geometric mean is helpful. It is defined as the n th root of the
product of n items of a series with following relation.
It is useful in ratio and Proportion data
Raw data or ungrouped data
n
log X i
GM n X 1 , X 2, .....X n = X 1 , X 2 ,..... X n
1
i 1
n = Antilog ( )
n
Grouped data
GM n X 1f1 , X 2f 2 ,..... X kf k = X1f1 , X 2f 2 ,.....X kf k 1
n
k
=
1
f1 log X1 f 2 log X 2 ..... f k log X k = 1 fi log X i
n n 1
15 | P a g e
Harmonic Mean
HM is the reciprocal of the arithmetic mean of the reciprocal of the values of a
variable or series.
n n
HM
1 1 1
1
..... x
x1 x 2 x n i
Grouped data
HM
f i
f i
f1 f2 f f
..... k x i
x1 x 2 xn i
2) GM= AM HM
16 | P a g e
Median: The median is the middle most items that divide the series into two equal parts
when ith items are arranged in ascending or descending order.
n 1
th
In case of raw data, the median is the term where n (odd) is the total no. of
2
observations whereas in case of n even number medial is the average between
th th
n n
and 1 terms, in case of grouped data it is given by the formula,
2 2
n/2 cf
Median l xCI
f
where, l = lower limit of the median class in which (n/2)th item falls
cf= cumulative frequency of the class preceding the median class
f = frequency of median class
CI= Class interval
Uses: It is useful when the extreme values of the series are either not available or
impossible to be obtained or abnormal. When in a group, the individual is
denoted by better than half the individual’s, median is used. It is also useful
when the items are not susceptible to measurement in definite units e.g.
intelligence, ability, efficiency etc.
Uses: It is useful when the extreme values of the series are either not available or
impossible to be obtained or abnormal. When in a group, the individual is
denoted by better than half the individual’s, median is used. It is also useful
when the items are not susceptible to measurement in definite units e.g.
intelligence, ability, efficiency etc.
** Median is the best method if data is suffer from outlier
Mode: The value of the variable which occurs most frequently or whose frequency is
maximum is known as mode.
Uses: Business forecasting is particularly based on the modal values. Meteorological
forecasting is also based on modal values.
Relation between Mean, Median and Mode:
Mode = 3Median – 2Mean
17 | P a g e
Lecture No. 4: Measures of Dispersion
Mean is though an important concept in statistics, it does not give a clear picture
as to how the different observations are distributed in a given distribution or the series
under study. Consider the following series.
Series Observations Mean
1 2, 3, 4, 7 4
2 4, 4, 4, 4 4
3 1, 1, 2, 12 4
4 3, 4, 4, 5 4
In the above series, the mean is same i.e. 4 but the spread of the observations
about the mean is in different manner. Hence after locating the measures of central
tendency, the next point is to find out the center. This can be done by measuring the
spread. The spread is also called a Scatter, Variation OR dispersion of the variate values.
Definition
Dispersion may be defined as the extend of the scatterness of observations around
a measure of central tendency and a measure of such scatter is called measures of
dispersion.
Different measures of dispersion
1) Range.
2) Absolute mean deviation or Absolute deviation (A.M.D.)
3) Standard deviation (S)
4) Variance (S2)
5) Standard error of mean ([Link].)
6) Coefficient of variation (C.V. %)
Requisite/Characteristics of an ideal measure of dispersion
(X i X) 2
S i 1
n 1
(2) Variable square method
n
n
( X i ) 2
X 2
i i 1
n
S i 1
n 1
n
( d i ) 2
d 2
i i 1
n
S i 1
di = Xi – A A = Assumed mean
n 1
f (X i i X) 2 k
S i 1
n 1
where n f
i 1
i ; fi= Frequency of ith class
19 | P a g e
k
k
( f i X i ) 2
f X i
2
i i 1
n
S i 1
n 1
(3) Assumed mean method
k
k
( f i d i ) 2
f d i
2
i i 1
n
S i 1
di = (Xi – A)
n 1
k
( f i d Xi ) 2
f d i
2
Xi i 1
n
S i 1
dxi = (Xi - A)/I
n 1
N1S12 N2 S 22 N1d12 N2 d 22
S12 =
N1 N2
where, S12 = Combined standard deviation; S1 = Standard deviation of first group;
N2 are the numbers of observation for series one and two; X 1 = A.M. for first
(2) The sum of squares of the deviations of items in the series from their arithmetic
mean is minimum. This is the reason why standard deviation is always computed
from the arithmetic mean.
(3) Addition or subtraction of a constant from the grouped of an observation will not
change the value of S.D.
(4) Multiplying or dividing each observation of a given series by a constant value will
multiply or divide the std. deviation by the same constant.
Variance: Variance is the square of standard deviation. It is also called the “Mean square
deviation". It is being used very extensively in analysis of variance of results from field
experiment. Symbolically denoted by
20 | P a g e
S2 = Sample variance and 2 = Population variance.
Method of computation
Raw data or ungrouped data
(1) Deviation method
(Xi 1
i - X )2
S2 =
( n 1)
(2) Variable square method
X X i / n
2 2
S .S .
2 i
S =
n 1 d. f .
where : Xi = Variate value S S = Sum of square
n = No. of observations. df = Degrees of freedom
d d i / n
2 2
2 i
S
n 1
where di = Xi - A
A = Assumed mean
Grouped data or frequency distribution
(1) Deviation method
f (X
i 1
i i - X )2
S2 =
(n 1)
2 i i
S
n 1
2 i i
S where, di = (Xi - A)
n 1
A = Assumed mean
fI = Frequency of ith class
21 | P a g e
1) If V(x) represent the variance of X series and V(y) represent the variance of Y series
then V (xy) = V(x) + V(y)
i.e. V(x+y) = V(x) + V (y) and V(x-y) = V(x) + V(y)
2) Multiplying or dividing each observation by a constant will multiply or divide the
variance by square of that constant.
e.g. V(ax) = a2 V(x).
3) Addition or subtraction a constant from the groups of each observation will not
change the value of variance.
Standard error of mean ( [Link].)
The standard deviation is the standard error of a single variate where as standard
error of mean is the standard deviation of sampling distribution of the sample mean OR
it refers to the average magnitude of difference between the sample estimate and
population parameter taken over all possible samples from the population.
Definition : It is defined as square root of the ratio of the variance to the total no. of
observations in a given set of data.
Symbolically, it is written as S X for sample and X for population.
S
SX where, S = Standard deviation; n = No. of observations
n
For statistical analysis work, the use of S X is common. It is also used to provide
22 | P a g e
The series for which the C.V.% is greater is said to be more variable or we say
less consistence, less homogeneous or less stable while the series having lower C.V.%
is called more consistence or more homogeneous.
(Measure of dispersion)
The yield/plot of two varieties X and Y in eleven fields in a certain season are as
under.
Identify, which variety is more consistent in the performance.
X 8 6 7 6 8 7 5 8 9 6 7
Y 9 4 3 8 9 9 8 4 9 6 8
Solution:
Sl. Xi xi- X (xi- X )2 Yi Y Y Y Y
i i
2
X i
(8 6 7 ........ 6 7) 77
Now, X i1
= 7
n 11 11
X
n
2
X
i
(1 1 0 ...... 4 1 0) 14
i
1.4
n 1 11 1 10
X
n
2
i X
SDX i
1.4 1.183
n 1
Variance (S2 x)= (1.183)=1.4
SD S 1.183
Coefficient of Variance of X (CV)% = 100 100 100 16 .90
Mean X 7
n
Y i
(9 4 3 ......... 6 8) 77
Now, Y i 1
7
n 11 11
23 | P a g e
Y Y
n
2
i
(4 9 16 ....1 1) 54
i
5.4
n 1 11 1 10
Y Y
n
2
i
SDy i
5.4 2.323
n 1
Variance of Y (S2 y)= (1.183)=5.4
SD S 2.323
Coefficient of Variance (CV)% = 100 100 100 33 .197
Mean Y 7
Standard deviation (S), Variance (S2) and Coefficient of Variation (C.V) of X variety are
smaller than Y variety, Hence, variety X is more stable than Y
24 | P a g e
Lecture 5: Probability
Statistics concerns itself with inductive reasoning/inference based on the
mathematics of probability.
Sampling variation needs support in terms of probability / reliability of inference.
Introduction:
In the study of a population, one can not make any firm statements about the
population concerned or its parameters, when only a sample investigation of the
population is available for scrutiny. Due to the sampling variation there is doubt about
sample investigation and hence it is practice to make statements of less definite nature in
terms of probability or chance. The probability or chance for any statement depends on
the number of favourable, unfavourable and total possible cases. e.g. in a tossing a coin
for getting a head, one has to consider that there are two equally likely cases, head and
tail, one is in favour of the statement and the other is against it.
The theory of probability aims to generalize the laws of chance, to discover the
regularities in the pattern in which events, depending on chance, repeat themselves. It
may be the tossing of a coin, a game of cards or the genetical ratios which may be the
object of our investigation.
Jacob Bernoullis, an Italian mathematician was the first to give concept and
definition of probability in 1713. The work of Gregor Mendel in Genetics showed that the
theory of probability could be applied to biological investigations.
Before going to study Probability, we should have clear Knowledge about set
theory.
Set:
A set is the collection of all individual numbers in a well defined space. In a set all the
member belonging to the defined particular space is counted once only. A member can
not be counted twice or more in a set. It is usually represented in flower braces.
For example:
Set of natural numbers = {1,2,3,…..}
Set of whole numbers = {0,1,2,3,…..}
The set that contains all the elements of a given collection is called the universal set and is
represented by the symbol ‘U’.
25 | P a g e
Subset
A set A is said to be subset of another set B if and only if every element of set A is also a
part of other set B. Denoted by ‘⊆‘. ‘A ⊆ B ‘ denotes A is a subset of B.
To prove A is the subset of B, we need to simply show that if x belongs to A then x also
belongs to B. To prove A is not a subset of B, we need to find out one element which is
part of set A but not belong to set B.
Union
Union of the sets A and B, denoted by A ∪ B, is the set of distinct element belongs to set
A or set B, or both.
Intersection
The intersection of the sets A and B, denoted by A ∩ B, is the set of elements belongs to
both A and B i.e. set of the common element in A and B.
26 | P a g e
Above is the Venn Diagram of A ∩ B.
Example: Consider the previous sets A and B. Find out A ∩ B. Solution : A ∩ B = {3, 4}.
Disjoint
Two sets are said to be disjoint if their intersection is the empty set .i.e sets have no
common elements.
For Example
Let A = {1, 3, 5, 7, 9} and B = { 2, 4 ,6 , 8}. A and B are disjoint set both of them have no
common elements.
Set Difference
Difference between sets is denoted by ‘A – B’, is the set containing elements of set A but
not in B. i.e all elements of A except the element of B.
27 | P a g e
Addition of sets A and B, referred to as Minknowski Addition, is the set in whose
elements are the sum of each possible pair of elements from the 2 sets (that is one element
is from set A and other is from set B).
Set subtraction follows the same rule, but with the subtraction operation on the elements.
It is to be observed that these operations are operable only on numeric data types. Even if
operated otherwise, it would only be a symbolic representation without any significance.
Further, it can be seen easily that set addition is commutative, while subtraction is not.
Definition of Probability:
Probability is a ratio of the number of “favourable” cases to the total number of
equally likely cases. If probability is denoted by P then
Suppose a coin is tossed, the possible outcomes (events) are head and tail. These
are equally likely and mutually exclusive events. The probability (P) of event head is 1/2.
Range of the P is 0 to 1
28 | P a g e
If n is the number of equally likely and mutually exclusive events for an event A,
of which m is the favourable to its occurrence, then the probability of A is the fraction
m/n."
P (A) = m/n
29 | P a g e
If two events A and B are mutually exclusive with probabilities P1 and P2
respectively, then the probability of occurrence of either of them (A or B) is equal to the
sum of the individual probabilities (A and B).
Proof: If an event A can happen in m1 ways and B in m2 ways, then the number of ways
in which either event can happen is m1 + m2. If the number of possibilities is n, then by
definition the probability of either the first or the second event happening is
m1 m2 m1 m2
P(A or B) = = P(A) + P(B) = P1 +P2
n n n
m1 m2
where P(A) P1 ; P(B) P2
n n
If K events are mutually exclusive with individual probabilities P1, P2, ... ,Pk then
P(anyone among K mutually exclusive events) = P1 + P2 + ...+ Pk.
Example: If A is the event drawing an ace from a pack of cards and B is the event
drawing a king, then P(ace = A) = 4/52 and P (king = B) = 4/52. The probability of
drawing either an ace or a king in a single draw is
Since both ace and king can not be drawn in a single draw and are thus mutually
exclusive events.
From the above explanation, one can point out two facts. They are:
Similarly we can generalize the rule for more than two events also.
i.e. P(A + B + C) = P(A) + P(B) + P(C) – P(AB) - P(AC) – P(BC) + P(ABC)
If two events A and B are dependent then the probability of both happening at a time is
given as follows:
P(AB) = P(A) . P (B/A) This is the conditional probability
or = P(B) . P(A/B)
Where P(B/A) means the probability of second event B dependent on the probability of
first event A.
In above, if P(B/A) = P(B) then A and B are independent events.
Example: Suppose a box contains 3 white balls and 2 black balls. Let A be the event “first
ball drawn is black” and B the event “second ball drawn is black”, where the balls are not
replaced after being drawn. Here A and B are dependent events.
2 2
P(A) probability of drawing first black ball.
32 5
32 | P a g e
1 1
P(B) P(B/A ) the probability of second black ball given the first
3 1 4
ball drawn is black
Then P(A.B) = P(both black) = 2/5 . 1/4
= 2/20 = 1/10
Similarly we can generalize the rule for more than one dependent event.
P(A.B.C) = P(A).P(B/A).P(C/AB)
33 | P a g e
Lecture No. 6. Binomial &Poisson Distributions
Probability Distribution
It is also called parent population distribution, theoretical distributions or
theoretical frequency distribution. In previous chapter, the probability of the occurrence
of a single event is obtained. In scientific research using statistical methodology, it is
often required to obtain the probabilities of occurrence of all possible events. A table of
the possible values (Xi) which a chance event may assume with a corresponding
probability distribution for each value is called a probability distribution for the
population. Following table gives the probability distribution of sum of two unbiased
dice.
Table: Probability distribution of sum of two dice.
____________________________________________________________________________
Xi 2 3 4 5 6 7 8 9 10 11 12
____________________________________________________________________________
fi 1 2 3 4 5 6 5 4 3 2 1
___________________________________________________________________________
pi 1/36 2/36 3/36 4/36 5/36 6/36 5/36 4/36 3/36 2/36 1/36
___________________________________________________________________________
k
pi = 1 pi = f(x) = f(xi)
i=1
Instead of a table of values such as above, one can represent the outcomes (pi) by
proper mathematical function over a range of Xi. In this chapter we would like to
describe three theoretical distribution
(i) Binomial distribution - James Bernoulli (1700)
(ii) Poisson distribution - S.D. Poisson (1857) and
BINOMIAL DISTRIBUTION
Binomial distribution was discovered by James Bernoulli in 1713.
This is very important distribution dealing with discrete variable. The binomial
distribution has two parameters viz. n and p. In other words, it is completely determined
by the values of n and p.
Let a random experiment be performed repeatedly and let the occurrence of an event
in any trial be called a success and its non-occurrence a failure. Consider a series of n
34 | P a g e
independent Burnoullian traials(n being finite, in which the probability ‘p’ of success in
any trial is constant for each trial. Then q=1-p is the probability of failure in any trial.
The probability of x successes and consequently (n-x) failures in n independent
trials, in a specified order (say) SSFSFFFS ….. FSF (where S represent success and F
failure) is given by the compound probability theorem by the expression:
P(SSFSFFFS ….. FSF)=P(S)P(S)P(F)P(S)P(F)P(F)P(F)P(S)…..P(F) P(S)P(F)
=p.p.q.p.q.q.q.p…..q.p.q
= p.p…..p… … ...q.q.q…….q
35 | P a g e
Conditions for using Binomial distribution
(1) The outcome or results of each trial in the process are characterized as one of two
types of possible outcomes.
(2) The possibility of outcome of any trial does not change and is independence of the
results of previous trials.
Use
It is useful in describing an enormous variety of real life events.
POISSON DISTRIBUTION
The Poisson distribution is the limiting form of the binomial probability
distributions n become infinitely large and p approaches 0 in such a way that np = m
remained constant. Such situation are fairly common. That is to say, a Poisson
distribution may be expected in cases were the chance of any individual event being
a success is small. e.g. occurrence of comparatively rare event, such as serious floods,
percentage infestation of any diseases etc.
Like binomial distribution, the variate of the Poisson distribution is also a
discrete one. The probability functions is
e-m mx
P(x) = ---------
x!
36 | P a g e
μ32 m 2 1
(1) β1 3
μ2 m
3
m
μ 4 m 3m 2 1
(2) β 2 3
μ22
m 3
m
Use
Poisson distribution is used in practice in wide variety of problems where there are
infrequently occurring events with respect to time, area, volume or similar unit. For
example it is used in quality control statistics to count the number of defects of an
item, or in biology to count number of bacteria, insects etc.
37 | P a g e
Lecture 7. Correlation Analysis
So far we have studied problems relating to one variable only. In practice we
come across a large number of problems involving the use of two or more than two
variables.
Univariate population
A population that is characterized by a single variable is termed as univariate
population e.g. population of height of students, weight, yield etc.
Bivariate population
When two variables are simultaneously studied in a single population is termed
as bivariate population e.g. the height and weight of the students, rainfall and yield, the
amount of fertilizer used and the crop yield.
If two quantities vary in such a way that movement in one are accompanied by
movements in the other, these quantities are said to be correlated e.g. price of
commodities and amount demanded, increase in rainfall up to a point and production of
crop. The degree of relationship between the variables under consideration is measured
through the correlation analysis.
Correlation
It indicates the association between the two or more variables in a bivariate
distribution or an analysis of the covariation of two or more variables is usually called
correlation.
Types of correlation
Correlation is described or classified in several different ways. Three of the most
important ways of classifying correlation are:
i) Positive or negative
ii) Simple, partial and multiple
iii) Linear and non-linear
Positive and negative correlation
Whether correlation is positive or negative would depend upon the direction of
change of the variable. If both the variables are varying in the same direction i.e. if as one
variable is increasing the other on an average is also increasing, correlation is said to be
positive. Eg: a) Hight and Weight b) Yield and fertilizer, c) Income and Expenditure. If,
on the other hand the variable is varying in opposite directions, i.e. as one variable is
38 | P a g e
increasing the other is decreasing or vice-versa, correlation is said to be negative. Eg:
Yield and Disease Incident, b) Price and supply
Positive correlation
X: 10 12 15 18 20 X: 80 70 60 40 30
Y: 15 20 22 25 37 Y: 50 45 30 20 10
Negative correlation
X: 20 30 40 60 80 X: 100 90 60 40 30
Y: 40 30 22 15 10 Y: 10 20 30 40 50
If the plotted points fall in a narrow band there would be a high degree of
correlation between the variables.
Y x Y x
x x x x
x x xx
xx x x
x x x x
x x xx
X X
High degree +ve r High degree -ve r
Rainfall and yield Intensity of diseases and yield
If the points are widely scattered over the diagram, it is the indication of very little
relationship between the variables.
Y x x Y x x x
x x x x xx
x x x x x x
x x x x x x
x x x x x x
x x x x
X X
Imperfect +ve Imperfect -ve
(Low degree of positive correlation) (Low degree of negative correlation)
If the points lie on a straight line parallel to the X-axis or in a haphazard manner it
shows absence of any relationship between the variables e.g. height of students and
marks.
Simple, Partial and Multiple correlation
When only two variables are studied it is a problem of simple correlation. When
three or more variables are studied it is a problem of either multiple or partial correlation.
39 | P a g e
In multiple correlation, three or more variables are studied simultaneously. In partial
correlation, we recognize more than two variables, but consider only two variables to be
influencing each other, the effect of other influencing variables being kept constant.
Linear and Non-linear (curvilinear) correlation
If the amount of change in one variable tends to bear a constant ratio to the
amount of change in the other variable than the correlation is said to be linear e.g.
X: 10, 20, 30, 40, 50
Y: 70,140, 210, 280, 350
Cov( XY )
=
XY
r
Cov( XY
=
xy =
SP( xy)
S X SY x . y2 2 SS X .SSY
X Y
where, xy XY n
X 2
x 2
X 2
n
Y 2
y 2
Y 2
n
1. A change in an origin does not affect the value of the correlation coefficient.
2. A change in a scale does not affect the value of correlation coefficient.
3. The value of correlation coefficient lies between -1 to +1.
4. Correlation coefficient is unit free.
40 | P a g e
5. Geometric mean of two-regression coefficient is equal to correlation coefficient.
Test of significance of correlation coefficient
Rank Correlation
The Karl Pearson’s method is based on the assumption that the population being
studied is normally distributed. When it is known that the population is not normal, or
when the shape of the distribution is not known there is a need for a measure of
correlation that involves no assumption about the parameters of the population.
This method was developed by Charles Spearman in 1904. This measure is
especially useful when quantitative measures for certain factors can not be fixed e.g. (i)
correlation between marks obtains in two different subjects by the same group of
students. (ii) Correlation of height and weight of the students can be worked out without
making exact measurement. We shall first stand the students according to height; the
same procedure can be utilized for weight for giving ranks. When there are two or more
items are of equal magnitude, their ranks are to be calculated by taking the average of
their ranks.
R 1
6 d i2 1 12 P 3 p or 1
6 d i2
n n2 1 nn 2 1
Where,
yx = Reg. coefficient of Y on X ; xy = Reg. coefficient of X on Y
and c is the intercept
42 | P a g e
Fitting of the regression lines
Intercept Y YX X
In case of the sample data the estimates of yx i.e. byx and the estimate of ‘’ as a
are obtained and placed in the equation.
a Y bYX X
Regression coefficient
Method of computation
YX
X Y
X Y
Cov( XY )
X
2
X
V (X )
bYX
X X Y Y
xy
XY X Y n
X X x X X n
2 2 2
2
43 | P a g e
Similarly,
XY
X Y
X Y
Cov( XY )
Y
2
Y
V (Y )
bYX
X X Y Y
xy
XY X Y n
Y Y y Y Y n
2 2 2
2
1) When our interest is to ascertain whether the effect of the independent variable
on the dependent variable is appreciable or not, we employ 't' test.
Ho : yx = 0
Ha : yx = 0
y xy x
2
b YX 2 2
t SE of b YX
SE of b YX n 2 x 2
The calculated t value is to be compared with the table t value at the desired level of
significance with (n-2) d.f. and conclusion is to be drawn.
SY SX
(i) r b YX .b XY (ii) bYX r (iii) b XY r
SX SY
45 | P a g e
Lecture 9: Introduction to test of Significance
The subject of statistics deals with statistical estimation and testing of
statistical hypothesis. These are the two important functions for drawing inference about
the population parameters. Statistical estimation is the technique of estimating the
population parameter values on the basis of information obtained from the sample.
Suppose we wish to know the yield of a crop. To know this figure, it is not necessary to
harvest entire field of that crop or all the fields of that crop grown in the region. One may
collect the sample from the fields by appropriate sampling procedure and on the basis of
sample information, one may estimate the average yield of the crop of entire area. The
estimate thus obtained is not the final form for drawing valid conclusion regarding
population from which the samples are drawn. It needs to be tested by applying an
appropriate test or method. Such test is known as the test of significance. Thus, test of
significance can be defined as “The statistical procedure for deciding whether the
observed difference between sample estimate and population Parametric value is
significant or not at specified level of significance".
Hypothesis: It is the statement specifying the parametric value of a distribution from
which the sample/s is/are drawn.
Null Hypothesis: It is a hypothesis of no difference between different populations
parametric values from which samples are drawn OR It is the
hypothesis of equality of population parametric values from which
sample/s is/are drawn.
Procedure for testing a hypothesis
Step I: Set appropriate null hypothesis
Let us consider that there are two methods for preparing compost.
Method A: standard method and
Method B: new method
Now to test which method is better, the hypothesis can be
1) B is better than A B > A
2) A is better than B A > B
3) B is not different from A A = B
The first two statements indicate a preferential attitude to one or the other of the
two methods. Hence, it is better to adopt the third statement and make the test. This third
46 | P a g e
statement is called the null hypothesis, which is denoted as Ho: symbolically Ho : 1 = 2
or 1 - 2 = 0 where 1 and 2 are the population parametric values.
In the above examples, suppose in first method the average nitrogen content is 1
and in the second method the average nitrogen is 2. Ho : 1 = 2 can be tested by the
appropriate test. As against the null hypothesis, the alternative hypothesis should also be
set up, which specifies those values, the researcher believes to hold true. Since one is
going to accept or reject the null hypothesis one has to set the alternative hypothesis also.
It is denoted by Ha : 1 2 or Ha : 1 < 2, 1 > 2 .
Step II: Fix appropriate level of significance
The confidence with which an experimenter reject or accept the null hypothesis
depends upon the significance level adopted. It is expressed in percentage such as 5 per
cent, 1 per cent etc. When the hypothesis in question, is accepted at 5 per cent level of
significance, the experimenter is running the risk that in the repeated cases
(experiments/trials), he will be making the wrong decision in about 5 per cent of the
cases. By rejecting the hypothesis at the same level, he runs the risk of rejecting a true
hypothesis in 5 out of every 100 occasions. Thus, level of significance is defined as:
"It is the maximum probability at which one would like to reject the null
hypothesis when it is true OR The level of significance is the average proportion of
incorrect statements made when the null hypothesis is true."
Step III: Set suitable test criterion
To construct a test criterion, one has to select the appropriate probability
distribution for the particular test viz. Z, t, F, 2 etc.
Step IV: Computation
This step involves the calculations of various statistics from sample data such as
mean and standard error of mean.
Step V: Conclusion
After doing the necessary calculations one has to decide whether to accept or
reject the null hypothesis at a certain level of significance. Therefore, the computed value
of the test criterion is compared with the table value. If the computed value is greater
table value, the observed difference is significant and Ho is not accepted. If calculated
value is less than or equal to table value the Ho is accepted at a given level of significance.
Not acceptance of Ho means the difference between sample estimate and the hypothetical
parametric value is a real difference, while acceptance of Ho means the difference
47 | P a g e
between sample estimate and population/hypothetical parametric value can be
explained due to chance variation (sampling variation).
Type - I and Type - II errors
While testing the hypothesis one is liable to commit two kinds of errors.
An error of first kind is made by rejecting the true null hypothesis. The
probability of committing a Type-I error is denoted by (alpha).
Type-II error is committed by accepting the null hypothesis when it is false. The
probability of Type-II error is denoted by (Beta).
Type-I error depends on the level of significance. When 5 per cent level of
significance is fixed, we fixed the probability of committing Type-I error at 5 per cent.
It is possible to control Type-I error by shifting the level of significance. Type-II
error increases as the Type-I error decreases. Therefore, the common practice is to keep
the Type-I error at five percent or one percent fixed and try to decrease Type-II error by
increasing sample size and following refined technique of conducting experiment.
Degrees of freedom
For testing any hypothesis the estimated statistic is compared with table value.
The knowledge of degrees of freedom is essential for referring the table value. With X1,
X2,.....Xn having constant sum, (n -1)X values can be given freely, but the nth X value will
be determined by the condition that the sum of all `X` is equal to the given constant
quantity i.e. one degree of freedom is lost. So in one way classification, number of
observations - 1 is called degrees of freedom and in general number of observation minus
number of independent constraints or restrictions is called degrees of freedom.
48 | P a g e
Lecture No. 10: One sample and two sample t-test for Mean
49 | P a g e
4) It is flatter than the normal distribution i.e. the area near the tail is large for t
distribution compared to normal distribution. Value of coefficient of kurtosis is less
than 3.
5) As sample size increases, the t distribution approaches to normal distribution.
6) There is need to know the d.f. to obtained the probability value from the table.
S
SX = (If population standard deviation is known)
n n
50 | P a g e
X O
t
SX
Step V : Conclusion
If calculated t < table t0.05,(n-1) d.f. observed difference is not significant. Ho: = o is
accepted. Acceptance of Ho: = o means the given small random sample has come from
the hypothetical population having mean o. If calculated t table t0.05,(n-1) d.f. observed
difference is significant. Ho: = o is rejected. If calculated t table t0.01, (n-1) d.f., observed
difference is highly significant. Ho: = o is rejected. Rejection of Ho: = o means the
given random sample does not come from the hypothetical population having mean o.
Where 1 is the population mean from which sample one is drawn and 2 is
the population mean from which the second sample is drawn.
Step II: Fix the level of significance. Usually 5 per cent and 1 per cent levels of
significance are fixed.
51 | P a g e
Step III: Calculate the following estimates.
Sample - I Sample - II
n1 n2
Xi Yi
i1 i1
i) Mean X Y
n1 n2
ii) Variance
n1 n2
( X i - X) 2 (Yi - Y)2
i1 i1
S2x = S 2y =
(n1 1) (n2 1)
1 1
Sp2 ( )
S n1 n2
( X Y)
( X 1) (Y 2 )
t
S( X Y )
Step V: If cal t < Table t 0.05, (n1+n2-2) d.f. difference is non significant at 5% level of
significance Ho: 1 = 2 accepted. Acceptance of Ho: 1 = 2 means both the
samples have came from the same population
If cal t table t0.05,(n1+n2-2) d.f. difference is significant at 5 % level of significant
Ho : 1 = 2 rejected at 5 % level of significance.
If cal t table t 0.01,(n1+n2-2) d.f. difference is highly significant 1 % level of
significance. Ho:1 = 2 rejected at 1% level of significance. Rejection of Ho : 1 =
2 means both the sample are drawn from two different populations.
52 | P a g e
Two sample 't' test ( Dependent sample) : Paired 't' test
Objective: To test whether the two small related random samples have come from the
same population.
(i) di = Xi - Yi i = 1,... ,n
(ii) d d n i
d d
2
2 i
(iii) S
n 1
d
2
S2 d
(iv) S.E. of d
i
n nn 1
d d
t (Under Ho : d = 0 )
Sd
Step V: Conclusion
If cal. t < table t 0.05, (n-1) d.f., observed difference is non-significant at 5% level of
significance and null hypothesis (Ho) is accepted.
Acceptance of null hypothesis (Ho: d = 0) means the given two related small
samples have come from the same population.
If cal. t table t 0.05, (n-1) d.f., observed difference is significant at 5% level of
significance and null hypothesis (Ho) is rejected at 5% level of significance. If
cal. t table t 0.01, (n-1) d.f., observed difference is highly significant at 1% level of
significance and null hypothesis (Ho) is rejected at 1% level of significance.
53 | P a g e
Rejection of null hypothesis (Ho : d = 0) means the given two related samples
does not come from the same population.
54 | P a g e
Step IV : Calculate normal deviate
X O
Z
SX
Sample II : Y1, Y2, ... ... ,Yn2 are random samples drawn from normally distributed
populations
OR
Let Sample - I Sample –II
Class value : X1, X2, ... , Xk1 Y1, Y2, ... , Yk2
Frequency : f1, f2, ... , fk1 = n1 f1, f2, ... , fk2 = n2
are two frequency distributions of two samples drawn from a normal populations.
55 | P a g e
Procedure :
Step I: Set the null hypothesis that both the samples have come from the same
population having mean and standard deviation or its estimate S.
i.e. Ho : 1 = 2 = against Ha : 1 2 or
1 > 2 or 1 < 2
Where 1 is the population mean from which sample one is drawn and 2 is the
population mean from which the second sample is drawn.
Step II : Fix the level of significance. Usually 5 per cent and 1 per cent levels of
significance are fixed.
Step III : Calculate the following estimates.
Sample-I Sample-II
(i) Mean :
n1 n2
Xi Yi
i1 i1
X Y
n1 n2
(ii) Variance :
n1 n2
( X i - X) 2 (Yi - Y)2
i1 i1
S2x = S 2y =
(n1 1) (n2 1)
n1 n2
(X i - X) 2 (Yi - Y) 2
i1 i1
Sp2
n1 n 2 - 2
1 1
Sp2 ( )
S n1 n2
( X Y)
( X 1 ) (Y 2 )
Z
S( X Y )
Step V: Conclusion:
56 | P a g e
If calculated Z 1.96 the observed difference is non significant at 5% level of
significance and Ho: 1 = 2 accepted. Acceptance of Ho: 1 = 2 means both the samples
have came from the same population. If calculated Z > 1.96, the observed difference is
significant at 5% level of significant and Ho : 1 = 2 rejected at 5 percent level of
significance and if calculated Z > 2.58, the observed difference is highly significant at
1% level of significance hence Ho : 1 = 2 rejected at 1% level of significance. Rejection of
Ho: 1 = 2 means both the sample are drawn from two different populations.
57 | P a g e
Lecture 11: Chi-Square Test of Independence of Attributes in 2x 2 Contingency Table
2 - test ( Chi-square test)
Chi-square was introduced by Karl Pearson in the year 1899. It is calculated by
K
( i - E i ) 2
2 =
i 1 Ei
Where, Oi = Observed frequency of ith class
Ei = Expected frequency of ith class
k = number of classes , i = 1,2,..,k
Definition : "It is the sum of the ratio of the square of deviations obtained between
observed and expected frequency to the expected frequency of the
respective class of the frequency distribution."
58 | P a g e
Therefore, the Z value can be worked out by using the following formula and it
should be compared with table Z value at 5 per cent or 1 per cent level of significance.
Z = 2 2 2n 1
Conditions for application of Chi-Square
1) Deviations (Oi - Ei) should be normally distributed.
2) Number of observations should be sufficiently large. It should be at least 50.
3) Expected frequency of any cell should not be very small. It should be at least 5 and
better if it is 10.
Uses
1) Testing goodness of fit.
2) Testing the independence of attributes for 2 x 2, 2 x c, r x 2 and r x c contingency table.
3) Testing the agreement of genetic ratio with the observed ratio.
4) Test of homogeneity of the families.
5) Test for the detection of linkage.
6) Testing of homogeneity of various variances (Bartlett's test of homogeneity)
7) Testing the heterogeneity among correlation coefficients.
1) Testing goodness of fit
When chi-square test is used to know whether the given sampling distribution is
in agreement with the theoretical or expected frequency distribution the test is
known as test of goodness of fit.
Procedure:
Step I : Set the appropriate null hypothesis.
Ho : Given sampling distribution is in the agreement with theoretical or
expected frequency distribution.
Ha : Given sampling distribution is not in agreement with theoretical or
expected frequency distribution.
Step II : Fix the level of significance.
Step III: Work out expected frequency according to given ratio or expectation.
Step IV: Calculate Chi square as
K
( i - E i ) 2
2 =
i 1 Ei
Where, Oi = Observed frequency of ith class
59 | P a g e
Ei = Expected frequency of ith class
k = number of classes , i = 1,2,..,k
Step V: Compare cal 2 with table value at 5% level of significance and (k-1) degree of
freedom.
Step VI: If cal 2 > table 2 0.05, (k-1)d.f. observed difference is significant at 5% level of
significance. Ho rejected.
If cal 2 < table 2 0.05, (k-1) d.f. observed difference is not significant at 5% level
of significance. Ho accepted.
Step VII: Conclusion: Non significance difference indicates that the given sampling
distribution is in agreement with theoretical distribution and the fit is good.
Significant difference indicates that the given sampling distribution is not in
agreement with theoretical distribution and the fit is poor.
2) Test of Independence
Another common use of the chi square test is in testing independence of
classifications.
Independence: The two attributes A and B are said to be independent to each other if the
proportion of A's among B's is the same as that in not - B's.
Variable: Any character which varying from individual to individual is termed as
variable.
Attribute: Attribute is that which is not capable of being described numerically e.g. sex,
blindness, colour, shape.
Contingency table: When the individuals in a sample have two characters or attributes
and a frequency distribution is made classifying them according to both so as
to show the relation between the characters, the resulted table is termed as
contingency table.
2
ad bc .N 2
2
R1 .R2 C1 .C 2
Where, a, b, c and d are the observed frequency of the respective cell R1, R2, C1, and
C2 are the rows and column totals. N is the grand total.
Step V: Compare calculated 2 with table value at 5% level of significance and (r-1) (c-1)
degree of freedom.
Step VI: If cal 2 > table 2 0.05, (r-1)(c-1) d.f., observed difference is significant at 5% level of
significance. Ho rejected.
If cal 2 < table 20.05, (r-1)(c-1) d.f., observed difference is not significant at 5%
level of significance. Ho accepted.
Step VII: Acceptance of Ho means the two characters are independent to each other.
Rejection of Ho means the two characters are not independent to each
other
Step IV: Work out expected frequency of each cell as follows.
R 1C1 R 2 C1
E(a11) = E(a21) =
N N
R1C 2 R 2C2
E(b12) = E(b22) =
N N
In general,
R iC j
E(Xij) =
N
61 | P a g e
Step VI: Compare calculated 2 with table value at 5% level of significance and (r-1) (c-
1) degree of freedom.
Step VII: If cal 2 > table 20.05, (r-1)(c-1) d.f., observed difference is significant at 5% level
of significance. Ho rejected.
If cal 2 < table 20.05, (r-1)(c-1) d.f., observed difference is not significant at 5%
level of significance. Ho accepted.
Step VIII: Acceptance of Ho means the two characters are independent to each other.
Rejection of Ho means the two characters are not independent to each
other
F test
t and Z tests are used for comparing two populations mean. When the population
is to be compared with respect to their variances the F test is used.
Definition: It is the ratio of the estimates of greater mean square or variance to smaller
variance of two different populations.
S12
F ; S12 > S22
S 22
Compare cal F with table F (n1 - 1) and (n2 - 1) d.f. and draw the conclusion.
62 | P a g e
Lecture 12: Introduction to Analysis of Variance and One Way ANOVA
Analysis of Variance (ANOVA):
Analysis of Variance (ANOVA) is a hypothesis-testing technique used to test the equality
of three or more population (or treatment) means by examining the variances of samples
that are taken.
ANOVA allows one to determine whether the differences between the samples are
simply due to random error (sampling errors) or whether there are systematic treatment
effects that causes the mean in one group to differ from the mean in another. Most of the
time ANOVA is used to compare the equality of three or more means, however when the
means from two samples are compared using ANOVA it is equivalent to using a t-test to
compare the means of independent samples.
ANOVA is based on comparing the variance (or variation) between the data samples to
variation within each particular sample. If the between variation is much larger than the
within variation, the means of different samples will not be equal. If the between and
within variations are approximately the same size, then there will be no significant
difference between sample means. There are two types of ANOVA that are commonly
used, the One-Way ANOVA and the Two-Way ANOVA.
One-way ANOVA
A one-way ANOVA has one independent variable and a two-way ANOVA was two
independent variables.
Statistical Model
Yij μ τ i ε ij
:
63 | P a g e
Where,Yij = Response of yield from the jth unit receiving the ith treatment
= General mean
i = Effect of ith treatment
ij
= Uncontrolled variation associated with jth unit receiving ith treatments.
Analysis:
H0: All the treatments are equal or T1= T2=T3=…..=Tt
H1: At least one variety is different from others
GT 2
1) Correction factor (C.F) =
n
2) Total SS= ( y ij ) 2 -C.F
yi . 2
3) SS for varieties = CF
k
4) SS for error = Total S. S – Treatment S. S
One way Analysis of variance
Source of M. S.
DF Sum of Squares (SS) Cal. F
variation (SS/DF)
Between the (t-1) t t r MST MST MSE
Yi.2 ( Y ij )
2
Treatment i 1 j 1
i 1
k kt
Between the By subtraction MSE
t(k-1)
Treatment
t r
( Y ij )2
Total (kt-1) t k
Y
i 1 j 1
2
ij
i 1 j 1
kt
Since, the Calculated F ˃ Table F0.05 at (4, 15) d.f, reject H0 and accept H1, there are
significant differences between the treatment means.
1 1
SEm
MS E
or SEd MS E
r or r0 r r
i j
CD (Critical Difference): It is such a significant difference that all the differences of two
mean greater than or equal to it are considered to be significant. Therefore it is the least
significant difference. It’s value is given by
64 | P a g e
Lecture 13: Introduction to sampling and Sampling versus Complete Enumeration
Introduction
Our knowledge, our actions and our attitudes are based to a very large extent on
samples. This is equally true in everyday life and in scientific research, whenever, a
person wants to buy a large quantity of a commodity say wheat, rice, fruits etc. be
decided about total lot of just by simply examining a small fraction of it . A doctor's
opinion about the state of health is determined by only examining one or two drops of
blood. We study the soil about its nutrient status, salinity status etc by using only 5g. of
soil. It has been experienced that the sample survey if planned properly, can give precise
information.
A sample is a part of population (total or aggregate). It is to be selected at
random for unbiased estimate and can be used as a basis for inferring about the
population. Most populations (a crop plant population, groundnut growers, pest
population pump-set owners etc.) are so large that it is rather next to impossible to
contact each of them during specified time, only a fraction of it is investigated and the
inference is drawn from it for the population. Consequently many generalizations which
arise from ordinary experience are likely to be unwarranted.
The sample has many advantages over a census or complete enumeration of a
finite population. If carefully designed, the sample is not only cheaper but may give
results which are as accurate or sometimes, more accurate than those of census. The
reason is that a census is subjected to biased error than that due to any sampling method
employed. The smallness of samples makes possible the use of the more care in the
design and execution of each step in the inquiry. In this way sources of error can be
investigated and reduced, eliminated or measured and corrected for. The other
advantages of sample survey are that it is less time consuming, involves less cost, has
greater scope and has greater operational facilities. It is for these reasons that sample
surveys are being preferred by the research scientists to complete enumeration.
65 | P a g e
Terminology
1) Population: A population is the aggregate of individuals or units or objects having at
least one common characteristics.
2) Sample: A sample is a part or fraction of large aggregate (Population) about which
some information is required.
3) Sampling: It is the method/process of selection of sample from the population.
4) Sampling Unit: It is the individual element or a group of elements on which
observations are to be made.
5) Parameter: Any measurable characteristic of population which is to be worked out by
utilizing each and every observation of population parameters help to
characterize the population mean and are the example of the
parameter. Parameters are generally unknown and constant.
6) Statistic: Any number estimated from sample (X and S are the statistic)
Types of population
1) Finite population: One can count rose plants in a garden i.e. counting the individuals
in the population is possible & hence it is finite population. e.g. Rose
plants in garden, No. of animals in herd, Fruits on tree etc.
2) Infinite population: The rose plants on the earth can not be counted. Thus it becomes
infinite population.
3) Real population: This is the population in which the members do exist in reality. e.g. a
heap of food grains, herd of cows etc.
4) Hypothetical population: The member of population does not exist in reality. e.g.
Possible throws of a die. Similarly, yield of a new variety to be evolved.
Possible results of an experiments etc. are the examples of this type of
population.
Advantages of sample study
1) Less expensive.
2) In a limited specific time one can complete the project work and collect the
information (i.e. greater speed)
3) With minimum technical persons, one can study the problem
66 | P a g e
4) More precise and accurate information regarding population (greater accuracy) may be
obtained.
5) Sample investigation can be carryout with minimum facilities.
6) Sample investigation is to be done when the unit is supposed to be destroyed while
taking observation.
Sampling plans/Designs
1) Simple random sampling.
2) Stratified random sampling.
3) Multistage sampling.
4) Cluster sampling.
5) Systematic sampling.
6) Purposive sampling and so on.
(1) Simple random sampling (SRS)
In a sample random sampling, the sample is selected in such a way that every
member of the population has an equal and independent chance of being selected in
the sample. It implies that selection of a sample from all possible samples that could be
chosen is equally likely.
Random selection of units is done using any one of the following methods.
a) Using Tickets, tags etc.
b) Random number tables.
a) Using Tickets, tags etc
To give an example, suppose that an experimenter wishes to draw a random
sample of size 10 individuals (say 10 plants for measuring the heights) from a finite
population, say 200 plants. A method of doing this would be assign a number to each
member of the population put the individuals into a box and mix them thoroughly. Draw
ten tickets or tags from the box. The number on these ten tickets or tags correspond the
plants to be selected. This method is time and labor consuming and also costly.
b) Use of Random Number table
The procedure (a) can be shortened by the use of a table of random number. Such
a table consists of numbers chosen in a fashion similar to drawing numbered tickets or
tags out of a box. This table is so made that all numbers 0, 1, 2.... appear with
approximately the same frequency. By combining the numbers in pairs we have two-
67 | P a g e
digit numbers (i.e. from 00 to 99) By using the number three at a time we have three digit
numbers from 000 to 999 etc.
The table should be entered in a random manner. Put a pencil aimlessly on a page
of the table. The point thus obtained on page is the starting point for selecting the
numbers, record the numbers until the required number of random digits is obtained.
Examples of SRS
1) Impact of T & V system on Socio-economic status of summer groundnut growers in a
specific Taluka.
Here population is “Summer groundnut growers registered under T & V system
of a given taluka. One can easily prepare the frame for this finite population and select
the random sample of summer groundnut growers.
2) Constraint analysis for milk: Productivity in a village
Here population is “milk producers in a given village. The frame for which can be
prepared easily and a random sample can be taken to fine out the factors responsible for
low productivity.
(2) Stratified Random sampling
In this scheme the heterogeneous population is sub-divided into homogeneous
several groups called stratum and then samples are drawn independently (at random)
from each stratum.
As the sampling variance of the estimate of mean depends on the within strata
variation, the stratification of heterogeneous population into homogeneous strata helps in
increasing precision of the estimates e.g. while studying average income of the staff
members of the Gujarat Agricultural University one has to employ stratified random
sampling. Because, simple random sampling may result into under or over estimation of
income of the staff member. Suppose only professors are selected for the sample then the
average income will be above the true average value, similarly if only helpers are selected
in the sample then the results will be on lower side.
When such heterogeneous population is required to be sampled then one has to
utilize the stratified random sampling in spite of simple random sampling.
68 | P a g e
Examples
1) Adoption level of improved agro technology by the cultivators.
2) A survey for nutritional status of school going students.
3) Socio-economic survey for rural and urban people of Valsad district.
(3) Multistage sampling
In this method, the selection is done in stages, called sub sampling. For example
in estimating the yield of wheat in a district. Talukas may be considered as primary
samplings unit (1st stage), Villages within taluka as secondary sampling units (2nd stage)
the cultivators within village within taluka as third stage sampling units and so on.
The advantage of this type of sampling is that at the first stage the frame of
primary sampling units is required which is easy to have and at the second stage the
frame of secondary sampling units is required for the selected primary sampling units
only and so on. Moreover, this method allows the use of different selection procedure in
different stages. It is because of this consideration that multi-stage sampling is used in
most of the large scale surveys.
(4) Systematic sampling
Systematic sampling is slight varying compared to the simple random sampling
in which only the first sample units is selected at random and the remaining units are
automatically selected in a definite sequence at equal spacing from one another. This
technique of drawing samples is usually recommended if the complete and up to date list
of the sampling units is available and the units are arranged in some systematic order
such alphabetical, chronological, geographical order etc. This requires the sampling units
in the population to be ordered in such a way that each item in the population is
uniquely identified by its order, for example the name of persons in a telephone
directory, the list of voters etc.
Sampling and Non-sampling Errors.
The inaccuracies or errors in any statistical investigation i.e. in the collection,
processing, analysis and interpretation of the data may be broadly classified as follows:
1) Sampling errors 2) Non sampling errors.
69 | P a g e
(1) Sampling Errors
In a sample survey, since only a small portion of population is studied and hence
its results are bound to differ from the census results and thus has a certain amount of
error. This error would always be there, no matter that the sample is drawn at random
and that it is highly representative. This error is attributed to fluctuations of sampling
and is called sampling error. Sampling error is due to the fact that only a subset of the
population (i.e. sample) has been used to estimate the population parameters and draw
inferences about the population. Thus, sampling error is present only in a sample survey
while it is completely absent in census surveys.
Sampling error may be due to following reasons.
(i) Faulty selection of the sample
(ii) Substitution
(iii) Faulty demarcation of sampling units
(iv) Error due to bias in the estimation method
(v) Variability of the population.
(2) Non Sampling Errors
Non-sampling errors are not attributed to chance and are a consequence of certain
factors which are within human control. In other words they are due to certain causes
which can be traced and may arise at any stage of the inquiry viz. panning and execution
of the survey and collection, processing and analysis of the data. This error is present in
sample and census.
Some of the important factors responsible for non sampling errors are as under
(i) Faulty planning including vague and faulty definitions of the population of the
statistical units to be used, incomplete list of population members.
(ii) Vague and imperfect questionnaire which might result in incomplete or wrong
information.
(iii) Defective methods of interviewing and asking questions.
(iv) Vagueness about the type of the data to be collected.
(v) Personal bias of the investigator.
(vi) Lack of trained and qualified investigators and lack of supervisory staff.
(vii) Failure of respondents memory to recall the events or happenings in the past.
(viii) Non response and inadequate or incomplete response.
(ix) Improper coverage.
70 | P a g e
(x) Compiling errors.
(xi) Publication errors.
71 | P a g e
Lecture 14: SRS with replacement (SRSWR) SRS, without replacement (SRSWOR)and
Use of Random Number Tables for selection of Simple Random Sample
There are wo type of SRS. They are
• SRS with replacement (SRSWR)
• SRS without replacement (SRSWOR)
SRS with replacement (SRSWR)
It is the random process in which a unit is selected and noted and then returned to the
population before the next drawing is made and this process is repeated n times, it give
rise to a simple random sample of n units.. This process is generally known as simple
random random sampling with replacement (SRSWR)
The probability of selection of an element remains unchanged after each draw
The same units could be selected more than once
Numbers on possible sample of size n from the Population of size N is given by
Nn
Example: 2 elements from 4 (ABCD)
(How many ways we can draw a sample of size 2 elements from a population of
size 4?) i.e. n=2 and N=4
• With SRSWR: Nn = 42 = 16
• AA, AB, AC, AD, BA, BB, BC, BD, CA, CB, CC, CD, DA, DB, DC, DD = 16
samples
SRS without replacement (SRSWOR)
It is a SRS process in which, once an element is selected as a sample unit, it will not be
replaced in the population pool.
The probability of selection of an element remains changed after each draw
The same units could not be selected more than once
Numbers on possible sample of size n from the Population of size N is given by
N
NCn =
n
N! 4!
6
Mathematically, ( N n )! n! 2!2 !
72 | P a g e
DA, DB, DC, DD
74 | P a g e
If the calculated t-value is less than the critical t-value at a given significance level, it indicates the observed difference is not statistically significant. This leads to the acceptance of the null hypothesis, suggesting that both samples come from the same population .
The Z-test is applicable when data follows a normal distribution, and either the sample size is large (n > 30) or the population standard deviation is known. If the calculated Z-value is less than 1.96, the mean difference is not significant at the 5% level, accepting the null hypothesis. If Z ≥ 1.96, the difference is significant, rejecting the null hypothesis .
The arithmetic mean is appreciated for several merits: it is rigidly defined, based on all observations, easily comprehensible, simple to calculate, facilitates easy mathematical treatment, and is minimally affected by sampling fluctuations . However, its demerits include being influenced by extreme values, becoming meaningless with large data variation, and not directly measuring growth or speed rates .
A paired t-test evaluates whether the mean difference between two sets of paired observations is zero. The null hypothesis (Ho: μd = 0) is tested against an alternative hypothesis at a predetermined significance level. If the calculated t-value is significant, it suggests that the mean differences observed are not by random chance, indicating dependency between the samples .
Simple random sampling ensures each member of the population has an equal chance of being selected, eliminating selection bias. It is straightforward and its random nature facilitates easy probability assessments for statistical analysis. However, it may require more effort if populations are vast .
The geometric mean is preferred when there is a need for an average that tends to lower values, avoiding the upward bias of the arithmetic mean which gives equal weight to all data points and has a tendency towards higher values. It is especially useful for ratio and proportion data .
In scientific experimentation, a hypothetical population represents outcomes that do not exist in reality, allowing researchers to explore potential scenarios. For example, possible dice throws or hypothetical crop yields in genetic engineering can be considered, giving a framework to analyze experimental results or forecasts .
The weighted mean is preferred when different observations are given different weights, which the arithmetic mean cannot adequately adjust for. It is useful when there are widely varying classes, differences in item importance, or when averaging ratios, percentages, or rates such as rupees per kilogram. It is particularly employed in contexts like calculating birth rates, death rates, and index numbers .
Rank correlation, developed by Charles Spearman, does not rely on the assumption of normal distribution, which is necessary for Karl Pearson’s method. It is used when quantitative measures cannot be fixed or when distribution shape is unknown, such as ranking students by height and weight without exact measurements .
A correlation coefficient of zero implies no linear relationship between the two variables under study; however, this does not necessarily imply that the variables are completely independent, as there may be nonlinear relationships not captured by the correlation coefficient .