0% found this document useful (0 votes)
39 views74 pages

Statistical Methods and Applications Overview

The document is a lecture note for a statistical methods course (STAT-211) covering various topics such as introduction to statistics, graphical representation, measures of central tendency, and probability distributions. It outlines the aims, limitations, and applications of statistics, particularly in agricultural research, and provides methods for data analysis including frequency distribution and graphical representation techniques. Additionally, it includes examples and procedures for constructing frequency tables and graphs.

Uploaded by

Suresh Sharma
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
39 views74 pages

Statistical Methods and Applications Overview

The document is a lecture note for a statistical methods course (STAT-211) covering various topics such as introduction to statistics, graphical representation, measures of central tendency, and probability distributions. It outlines the aims, limitations, and applications of statistics, particularly in agricultural research, and provides methods for data analysis including frequency distribution and graphical representation techniques. Additionally, it includes examples and procedures for constructing frequency tables and graphs.

Uploaded by

Suresh Sharma
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Lecture Note

For
STAT-211 (Statistical Methods)
Year-II, Sem-I, Academic Year: 2020-21

Lecture Topics Page


No.
1 Introduction to Statistics and Its Application 1-4
2 Graphical Representation: 5-10
3 Measures of Central Tendency 11-17
4 Measures of Dispersion 18-24
5 Probability 25-33
6 Binomial &Poisson Distributions 34-37
7 Correlation Analysis 38-41
8 Linear Regression Analysis 41-45
9 Introduction to test of Significance 46-48
10 One sample and two sample t-test for Mean 49-57
11 Chi-Square Test of Independence of Attributes in 2x 2 58-62
Contingency Table
12 Introduction to Analysis of Variance and One Way 63-62
ANOVA
13 Introduction to sampling and Sampling versus Complete 63-71
Enumeration
14 SRS with replacement (SRSWR) SRS, without 72-74
replacement (SRSWOR)and Use of Random Number
Tables for selection of Simple Random Sample

1|Page
Lecture No. 1: Introduction to Statistics and Its Application
Statistics
Statistics is the science of application of mathematics on collected data.
OR Meaning of statistics Statistics is concerned with scientific methods for collecting,
organising, summarising, presenting and analysing data as well as deriving valid
conclusions and making reasonable decisions on the basis of this analysis.
Father of Statistics: R. A. Fisher
Father of Indian Statistics: P.C. Mahalonobis
The word statistics is generally used in two different ways.
(1) When it is used in plural, it means the quantitative data affected to a marked extent
by a multiplicity of causes. When we say `collect statistics’ means collect the
numerical data which are to be analyzed and interpreted. e.g.
(i) Wheat production affected by various causes.
(ii) Collect the data of height of the students of second semester.
(2) When it is used in singular, it means "the science of collecting, classifying and using
the data for further statistical treatments". It involves the methods of analysis used in
the analysis and interpretation of data and they are known as statistical methods.
Definitions:
(1):It is a study of population, variation and the methods for reduction of the data.
Biometry: When the principles of statistics are applied on living thing or organisms, the
science is called biometry.
Statistical methods: The methods by which statistical data are analyzed are called
statistical methods.
Aims of studying statistics
(1) To study the population: The study of population of any kind agricultural activities on
the basis of sample data.
(2) To understand the nature of variability e.g. Height of plants. The biological
phenomena observed under one set of conditions are never duplicated exactly under
another set of similar conditions. Therefore, repetition of experiment is necessary to
account all the factors causing variation. In biological phenomena were variation is a
rule rather than exception it is this function that has wide application.

2|Page
(3) To express the facts in summary form (the facts that are based on large number of
observations) e. g. It is not possible for one to form a precise idea about the income
position of the population of India from the records of individuals.
(4) To provide correct method(s) for taking sample (sampling).
(5) To provide proper method for comparison of two or more things.
(6) It helps in prediction/forecasting the yield of a particular crop for a particular year on
the basis of the past records.
Limitations of Statistics
Statistics with its wide applications in almost every sphere of human activity is
not without limitations. The following are the important limitations.
1) It does not deal with individual.
2) It deals only with quantitative characters.
3) Statistical results are true only on an average.
4) Statistics can be misused.
5) It does not reveal the entire story.
6) Expert knowledge is must to handle the statistical data.
(1) It does not study individuals
Statistics deals with an aggregate of objects and does not give any specific
recognition to the individual items of a series. e. g. the individual figures of
agricultural production of any country for a particular year are meaningless unless,
to facilitate comparison, similar figures of other countries or of the same country of
different years are given. Height of Mr. X is 5' 8" does not constitute statistical
statement. The average height of an Indian is 5' 8".
(2) It deals only with quantitative characters.
Efficiency, honesty, intelligence factors can be measure indirectly e. g. efficiency of
selling agent can be judged by studying the no. of articles sold by him.
(3) Statistical results are true only on an average.
Average consumption of milk per head in a certain locality is 1/2 liter but it does
not give any idea of the shortage of milk faced by the poor. The conclusions
obtained statistically are not universally true; they are true only under certain
conditions. This is because statistics as a science less exact as compared to natural
sciences.

3|Page
(4) Statistics can be misused.
Because if conclusions are based on incomplete information, statistics can prove
anything. There are three types lies: lies, dammed lies and statistics. Statistics are
like clay of which one can make a God or Devil as he please.
Importance/Application in agricultural research
(1) It helps to understand nature of variability or differences.
(2) To arrive at the meaningful conclusion on the basis of sample study in the field.
(3) Express the data/result of the field experiment in summary form.
(4) Sampling
a) In state agricultural survey, for estimation of areas and yield of crops.
b) In price fixation policy of various agricultural commodities.
c) In agricultural extension survey, to study the impact of programs.
d) In agricultural economics survey, to study the demand-supply policy, the
growth rate of population and cost of production of various crops.
(5) In agricultural meteorology for weather forecasting and to correlate weather
parameters with crop production.
Variable : The characteristics which show variation or variability are called variables or
variats e.g. cabbage yield, wheat yield per hectares of the growers. Variable
can of two types,
(i) Qualitative : The characteristics which can not be measured numerically or in terms
of magnitude e.g. flower color, nature of surface.
(ii) Quantitative: The characteristics which can be measured in terms of magnitude e.g.
yield of crop, height, weight. The quantitative characteristics are of
two types.
(a) Discrete : Character which takes only integer values/or whole value. There is a
definite gap between two values. e.g. No. of students in a class, No. of
bacteria in given area.
(b) Continuous: The quantity which can take any numerical value within a certain range.
Height, weight (They are in fraction and there is no definite gap between values).

4|Page
Lecture 2. Graphical Representation:
Graph: A display of point and line. In a graph of pair value (x, y), so called the co-
ordinate of o point, are plotted on a graph pepper suitably choosing the scales along X-
axis (abscissa) and Y-axis (ordinate). The plotted points are joined by the straight line in
their sequence of occurrence. The figure so obtain are called graph.
Advantage of Graphical or Diagramatics representation:
1. Diagram give as bird’s-eye view of complex data.
2. They have long lasting impression
3. Easy to understand even by common man
4. They save time and Labour.
5. They facilitate comparison
One dimensional diagram or graph
Bar diagram and Line Diagram
Different Types of Bar diagram are.
1. Simple bar chart
2. Multiple bar chart
3. Sub-divided bar chart
Two dimensional diagram or graph
Circle, Rectangle, Pie Chart
Three Dimensional Diagram
1. Cubes, Cylinder and spheres

Frequency Distribution
Objectives:
1. To condense the mass of data in such a manner that similarities and dissimilarities can
be easily understand.
2. To enable statistical treatment to the data collected.
Frequency: The no. or individual of items occurring in each class is termed as frequency.
Frequency distribution: The manner in which the frequencies are distributed over the
different class is called frequency distribution of the character under study and the table
indicating frequency distribution is called frequency table.

5|Page
Class limit: It is the lowest and highest values of the distribution that can be included in
the class e.g 10-20, 20-30 etc. Two boundaries of a class are known as the lower limit and
upper limit of a class.
Class interval: The width of a class that is the difference of upper and lower limit of the
class is known as class interval.
Class mid point: It is the value lying half way between the lower limit (LL) and upper
limit (UL) of a class interval i.e. (LL + UL)/2.
Points while deciding class interval/classes:
1. It should be of uniform width which facilitates the statistical computation.
2. Range of the class should cover the data and should be continuous.
3. It should be convenient to make the mid-point of a class.
4. It should not be over lapping.
Types of frequency distribution:
(1) Discrete frequency distribution
(2) Continuous frequency distribution
Methods of classifying the data according to class interval:
Exclusive method: When the class intervals are so fixed that the upper limit of one class
is the lower limit of the next class. This method is known as exclusive method e.g. 0-10,
10-20. Usually this method is preferred for continuous type of data. The data observed
up to 9.99 would be included in 0-10 class while 10 or greater than 10 will be included in
10-20 class.
Inclusive method: In this method of classification, the upper limit of one class is
included in that class itself e.g. 100-199, 200-299. The value of 100 and 199 will be
included in the class of 100-199. This method is preferred for discrete type of data.
**Procedure to form frequency distribution:
Step 1: Find range of the data. Range = Highest value – Lowest value.
Step 2: Fix the number of classes. Number of classes should preferably between 5 to 15
and should not be less than 5 and more than 30.
Approximate no. of classes = K =1 + 3.322 log N (Sturge’s rule) where N = no. of
observations under study.
Step 3: Fix the class interval = CI = Range/No. of classes or (L-S)/K where L = largest
value and S = smallest value
Step 4: Arrange different classes in ascending order of magnitude

6|Page
Step 5: Pick up the values of observation and make tally mark against respective classes.
Step 6: Find total tally mark of each class which will give the no. of frequencies in the
respective classes.

Graphical Representation
Graphical representation is used when we have to represent the data of a
frequency distribution and of a time series. It is represented by points which are plotted
on a graph paper.
Advantages of graphical representation
1. Easy to understand and interpret data at a glance.
2. It facilitates comparisons.
3. It gives eye view of complex data.
4. It has long lasting impression.
5. It gives an attractive and interesting view.
Limitations of graphical representation
1. It cannot show all those facts which are there in the tables.
2. It shows tendency and fluctuations, actual values are not known.
3. The charts take more time to be drawn than the tables.
Graphs of Frequency Distribution
Histogram: It is a bar diagram which is suitable for frequency distributions with
continuous classes. The width of all bars is equal to class interval and heights of the bars
are in proportion to the frequencies of the respective classes. In this diagram bars touch
each other but one bar never overlaps the other.

Frequency polygon: When the mid points of the tops of the adjacent bars of a histogram
are joined in order by a straight line, then the graph of lines so obtained is called a
frequency polygon.

7|Page
Frequency curve: A frequency curve is a graphical representation of frequencies
corresponding to their variates values by a smooth curve. A smoothened frequency
polygon represents a frequency curve.

Ogive or Cumulative frequency curve: it is a graph plotted for the variates values and
their corresponding cumulative frequencies and joined by a free hand smooth curve. The
curve is ‘S’ shaped. There are two methods of constructing ogive viz. (i) “less than”
method and (ii) “more than” method. In the “less than” method, we start with the upper
limit of classes and go on adding the frequencies however, in the “more than” methods,
we start with the lower limit of classes. The first method gives a rising curve whereas
second method shows a declining curve.

Example1:
Following are the make obtain by 60 students in an examination of 100 mark
85 38 55 49 52 60 65 42 31 30
29 18 60 32 13 80 9 15 51 40
54 59 90 69 36 95 5 59 66 70
69 28 50 59 61 41 85 76 30 45
58 47 20 75 56 52 60 23 18 48
45 72 86 66 42 28 35 37 60 80

8|Page
Prepare frequency table of the data and draw Histogram, Frequency polygon and ogive.
[no. of
Soln:
Step 1. Range= Highest –lowest value = 95-5 = 90
Here, Highest value = 95, Lowest value= 5;
Step 2. Number of class, K=1+3.322* log60 = 6.90 ~7
Range 90
Step 3. Class interval=   12.86  13
Numberof class 7
For convenience, the class interval is taken as 10 instead of 9. The lowest class may be
started from 0 instead of 5 and similarly to include the highest value (95) the highest class
is taken with a upper limit 100.
For For Polygon For Ogive
Histogram
Class Tally mark Frequency Mid-value of Cumulative
(X) X Frequency
5-18 1111 4 11.5 4
18-31 1111 1111 1 9 24.5 13
31-44 1111 1111 11 10 37.5 23
44-57 1111 1111 1111 12 50.5 35
57-70 11111111111111 14 63.5 49
70-83 1111 11 6 80.5 55
83-96 1111 11 5 93.5 60
TOTAL N=60

Histogram
16
Number of student (Frequency)

14
12
10
8
6
4
2
0
5 18 31 44 57 70 83 96

Classe
s

9|Page
Frequency Polygon

16

Number of student (Frequency)


14
12
10
8 Series 1
6
4
2
0
11.5 24.5 37.5 50.5 63.5 80.5 93.5
Class Mid Value

ogive
70

60
Cummulative Frequency

50

40

30
Y-Values
20

10

0
0 20 40 60 80 100
Class mid Values

10 | P a g e
Lecture 3: Measures of Central Tendency
Different groups of data or statistical series or frequency distributions differ in
four characteristics.
1) Central tendency or location 3) Skewness or symmetry
2) Dispersion or variation 4) Kurtosis or peaked ness
Central tendency: A values of the variable tend to cluster around a central value or
centrally located observation of the distribution. This characteristic is known as central
tendency.
This centrally located value which represents the group of values is termed as the
measure of central tendency e.g. an average is called measure of central tendency.
Objectives
1) To get one single value that describes the characteristics of the entire
series/group.
2) To compare two or more distributions.

Requisite/Characteristics of ideal measures of central tendency


Since an average is a single value representing a group of values, it is expected
that such a value should satisfied the following properties.
1) It should be rigidly defined.
2) It should be based on all the observations.
3) It should be easy to understand or comprehensible, otherwise its use will be limited
4) It should be easy to calculate.
5) It should be amenable to further mathematical treatment.
6) It should be least affected by fluctuation of sampling.
7) It should be least affected by the extreme values.
Different measures of Central tendency
1) Arithmetic mean (A.M.) Algebraic average
2) Median
3) Mode Positional average

4) Geometric mean (GM)


5) Harmonic mean (H.M.) Algebraic average
6) Weighted mean (W.M.)

11 | P a g e
(1) Arithmetic mean or Mean
It is the most common and ideal measure of central tendency. It is defined as the
sum of the observed values of the character (or variable) divided by the number of
observations considered in obtained sum (total).
_
Symbolically X: Sample mean
 : Population mean

X i
X  i1

n
Raw data or ungrouped data

(i) Direct method:


n

X i
X  i1

n
Grouped data

(i) Direct method


k

fX i i
X i 1
k

f i 1
i n

where, f = freq. of k th class, X = class mid value, k = no. of classes


Sum
Xi 3 4 5 4 3 5 4 28
Xi-Mean -1 0 1 0 -1 1 0 0
(Xi-5) -2 -1 0 -1 -2 0 -1 -7
(Xi- 1 0 1 0 1 1 0 4
Mean) 2

(X-5)2 4 1 0 1 4 0 1 11
x 3 4 5 4 3 5 4 Product=14400
log x 0.477 0.602 0.698 0.602 0.4771 0.698 0.602 Sum=4.158
1/x 0.333 0.25 0.2 0.25 0.333 0.2 0.25 Sum=1.8166

(3  4  .....  4) 28
Mean ( X ) =  4
7 7
GM= (3x4……x4)1/7 = (14400)1/7 = 3.926 or

12 | P a g e
 n 
  log xi 
Anti log(  i 1  = Antilog(  log 3  log 4......  log 4 
GM=
   
n  n 
 
 
 0.477  0.602......  0.602 
Anti log    Anti log( 0.594)  3.926
 7 

(0.33  0.25 ........  0.25) 1.816


1/HM=   0.2595
7 7
1
HM=  3.853
0.2595

AM˃GM˃HM

Properties of Arithmetic mean


1) The algebraic sum of the deviations of a set of values (observed values) from their
arithmetic mean is zero.
2) The sum of squares of the deviations of a set of values from their arithmetic mean is
always minimum
3) Amenability of arithmetic mean to further mathematical calculation.
_ _
(a) Let X1 be the mean of n1 observations, X2 be the mean of n2 observations, Xk be
the mean of nk observations then the combined mean of N observations is given
by (where N = n1 + n2 +...,+nk )
n1X1  n 2 X 2  ......  nk X k
X
n1  n 2  ......  nk

This is also called weighted mean (W.M.)


(b) Adding or subtracting a constant from each observation of a given series will
add or subtract the same constant to the arithmetic mean.
(c) Multiplying or diving each observation by a constant will multiply or divide the
arithmetic mean by the same constant.

13 | P a g e
Merits and demerits of Arithmetic mean
Merits
1) It is rigidly defined.
2) It is based on all the observations.
3) It is readily comprehensible.
4) It is easy to calculate.
5) Its algebraic (Mathematical) treatment is especially easy and definitely possible.
6) It is also least affected by the fluctuation of sampling.
Demerits
1) It is affected by the extreme values.
2) If there is large variation in the data then A.M. becomes some times meaning less.
3) It is not used to measure rate of growth or rate of speed directly.
Uses: It is most popular and simple estimate and used widely in almost all the fields
of studies such as social science, economics, business, agriculture, medical sciences,
engineering and such other sciences.

Weighted Mean
When different observations are to be given different weights, arithmetic mean
does not prove to be a good measure of central tendency. In such cases weighted mean is
to be calculated.
If X1, X2, X3,... Xn are different observation and W1, W2, W3,... Wn are their
respective weights then,

W.M. 
W1X1  W2 X 2  ......  Wn Xn
=
W X i i

W1  W2  ......  Wn W i

Merits and demerits of Weighted mean (W. M.)


W.M. is the A.M., hence merits and demerits are the same as there for the
arithmetic mean.

Uses
1) Used when the number of individuals in different classes of grouped widely
varying.
2) Used when the importance of all the items in a series is not same.

14 | P a g e
3) Used when the ratios, percentages or rates e.g. rupees per kilogram, rupees per
meter etc. are to be averaged.
4) Weighted mean is particularly used in calculating birth rates, death rates, index
numbers, average yield etc.

Geometric Mean
AM gives equal weightage to all the items and has got a tendency towards the
higher values. Sometimes it is necessary to get average having a tendency towards the
lower values. In such case, geometric mean is helpful. It is defined as the n th root of the
product of n items of a series with following relation.
It is useful in ratio and Proportion data
Raw data or ungrouped data
n

 log X i
GM  n X 1 , X 2, .....X n =  X 1 , X 2 ,..... X n 
1
i 1
n = Antilog ( )
n

Grouped data

GM  n X 1f1 , X 2f 2 ,..... X kf k = X1f1 , X 2f 2 ,.....X kf k 1
n

k
=
1
 f1 log X1  f 2 log X 2  ..... f k log X k  = 1  fi log X i
n n 1

Merits and demerits of Geometric mean


Merits
1) It is rigidly defined.
2) It is based on all the observations.
3) It is not much affected by the fluctuation of sampling.
4) It gives less weightage to large items and more to small items.
5) It is suitable for averaging ratios, average rate of change, index nos.
Demerits
1) It is difficult to understand.
2) It cannot be calculated when there are negative values.
3) If any item of the series is zero, it would be also zero.

15 | P a g e
Harmonic Mean
HM is the reciprocal of the arithmetic mean of the reciprocal of the values of a
variable or series.

Raw data or ungrouped data

n n
HM  
 1 1 1
 
1
 .....   x
 x1 x 2 x n  i

Grouped data

HM 
f i

f i

 f1 f2 f  f
   .....  k  x i

 x1 x 2 xn  i

Merits and demerits of Harmonic mean


Merits
1) It is rigidly defined.
2) It is based on all the observations.
3) It is not much affected by the fluctuation of sampling.
4) It gives greater weightage to smaller values.
5) It is useful in average price, speed, time, distance, quantity etc.
Demerits
1) It is not easy to calculate and understand.
2) It cannot be calculated if any value is zero or negative.
3) It gives large weightage to smaller values.
Uses: Time series data, units purchased per rupee, kilometers covered per hour,
problems solved per time.
Relation between AM, GM and HM:
1) AM > GM > HM

2) GM= AM  HM

16 | P a g e
Median: The median is the middle most items that divide the series into two equal parts
when ith items are arranged in ascending or descending order.

 n 1
th

In case of raw data, the median is the   term where n (odd) is the total no. of
 2 
observations whereas in case of n even number medial is the average between
th th
n n 
  and   1 terms, in case of grouped data it is given by the formula,
2 2 

In case of grouped data, it is given by the formula,

 n/2   cf 
Median  l   xCI
 f 
where, l = lower limit of the median class in which (n/2)th item falls
cf= cumulative frequency of the class preceding the median class
f = frequency of median class
CI= Class interval
Uses: It is useful when the extreme values of the series are either not available or
impossible to be obtained or abnormal. When in a group, the individual is
denoted by better than half the individual’s, median is used. It is also useful
when the items are not susceptible to measurement in definite units e.g.
intelligence, ability, efficiency etc.
Uses: It is useful when the extreme values of the series are either not available or
impossible to be obtained or abnormal. When in a group, the individual is
denoted by better than half the individual’s, median is used. It is also useful
when the items are not susceptible to measurement in definite units e.g.
intelligence, ability, efficiency etc.
** Median is the best method if data is suffer from outlier
Mode: The value of the variable which occurs most frequently or whose frequency is
maximum is known as mode.
Uses: Business forecasting is particularly based on the modal values. Meteorological
forecasting is also based on modal values.
Relation between Mean, Median and Mode:
Mode = 3Median – 2Mean

17 | P a g e
Lecture No. 4: Measures of Dispersion
Mean is though an important concept in statistics, it does not give a clear picture
as to how the different observations are distributed in a given distribution or the series
under study. Consider the following series.
Series Observations Mean
1 2, 3, 4, 7 4
2 4, 4, 4, 4 4
3 1, 1, 2, 12 4
4 3, 4, 4, 5 4

In the above series, the mean is same i.e. 4 but the spread of the observations
about the mean is in different manner. Hence after locating the measures of central
tendency, the next point is to find out the center. This can be done by measuring the
spread. The spread is also called a Scatter, Variation OR dispersion of the variate values.
Definition
Dispersion may be defined as the extend of the scatterness of observations around
a measure of central tendency and a measure of such scatter is called measures of
dispersion.
Different measures of dispersion
1) Range.
2) Absolute mean deviation or Absolute deviation (A.M.D.)
3) Standard deviation (S)
4) Variance (S2)
5) Standard error of mean ([Link].)
6) Coefficient of variation (C.V. %)
Requisite/Characteristics of an ideal measure of dispersion

Measures of dispersion should possess all those characteristics which are


considered essential for measures of central tendency viz.
1) It should be based on all observations.
2) It should be readily comprehensible.
3) It should be fairly easily calculated.
4) It should be simple to understand.
5) It should not be affected by sampling fluctuations.
6) It should be amenable to algebraic treatment.
18 | P a g e
Standard deviation (S)
The standard deviation or "root of mean square deviation" is the most common
and efficient estimator used in statistics. It is based on deviation from arithmetic mean
and is denoted by S or . S = Standard deviation for sample.  = Standard deviation
for population.
Definition
"It is a square root of a ratio of sum of square of deviation calculated from
arithmetic mean to the total number of observations minus one”.
Method of computation

Raw data or ungrouped data

(1) Deviation method


n

 (X i  X) 2
S i 1
n 1
(2) Variable square method
n

n
( X i ) 2
X 2
i  i 1
n
S i 1
n 1

where: Xi = Variate value n = No. of observations

(3) Assumed mean method


n

n
( d i ) 2
d 2
i  i 1
n
S i 1
di = Xi – A A = Assumed mean
n 1

Grouped data or frequency distribution

(1) Deviation method


k

 f (X i i  X) 2 k
S i 1
n 1
where n  f
i 1
i ; fi= Frequency of ith class

(2) Variable square method

19 | P a g e
k

k
( f i X i ) 2
f X i
2
i  i 1
n
S i 1
n 1
(3) Assumed mean method
k

k
( f i d i ) 2
f d i
2
i  i 1
n
S i 1
di = (Xi – A)
n 1

(4) Step deviation method


k

k
(  f i d Xi ) 2
f d i
2
Xi  i 1
n
S i 1
dxi = (Xi - A)/I
n 1

Properties of Standard deviation


(1) Combined standard deviation can be calculated using following formula when
two series are given under study. It is symbolically denoted by S12.

N1S12  N2 S 22  N1d12  N2 d 22
S12 =
N1  N2
where, S12 = Combined standard deviation; S1 = Standard deviation of first group;

S2 = Standard deviation of second group; d1 = ( X1 - X12 ) ; d2 = ( X 2 - X12 ) ; N1 and

N2 are the numbers of observation for series one and two; X 1 = A.M. for first

series; X 2 = A.M. for second series; X 12 = weighted or combined mean.

(2) The sum of squares of the deviations of items in the series from their arithmetic
mean is minimum. This is the reason why standard deviation is always computed
from the arithmetic mean.
(3) Addition or subtraction of a constant from the grouped of an observation will not
change the value of S.D.
(4) Multiplying or dividing each observation of a given series by a constant value will
multiply or divide the std. deviation by the same constant.
Variance: Variance is the square of standard deviation. It is also called the “Mean square
deviation". It is being used very extensively in analysis of variance of results from field
experiment. Symbolically denoted by

20 | P a g e
S2 = Sample variance and 2 = Population variance.

Method of computation
Raw data or ungrouped data
(1) Deviation method

(Xi 1
i - X )2
S2 =
( n  1)
(2) Variable square method

X   X i  / n
2 2
S .S .

2 i
S =
n 1 d. f .
where : Xi = Variate value S S = Sum of square
n = No. of observations. df = Degrees of freedom

(3) Assumed mean method

d   d i  / n
2 2


2 i
S
n 1
where di = Xi - A
A = Assumed mean
Grouped data or frequency distribution
(1) Deviation method

 f (X
i 1
i i - X )2
S2 =
(n  1)

(2) Variable square method


fX   f i X i  / n
2 2


2 i i
S
n 1

(3) Assumed mean method :


fd   f i d i  / n
2 2


2 i i
S where, di = (Xi - A)
n 1
A = Assumed mean
fI = Frequency of ith class

(4) Step deviation method :


 f dx   f dx  2 2
/n

2 i i i i
S xI 2 dxi = (Xi - A)/I
n 1
Properties of variance

21 | P a g e
1) If V(x) represent the variance of X series and V(y) represent the variance of Y series
then V (xy) = V(x) + V(y)
i.e. V(x+y) = V(x) + V (y) and V(x-y) = V(x) + V(y)
2) Multiplying or dividing each observation by a constant will multiply or divide the
variance by square of that constant.
e.g. V(ax) = a2 V(x).
3) Addition or subtraction a constant from the groups of each observation will not
change the value of variance.
Standard error of mean ( [Link].)
The standard deviation is the standard error of a single variate where as standard
error of mean is the standard deviation of sampling distribution of the sample mean OR
it refers to the average magnitude of difference between the sample estimate and
population parameter taken over all possible samples from the population.
Definition : It is defined as square root of the ratio of the variance to the total no. of
observations in a given set of data.
Symbolically, it is written as S X for sample and  X for population.
S
SX  where, S = Standard deviation; n = No. of observations
n
For statistical analysis work, the use of S X is common. It is also used to provide

confidence limit on population mean and for test of significance.


Coefficient of variation (C.V. %)
It is a relative measure of variation and widely used to compare two or more
statistical series.
The statistical series may differ from one another with respect to their mean or
standard deviation or both. Some times they may also differ with respect to their units
and then their comparison is not possible. To have a comparable idea about the
variability present in them, C.V. % is used. It was developed by Karl Pearson".
Definition : "It is a percentage ratio of standard deviation to the arithmetic mean of a
given series". It is without unit or unit less.
S
CV%  x100
X

22 | P a g e
The series for which the C.V.% is greater is said to be more variable or we say
less consistence, less homogeneous or less stable while the series having lower C.V.%
is called more consistence or more homogeneous.

(Measure of dispersion)

The yield/plot of two varieties X and Y in eleven fields in a certain season are as
under.
Identify, which variety is more consistent in the performance.
X 8 6 7 6 8 7 5 8 9 6 7
Y 9 4 3 8 9 9 8 4 9 6 8
Solution:
Sl. Xi xi- X (xi- X )2 Yi Y  Y  Y  Y 
i i
2

1 8.00 1.00 1.00 9.00 2.00 4.00


2 6.00 -1.00 1.00 4.00 -3.00 9.00
3 7.00 0.00 0.00 3.00 -4.00 16.00
4 6.00 -1.00 1.00 8.00 1.00 1.00
5 8.00 1.00 1.00 9.00 2.00 4.00
6 7.00 0.00 0.00 9.00 2.00 4.00
7 5.00 -2.00 4.00 8.00 1.00 1.00
8 8.00 1.00 1.00 4.00 -3.00 9.00
9 9.00 2.00 4.00 9.00 2.00 4.00
10 6.00 -1.00 1.00 6.00 -1.00 1.00
11 7.00 0.00 0.00 8.00 1.00 1.00
Sum 77.00 0.00 14.00 Sum 77.00 0.00 54.00
Mean 7.00 Mean 7.00
n

X i
(8  6  7  ........  6  7) 77
Now, X  i1
=  7
n 11 11

 X 
n
2
X
i
(1  1  0  ......  4  1  0) 14
i
   1.4
n 1 11  1 10

 X 
n
2
i X
SDX  i
 1.4  1.183
n 1
Variance (S2 x)= (1.183)=1.4
SD S 1.183
Coefficient of Variance of X (CV)% =  100   100   100  16 .90
Mean X 7
n

Y i
(9  4  3  .........  6  8) 77
Now, Y  i 1
  7
n 11 11

23 | P a g e
 Y  Y 
n
2
i
(4  9  16  ....1  1) 54
i
   5.4
n 1 11  1 10

 Y  Y 
n
2
i
SDy  i
 5.4  2.323
n 1
Variance of Y (S2 y)= (1.183)=5.4
SD S 2.323
Coefficient of Variance (CV)% =  100   100   100  33 .197
Mean Y 7
Standard deviation (S), Variance (S2) and Coefficient of Variation (C.V) of X variety are
smaller than Y variety, Hence, variety X is more stable than Y

24 | P a g e
Lecture 5: Probability
Statistics concerns itself with inductive reasoning/inference based on the
mathematics of probability.
Sampling variation needs support in terms of probability / reliability of inference.
Introduction:

In the study of a population, one can not make any firm statements about the
population concerned or its parameters, when only a sample investigation of the
population is available for scrutiny. Due to the sampling variation there is doubt about
sample investigation and hence it is practice to make statements of less definite nature in
terms of probability or chance. The probability or chance for any statement depends on
the number of favourable, unfavourable and total possible cases. e.g. in a tossing a coin
for getting a head, one has to consider that there are two equally likely cases, head and
tail, one is in favour of the statement and the other is against it.
The theory of probability aims to generalize the laws of chance, to discover the
regularities in the pattern in which events, depending on chance, repeat themselves. It
may be the tossing of a coin, a game of cards or the genetical ratios which may be the
object of our investigation.
Jacob Bernoullis, an Italian mathematician was the first to give concept and
definition of probability in 1713. The work of Gregor Mendel in Genetics showed that the
theory of probability could be applied to biological investigations.
Before going to study Probability, we should have clear Knowledge about set
theory.

Set:
A set is the collection of all individual numbers in a well defined space. In a set all the
member belonging to the defined particular space is counted once only. A member can
not be counted twice or more in a set. It is usually represented in flower braces.

For example:
Set of natural numbers = {1,2,3,…..}
Set of whole numbers = {0,1,2,3,…..}
The set that contains all the elements of a given collection is called the universal set and is
represented by the symbol ‘U’.

25 | P a g e
Subset
A set A is said to be subset of another set B if and only if every element of set A is also a
part of other set B. Denoted by ‘⊆‘. ‘A ⊆ B ‘ denotes A is a subset of B.
To prove A is the subset of B, we need to simply show that if x belongs to A then x also
belongs to B. To prove A is not a subset of B, we need to find out one element which is
part of set A but not belong to set B.

Union
Union of the sets A and B, denoted by A ∪ B, is the set of distinct element belongs to set
A or set B, or both.

Above is the Venn Diagram of A U B.

Example : Find the union of A = {2, 3, 4} and B = {3, 4, 5};


Solution : A ∪ B = {2, 3, 4, 5}.

Intersection
The intersection of the sets A and B, denoted by A ∩ B, is the set of elements belongs to
both A and B i.e. set of the common element in A and B.

26 | P a g e
Above is the Venn Diagram of A ∩ B.

Example: Consider the previous sets A and B. Find out A ∩ B. Solution : A ∩ B = {3, 4}.

Disjoint
Two sets are said to be disjoint if their intersection is the empty set .i.e sets have no
common elements.

Above is the Venn Diagram of A disjoint B.

For Example
Let A = {1, 3, 5, 7, 9} and B = { 2, 4 ,6 , 8}. A and B are disjoint set both of them have no
common elements.

Set Difference
Difference between sets is denoted by ‘A – B’, is the set containing elements of set A but
not in B. i.e all elements of A except the element of B.

Above is the Venn Diagram of A-B.

Addition & Subtraction

27 | P a g e
Addition of sets A and B, referred to as Minknowski Addition, is the set in whose
elements are the sum of each possible pair of elements from the 2 sets (that is one element
is from set A and other is from set B).
Set subtraction follows the same rule, but with the subtraction operation on the elements.
It is to be observed that these operations are operable only on numeric data types. Even if
operated otherwise, it would only be a symbolic representation without any significance.
Further, it can be seen easily that set addition is commutative, while subtraction is not.

For Additions and consequently Subtraction, please refer this answer.


Formula:
1. A∪B=n(A)+n(B)-n(A∩B)
Properties of Union and Intersection of sets:
1. Associative Properties: A ∪ (B ∪ C) = (A ∪ B) ∪ C and A ∩ (B ∩ C) = (A ∩ B) ∩ C
2. Commutative Properties: A ∪ B = B ∪ A and A ∩ B = B ∩ A
3. Identity Property for Union: A ∪ φ = A
4. Intersection Property of the Empty Set: A ∩ φ = φ
5. Distributive Properties: A ∪(B ∩ C) = (A ∪ B) ∩ (A ∪ C) similarly for intersection.

Definition of Probability:
Probability is a ratio of the number of “favourable” cases to the total number of
equally likely cases. If probability is denoted by P then

Number of favourable cases


P
Total number of equally likely cases

Suppose a coin is tossed, the possible outcomes (events) are head and tail. These
are equally likely and mutually exclusive events. The probability (P) of event head is 1/2.
Range of the P is 0 to 1

Number of favourable event(s) 1


i. e. P(Head)  
Total no. of events 2

28 | P a g e
If n is the number of equally likely and mutually exclusive events for an event A,
of which m is the favourable to its occurrence, then the probability of A is the fraction
m/n."
P (A) = m/n

Random Experiment: A happening with two or more outcomes is called an experiment.


If the outcomes are associated with uncertainties, the experiment is called random.
Trial and event: An experiment which repeated under essentially identical conditions,
possible out comes is called events and experiment is known as trial. (Any out come or
results of an experiment is termed as event) e.g. throwing a die is a trial and getting 1 or 2
… is an event.
Simple event: The occurrence of a single event is known as simple event.
Compound events: The occurrence of two or more in connection with each other, the
joint occurrence is called the compound events.
Exhaustive events: All possible out comes of any trial / experiment are known as
exhaustive events. e.g. (i) tossing a coin there are two exhaustive events viz. head and
tail. (Possibility of the coin standing on an edge being ignored) (ii) throwing of a die,
there are six exhaustive events.
Mutually exclusive events: Events are said to be mutually exclusive if happening of any
one of them precludes the happening of other OR two events are said to be mutually
exclusive when both can not happen simultaneously in a single trial.
Independent events: Events are said to be independent if occurrence of any event is not
affected by the occurrence the remaining events e.g. in tossing an unbiased coin event of
getting head in the first toss is independent of getting a head in second, third and
subsequent throws.
Dependent events: Dependent events are those in which the occurrence or non-
occurrence of one event in anyone trial affects the probability of other events in other
trial.
Equally likely events: Events are said to be equally likely when one does not occur more
often than the others e.g. in throwing a die, all the six faces are equally likely to come.
(i) Law of addition:
Rule (i) A: When events are mutually exclusive

29 | P a g e
If two events A and B are mutually exclusive with probabilities P1 and P2
respectively, then the probability of occurrence of either of them (A or B) is equal to the
sum of the individual probabilities (A and B).

In symbols P(A or B) = P (A) + P(B) = P1 + P2

Proof: If an event A can happen in m1 ways and B in m2 ways, then the number of ways
in which either event can happen is m1 + m2. If the number of possibilities is n, then by
definition the probability of either the first or the second event happening is
m1  m2 m1 m2
P(A or B)  =  = P(A) + P(B) = P1 +P2
n n n
m1 m2
where P(A)   P1 ; P(B)   P2
n n

Similarly, in general, P(A or B or C) = P(A) + P(B) + P(C)

If K events are mutually exclusive with individual probabilities P1, P2, ... ,Pk then
P(anyone among K mutually exclusive events) = P1 + P2 + ...+ Pk.
Example: If A is the event drawing an ace from a pack of cards and B is the event
drawing a king, then P(ace = A) = 4/52 and P (king = B) = 4/52. The probability of
drawing either an ace or a king in a single draw is

P(ace or king) = P(A or B) = P(ace ) + P(king)


= P(A) + P(B)
= 4/52 + 4/52 = 8/52

Since both ace and king can not be drawn in a single draw and are thus mutually
exclusive events.
From the above explanation, one can point out two facts. They are:

(i) The probability, P1 of an event lies between zero and one.


(ii) The sum of the probabilities of mutually exclusive events is one.

Rule (i) B: When events are not mutually exclusive


30 | P a g e
If A and B are not mutually exclusive events, then the probability of either of them
is equal to the sum of their probabilities less the probability of their simultaneous
occurrence.
Symbolically
P(A or B) = P(A) + P(B) - P(AB)

where P(AB) is the probability of joint occurrence of A and B

Example: This will serve the proof also.


If A is the event “drawing an ace” from a pack of cards, and B is the event
“drawing a spade card”; then A and B are not mutually exclusive events, since the ace of
spade can be drawn. Thus the probability of drawing either ace or a spade or both is
P(ace or spade) = P(ace) + P(spade) – P(ace of spade)
= 4/52 + 13/52 - 1/52
= 16/52

Similarly we can generalize the rule for more than two events also.
i.e. P(A + B + C) = P(A) + P(B) + P(C) – P(AB) - P(AC) – P(BC) + P(ABC)

(ii) Law of multiplication:

Independent and dependent events: Events are said to be dependent or independent


accordingly as the occurrence of one does or does not affect the occurrence of the others.
Two events, drawing of a king and queen will be independent if the drawing of the card
is replaced after the first draw but if the card after first draw is not replaced and another
card is drawn for the second event, the probability of occurrence of the second event will
depend on the probability of the occurrence of the first. Hence in the latter case the
second event will be dependent on the first.
Rule A: When events are independent
If A and B are two independent events, with individual probabilities P1 and P2
respectively, then the probability of both happening at a time is the product of their
respective probability (P1.P2)
i.e. P(AB) = P(A) . P(B)
= Pl . P2
31 | P a g e
Proof: Let n1 and m1 be the possible and favourable numbers of cases for the event A and
n2 and m2 for the event B then
P(A) = m1/nl and P(B) = m2/n2
Since two events are independent, we can associate n2 possible cases for B with each of
the n1 possible cases for A, so that the total number of possible cases is n1.n2.
Similarly the total number of favourable cases for “A and B” is m.1m2
Thus,
m1.m2 m1 m2
P(A and B, both at a time) =  .  P1.P2
n1.n2 n1 n2
i .e. P(A and B) = P(A).P(B)
Similarly, the probability of occurrence of several independent events is the product of
their separate probabilities.
P(A.B.C. .... K) = P(A).P(B).P(C). .... .P(K)
Example: One urn contains 6 white flowers and 10 red flowers; second urn contains 8
white flowers and 12 red flowers. One flower is taken out from each of the urn. What is
the probability that the flowers drawn are white?
The probability of a white flower from the 1st urn is 6/16 and from 2nd urn is 8/20. Both
events are independent & hence required probability is the product 6/16 x 8/20 = 3/20.
Rule B: When events are dependent

If two events A and B are dependent then the probability of both happening at a time is
given as follows:
P(AB) = P(A) . P (B/A) This is the conditional probability
or = P(B) . P(A/B)
Where P(B/A) means the probability of second event B dependent on the probability of
first event A.
In above, if P(B/A) = P(B) then A and B are independent events.
Example: Suppose a box contains 3 white balls and 2 black balls. Let A be the event “first
ball drawn is black” and B the event “second ball drawn is black”, where the balls are not
replaced after being drawn. Here A and B are dependent events.
2 2
P(A)   probability of drawing first black ball.
32 5

32 | P a g e
1 1
P(B)    P(B/A ) the probability of second black ball given the first
3 1 4
ball drawn is black
Then P(A.B) = P(both black) = 2/5 . 1/4
= 2/20 = 1/10
Similarly we can generalize the rule for more than one dependent event.
P(A.B.C) = P(A).P(B/A).P(C/AB)

33 | P a g e
Lecture No. 6. Binomial &Poisson Distributions
Probability Distribution
It is also called parent population distribution, theoretical distributions or
theoretical frequency distribution. In previous chapter, the probability of the occurrence
of a single event is obtained. In scientific research using statistical methodology, it is
often required to obtain the probabilities of occurrence of all possible events. A table of
the possible values (Xi) which a chance event may assume with a corresponding
probability distribution for each value is called a probability distribution for the
population. Following table gives the probability distribution of sum of two unbiased
dice.
Table: Probability distribution of sum of two dice.
____________________________________________________________________________
Xi 2 3 4 5 6 7 8 9 10 11 12
____________________________________________________________________________
fi 1 2 3 4 5 6 5 4 3 2 1
___________________________________________________________________________
pi 1/36 2/36 3/36 4/36 5/36 6/36 5/36 4/36 3/36 2/36 1/36
___________________________________________________________________________

k
 pi = 1 pi = f(x) = f(xi)
i=1

Instead of a table of values such as above, one can represent the outcomes (pi) by
proper mathematical function over a range of Xi. In this chapter we would like to
describe three theoretical distribution
(i) Binomial distribution - James Bernoulli (1700)
(ii) Poisson distribution - S.D. Poisson (1857) and

BINOMIAL DISTRIBUTION
Binomial distribution was discovered by James Bernoulli in 1713.
This is very important distribution dealing with discrete variable. The binomial
distribution has two parameters viz. n and p. In other words, it is completely determined
by the values of n and p.
Let a random experiment be performed repeatedly and let the occurrence of an event
in any trial be called a success and its non-occurrence a failure. Consider a series of n

34 | P a g e
independent Burnoullian traials(n being finite, in which the probability ‘p’ of success in
any trial is constant for each trial. Then q=1-p is the probability of failure in any trial.
The probability of x successes and consequently (n-x) failures in n independent
trials, in a specified order (say) SSFSFFFS ….. FSF (where S represent success and F
failure) is given by the compound probability theorem by the expression:
P(SSFSFFFS ….. FSF)=P(S)P(S)P(F)P(S)P(F)P(F)P(F)P(S)…..P(F) P(S)P(F)
=p.p.q.p.q.q.q.p…..q.p.q
= p.p…..p… … ...q.q.q…….q

x factors (n-x) factors


But successes in n trials can occur in nCx ways and probability for each of these ways is px
qn-x. Hence the probability of x successes in n trials in any order whatsoever is given by
the addition theorem of probability by the expression :
p(X=x)= p(x)= nCx px qn-x

Where p(x) denotes the probability of getting exactly x successes.


Properties of Binomial distribution
(1) The shape of the distribution depends on the values of q and p. If p = q, the
shape if it is symmetrical. If p = q the shape of it is, asymmetrical but the
asymmetry decreases as n increases.
(2) Arithmetic mean = np
(3) Standard deviation = npq
(4) Variance = npq
(5) Central moments value
(a) First moment 1 = 0
(b) Second moment 2 = npq
(c) Third moment 3 = npq (q-p)
(d) Fourth moment 4= 3n2p2q2 + npq (1-6pq)
(6) -coefficients
(q-p)2 1- 6pq
(a) 1 = -------- (b) 2 = 3 + --------
npq npq

35 | P a g e
Conditions for using Binomial distribution
(1) The outcome or results of each trial in the process are characterized as one of two
types of possible outcomes.
(2) The possibility of outcome of any trial does not change and is independence of the
results of previous trials.
Use
It is useful in describing an enormous variety of real life events.

POISSON DISTRIBUTION
The Poisson distribution is the limiting form of the binomial probability
distributions n become infinitely large and p approaches 0 in such a way that np = m
remained constant. Such situation are fairly common. That is to say, a Poisson
distribution may be expected in cases were the chance of any individual event being
a success is small. e.g. occurrence of comparatively rare event, such as serious floods,
percentage infestation of any diseases etc.
Like binomial distribution, the variate of the Poisson distribution is also a
discrete one. The probability functions is
e-m mx
P(x) = ---------
x!

Where P(x) represents the number of successes


m represents the average number of successes (m = np)
e is a constant (e = 2.7183)
Properties of Poisson distribution
(1) Arithmetic mean = m
(2) Variance =m
(3) Standard deviation = m
(4) Central moment value : (1) First moment = 1 = 0
(2) Second moment = 2 = m
(3) Third moment = 3 = m2
(4) Fourth moment = 4 = m+3m2
(5)  Coefficients

36 | P a g e
μ32 m 2 1
(1) β1   3 
μ2 m
3
m

μ 4 m  3m 2 1
(2) β 2    3
μ22
m 3
m
Use
Poisson distribution is used in practice in wide variety of problems where there are
infrequently occurring events with respect to time, area, volume or similar unit. For
example it is used in quality control statistics to count the number of defects of an
item, or in biology to count number of bacteria, insects etc.

37 | P a g e
Lecture 7. Correlation Analysis
So far we have studied problems relating to one variable only. In practice we
come across a large number of problems involving the use of two or more than two
variables.
Univariate population
A population that is characterized by a single variable is termed as univariate
population e.g. population of height of students, weight, yield etc.
Bivariate population
When two variables are simultaneously studied in a single population is termed
as bivariate population e.g. the height and weight of the students, rainfall and yield, the
amount of fertilizer used and the crop yield.
If two quantities vary in such a way that movement in one are accompanied by
movements in the other, these quantities are said to be correlated e.g. price of
commodities and amount demanded, increase in rainfall up to a point and production of
crop. The degree of relationship between the variables under consideration is measured
through the correlation analysis.
Correlation
It indicates the association between the two or more variables in a bivariate
distribution or an analysis of the covariation of two or more variables is usually called
correlation.
Types of correlation
Correlation is described or classified in several different ways. Three of the most
important ways of classifying correlation are:
i) Positive or negative
ii) Simple, partial and multiple
iii) Linear and non-linear
Positive and negative correlation
Whether correlation is positive or negative would depend upon the direction of
change of the variable. If both the variables are varying in the same direction i.e. if as one
variable is increasing the other on an average is also increasing, correlation is said to be
positive. Eg: a) Hight and Weight b) Yield and fertilizer, c) Income and Expenditure. If,
on the other hand the variable is varying in opposite directions, i.e. as one variable is

38 | P a g e
increasing the other is decreasing or vice-versa, correlation is said to be negative. Eg:
Yield and Disease Incident, b) Price and supply
Positive correlation

X: 10 12 15 18 20 X: 80 70 60 40 30
Y: 15 20 22 25 37 Y: 50 45 30 20 10

Negative correlation

X: 20 30 40 60 80 X: 100 90 60 40 30
Y: 40 30 22 15 10 Y: 10 20 30 40 50

If the plotted points fall in a narrow band there would be a high degree of
correlation between the variables.
Y x Y x
x x x x
x x xx
xx x x
x x x x
x x xx
X X
High degree +ve r High degree -ve r
Rainfall and yield Intensity of diseases and yield

If the points are widely scattered over the diagram, it is the indication of very little
relationship between the variables.
Y x x Y x x x
x x x x xx
x x x x x x
x x x x x x
x x x x x x
x x x x
X X
Imperfect +ve Imperfect -ve
(Low degree of positive correlation) (Low degree of negative correlation)

If the points lie on a straight line parallel to the X-axis or in a haphazard manner it
shows absence of any relationship between the variables e.g. height of students and
marks.
Simple, Partial and Multiple correlation

When only two variables are studied it is a problem of simple correlation. When
three or more variables are studied it is a problem of either multiple or partial correlation.
39 | P a g e
In multiple correlation, three or more variables are studied simultaneously. In partial
correlation, we recognize more than two variables, but consider only two variables to be
influencing each other, the effect of other influencing variables being kept constant.
Linear and Non-linear (curvilinear) correlation

If the amount of change in one variable tends to bear a constant ratio to the
amount of change in the other variable than the correlation is said to be linear e.g.
X: 10, 20, 30, 40, 50
Y: 70,140, 210, 280, 350

Correlation would be called non-linear if the amount of change in one variable


does not bear a constant ratio to the amount of change in the other variable.
Algebraic method (Karl Pearson Coefficient of correlation)
 (population) and its estimate as ‘r’ (sample) indicate Karl Pearson coefficient of
correlation.
Definition: It is a measure of intensity of association between two variables in a bivariate
population.
Computational formula:

Cov( XY )
 =
 XY

r
Cov( XY
=
 xy =
SP( xy)
S X SY  x . y2 2 SS X .SSY

 X Y 
where,  xy   XY  n
 X  2

x 2
X 2

n
 Y  2

y 2
 Y 2

n

Properties of correlation coefficient:

1. A change in an origin does not affect the value of the correlation coefficient.
2. A change in a scale does not affect the value of correlation coefficient.
3. The value of correlation coefficient lies between -1 to +1.
4. Correlation coefficient is unit free.

40 | P a g e
5. Geometric mean of two-regression coefficient is equal to correlation coefficient.
Test of significance of correlation coefficient

Comparison of sample 'r' with population value

Ho:  = 0 (both the variables are not linearly associated)


H a:   0
1 r2
t(n-2) = r -  / SE of r SE of r =
n2

r
n  2 under Ho :  = 0
1 r2

If cal. t  table t0.05, (n-2) d.f. Ho: rejected


Rejection of Ho: Means there is an association between two variables under study.
If cal. t < table t0.05 (n-2) d.f. Ho: accepted
Acceptance of Ho: indicates that there is no association between two variables in
the population.

Rank Correlation
The Karl Pearson’s method is based on the assumption that the population being
studied is normally distributed. When it is known that the population is not normal, or
when the shape of the distribution is not known there is a need for a measure of
correlation that involves no assumption about the parameters of the population.
This method was developed by Charles Spearman in 1904. This measure is
especially useful when quantitative measures for certain factors can not be fixed e.g. (i)
correlation between marks obtains in two different subjects by the same group of
students. (ii) Correlation of height and weight of the students can be worked out without
making exact measurement. We shall first stand the students according to height; the
same procedure can be utilized for weight for giving ranks. When there are two or more
items are of equal magnitude, their ranks are to be calculated by taking the average of
their ranks.

R  1
 
6  d i2  1 12  P 3 p  or 1
6 d i2

n n2 1  nn 2  1

Where di2 = square of difference of rank


n = number of pairs
P = number of items where ranks are common
41 | P a g e
Lecture 8: Linear Regression Analysis
The word regression was first used by Sir Fransis Galton and he introduces
functional relationships between two variables. Many a times it is observed that change
in one variable from a bivariate population causes change in the other variable, indicating
a cause and effect relationship between the two variables. The former variable is termed
as independent variable whereas, the later as dependent variable. Quantity of fertilizer
and the crop will have this type of cause and effect relationship, where as quantity of
fertilizer could be termed as independent variable and crop yield as dependent variable.
The functional relationship between this independent and dependent variable is known
as regression relationship.
Definition: Regression is a study of average relationship between two or more variables
in terms of original units of the data.
Regression lines
In a scatter diagram if the points are scattered around a line than the relationship
between two variables can be considered as linear. The resulting line is termed as
regression line or line of best fit. For any pair of two variables that are related with each
other linearly a set of two regression lines could be observed and they can be represented
by two equations which are called regression equations. Let X and Y are the two
variables. Then the two regression lines can be given by the following two equations.

𝑌 = 𝛽𝑌𝑋 𝑋 + 𝐶 ...... (i)

𝑋 = 𝛽𝑋𝑌 + 𝐶 ...... (ii)

Where,
yx = Reg. coefficient of Y on X ; xy = Reg. coefficient of X on Y
and c is the intercept

We may observe that in first regression equation Y is considered as the dependent


and X as independent where as in the second it is the reverse case.
These lines have been shown in the following diagram.

42 | P a g e
Fitting of the regression lines

A regression equation which represents a straight line is of the following form.



Y     YX X Y =

Here Y is the dependent variable and X is the independent variable.  yx is


population regression coefficient of Y on X

Intercept   Y   YX X

In case of the sample data the estimates of yx i.e. byx and the estimate of ‘’ as a
are obtained and placed in the equation.

a  Y  bYX X

In a similar fashion the regression equation of the straight line where X is


considered as dependent variable and Y as the independent variable the form of the
equation would be

X   '   XY Y

The estimate of xy is bxy and ' is a '  X  b XY Y

Regression coefficient

Regression coefficient can be defined as the average increase or decrease in the


dependent variable for a unit change in the independent variable or it is the average rate
of change in dependent variable with a unit change in independent variable. It is
represented by yx and xy for the population regression coefficient. In practice they are
estimated with the help of the sample from the bivariate population under consideration
and these estimates are generally represented as byx and bxy respectively.

Method of computation

 YX 
  X   Y   
X Y

Cov( XY )
 X   
2
X
V (X )

bYX 
 X  X Y  Y  
 xy 
 XY   X Y  n
 X  X  x  X   X  n
2 2 2
2

43 | P a g e
Similarly,

 XY 
  X   Y   
X Y

Cov( XY )
 Y   
2
Y
V (Y )

bYX 
 X  X Y  Y  
 xy 
 XY   X Y  n
 Y  Y  y Y  Y  n
2 2 2
2

Test of significance of regression coefficient

1) When our interest is to ascertain whether the effect of the independent variable
on the dependent variable is appreciable or not, we employ 't' test.

Ho : yx = 0
Ha : yx = 0

 y   xy   x
2
b YX 2 2
t SE of b YX 
SE of b YX n  2 x 2

Where, n = size of the sample

The calculated t value is to be compared with the table t value at the desired level of
significance with (n-2) d.f. and conclusion is to be drawn.

Properties of regression coefficient

1) Geometric mean between regression coefficients is correlation coefficient i.e. r =


b yx . b xy
a) Arithmetic mean of byx & bxy is equal to or greater than correlation coefficient
b Y X  b XY
i.e. r
2
b) If one regression coefficient is greater than unity than other regression
coefficient must be less than unity.
2) Regression coefficient is independent of origin but not scale
2) Regression coefficient lies between -  to + 
3) Regression coefficient has unit
5) Regression coefficient has one way relationship
Uses of regression
1) To predict the value of Y for a given value of X with the help of regression equation.
2) To know the rate of change in Y for a unit change in X with the help of regression
coefficient.
44 | P a g e
Relations among r, byx, bxy, Sx and Sy

SY SX
(i) r  b YX .b XY (ii) bYX  r (iii) b XY  r
SX SY

Differences between Correlation and Regression


Correlation Regression
1 It deals with mutual association It deals with cause and effect
relationship
2 It is two way relationship It is one way relationship
3 Correlation coefficient is unit free Regression coefficient is in the units
of dependent variable
4 Correlation coefficient lies Regression coefficient lies between
between - 1 to + 1 -  and + 
5 For a given value of one variable For a given value of independent
other variable can not be variable the value of the dependent
predicted variable can be predicted.

45 | P a g e
Lecture 9: Introduction to test of Significance
The subject of statistics deals with statistical estimation and testing of
statistical hypothesis. These are the two important functions for drawing inference about
the population parameters. Statistical estimation is the technique of estimating the
population parameter values on the basis of information obtained from the sample.
Suppose we wish to know the yield of a crop. To know this figure, it is not necessary to
harvest entire field of that crop or all the fields of that crop grown in the region. One may
collect the sample from the fields by appropriate sampling procedure and on the basis of
sample information, one may estimate the average yield of the crop of entire area. The
estimate thus obtained is not the final form for drawing valid conclusion regarding
population from which the samples are drawn. It needs to be tested by applying an
appropriate test or method. Such test is known as the test of significance. Thus, test of
significance can be defined as “The statistical procedure for deciding whether the
observed difference between sample estimate and population Parametric value is
significant or not at specified level of significance".
Hypothesis: It is the statement specifying the parametric value of a distribution from
which the sample/s is/are drawn.
Null Hypothesis: It is a hypothesis of no difference between different populations
parametric values from which samples are drawn OR It is the
hypothesis of equality of population parametric values from which
sample/s is/are drawn.
Procedure for testing a hypothesis
Step I: Set appropriate null hypothesis
Let us consider that there are two methods for preparing compost.
Method A: standard method and
Method B: new method
Now to test which method is better, the hypothesis can be
1) B is better than A B > A
2) A is better than B A > B
3) B is not different from A A = B
The first two statements indicate a preferential attitude to one or the other of the
two methods. Hence, it is better to adopt the third statement and make the test. This third

46 | P a g e
statement is called the null hypothesis, which is denoted as Ho: symbolically Ho : 1 = 2
or 1 - 2 = 0 where 1 and 2 are the population parametric values.
In the above examples, suppose in first method the average nitrogen content is 1
and in the second method the average nitrogen is 2. Ho : 1 = 2 can be tested by the
appropriate test. As against the null hypothesis, the alternative hypothesis should also be
set up, which specifies those values, the researcher believes to hold true. Since one is
going to accept or reject the null hypothesis one has to set the alternative hypothesis also.
It is denoted by Ha : 1  2 or Ha : 1 < 2, 1 > 2 .
Step II: Fix appropriate level of significance
The confidence with which an experimenter reject or accept the null hypothesis
depends upon the significance level adopted. It is expressed in percentage such as 5 per
cent, 1 per cent etc. When the hypothesis in question, is accepted at 5 per cent level of
significance, the experimenter is running the risk that in the repeated cases
(experiments/trials), he will be making the wrong decision in about 5 per cent of the
cases. By rejecting the hypothesis at the same level, he runs the risk of rejecting a true
hypothesis in 5 out of every 100 occasions. Thus, level of significance is defined as:
"It is the maximum probability at which one would like to reject the null
hypothesis when it is true OR The level of significance is the average proportion of
incorrect statements made when the null hypothesis is true."
Step III: Set suitable test criterion
To construct a test criterion, one has to select the appropriate probability
distribution for the particular test viz. Z, t, F, 2 etc.
Step IV: Computation
This step involves the calculations of various statistics from sample data such as
mean and standard error of mean.
Step V: Conclusion
After doing the necessary calculations one has to decide whether to accept or
reject the null hypothesis at a certain level of significance. Therefore, the computed value
of the test criterion is compared with the table value. If the computed value is greater
table value, the observed difference is significant and Ho is not accepted. If calculated
value is less than or equal to table value the Ho is accepted at a given level of significance.
Not acceptance of Ho means the difference between sample estimate and the hypothetical
parametric value is a real difference, while acceptance of Ho means the difference

47 | P a g e
between sample estimate and population/hypothetical parametric value can be
explained due to chance variation (sampling variation).
Type - I and Type - II errors
While testing the hypothesis one is liable to commit two kinds of errors.
An error of first kind is made by rejecting the true null hypothesis. The
probability of committing a Type-I error is denoted by  (alpha).
Type-II error is committed by accepting the null hypothesis when it is false. The
probability of Type-II error is denoted by  (Beta).
Type-I error depends on the level of significance. When 5 per cent level of
significance is fixed, we fixed the probability of committing Type-I error at 5 per cent.
It is possible to control Type-I error by shifting the level of significance. Type-II
error increases as the Type-I error decreases. Therefore, the common practice is to keep
the Type-I error at five percent or one percent fixed and try to decrease Type-II error by
increasing sample size and following refined technique of conducting experiment.
Degrees of freedom
For testing any hypothesis the estimated statistic is compared with table value.
The knowledge of degrees of freedom is essential for referring the table value. With X1,
X2,.....Xn having constant sum, (n -1)X values can be given freely, but the nth X value will
be determined by the condition that the sum of all `X` is equal to the given constant
quantity i.e. one degree of freedom is lost. So in one way classification, number of
observations - 1 is called degrees of freedom and in general number of observation minus
number of independent constraints or restrictions is called degrees of freedom.

48 | P a g e
Lecture No. 10: One sample and two sample t-test for Mean

Small sample or Student's ‘t’ test


When the sample is large and if r is not known, we estimate the same and can be
used in Z test. But if 'n' is small error will be more for replacing r by S and under that
situation the Z remain no longer normal, but changes to another distribution named "t".
The "t" distribution was found out by W.S. Gossett in the name of 'Student' in 1908.
Values of 't' depends on degree of freedom and is always greater than its limiting
value of Z for any unit degree of freedom. When d.f. is large tZ. Difference between t
and Z becomes more and more marked as n become smaller and smaller.
Definition: It is the ratio of the deviation between sample mean and hypothetical mean
to the standard error of mean estimated from the small sample.
Conditions for applying 't' test
1) Data follow normal distribution.
2) The sample is small (n < 30) and the standard deviation of the population is estimated
from the sample.
Uses
1) Comparing sample mean with hypothetical mean or population mean.
2) Comparing two sample means.
(a) When the number of observation of both the samples are unequal
(n1  n2).
(b) When number of observations of both the samples are equal (n1 = n2).
(c) When the observations are paired.
3) Comparing the regression coefficient of sample with the hypothetical or population
regression coefficient.
4) Comparing the correlation coefficient with the correlation coefficient of population.
5) Comparing two regression coefficients.
Characteristics of "t" distribution
1) It is the exact distribution and not approximate.
2) t value ranges from -  to + 
3) The distribution is symmetrical one.

49 | P a g e
4) It is flatter than the normal distribution i.e. the area near the tail is large for t
distribution compared to normal distribution. Value of coefficient of kurtosis is less
than 3.
5) As sample size increases, the t distribution approaches to normal distribution.
6) There is need to know the d.f. to obtained the probability value from the table.

One Sample 't' test


Objective: To test whether the given small sample (n < 30) has come from the population
having mean .
Procedure :
If (i) X1, X2, X3, ... , Xn is the given sample (n < 30) or
(ii) class value : X1, X2, ... ,Xk with corresponding
frequencies : f1, f2, ... ,fk (fi = n) of a given sample,
Step I : Set the null hypothesis : Ho :  = o or Ho :  - o = 0
Ha :   o (two tailed test) or
 < o,  > o (one tailed test)
Where,  is the population mean from which the random sample has been
drawn and o is the mean of the hypothetical population.
Step II : Fix the level of significance .Usually 5 and 1 per cent levels of significance are
fixed.
Step III : Calculate the following estimates.
If (i) X1, X2, X3, ... , Xn is the given sample (n < 30) or
(ii) class value : X1, X2,...,Xk with corresponding frequencies
f1, f2, ... ,fk (fi = n) of a given sample, Calculate,
n
 Xi
i1
Sample mean X 
n
n
 (X i - X) 2
i1
Variance : S2 =
(n  1)

Standard error of mean:

S 
SX  = (If population standard deviation is known)
n n

Step IV : Compute the student 't' with (n-1) degree of freedom

50 | P a g e
X  O
t 
SX

Step V : Conclusion
If calculated t < table t0.05,(n-1) d.f. observed difference is not significant. Ho:  = o is
accepted. Acceptance of Ho:  = o means the given small random sample has come from
the hypothetical population having mean o. If calculated t  table t0.05,(n-1) d.f. observed
difference is significant. Ho:  = o is rejected. If calculated t  table t0.01, (n-1) d.f., observed
difference is highly significant. Ho:  = o is rejected. Rejection of Ho:  = o means the
given random sample does not come from the hypothetical population having mean o.

Two sample 't' test (Independent sample)


Objective : To test whether the given two small random samples have come from the
same population having mean o.
Let sample - I : X1, X2, ... ,Xn1
sample - II : Y1, Y2, ... ,Yn2 are two random samples drawn from a
population.
Procedure:
Step I: Set the null hypothesis that both the samples have come from the same
population having mean  and standard deviation S.
i.e. Ho : 1 = 2 =  against Ha : 1  2 or
1 > 2 or 1 < 2

Where 1 is the population mean from which sample one is drawn and 2 is
the population mean from which the second sample is drawn.
Step II: Fix the level of significance. Usually 5 per cent and 1 per cent levels of
significance are fixed.

51 | P a g e
Step III: Calculate the following estimates.
Sample - I Sample - II
n1 n2
 Xi  Yi
i1 i1
i) Mean X  Y 
n1 n2

ii) Variance
n1 n2
 ( X i - X) 2  (Yi - Y)2
i1 i1
S2x = S 2y =
(n1  1) (n2  1)

(iii) Pooled sample variance


n1 n2
 ( X i - X) 2   (Yi - Y) 2
i1 i1
Sp2 
n1  n 2 - 2

(iv) Standard error of mean of differences

1 1
Sp2 (  )
S  n1 n2
( X Y)

Step IV: Calculate student 't' with n1 + n2 - 2 d.f.

( X  1)  (Y   2 )
t 
S( X  Y )

Step V: If cal t < Table t 0.05, (n1+n2-2) d.f. difference is non significant at 5% level of
significance Ho: 1 = 2 accepted. Acceptance of Ho: 1 = 2 means both the
samples have came from the same population 
If cal t  table t0.05,(n1+n2-2) d.f. difference is significant at 5 % level of significant
Ho : 1 = 2 rejected at 5 % level of significance.
If cal t  table t 0.01,(n1+n2-2) d.f. difference is highly significant 1 % level of
significance. Ho:1 = 2 rejected at 1% level of significance. Rejection of Ho : 1 =
2 means both the sample are drawn from two different populations.

52 | P a g e
Two sample 't' test ( Dependent sample) : Paired 't' test

Objective: To test whether the two small related random samples have come from the
same population.

Let Sample-I : X1, X2, ... ,Xn and


Sample-II: Y1, Y2, ... ,Yn be two related sample such that (X1, Y1), (X2, Y2), ... , ( Xn, Yn)
are the pairs of related observations.
Procedure :
Step I : Set the null hypothesis : Ho : d = 0 ; Ha : d  0
Where, d is the average difference between Xi - Yi in the population.
Step II: Fix the level of significance. Usually 5 per cent and 1 per cent levels of
significance are fixed.
Step III: Calculate the following estimates.

(i) di = Xi - Yi i = 1,... ,n

(ii) d  d n i

 d  d
2


2 i
(iii) S
n 1

 d 
2
S2 d
(iv) S.E. of d  
i

n nn  1

Step IV: Calculate the student t with n-1 d.f.

d  d
t  (Under Ho : d = 0 )
Sd
Step V: Conclusion
If cal. t < table t 0.05, (n-1) d.f., observed difference is non-significant at 5% level of
significance and null hypothesis (Ho) is accepted.
Acceptance of null hypothesis (Ho: d = 0) means the given two related small
samples have come from the same population.
If cal. t  table t 0.05, (n-1) d.f., observed difference is significant at 5% level of
significance and null hypothesis (Ho) is rejected at 5% level of significance. If
cal. t  table t 0.01, (n-1) d.f., observed difference is highly significant at 1% level of
significance and null hypothesis (Ho) is rejected at 1% level of significance.
53 | P a g e
Rejection of null hypothesis (Ho : d = 0) means the given two related samples
does not come from the same population.

Large sample test: Z - test


It is a large sample test and can be utilized for testing the hypothesis if the
following conditions are satisfied.
(1) Data follow normal distribution.
(2) Sample size should be large ( n > 30 ) or
(3) The standard deviation of population should be known if sample is not large.
Z test can be defined as "It is the ratio of the difference between the estimated
population mean and hypothetical mean to the standard error of mean based on
population standard deviation or its estimate from large sample.

One sample Z test


Objective: To test whether the given large random sample has come from the given
population with mean  and variance 2 or its estimate S2.
Procedure :
Step I : Set the null hypothesis : Ho :  = o or Ho :  - o = 0
Ha :   o or (two tailed)
 < o,  > o ( one tailed test)
Where,  is the population mean from which the random sample has been
drawn and o is the mean of the hypothetical population.
Step II : Fix the level of significance:
Step III: Computation
If (i) X1, X2, X3, ... , Xn is the given sample (n > 30) or
(ii) class value : X1, X2,...,Xk with corresponding frequencies: f1, f2, ... ,fk
(Total = n) of given sample, calculate,
n n
 Xi  (X i - X) 2
i1 i1
Sample mean X  Variance : S2 =
n (n  1)

Standard error of mean :


S 
SX  = (If population standard deviation is known)
n n

54 | P a g e
Step IV : Calculate normal deviate
X O
Z 
SX

Step - V : Conclusion: If calculated Z  1.96 the difference is non significant at 5% level


of significant. Ho:  = o is accepted. Acceptance of hypothesis revealed that
the given sample has come from the population having mean o. If calculated
Z  1.96 the difference is significant at 5% level of significant. Ho:  = o is
rejected. If calculated Z  2.58 the difference is highly significant at 1% level of
significant. Ho:  = o rejected Rejection of hypothesis indicates that the given
sample does not come from the population having mean o. The confidence of
rejection being 95 per cent. If the difference is highly significant, the
confidence of rejection is 99 per cent.

Two sample Z test


Objective: To test whether two randomly selected samples have come from the same
population having mean  and a standard deviation  or its estimate S.
Generally there is little interest in comparing sample mean with population mean.
A more frequent problem usually met with in agriculture is involved in comparison of 2
samples i.e. means of 2 samples. e.g.
i) We may require comparing variety A with variety B of a crop.
ii) Comparison of two manure for their effect on yield.
iii) Comparison of two rations for milk yield.
iv) Comparison of yield of the same varieties on two different farms etc. need to be
tested for their difference in means (for location specificity).
Let, Sample I : X1, X2, ... ... , Xn1 and

Sample II : Y1, Y2, ... ... ,Yn2 are random samples drawn from normally distributed
populations
OR
Let Sample - I Sample –II
Class value : X1, X2, ... , Xk1 Y1, Y2, ... , Yk2
Frequency : f1, f2, ... , fk1 = n1 f1, f2, ... , fk2 = n2
are two frequency distributions of two samples drawn from a normal populations.

55 | P a g e
Procedure :
Step I: Set the null hypothesis that both the samples have come from the same
population having mean  and standard deviation  or its estimate S.
i.e. Ho : 1 = 2 =  against Ha : 1  2 or
1 > 2 or 1 < 2
Where 1 is the population mean from which sample one is drawn and 2 is the
population mean from which the second sample is drawn.

Step II : Fix the level of significance. Usually 5 per cent and 1 per cent levels of
significance are fixed.
Step III : Calculate the following estimates.

Sample-I Sample-II
(i) Mean :
n1 n2
 Xi  Yi
i1 i1
X  Y 
n1 n2
(ii) Variance :
n1 n2
 ( X i - X) 2  (Yi - Y)2
i1 i1
S2x = S 2y =
(n1  1) (n2  1)

(v) Pooled sample variance

n1 n2
 (X i - X) 2   (Yi - Y) 2
i1 i1
Sp2 
n1  n 2 - 2

(vi) Standard error of mean of differences

1 1
Sp2 (  )
S  n1 n2
( X Y)

Step IV : Calculate normal deviate (Z)

( X   1 )  (Y   2 )
Z 
S( X  Y )

Step V: Conclusion:

56 | P a g e
If calculated Z  1.96 the observed difference is non significant at 5% level of
significance and Ho: 1 = 2 accepted. Acceptance of Ho: 1 = 2 means both the samples
have came from the same population. If calculated Z > 1.96, the observed difference is
significant at 5% level of significant and Ho : 1 = 2 rejected at 5 percent level of
significance and if calculated Z > 2.58, the observed difference is highly significant at
1% level of significance hence Ho : 1 = 2 rejected at 1% level of significance. Rejection of
Ho: 1 = 2 means both the sample are drawn from two different populations.

57 | P a g e
Lecture 11: Chi-Square Test of Independence of Attributes in 2x 2 Contingency Table
2 - test ( Chi-square test)
Chi-square was introduced by Karl Pearson in the year 1899. It is calculated by

K
( i - E i ) 2
2 = 
i 1 Ei
Where, Oi = Observed frequency of ith class
Ei = Expected frequency of ith class
k = number of classes , i = 1,2,..,k
Definition : "It is the sum of the ratio of the square of deviations obtained between
observed and expected frequency to the expected frequency of the
respective class of the frequency distribution."

Properties of Chi-Square distribution


1) Chi-square distribution is not exact distribution as “t" distribution.
2) It is not symmetrical distribution but it is positively skewed distribution.
3) The value of its varies from 0 to . When there is a perfect agreement of observed
frequency distribution with hypothetical frequency distribution, the value of chi-
square will be zero, while the value of its increases as there is a departure from the
agreement and will increased up to infinity.
4) If 21, 22,... , 2k are chi-square values of different samples with n1, n2,...,nk degrees of
freedom respectively, the pooled chi-square value will be equal to
K
( Οi - Ei )2
2 p =  with n =  ni degrees of freedom.
i 1 Ei
5) The different central moments are 2 = 2n;  3 = 8n, 4 = 48 n + 12n2.
6) As the number of observation tends to infinity, the chi-square distribution tends to
normality.
7) The table chi-square value depends upon degrees of freedom. The table chi-square
values can be obtained for 1 to 30 d.f., then it is not available from the table. As the

number of degrees of freedom exceeds 30, it is found that  2 will be distributed

approximately normal about the mean 2n  1 with a unit standard deviation.

58 | P a g e
Therefore, the Z value can be worked out by using the following formula and it
should be compared with table Z value at 5 per cent or 1 per cent level of significance.

Z = 2  2  2n  1
Conditions for application of Chi-Square
1) Deviations (Oi - Ei) should be normally distributed.
2) Number of observations should be sufficiently large. It should be at least 50.
3) Expected frequency of any cell should not be very small. It should be at least 5 and
better if it is 10.
Uses
1) Testing goodness of fit.
2) Testing the independence of attributes for 2 x 2, 2 x c, r x 2 and r x c contingency table.
3) Testing the agreement of genetic ratio with the observed ratio.
4) Test of homogeneity of the families.
5) Test for the detection of linkage.
6) Testing of homogeneity of various variances (Bartlett's test of homogeneity)
7) Testing the heterogeneity among correlation coefficients.
1) Testing goodness of fit
When chi-square test is used to know whether the given sampling distribution is
in agreement with the theoretical or expected frequency distribution the test is
known as test of goodness of fit.
Procedure:
Step I : Set the appropriate null hypothesis.
Ho : Given sampling distribution is in the agreement with theoretical or
expected frequency distribution.
Ha : Given sampling distribution is not in agreement with theoretical or
expected frequency distribution.
Step II : Fix the level of significance.
Step III: Work out expected frequency according to given ratio or expectation.
Step IV: Calculate Chi square as
K
( i - E i ) 2
2 = 
i 1 Ei
Where, Oi = Observed frequency of ith class

59 | P a g e
Ei = Expected frequency of ith class
k = number of classes , i = 1,2,..,k
Step V: Compare cal 2 with table value at 5% level of significance and (k-1) degree of
freedom.
Step VI: If cal 2 > table 2 0.05, (k-1)d.f. observed difference is significant at 5% level of
significance. Ho rejected.
If cal 2 < table 2 0.05, (k-1) d.f. observed difference is not significant at 5% level
of significance. Ho accepted.
Step VII: Conclusion: Non significance difference indicates that the given sampling
distribution is in agreement with theoretical distribution and the fit is good.
Significant difference indicates that the given sampling distribution is not in
agreement with theoretical distribution and the fit is poor.
2) Test of Independence
Another common use of the chi square test is in testing independence of
classifications.
Independence: The two attributes A and B are said to be independent to each other if the
proportion of A's among B's is the same as that in not - B's.
Variable: Any character which varying from individual to individual is termed as
variable.
Attribute: Attribute is that which is not capable of being described numerically e.g. sex,
blindness, colour, shape.
Contingency table: When the individuals in a sample have two characters or attributes
and a frequency distribution is made classifying them according to both so as
to show the relation between the characters, the resulted table is termed as
contingency table.

Procedure for test of Independence of attribute in case of 2 x 2 contingency table:

Step I: Set the appropriate null hypothesis.


Ho: The given classification of group of individuals independent to each other.
Ha: The given classification of group of individuals is not independent to each
other.

Step II: Fix the level of significance.


Step III: Let group A and B are classified in two ways, the results of the classification
can be set out the following table.
60 | P a g e
Class A1 A2 Total
B1 a b R1
B2 c d R2
Total C1 C2 N

Step IV: Calculate Chi square as

 2

ad  bc  .N 2
2

R1 .R2 C1 .C 2
Where, a, b, c and d are the observed frequency of the respective cell R1, R2, C1, and
C2 are the rows and column totals. N is the grand total.

Step V: Compare calculated 2 with table value at 5% level of significance and (r-1) (c-1)
degree of freedom.

Step VI: If cal 2 > table 2 0.05, (r-1)(c-1) d.f., observed difference is significant at 5% level of
significance. Ho rejected.
If cal 2 < table 20.05, (r-1)(c-1) d.f., observed difference is not significant at 5%
level of significance. Ho accepted.

Step VII: Acceptance of Ho means the two characters are independent to each other.
Rejection of Ho means the two characters are not independent to each
other
Step IV: Work out expected frequency of each cell as follows.
R 1C1 R 2 C1
E(a11) = E(a21) =
N N
R1C 2 R 2C2
E(b12) = E(b22) =
N N
In general,
R iC j
E(Xij) =
N

Step V: Calculate Chi square as


K
( i - E i ) 2
2 = 
i 1 Ei
Where, Oi = Observed frequency of ith class
Ei = Expected frequency of ith class
k = number of classes , i = 1,2,..,k

61 | P a g e
Step VI: Compare calculated 2 with table value at 5% level of significance and (r-1) (c-
1) degree of freedom.
Step VII: If cal 2 > table 20.05, (r-1)(c-1) d.f., observed difference is significant at 5% level
of significance. Ho rejected.
If cal 2 < table 20.05, (r-1)(c-1) d.f., observed difference is not significant at 5%
level of significance. Ho accepted.
Step VIII: Acceptance of Ho means the two characters are independent to each other.
Rejection of Ho means the two characters are not independent to each
other

F test
t and Z tests are used for comparing two populations mean. When the population
is to be compared with respect to their variances the F test is used.
Definition: It is the ratio of the estimates of greater mean square or variance to smaller
variance of two different populations.

S12
F  ; S12 > S22
S 22

Compare cal F with table F (n1 - 1) and (n2 - 1) d.f. and draw the conclusion.

62 | P a g e
Lecture 12: Introduction to Analysis of Variance and One Way ANOVA
Analysis of Variance (ANOVA):
Analysis of Variance (ANOVA) is a hypothesis-testing technique used to test the equality
of three or more population (or treatment) means by examining the variances of samples
that are taken.
ANOVA allows one to determine whether the differences between the samples are
simply due to random error (sampling errors) or whether there are systematic treatment
effects that causes the mean in one group to differ from the mean in another. Most of the
time ANOVA is used to compare the equality of three or more means, however when the
means from two samples are compared using ANOVA it is equivalent to using a t-test to
compare the means of independent samples.
ANOVA is based on comparing the variance (or variation) between the data samples to
variation within each particular sample. If the between variation is much larger than the
within variation, the means of different samples will not be equal. If the between and
within variations are approximately the same size, then there will be no significant
difference between sample means. There are two types of ANOVA that are commonly
used, the One-Way ANOVA and the Two-Way ANOVA.

Definition of ANOVA: It is a mathematical process of partitioning the total sum of


squares into various recognized sources of variation.
Assumptions underlying One Way ANOVA
For valid use of ANOVA certain assumption should be satisfied.
1. The population from which each sample mean has been drawn is normally
distributed.
2. The treatment effects and error are additive in nature.
3. Errors are normally and independently distributed with mean zero and common
variance σ2.

One-way ANOVA
A one-way ANOVA has one independent variable and a two-way ANOVA was two
independent variables.
Statistical Model
Yij  μ  τ i  ε ij
:

63 | P a g e
Where,Yij = Response of yield from the jth unit receiving the ith treatment
 = General mean
 i = Effect of ith treatment
 ij
= Uncontrolled variation associated with jth unit receiving ith treatments.

Analysis:
H0: All the treatments are equal or T1= T2=T3=…..=Tt
H1: At least one variety is different from others

GT 2
1) Correction factor (C.F) =
n
2) Total SS= ( y ij ) 2 -C.F

  yi . 2 
3) SS for varieties =    CF
 k 
 
4) SS for error = Total S. S – Treatment S. S
One way Analysis of variance
Source of M. S.
DF Sum of Squares (SS) Cal. F
variation (SS/DF)
Between the (t-1) t t r MST MST MSE
Yi.2 ( Y ij )
2
Treatment i 1 j 1
i 1

k kt
Between the By subtraction MSE
t(k-1)
Treatment
t r
( Y ij )2
Total (kt-1) t k

 Y
i 1 j 1
2
ij 
i 1 j 1

kt

Since, the Calculated F ˃ Table F0.05 at (4, 15) d.f, reject H0 and accept H1, there are
significant differences between the treatment means.

1 1
SEm 
MS E
or SEd  MS E   
r or r0 r r 
 i j 

CD (Critical Difference): It is such a significant difference that all the differences of two
mean greater than or equal to it are considered to be significant. Therefore it is the least
significant difference. It’s value is given by

CD (Critical Difference) at 5 % level of significance = SEm x 2 to.o5 at Error d.f.

64 | P a g e
Lecture 13: Introduction to sampling and Sampling versus Complete Enumeration
Introduction
Our knowledge, our actions and our attitudes are based to a very large extent on
samples. This is equally true in everyday life and in scientific research, whenever, a
person wants to buy a large quantity of a commodity say wheat, rice, fruits etc. be
decided about total lot of just by simply examining a small fraction of it . A doctor's
opinion about the state of health is determined by only examining one or two drops of
blood. We study the soil about its nutrient status, salinity status etc by using only 5g. of
soil. It has been experienced that the sample survey if planned properly, can give precise
information.
A sample is a part of population (total or aggregate). It is to be selected at
random for unbiased estimate and can be used as a basis for inferring about the
population. Most populations (a crop plant population, groundnut growers, pest
population pump-set owners etc.) are so large that it is rather next to impossible to
contact each of them during specified time, only a fraction of it is investigated and the
inference is drawn from it for the population. Consequently many generalizations which
arise from ordinary experience are likely to be unwarranted.
The sample has many advantages over a census or complete enumeration of a
finite population. If carefully designed, the sample is not only cheaper but may give
results which are as accurate or sometimes, more accurate than those of census. The
reason is that a census is subjected to biased error than that due to any sampling method
employed. The smallness of samples makes possible the use of the more care in the
design and execution of each step in the inquiry. In this way sources of error can be
investigated and reduced, eliminated or measured and corrected for. The other
advantages of sample survey are that it is less time consuming, involves less cost, has
greater scope and has greater operational facilities. It is for these reasons that sample
surveys are being preferred by the research scientists to complete enumeration.

65 | P a g e
Terminology
1) Population: A population is the aggregate of individuals or units or objects having at
least one common characteristics.
2) Sample: A sample is a part or fraction of large aggregate (Population) about which
some information is required.
3) Sampling: It is the method/process of selection of sample from the population.
4) Sampling Unit: It is the individual element or a group of elements on which
observations are to be made.
5) Parameter: Any measurable characteristic of population which is to be worked out by
utilizing each and every observation of population parameters help to
characterize the population mean  and  are the example of the
parameter. Parameters are generally unknown and constant.
6) Statistic: Any number estimated from sample (X and S are the statistic)

Types of population
1) Finite population: One can count rose plants in a garden i.e. counting the individuals
in the population is possible & hence it is finite population. e.g. Rose
plants in garden, No. of animals in herd, Fruits on tree etc.
2) Infinite population: The rose plants on the earth can not be counted. Thus it becomes
infinite population.
3) Real population: This is the population in which the members do exist in reality. e.g. a
heap of food grains, herd of cows etc.

4) Hypothetical population: The member of population does not exist in reality. e.g.
Possible throws of a die. Similarly, yield of a new variety to be evolved.
Possible results of an experiments etc. are the examples of this type of
population.
Advantages of sample study
1) Less expensive.
2) In a limited specific time one can complete the project work and collect the
information (i.e. greater speed)
3) With minimum technical persons, one can study the problem

66 | P a g e
4) More precise and accurate information regarding population (greater accuracy) may be
obtained.
5) Sample investigation can be carryout with minimum facilities.
6) Sample investigation is to be done when the unit is supposed to be destroyed while
taking observation.
Sampling plans/Designs
1) Simple random sampling.
2) Stratified random sampling.
3) Multistage sampling.
4) Cluster sampling.
5) Systematic sampling.
6) Purposive sampling and so on.
(1) Simple random sampling (SRS)
In a sample random sampling, the sample is selected in such a way that every
member of the population has an equal and independent chance of being selected in
the sample. It implies that selection of a sample from all possible samples that could be
chosen is equally likely.
Random selection of units is done using any one of the following methods.
a) Using Tickets, tags etc.
b) Random number tables.
a) Using Tickets, tags etc
To give an example, suppose that an experimenter wishes to draw a random
sample of size 10 individuals (say 10 plants for measuring the heights) from a finite
population, say 200 plants. A method of doing this would be assign a number to each
member of the population put the individuals into a box and mix them thoroughly. Draw
ten tickets or tags from the box. The number on these ten tickets or tags correspond the
plants to be selected. This method is time and labor consuming and also costly.
b) Use of Random Number table
The procedure (a) can be shortened by the use of a table of random number. Such
a table consists of numbers chosen in a fashion similar to drawing numbered tickets or
tags out of a box. This table is so made that all numbers 0, 1, 2.... appear with
approximately the same frequency. By combining the numbers in pairs we have two-

67 | P a g e
digit numbers (i.e. from 00 to 99) By using the number three at a time we have three digit
numbers from 000 to 999 etc.
The table should be entered in a random manner. Put a pencil aimlessly on a page
of the table. The point thus obtained on page is the starting point for selecting the
numbers, record the numbers until the required number of random digits is obtained.
Examples of SRS
1) Impact of T & V system on Socio-economic status of summer groundnut growers in a
specific Taluka.
Here population is “Summer groundnut growers registered under T & V system
of a given taluka. One can easily prepare the frame for this finite population and select
the random sample of summer groundnut growers.
2) Constraint analysis for milk: Productivity in a village
Here population is “milk producers in a given village. The frame for which can be
prepared easily and a random sample can be taken to fine out the factors responsible for
low productivity.
(2) Stratified Random sampling
In this scheme the heterogeneous population is sub-divided into homogeneous
several groups called stratum and then samples are drawn independently (at random)
from each stratum.
As the sampling variance of the estimate of mean depends on the within strata
variation, the stratification of heterogeneous population into homogeneous strata helps in
increasing precision of the estimates e.g. while studying average income of the staff
members of the Gujarat Agricultural University one has to employ stratified random
sampling. Because, simple random sampling may result into under or over estimation of
income of the staff member. Suppose only professors are selected for the sample then the
average income will be above the true average value, similarly if only helpers are selected
in the sample then the results will be on lower side.
When such heterogeneous population is required to be sampled then one has to
utilize the stratified random sampling in spite of simple random sampling.

68 | P a g e
Examples
1) Adoption level of improved agro technology by the cultivators.
2) A survey for nutritional status of school going students.
3) Socio-economic survey for rural and urban people of Valsad district.
(3) Multistage sampling
In this method, the selection is done in stages, called sub sampling. For example
in estimating the yield of wheat in a district. Talukas may be considered as primary
samplings unit (1st stage), Villages within taluka as secondary sampling units (2nd stage)
the cultivators within village within taluka as third stage sampling units and so on.
The advantage of this type of sampling is that at the first stage the frame of
primary sampling units is required which is easy to have and at the second stage the
frame of secondary sampling units is required for the selected primary sampling units
only and so on. Moreover, this method allows the use of different selection procedure in
different stages. It is because of this consideration that multi-stage sampling is used in
most of the large scale surveys.
(4) Systematic sampling
Systematic sampling is slight varying compared to the simple random sampling
in which only the first sample units is selected at random and the remaining units are
automatically selected in a definite sequence at equal spacing from one another. This
technique of drawing samples is usually recommended if the complete and up to date list
of the sampling units is available and the units are arranged in some systematic order
such alphabetical, chronological, geographical order etc. This requires the sampling units
in the population to be ordered in such a way that each item in the population is
uniquely identified by its order, for example the name of persons in a telephone
directory, the list of voters etc.
Sampling and Non-sampling Errors.
The inaccuracies or errors in any statistical investigation i.e. in the collection,
processing, analysis and interpretation of the data may be broadly classified as follows:
1) Sampling errors 2) Non sampling errors.

69 | P a g e
(1) Sampling Errors
In a sample survey, since only a small portion of population is studied and hence
its results are bound to differ from the census results and thus has a certain amount of
error. This error would always be there, no matter that the sample is drawn at random
and that it is highly representative. This error is attributed to fluctuations of sampling
and is called sampling error. Sampling error is due to the fact that only a subset of the
population (i.e. sample) has been used to estimate the population parameters and draw
inferences about the population. Thus, sampling error is present only in a sample survey
while it is completely absent in census surveys.
Sampling error may be due to following reasons.
(i) Faulty selection of the sample
(ii) Substitution
(iii) Faulty demarcation of sampling units
(iv) Error due to bias in the estimation method
(v) Variability of the population.
(2) Non Sampling Errors
Non-sampling errors are not attributed to chance and are a consequence of certain
factors which are within human control. In other words they are due to certain causes
which can be traced and may arise at any stage of the inquiry viz. panning and execution
of the survey and collection, processing and analysis of the data. This error is present in
sample and census.
Some of the important factors responsible for non sampling errors are as under
(i) Faulty planning including vague and faulty definitions of the population of the
statistical units to be used, incomplete list of population members.
(ii) Vague and imperfect questionnaire which might result in incomplete or wrong
information.
(iii) Defective methods of interviewing and asking questions.
(iv) Vagueness about the type of the data to be collected.
(v) Personal bias of the investigator.
(vi) Lack of trained and qualified investigators and lack of supervisory staff.
(vii) Failure of respondents memory to recall the events or happenings in the past.
(viii) Non response and inadequate or incomplete response.
(ix) Improper coverage.

70 | P a g e
(x) Compiling errors.
(xi) Publication errors.

Sampling versus Complete Enumeration


BASIS FOR CENSUS SAMPLING
COMPARISON (Complete Enumeration)
Meaning A systematic method that collects Sampling refers to a portion of
and records the data about all the the population selected to
members of the population is represent the entire group, in
called Census. all its characteristics.
Enumeration Complete Partial
Study of Each and every unit of the Only a handful of units of the
population. population.
Time required It is a time consuming process. It is a fast process.
Cost Expensive method Economical method
Results Reliable and accurate Less reliable and accurate, due
to the margin of error in the
data collected.
Error Not present. Depends on the size of the
population
Appropriate for Population of heterogeneous Population of homogeneous
nature. nature.

71 | P a g e
Lecture 14: SRS with replacement (SRSWR) SRS, without replacement (SRSWOR)and
Use of Random Number Tables for selection of Simple Random Sample
There are wo type of SRS. They are
• SRS with replacement (SRSWR)
• SRS without replacement (SRSWOR)
SRS with replacement (SRSWR)
It is the random process in which a unit is selected and noted and then returned to the
population before the next drawing is made and this process is repeated n times, it give
rise to a simple random sample of n units.. This process is generally known as simple
random random sampling with replacement (SRSWR)
 The probability of selection of an element remains unchanged after each draw
 The same units could be selected more than once
 Numbers on possible sample of size n from the Population of size N is given by
Nn
Example: 2 elements from 4 (ABCD)
(How many ways we can draw a sample of size 2 elements from a population of
size 4?) i.e. n=2 and N=4
• With SRSWR: Nn = 42 = 16
• AA, AB, AC, AD, BA, BB, BC, BD, CA, CB, CC, CD, DA, DB, DC, DD = 16
samples
SRS without replacement (SRSWOR)
It is a SRS process in which, once an element is selected as a sample unit, it will not be
replaced in the population pool.
 The probability of selection of an element remains changed after each draw
 The same units could not be selected more than once
 Numbers on possible sample of size n from the Population of size N is given by

N
 
NCn =  
n

N! 4!
 6
Mathematically, ( N  n )! n! 2!2 !

AA, AB, AC, AD,


BA, BB, BC, BD,
CA, CB, CC, CD,

72 | P a g e
DA, DB, DC, DD

AB, AC, AD,


BA, BC, BD,
CA, CB, CD,
DA, DB, DC, = 6 sample

Use of Random Number Tables for selection of Simple Random Sample


Random number table
A random number table is a series of digits (0 to 9) arranged randomly in rows and
columns, as demonstrated in the small sample shown below. The table usually contains
5-digit numbers, arranged in rows and columns, for ease of reading. Typically, a full table
may extend over as many as four or more pages. You will find random number tables in
most statistical textbooks. Random number tables have been in existence since 1927 and
are generated by a variety of methods.
Random Numbers Table
74656 01687 49537 94980 63873 71586 54363 93265
56310 55371 78981 12954 00606 73370 71597 07696
96341 74861 47669 08684 53648 60225 23246 13571
60911 29205 64466 63329 81148 92321 69714 79051
72353 75045 06980 99341 004813 08486 49384 71868
97496 89568 20279 85178 95569 14765 43435 78357
23337 04321 34254 86360 43985 50785 91724 37865
59195 15841 88251 42855 72307 03140 98573 51990
75254 65582 61527 13713 30245 53279 62057 24161
63788 75023 00769 21277 23315 03930 45525 85908
22572 50327 57994 78755 4212 72936 12288 24008
42476 22412 04823 92008 75372 88179 55569 94389
08839 05299 45136 23124 88665 11751 92365 03277
39853 56122 47465 21959 21029 08233 92665 36310

How to use a random number table:


1. Assume you have the test scores for a population of 200 students. Each student has
been assigned a number from 1 to 200. We want to randomly sample only 5 of the
students for this demo.
2. Since the population size is a three-digit number, we will use the first three digits of the
numbers listed in the table.
73 | P a g e
3. Without looking, point to a starting spot in the table. Assume we land on 78981 (3rd
column, 2nd entry).
4. This location gives the first three digits to be 789. This choice is too large (> 200), so we
choose the next number in that column. Keep in mind that we are looking for numbers
whose first three digits are from 001 to 200 (representing students).
5. The second choice gives the first three digits to be 476, also too large. Continue down
the column until you find 5 of the numbers whose first three digits are less than or equal
to 200.
6. From this table, we arrive at 069 (06980), 007 (00769), 048 (04823), 129 (12954), and 086
(08684).
7. RESULT: Students 69, 7, 48, 129, and 86 will be used for our random sample. Our
sample set of students has been randomly selected where each student had an equal
chance of being selected and the selection of one student did not influence the selection of
other students.

74 | P a g e

Common questions

Powered by AI

If the calculated t-value is less than the critical t-value at a given significance level, it indicates the observed difference is not statistically significant. This leads to the acceptance of the null hypothesis, suggesting that both samples come from the same population .

The Z-test is applicable when data follows a normal distribution, and either the sample size is large (n > 30) or the population standard deviation is known. If the calculated Z-value is less than 1.96, the mean difference is not significant at the 5% level, accepting the null hypothesis. If Z ≥ 1.96, the difference is significant, rejecting the null hypothesis .

The arithmetic mean is appreciated for several merits: it is rigidly defined, based on all observations, easily comprehensible, simple to calculate, facilitates easy mathematical treatment, and is minimally affected by sampling fluctuations . However, its demerits include being influenced by extreme values, becoming meaningless with large data variation, and not directly measuring growth or speed rates .

A paired t-test evaluates whether the mean difference between two sets of paired observations is zero. The null hypothesis (Ho: μd = 0) is tested against an alternative hypothesis at a predetermined significance level. If the calculated t-value is significant, it suggests that the mean differences observed are not by random chance, indicating dependency between the samples .

Simple random sampling ensures each member of the population has an equal chance of being selected, eliminating selection bias. It is straightforward and its random nature facilitates easy probability assessments for statistical analysis. However, it may require more effort if populations are vast .

The geometric mean is preferred when there is a need for an average that tends to lower values, avoiding the upward bias of the arithmetic mean which gives equal weight to all data points and has a tendency towards higher values. It is especially useful for ratio and proportion data .

In scientific experimentation, a hypothetical population represents outcomes that do not exist in reality, allowing researchers to explore potential scenarios. For example, possible dice throws or hypothetical crop yields in genetic engineering can be considered, giving a framework to analyze experimental results or forecasts .

The weighted mean is preferred when different observations are given different weights, which the arithmetic mean cannot adequately adjust for. It is useful when there are widely varying classes, differences in item importance, or when averaging ratios, percentages, or rates such as rupees per kilogram. It is particularly employed in contexts like calculating birth rates, death rates, and index numbers .

Rank correlation, developed by Charles Spearman, does not rely on the assumption of normal distribution, which is necessary for Karl Pearson’s method. It is used when quantitative measures cannot be fixed or when distribution shape is unknown, such as ranking students by height and weight without exact measurements .

A correlation coefficient of zero implies no linear relationship between the two variables under study; however, this does not necessarily imply that the variables are completely independent, as there may be nonlinear relationships not captured by the correlation coefficient .

You might also like