0% found this document useful (0 votes)
26 views21 pages

Introduction to Basic Statistics Concepts

The document outlines a Basic Statistics course created by Dr. Vinay Kumar, covering fundamental concepts such as the definition, scope, and limitations of statistics, as well as data types and frequency distributions. It also introduces measures of central tendency, including arithmetic mean, and discusses their merits and demerits. The course aims to provide a comprehensive understanding of statistical methods applicable across various fields.

Uploaded by

ucetprincipal607
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
26 views21 pages

Introduction to Basic Statistics Concepts

The document outlines a Basic Statistics course created by Dr. Vinay Kumar, covering fundamental concepts such as the definition, scope, and limitations of statistics, as well as data types and frequency distributions. It also introduces measures of central tendency, including arithmetic mean, and discusses their merits and demerits. The course aims to provide a comprehensive understanding of statistical methods applicable across various fields.

Uploaded by

ucetprincipal607
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

Basic Statistics

Course Information
Course Name: Basic Statistics

Content Creator: Dr. Vinay Kumar


Chaudhary Charan Singh Haryana Agricultural University, Hisar

Course Reviewer: Dr Dhaneshkumar V Patel


Unagadh Agricultural University, Junagadh

Lesson 1: Introduction of Statistics,


Scope and Types of Data
Objectives of the Lesson
1. Origin and growth of statistics
2. Importance and characteristics of statistics
3. Limitations and scope of statistics
4. Frequency distribution and its types

1.1 Introduction
Modern age is the age of science which requires that every aspect, whether it
pertains to natural phenomena, politics, economics or any other field, should be
expressed in an unambiguous and precise form. A phenomenon expressed in
ambiguous and vague terms might be difficult to understand in proper
perspective. Therefore, in order to provide an accurate and precise explanation
of a phenomenon or a situation, figures are often used.

The word 'Statistics' is probably derived from the Latin word 'status' or the
Italian word 'statista' or the German word 'statistik', each of which means a
'political state'. The word 'Statistics' is used in singular as well as in plural
sense. As a plural, statistics may be defined as the numerical data relating to an
aggregate of individuals and as a singular it is defined as the science of
collection, organization, presentation, analysis and interpretation of numerical
data.

1.2 Definition of Statistics


Statistics has been defined differently by various authors from time to time.
Important definitions include:

 Lovitt: The branch of science which deals with the collection,


classification and tabulation of numerical facts as the basis for
explanations, description and comparison of phenomenon.

 Corxton and Cowden: The science which deals with the collection,
analysis and interpretation of numerical data.

 Kings: The science of statistics is the method of judging collective,


natural or social phenomenon from the results obtained from the analysis
or enumeration or collection of estimates.

 Boddington: Statistics is a science of estimates and probabilities.

 Wallis and Roberts: Statistics is a branch of science, which provides


tools (techniques) for decision making in the face of uncertainty
(Probability). This is the modern definition of statistics which covers the
entire body of statistics.

 Sir R.A. Fisher: "The science of Statistics is essentially a branch of


applied mathematics and may be regarded as mathematics applied to
observational data". Fisher's definition is most exact in the sense that it
covers all aspects and fields of Statistics viz. Collection, Organization,
Presentation, Analysis and Interpretation of data. Sir R.A. Fisher (1890-
1962) is known as the "Father of Modern Statistics".

1.3 Scope of Statistics


During the last few decades, statistics has penetrated into almost all sciences
like agriculture, biology, business, social, engineering, medical, etc. Statistical
methods are commonly used for analyzing and interpreting experimental data.
Many new branches of statistics have emerged such as Industrial Statistics,
Biometrics, Biostatistics, Agricultural Statistics and Statistical Bioinformatics.

Statistics is the science that transforms data into information and the role of
statisticians is to serve science and society through the development,
understanding, and dissemination of state-of-the-art techniques for collecting,
presenting, analyzing, and drawing inferences from data.

Scope of Statistics includes:


 Physical and Natural Sciences: Statistics has great significance in
propounding and verifying scientific laws. It is used in agricultural and
biological research for efficient planning of experiments and for
interpreting experimental data.

 Economic Planning: Statistics is of vital importance in economic


planning. Priorities of planning are determined on the basis of statistics
related to the resource base of the country and its short-term and long-
term needs.
 Economic Phenomena: Statistical techniques are used to study wages,
price analysis, analysis of time series, and demand analysis.

 Business: Successful business executives make use of statistical


techniques for studying the needs and future prospects of their products.
The formulation of a production plan in advance requires relevant details
and their proper analysis.

 Industry: Statistical tools are very helpful in quality control and


assessment. Inspection plans and control charts are of immense
importance and widely used for quality control purposes.

1.4 Limitations of Statistics


 Statistical methods are best applicable to quantitative data.
 Statistical decisions are subject to a certain degree of error.
 Statistical laws do not deal with individual observations but with a group
of observations.
 Statistical conclusions are true on an average.
 Statistics is liable to be misused. The misuse of statistics may arise
because of the use of statistical tools by inexperienced and untrained
persons.
 Statistical results may lead to fallacious conclusions if quoted out of
context or manipulated.

1.5 Concepts, Definitions, Frequency


Distributions & Frequency Curves
1.5.1 Raw Data
The data collected by an investigator which have not been organized numerically
and used by anybody else.

1.5.2 Array
An arrangement of raw numerical data in ascending or descending order of
magnitude.

1.5.3 Primary Data


The data collected directly from the original source is called the primary data
i.e. the data collected for the first time. The primary data may be collected by:

1. Direct interview method


2. Through mail
3. Through designed experiments
[Link] Direct Interview Method
In this method the investigator contacts the units/individuals and has personal
interview. The information is recorded on the questionnaire or schedule. This
information will be more reliable and correct but more expenditure may be
involved.

[Link] Through Mail


The data may be collected through correspondence. The questionnaire or
schedules are sent by mail with the instructions for filling the same and return.
It is less costly to get the data by mail. The main drawback of this method is the
poor response.

[Link] Through Designed Experiments


Data are generated as outcome of the research conducted by the investigator
himself.

1.5.4 Secondary Data


Sometimes the data which we need had already been collected by some agencies
for their study or the data are available in the published records. The data,
which have already been collected by some agency and have been processed or
used at least once are called secondary data. Secondary data may be collected
from organizations or private agencies, government records, journals etc.

1.5.5 Variable
A quantitative and qualitative characteristic that varies from observation to
observation in the same group is called a variable. In case of quantitative
variables, observations are made using interval scales whereas in case of
qualitative variables nominal scales are used. Variables are of two types:
discrete variable and continuous variable.

1.5.6 Discrete Variable


A variable that takes only specific values in a given range, usually the integral
values e.g. number of students in a college, number of petals in a flower, number
of tillers in a plant etc.

1.5.7 Continuous Variable


A variable which can theoretically assume any value between two given values is
called a continuous variable. A continuous variable can take any value within a
certain range, for example yield of a crop, height of plants and birth rates etc.

1.6 Classification of Data on the Basis of


Scales
Four levels or scales of data measurement are:
 Nominal Scale: Lowest level where only names are meaningful
 Ordinal Scale: Ordinal adds an order to the names
 Interval Scale: Interval adds meaningful differences
 Ratio Scale: Ratio adds a zero so that ratios are meaningful

Frequency
The number of times an individual item is repeated in a series is called its
frequency. In case of grouped data, the number of observations lying in any
class is known as the frequency of that class.

Frequency Distribution
It is a tabular arrangement of data values along with their frequencies.

Cumulative Frequency (Less Than Type)


The cumulative frequency corresponding to any value or class is the number of
observations less than or equal to that value or upper limit of that class. It may
also be defined as the total of all frequencies up to the value or the class.

Relative Frequency
The relative frequency of a class is the frequency of the class divided by the total
frequency of all the classes and is generally expressed as a percentage.

Frequency of the class


Relative Frequency=
Total frequency of all classes

Rules for Constructing a Frequency


Distribution
The following points should be borne in mind while tabulating or classifying an
observed frequency distribution:

1. The classes should be well defined and non-overlapping.


2. As far as possible the class interval should be of equal width.
3. The classes should be exhaustive i.e. the range of the classes should cover
the entire range of the data.
4. As a general rule, the number of classes should be between 10 and 15 and
never more than 20 and not less than 5. However the exact number
depends upon the data in hand.
5. Open-ended classes should be avoided.

Sturge's Formula
A numerical formula as suggested by H.A. Sturge may be used for determining
approximately the class size and the number of classes. According to this
formula the number of classes (k) is given:

k =1+3.322 log 10 ⁡N

where N is the number of observations. Then class size is determined as:

Largest value − Smallest value Range


Class width (h)= =
Number of Classes k

Ungrouped or Discrete Frequency Distribution


When the number of observations in the data is small, then the listing of the
frequency of occurrence against the value of variable is called the discrete
frequency distribution.

Example: Raw data showing the number of children of 20 families: 2, 0, 3, 1, 1,


3, 4, 2, 0, 3, 4, 2, 2, 1, 0, 4, 1, 2, 2, 3

Number 0 1 2 3 4
of
Children
Frequenc 3 4 6 4 3
y

Grouped (Continuous) Frequency Distribution


When the data is very large it becomes necessary to condense the data into a
suitable number of class interval of the variable along with the corresponding
frequencies.

Exclusive Method
In this method, the upper limit of any class interval is kept the same as the lower
limit of the just higher class or there is no gap between upper limit of class and
lower limit of just class. It is continuous distribution.

Class Frequency
0-10 2
10-20 4
20-30 5
30-40 3
40-50 1
Inclusive Method
There will be a gap between the upper limit of any class and the lower limit of
just higher class. It is discontinuous distribution.

Class Frequency
0-9 2
10-19 4
20-29 5
30-39 3
40-49 1

One can convert discontinuous distribution to continuous distribution by


subtracting half of the gap (0.5 in this case) from lower limit and by adding the
same quantity to the upper limit.

Lesson 2: Measures of Central


Tendency
Objectives of the Lesson
1. Characteristics for ideal averages
2. Various measures of Central tendency
3. Merit and Demerits of Various measures of central tendency
4. Deciles and Percentiles

3.1 Types of Averages


There are five averages. Among them mean, median and mode are called simple
averages and the other two averages geometric mean and harmonic mean are
called special averages.

3.2 Characteristics for a Good or an Ideal


Average
The following properties should be possessed for an ideal average:

1. It should be rigidly defined.


2. It should be easy to understand and compute.
3. It should be based on all items in the data.
4. Its definition shall be in the form of a mathematical formula.
5. It should be capable of further algebraic treatment.
6. It should have sampling stability.
7. It should be capable of being used in further statistical computations or
processing.

3.3 Arithmetic Mean


The arithmetic mean (or, simply average or mean) of a set of numbers is
obtained by dividing the sum of numbers of the set by the number of numbers. If
the variable x assumes n values x 1 , x 2 … x n then the mean is given by:

x 1 + x 2+ x3 +⋯+ x n ∑ x
X́ = =
n n

Example 1: Calculate the Mean


Calculate the mean for 2, 4, 6, 8, and 10.

Solution:

2+ 4+6 +8+10 30
Mean = = =6
5 5

Direct Method
If the observations x 1 , x 2 … x n have frequencies f 1 , f 2 , f 3 , … , f n respectively, then
the mean is given by:

f 1 x 1 + f 2 x 2 +⋯+f n x n ∑ f i x i
Mean ( X́ )= =
f 1 +f 2+⋯+ f n ∑fi
Example 2: Frequency Distribution
Given the following frequency distribution, calculate the arithmetic mean.

Marks 50 55 60 65 70 75
(x)
No of 2 5 4 4 5 5
Student
s (f)

Solution:

Marks 50 55 60 65 70 75 Total
(x)
No of 2 5 4 4 5 5 25
Studen
ts (f)
fx 100 275 240 260 350 375 1600

Mean ( X́ )=
∑ f i x i = 1600 =64
∑ f i 25
Short Cut Method
In some problems, where the number of variables is large or the values of x i or f i
are larger, then the calculations become tedious. To overcome this difficulty, we
use short cut or deviation method in which an approximate mean, called
assumed mean is taken. This assumed mean is taken preferably near the middle,
say A, and the deviation d i=x i − A is calculated for each variable.

Then the mean is given by the formula:

Mean ( X́ )=A +
∑ f i di
∑ fi
Mean for Grouped Frequency Distribution
Find the class mark or mid-value x i of each class, as:

lower limit +upper limit


x i=class marks=
2
Then:

X́ =
∑ f i xi or X́ =A +
∑ f i d i , where d =x − A
∑ fi ∑ fi i i

Example 4: Distribution by Income Groups


Following is the distribution of persons according to different income groups.
Calculate arithmetic mean.

Incom 0-10 10-20 20-30 30-40 40-50 50-60 60-70


e Class
Interv
al
Numb 6 8 10 12 7 4 3
er of
Person
s
Solution:

X́ =A +
∑ f i d i ×h=35+ −20 ×10=35 − 4=31
∑ fi 50

3.3.1 Merits and Demerits of Arithmetic Mean


Merits
1. It is rigidly defined.
2. It is easy to understand and easy to calculate.
3. If the number of items is sufficiently large, it is more accurate and more
reliable.
4. It is a calculated value and is not based on its position in the series.
5. It is possible to calculate even if some of the details of the data are
lacking.
6. Of all averages, it is affected least by fluctuations of sampling.
7. It provides a good basis for comparison.

Demerits
1. It cannot be obtained by inspection nor located through a frequency
graph.
2. It cannot be used in the study of qualitative phenomena not capable of
numerical measurement i.e. Intelligence, beauty, honesty etc.
3. It can ignore any single item only at the risk of losing its accuracy.
4. It is affected very much by extreme values.
5. It cannot be calculated for open-end classes.
6. It may lead to fallacious conclusions, if the details of the data from which
it is computed are not given.

3.4 Harmonic Mean (H.M.)


Harmonic mean of a set of observations is defined as the reciprocal of the
arithmetic average of the reciprocal of the given values. If x 1 , x 2 ,… , x n are n
observations:

n
H M= n

∑ (1/ x i)
i=1

For a frequency distribution:


n
H M= n

∑ f (1/ x i)
i=1

Example 5: Calculate H.M.


From the given data calculate H.M. 5, 10, 17, 24, and 30.

Solution:

x 1/x
5 0.2000
10 0.1000
17 0.0588
24 0.0417
30 0.0333
Total 0.4338

n 5
H M= = =11.52
∑ (1/ x i) 0.4338
3.5 Geometric Mean (G.M.)
The geometric mean of a series containing n observations is the nth root of the
product of the values. If x 1 , x 2 ,… , x n are observations then:

G.M.= √ x 1 ⋅ x 2 ⋯ x n=¿
n

1
log ⁡G.M.= ( log ⁡x 1 + log ⁡x 2 +⋯+log ⁡x n )=
∑ log ⁡xi
n n

G.M.=Antilog
∑ log ⁡x i
n

Example 7: Calculate Geometric Mean


Calculate the geometric mean (G.M.) of the following series of monthly income
of a batch of families 180, 250, 490, 1400, 1050.

Solution:

x Log x
180 2.2553
250 2.3979
490 2.6902
1400 3.1461
1050 3.0212
Total 13.5107

G.M.=Antilog
∑ log ⁡x i =Antilog 13.5107 =Antilog 2.70=503.6
n 5

3.6 Positional Averages (Median and Mode)


These averages are based on the position of the given observation in a series,
arranged in an ascending or descending order. The magnitude or the size of the
values does not matter as was in the case of arithmetic mean.

3.6.1 Median
The median is the middle value of a distribution i.e., median of a distribution is
the value of the variable which divides it into two equal parts. It is the value of
the variable such that the number of observations above it is equal to the
number of observations below it.

Ungrouped or Raw Data


Arrange the given values in the increasing or decreasing order. If the number of
values are odd, median is the middle value. If the number of values are even,
median is the mean of middle two values.

By formula:

( )
th
N +1
Median, M d= item
2

When Odd Numbers of Values are Given


Example 10: Find median for the following data: 25, 18, 27, 10, 8, 30, 42, 20,
53

Solution:

Arranging the data in the increasing order: 8, 10, 18, 20, 25, 27, 30, 42, 53

Here, numbers of observations are odd (N = 9)

( ) ( ) item =¿
th th
N +1 9+1
Median, M d= item=
2 2
When Even Numbers of Values are Given
Example 11: Find median for the following data: 5, 8, 12, 30, 18, 10, 2, 22

Solution:

Arranging the data in the increasing order: 2, 5, 8, 10, 12, 18, 22, 30

10+12
Here median is the mean of the middle two items (i.e) mean of (10, 12) =
2
= 11

Grouped Data
In a grouped distribution, values are associated with frequencies. Grouping can
be in the form of a discrete frequency distribution or a continuous frequency
distribution. Whatever may be the type of distribution, cumulative frequencies
have to be calculated to know the total number of items.

For Continuous Series:

The steps given below are followed for the calculation of median in continuous
series:

Step 1: Find cumulative frequencies.

N
Step 2: Find
2
N
Step 3: See in the cumulative frequency the value first greater than . Then the
2
corresponding class interval is called the Median Class. Then apply the formula
for Median:

N
−c f
2
M d=l+ ×h
F
Where:

 l = lower limit of the median class


 ∑ f i=n = number of observations
 f = frequency of the median class
 h = size of the median class
 c f = cumulative frequency of the class preceding the median class
 N = total frequency

3.7 Quartiles
The quartiles divide the distribution in four parts. There are three quartiles. The
second quartile divides the distribution into two halves and therefore is the same
as the median. The first (lower) quartile (Q 1) marks off the first one-fourth, the
third (upper) quartile (Q 3) marks off the three-fourth.

Q3 −Q1
Q.D.=
2

( ) ( )
th th
N +1 N +1
Where Q 1= item and Q 3=3 item
4 4

3.10 Mode
The mode or modal value of a distribution is that value of the variable for which
the frequency is the maximum. It refers to that value in a distribution which
occurs most frequently. It shows the center of concentration of the frequency
around a given value. Therefore, where the purpose is to know the point of the
highest concentration it is preferred. It is, thus, a positional measure.

Ungrouped or Raw Data


For ungrouped data or a series of individual observations, mode is often found
by mere inspection.

Example 19: 2, 7, 10, 15, 10, 17, 8, 10, 2

Mode =M 0=10

Grouped Data
For Continuous Distribution:

See the highest frequency then the corresponding value of class interval is
called the modal class. Then apply the following formula:

f −f 1
Mode, M o =l+ ×h
2 f − f 1 −f 2

Where:

 l = lower limit of the modal class


 f = frequency of the modal class
 f 1 = frequency of the class preceding the modal class
 f 2 = frequency of the class following the modal class
 h = size of the modal class

3.11 Empirical Relationship between Averages


In a symmetrical distribution the three simple averages mean = median = mode.
For a moderately asymmetrical distribution, the relationship between them are
brought by Prof. Karl Pearson as:

Mode=3 Median −2 Mean


Example 21: If the mean and median of a moderately asymmetrical series are
26.8 and 27.9 respectively, what would be its most probable mode?

Solution:

Using the empirical formula:

Mode=3 median − 2 mean=3× 27.9 −2 ×26.8=30.1

Lesson 3: Measures of Dispersion


Objectives of the Lesson
1. Characteristics of good measure of Dispersion
2. Various absolute and relative measures of Dispersion
3. Mean deviation, Standard deviation and Coefficient of variation

4.1 Introduction
The measures of central tendency serve to locate the center of the distribution,
but they do not reveal how the items are spread out on either side of the center.
This characteristic of a frequency distribution is commonly referred to as
dispersion. In a series all the items are not equal. There is difference or variation
among the values. The degree of variation is evaluated by various measures of
dispersion. Small dispersion indicates high uniformity of the items, while large
dispersion indicates less uniformity.

Example: Consider the following marks of two students:

Student I 68 75 65 67 70
Student II 85 90 80 25 65

Both have got a total of 345 and an average of 69 each. The fact is that the
second student has failed in one paper. When the averages alone are considered,
the two students are equal. But first student has less variation than second
student. Less variation is a desirable characteristic.

4.2 Characteristics of a Good Measure of


Dispersion
An ideal measure of dispersion is expected to possess the following properties:

1. It should be rigidly defined


2. It should be based on all the items.
3. It should not be unduly affected by extreme items.
4. It should lend itself for algebraic manipulation.
5. It should be simple to understand and easy to calculate.

4.3 Absolute and Relative Measures


There are two kinds of measures of dispersion, namely:

1. Absolute measure of dispersion


2. Relative measure of dispersion
Absolute measure of dispersion indicates the amount of variation in a set of
values in terms of units of observations. On the other hand relative measures of
dispersion are free from the units of measurements of the observations. They are
pure numbers. They are used to compare the variation in two or more sets,
which are having different units of measurements of observations.

Absolute Measure Relative Measure


Range Co-efficient of Range
Quartile deviation Co-efficient of Quartile deviation
Mean deviation Co-efficient of Mean deviation
Standard deviation Co-efficient of variation

4.3.1 Range and Coefficient of Range


Range: This is the simplest possible measure of dispersion and is defined as the
difference between the largest and smallest values of the variable.

Range=L− S
Where: L = Largest value; S = Smallest value

Co-efficient of Range:

L−S
Co-efficient of Range=
L+ S

Example 1: Range and Coefficient


Find the value of range and its co-efficient for the following data: 7, 9, 6, 8, 11,
10, 4
Solution:

L=11, S=4
Range=L− S=11− 4=7
L − S 11− 4 7
Co-efficient of Range= = = =0.4667
L+ S 11+4 15

4.3.2 Quartile Deviation and Coefficient of


Quartile Deviation
Quartile Deviation (Q.D.): Quartile Deviation is half of the difference between
the first and third quartiles. Hence, it is called Semi Inter Quartile Range.

Q3 −Q1
Q.D.=
2

Among the quartiles Q 1, Q 2 and Q 3, the range Q 3 −Q 1 is called inter quartile


Q3 − Q1
range and , semi inter quartile range.
2
Co-efficient of Quartile Deviation:

Q 3 −Q 1
Co-efficient of Quartile Deviation=
Q3 +Q1

4.3.3 Mean Deviation and Coefficient of Mean


Deviation
Mean Deviation: The range and quartile deviation are not based on all
observations. They are positional measures of dispersion. They do not show any
scatter of the observations from an average. The mean deviation is a measure of
dispersion based on all items in a distribution.

Definition: Mean deviation is the arithmetic mean of the deviations of a series


computed from any measure of central tendency; i.e., the mean, median or
mode, all the deviations are taken as positive i.e., signs are ignored.

Coefficient of Mean Deviation:

Mean Deviation (M.D.)


Coefficient of Mean Deviation =
Mean or Median or Mode
If the result is desired in percentage:

Mean Deviation (M.D.)


Coefficient of Mean Deviation = × 100
Mean or Median or Mode
Computation of Mean Deviation - Individual Series:

1. Calculate the average mean, median or mode of the series.


2. Take the deviations of items from average ignoring signs and denote
these deviations by |D|.
3. Compute the total of these deviations, i.e., Σ∨D∨¿
4. Divide this total obtained by the number of items.

M.D.=∑ ∨D∨ ¿ ¿
n

Example 6: Mean Deviation


Calculate mean deviation from mean and median for the following data: 100,
150, 200, 250, 360, 490, 500, 600, and 671. Also calculate co-efficient of M.D.

Solution:

Mean =
∑ X = 100+150+200+ 250+360+ 490+500+600+671 = 3321 =369
n 9 9
Arranging data in ascending order: 100, 150, 200, 250, 360, 490, 500, 600, 671

( ) ( )
th th
N +1 9+1 th
Median, M d=Value of item= item=5 item=360
2 2

1570
M.D. from mean =∑ ∨D∨ ¿ = =174.44 ¿
n 9
Mean Deviation (M.D.) 174.44
Coefficient of M.D.= = =0.47
Mean 369
1561
M.D. from median=∑ ∨D∨ ¿ = =173.44 ¿
n 9
Mean Deviation (M.D.) 173.44
Coefficient of M.D.= = =0.48
Median 360
Mean Deviation - Discrete Series:

M.D.=∑ f ∨D∨ ¿ ¿
n
Mean Deviation - Continuous Series:

The method of calculating mean deviation in a continuous series is the same as


the discrete series. In continuous series we have to find out the mid points of the
various classes and take deviation of these points from the average selected.

M.D.=∑ f ∨D∨ ¿ ¿
n
Additional Content: Design of
Experiments
10.5 Design of Experiments
Choice of treatments, method of assigning treatments to experimental units and
arrangement of experimental units in different patterns are known as designing
an experiment. We study the effect of changes in one variable on another
variable. For example how the application of various doses of fertilizer affects
the grain yield.

Variable whose change we wish to study is known as response variable.


Variable whose effect on the response variable we wish to study is known
as factor.

Treatment
Objects of comparison in an experiment are defined as treatments. Examples are
Varieties tried in a trail and different chemicals.

Experimental Unit
The object to which treatments are applied or basic objects on which the
experiment is conducted is known as experimental unit. Example: piece of land,
an animal, etc

Experimental Error
Response from all experimental units receiving the same treatment may not be
same even under similar conditions. These variations in responses may be due to
various reasons. Other factors like heterogeneity of soil, climatic factors and
genetic differences, etc also may cause variations (known as extraneous factors).
The variations in response caused by extraneous factors are known as
experimental error.

Our aim of designing an experiment will be to minimize the experimental error.

10.6 Basic Principles


To reduce the experimental error we adopt certain principles known as basic
principles of experimental design. The basic principles are:

1. Replication
2. Randomization
3. Local control

10.6.1 Replication
Repeated application of the treatments is known as replication. When the
treatment is applied only once we have no means of knowing about the variation
in the results of a treatment. Only when we repeat several times we can estimate
the experimental error.

With the help of experimental error we can determine whether the obtained
differences between treatment means are real or not. When the number of
replications is increased, experimental error reduces.

10.6.2 Randomization
When all the treatments have equal chance of being allocated to different
experimental units it is known as randomization.

If our conclusions are to be valid, treatment means and differences among


treatment means should be estimated without any bias. For this purpose we use
the technique of randomization.

10.6.3 Local Control


Experimental error is based on the variations from experimental unit to
experimental unit. This suggests that if we group the homogenous experimental
units into blocks, the experimental error will be reduced considerably. Grouping
of homogenous experimental units into blocks is known as local control of error.

In order to have valid estimate of experimental error the principles of replication


and randomization are used. In order to reduce the experimental error, the
principles of replication and local control are used. In general to have precise,
valid and accurate result we adopt the basic principles.

10.7 Completely Randomized Design (CRD)


CRD is the basic single factor design. In this design the treatments are assigned
completely at random so that each experimental unit has the same chance of
receiving any one treatment. But CRD is appropriate only when the experimental
material is homogeneous. As there is generally large variation among
experimental plots due to many factors CRD is not preferred in field
experiments.

In laboratory experiments and greenhouse studies it is easy to achieve


homogeneity of experimental materials and therefore CRD is most useful in such
experiments.

10.7.1 Layout of a CRD


Completely randomized Design is the one in which all the experimental units are
taken in a single group which are homogeneous as far as possible. The
randomization procedure for allotting the treatments to various units will be as
follows:

Step 1: Determine the total number of experimental units.


Step 2: Assign a plot number to each of the experimental units starting from left
to right for all rows.

Step 3: Assign the treatments to the experimental units by using random


numbers.

Statistical Model for CRD:

Y i j =μ+t i + ei j

Where:

 μ = overall mean effect


 t i = true effect of the ith treatment
 e i j = error term of the jth unit receiving ith treatment
Null Hypothesis:

H 0 : μ1=μ2=⋯=μk or There is no significant difference between the treatments

Alternative Hypothesis:

H 1 : μ 1 ≠ μ2 ≠ ⋯ ≠ μ k

There is significant difference between the treatments.

This comprehensive document covers the fundamentals of statistics, measures of


central tendency, measures of dispersion, and design of experiments.

You might also like