Introduction to Basic Statistics Concepts
Introduction to Basic Statistics Concepts
Course Information
Course Name: Basic Statistics
1.1 Introduction
Modern age is the age of science which requires that every aspect, whether it
pertains to natural phenomena, politics, economics or any other field, should be
expressed in an unambiguous and precise form. A phenomenon expressed in
ambiguous and vague terms might be difficult to understand in proper
perspective. Therefore, in order to provide an accurate and precise explanation
of a phenomenon or a situation, figures are often used.
The word 'Statistics' is probably derived from the Latin word 'status' or the
Italian word 'statista' or the German word 'statistik', each of which means a
'political state'. The word 'Statistics' is used in singular as well as in plural
sense. As a plural, statistics may be defined as the numerical data relating to an
aggregate of individuals and as a singular it is defined as the science of
collection, organization, presentation, analysis and interpretation of numerical
data.
Corxton and Cowden: The science which deals with the collection,
analysis and interpretation of numerical data.
Statistics is the science that transforms data into information and the role of
statisticians is to serve science and society through the development,
understanding, and dissemination of state-of-the-art techniques for collecting,
presenting, analyzing, and drawing inferences from data.
1.5.2 Array
An arrangement of raw numerical data in ascending or descending order of
magnitude.
1.5.5 Variable
A quantitative and qualitative characteristic that varies from observation to
observation in the same group is called a variable. In case of quantitative
variables, observations are made using interval scales whereas in case of
qualitative variables nominal scales are used. Variables are of two types:
discrete variable and continuous variable.
Frequency
The number of times an individual item is repeated in a series is called its
frequency. In case of grouped data, the number of observations lying in any
class is known as the frequency of that class.
Frequency Distribution
It is a tabular arrangement of data values along with their frequencies.
Relative Frequency
The relative frequency of a class is the frequency of the class divided by the total
frequency of all the classes and is generally expressed as a percentage.
Sturge's Formula
A numerical formula as suggested by H.A. Sturge may be used for determining
approximately the class size and the number of classes. According to this
formula the number of classes (k) is given:
k =1+3.322 log 10 N
Number 0 1 2 3 4
of
Children
Frequenc 3 4 6 4 3
y
Exclusive Method
In this method, the upper limit of any class interval is kept the same as the lower
limit of the just higher class or there is no gap between upper limit of class and
lower limit of just class. It is continuous distribution.
Class Frequency
0-10 2
10-20 4
20-30 5
30-40 3
40-50 1
Inclusive Method
There will be a gap between the upper limit of any class and the lower limit of
just higher class. It is discontinuous distribution.
Class Frequency
0-9 2
10-19 4
20-29 5
30-39 3
40-49 1
x 1 + x 2+ x3 +⋯+ x n ∑ x
X́ = =
n n
Solution:
2+ 4+6 +8+10 30
Mean = = =6
5 5
Direct Method
If the observations x 1 , x 2 … x n have frequencies f 1 , f 2 , f 3 , … , f n respectively, then
the mean is given by:
f 1 x 1 + f 2 x 2 +⋯+f n x n ∑ f i x i
Mean ( X́ )= =
f 1 +f 2+⋯+ f n ∑fi
Example 2: Frequency Distribution
Given the following frequency distribution, calculate the arithmetic mean.
Marks 50 55 60 65 70 75
(x)
No of 2 5 4 4 5 5
Student
s (f)
Solution:
Marks 50 55 60 65 70 75 Total
(x)
No of 2 5 4 4 5 5 25
Studen
ts (f)
fx 100 275 240 260 350 375 1600
Mean ( X́ )=
∑ f i x i = 1600 =64
∑ f i 25
Short Cut Method
In some problems, where the number of variables is large or the values of x i or f i
are larger, then the calculations become tedious. To overcome this difficulty, we
use short cut or deviation method in which an approximate mean, called
assumed mean is taken. This assumed mean is taken preferably near the middle,
say A, and the deviation d i=x i − A is calculated for each variable.
Mean ( X́ )=A +
∑ f i di
∑ fi
Mean for Grouped Frequency Distribution
Find the class mark or mid-value x i of each class, as:
X́ =
∑ f i xi or X́ =A +
∑ f i d i , where d =x − A
∑ fi ∑ fi i i
X́ =A +
∑ f i d i ×h=35+ −20 ×10=35 − 4=31
∑ fi 50
Demerits
1. It cannot be obtained by inspection nor located through a frequency
graph.
2. It cannot be used in the study of qualitative phenomena not capable of
numerical measurement i.e. Intelligence, beauty, honesty etc.
3. It can ignore any single item only at the risk of losing its accuracy.
4. It is affected very much by extreme values.
5. It cannot be calculated for open-end classes.
6. It may lead to fallacious conclusions, if the details of the data from which
it is computed are not given.
n
H M= n
∑ (1/ x i)
i=1
∑ f (1/ x i)
i=1
Solution:
x 1/x
5 0.2000
10 0.1000
17 0.0588
24 0.0417
30 0.0333
Total 0.4338
n 5
H M= = =11.52
∑ (1/ x i) 0.4338
3.5 Geometric Mean (G.M.)
The geometric mean of a series containing n observations is the nth root of the
product of the values. If x 1 , x 2 ,… , x n are observations then:
G.M.= √ x 1 ⋅ x 2 ⋯ x n=¿
n
1
log G.M.= ( log x 1 + log x 2 +⋯+log x n )=
∑ log xi
n n
G.M.=Antilog
∑ log x i
n
Solution:
x Log x
180 2.2553
250 2.3979
490 2.6902
1400 3.1461
1050 3.0212
Total 13.5107
G.M.=Antilog
∑ log x i =Antilog 13.5107 =Antilog 2.70=503.6
n 5
3.6.1 Median
The median is the middle value of a distribution i.e., median of a distribution is
the value of the variable which divides it into two equal parts. It is the value of
the variable such that the number of observations above it is equal to the
number of observations below it.
By formula:
( )
th
N +1
Median, M d= item
2
Solution:
Arranging the data in the increasing order: 8, 10, 18, 20, 25, 27, 30, 42, 53
( ) ( ) item =¿
th th
N +1 9+1
Median, M d= item=
2 2
When Even Numbers of Values are Given
Example 11: Find median for the following data: 5, 8, 12, 30, 18, 10, 2, 22
Solution:
Arranging the data in the increasing order: 2, 5, 8, 10, 12, 18, 22, 30
10+12
Here median is the mean of the middle two items (i.e) mean of (10, 12) =
2
= 11
Grouped Data
In a grouped distribution, values are associated with frequencies. Grouping can
be in the form of a discrete frequency distribution or a continuous frequency
distribution. Whatever may be the type of distribution, cumulative frequencies
have to be calculated to know the total number of items.
The steps given below are followed for the calculation of median in continuous
series:
N
Step 2: Find
2
N
Step 3: See in the cumulative frequency the value first greater than . Then the
2
corresponding class interval is called the Median Class. Then apply the formula
for Median:
N
−c f
2
M d=l+ ×h
F
Where:
3.7 Quartiles
The quartiles divide the distribution in four parts. There are three quartiles. The
second quartile divides the distribution into two halves and therefore is the same
as the median. The first (lower) quartile (Q 1) marks off the first one-fourth, the
third (upper) quartile (Q 3) marks off the three-fourth.
Q3 −Q1
Q.D.=
2
( ) ( )
th th
N +1 N +1
Where Q 1= item and Q 3=3 item
4 4
3.10 Mode
The mode or modal value of a distribution is that value of the variable for which
the frequency is the maximum. It refers to that value in a distribution which
occurs most frequently. It shows the center of concentration of the frequency
around a given value. Therefore, where the purpose is to know the point of the
highest concentration it is preferred. It is, thus, a positional measure.
Mode =M 0=10
Grouped Data
For Continuous Distribution:
See the highest frequency then the corresponding value of class interval is
called the modal class. Then apply the following formula:
f −f 1
Mode, M o =l+ ×h
2 f − f 1 −f 2
Where:
Solution:
4.1 Introduction
The measures of central tendency serve to locate the center of the distribution,
but they do not reveal how the items are spread out on either side of the center.
This characteristic of a frequency distribution is commonly referred to as
dispersion. In a series all the items are not equal. There is difference or variation
among the values. The degree of variation is evaluated by various measures of
dispersion. Small dispersion indicates high uniformity of the items, while large
dispersion indicates less uniformity.
Student I 68 75 65 67 70
Student II 85 90 80 25 65
Both have got a total of 345 and an average of 69 each. The fact is that the
second student has failed in one paper. When the averages alone are considered,
the two students are equal. But first student has less variation than second
student. Less variation is a desirable characteristic.
Range=L− S
Where: L = Largest value; S = Smallest value
Co-efficient of Range:
L−S
Co-efficient of Range=
L+ S
L=11, S=4
Range=L− S=11− 4=7
L − S 11− 4 7
Co-efficient of Range= = = =0.4667
L+ S 11+4 15
Q3 −Q1
Q.D.=
2
Q 3 −Q 1
Co-efficient of Quartile Deviation=
Q3 +Q1
M.D.=∑ ∨D∨ ¿ ¿
n
Solution:
Mean =
∑ X = 100+150+200+ 250+360+ 490+500+600+671 = 3321 =369
n 9 9
Arranging data in ascending order: 100, 150, 200, 250, 360, 490, 500, 600, 671
( ) ( )
th th
N +1 9+1 th
Median, M d=Value of item= item=5 item=360
2 2
1570
M.D. from mean =∑ ∨D∨ ¿ = =174.44 ¿
n 9
Mean Deviation (M.D.) 174.44
Coefficient of M.D.= = =0.47
Mean 369
1561
M.D. from median=∑ ∨D∨ ¿ = =173.44 ¿
n 9
Mean Deviation (M.D.) 173.44
Coefficient of M.D.= = =0.48
Median 360
Mean Deviation - Discrete Series:
M.D.=∑ f ∨D∨ ¿ ¿
n
Mean Deviation - Continuous Series:
M.D.=∑ f ∨D∨ ¿ ¿
n
Additional Content: Design of
Experiments
10.5 Design of Experiments
Choice of treatments, method of assigning treatments to experimental units and
arrangement of experimental units in different patterns are known as designing
an experiment. We study the effect of changes in one variable on another
variable. For example how the application of various doses of fertilizer affects
the grain yield.
Treatment
Objects of comparison in an experiment are defined as treatments. Examples are
Varieties tried in a trail and different chemicals.
Experimental Unit
The object to which treatments are applied or basic objects on which the
experiment is conducted is known as experimental unit. Example: piece of land,
an animal, etc
Experimental Error
Response from all experimental units receiving the same treatment may not be
same even under similar conditions. These variations in responses may be due to
various reasons. Other factors like heterogeneity of soil, climatic factors and
genetic differences, etc also may cause variations (known as extraneous factors).
The variations in response caused by extraneous factors are known as
experimental error.
1. Replication
2. Randomization
3. Local control
10.6.1 Replication
Repeated application of the treatments is known as replication. When the
treatment is applied only once we have no means of knowing about the variation
in the results of a treatment. Only when we repeat several times we can estimate
the experimental error.
With the help of experimental error we can determine whether the obtained
differences between treatment means are real or not. When the number of
replications is increased, experimental error reduces.
10.6.2 Randomization
When all the treatments have equal chance of being allocated to different
experimental units it is known as randomization.
Y i j =μ+t i + ei j
Where:
Alternative Hypothesis:
H 1 : μ 1 ≠ μ2 ≠ ⋯ ≠ μ k