Module 1:
Primary data and secondary data:
Primary data:
They are the data originally collected for an investigation for the first time. They are collected
primary sources. They are in the form of raw material.
1. Observation method
2. Interview method
3. Questionnaire method
4. Schedule method
Secondary data:
Secondary data are those data that are collected from primary data which have been already
collected. Sources of secondary data are:
1. Official reports
2. Research papers
3. Newspapers
4. Year book
5. Documents
Techniques of data collecting (census and sampling):
Population: a population is the total objects under consideration.
1. Finite population
2. Infinite population
Census method:
It is a statistical investigation in which the data are collected for each and every unit of the
population. It is known as complete survey or complete enumeration.
For example, consider the college students as the population where each and every student’s
information are investigated.
Sampling method:
Sampling is the process of using a subset of a population to represent the whole population.
For example, college students are considered as a population where each departments of the
college are subsets of it. Sampling selects one of the departments as its main population.
Difference between census and sampling:
CENSUS METHOD SAMPLING METHOD
This method collects information This method collects data only from a
from all the units in the population. representative part of the population.
It gives more accurate results. It gives less accurate results.
This method cannot be applied in This method can be applied in all
cases like infinite population. cases.
It is more time consuming and is It less time consuming and less
expensive. expensive.
Detailed study of all items is not Detailed study of all items in the
possible. sample is possible.
Different methods for selecting samples:
Sampling are of two types:
1. Random sampling (probability sampling)
2. Non-random sampling (non-probability sampling)
Random sampling:
This type of sampling is also known as probability sampling because such sampling design is based
on probability for selection of each item.
Each item in the population has an equal chance of being selected or excluded. There are two types
of random sampling:
1. Simple sampling
2. Complex sampling
Simple sampling:
In this case samples are selected from the population in such a manner that every item of the
population has an equal chance of being selected.
Complex sampling:
In this case, the sample members do not gave equal probability of being selected.
Types of complex sampling:
1. Stratified sampling:
This method is applied when the population is heterogeneous. The entire universe is divided into
certain number of strata on the basis of certain factors such as age, sex, income, education etc.
2. Systematic sampling:
Under this method the initial unit of the sample is selected at random and the others are selected on
a predetermined basis of space interval.
3. Cluster sampling:
It is a type of sampling in which the whole universe is divided into a certain subgroups called
clusters. The clusters are selected at a random.
Classification and types of classification:
Classification is the process of arranging data into similar groups according to their characteristics.
There are 4 types of classifications:
1. Geographical classification:
Classification on the basis of place is called geographical classification. In this classification, data are
classified on the basis of states, districts, towns etc.
2. Chronological classification:
Classification on the basis of time or periods is called chronological classification. In this
classification, data are classified on the basis of year, month, day, week etc.
3. Qualitative classification:
Classification according to attributes is known as qualitative classification. It cannot be measured.
For example, education, sex, nationality, literacy etc. which cannot be measured.
4. Quantitative classification:
Classification on the basis of measurable characteristics of the units under study is called
quantitative classification or classification according to variables. For example, age, marks, income,
profit etc. which can be measured.
Statistical series:
A series is a systematic arrangement of the items into some logical order.
According to Canor, “if two quantities can be arranged side by side so that measurable differences in
the one corresponds with measurable differences in the other, the result is said to form a statisticall
series”.
Statistical series can be divided into 3:
1. Time series: shown timewise.
2. Spatial series: presented with reference to places.
3. Condition series: presented with reference to some condition.
Individual series, discrete series and continuous series:
They are divided on the basis of construction.
In individual series, the items are listed singly showing the observations relating to them. For
example, the marks obtained by students can be shown in an individual series by showing the class
number of the students against their marks.
Class no.
of Marks
students
1 60
2 75
3 62
4 59
In discrete series, the exact measurements of various items are shown. For each measurements is to
be treated as a class. Frequency relating to such measurements are also shown against them.
Daily No. of
wages(Rs) workers
15 2
16 5
20 8
24 13
In continuous series, measurements are approximate and therefore expressed in class intervals. The
class intervals are continuously shown from the beginning to end with corresponding frequencies.
Age No. of
people
0-10 3
10-20 5
20-30 8
30-40 10
Scales of measurement:
Scales of measurement shows how variables are defined and categorised.
There are 4 levels of measurement:
1. Nominal
2. Ordinal
3. Interval
4. Ratio
Nominal:
First scale of measurement or nominal level assigns labels or tags to identify and classify the objects,
for example, name, age, mark etc. (Naming of the variables).
Ordinal:
Second scale of measurement that reports the ordering and ranking of the data such as grades,
degree of education and 1st, 2nd and 3rd ranking etc.
Interval:
The third scale of measurement is interval scale which is defined as a quantitative measurement
scale in which the difference between the two variables. The difference between them cannot be
zero. For example, the difference of 60 and 50 degree is 10 degree and also the difference of 80 and
70 degree is 10 degree.
Ratio scale:
The ratio scale is the 4th level of measurement scale, which is also quantitative. This level allows zero
values unlike interval scale. For example, length, duration, profit, sales etc.
Module 2:
Measures of central tendency:
It is a value lying between the minimum and maximum value of the series and generally located at
the centre.
It means finding out the central value or average value of a statistical series.
An average is a single significant figure which sum up the characteristics of a group of figures. It
represents the whole series.
Characteristics of a good average:
It should be clearly defined.
It should be based on all observations of the data.
It should be easy to calculate and simple to follow.
It should not influence sampling fluctuations.
It should be suitable for further algebraic treatments.
Types of averages:
1. Mathematical averages:
Arithmetic mean, geometric mean and harmonic mean are the types of mathematical averages.
These are called mathematical averages because their values are computed through the
mathematical equations.
2. Positional averages:
Median, mode, quartiles, percentiles and deciles are the types of positional averages.
These are called positional averages because their values are determined by locating their position in
the series.
Mathematical averages:
Arithmetic mean:
Arithmetic mean may be defined as the value obtained by dividing the total of the values of a
variable(x) by the total number of observations.
Arithmetic mean can be of two types, simple arithmetic mean and weighted arithmetic mean.
Simple mean is mean of items which are given equal importance.
Simple arithmetic mean can be calculated as, x=
∑x
n
Weighted mean is the mean of items which are given different weights in accordance with their
relative importance.
Weighted mean can be calculated as, x w=
∑ Wx
∑w
Geometric mean:
Geometric mean is one of the mathematically average if there are ‘n’ values in a series then their
GM is defined as nth root of product of those n values.
GM is calculated as, GM = antilog
(∑ )
log x
n
Harmonic mean:
Harmonic mean of a set of n values I defined as the reciprocal of the mean of the reciprocal of these
values.
n
Harmonic mean can be calculated as, HM = 1
∑x
Positional averages:
Median:
Median is the middle value in a set of data. First, organize and order the data in ascending order
then, divide the number of observations by 2 to find the midpoint value.
Median = ( n+12 )
Partition values:
Quartiles: The values which divides the data into 4 equal parts are known as quartiles. There are 3
quartiles, Q1, Q2 and Q3.
The quartiles can be calculated as, Q1 = ( n+14 ) , Q = 3( n+14 ). The quartile Q is the same as median.
3 2
Deciles: The values which divides the data into 10 equal parts are known as deciles. There are 10
quartiles, D1, D2 … D10.
The deciles can be calculated as, D1 = ( n+1
10 )
, D = r(
r
10 )
n+1
. The decile D is the same as median.
5
Percentile: The values which divides the data into 100 equal parts are known as deciles. There are
100 percentiles, P1, P2 … P100.
The percentiles can be calculated as, P1 = ( n+1
100 ) , P = r(
r
100 )
n+1
. The percentile P is the same as
50
median.
Mode:
Mode is the value of that item of a series which occurs most frequently in the series.
When no items appear more number of times than others we say the mode is ill-defined.
Ill-defined mode can be calculated as, mode = 3median-2mean.
Measures of dispersion:
Dispersion refers to variability in the size of the item. It speaks about the spread or scatter of the
values in a series. It indicates lack of uniformity in the size of item.
Types of measures of dispersion:
Measures of dispersion are classified into relative measure and absolute measure.
Absolute measure of dispersion:
Absolute measures are expressed in the same unit in which data are collected. They measure
variability in series.
Types of absolute measure of dispersion:
1. Range
2. Quartile deviation
3. Mean deviation
4. Standard deviation
Relative measure of dispersion:
Relative measure of dispersion are useful comparing two series for their variability. Greater the
value of relative measure in series, greater is the variability in series.
Types of relative measure of dispersion:
1. Coefficient of range
2. Coefficient of QD
3. Coefficient of MD
4. Coefficient of variation
Range:
Range is the simplest possible measure of dispersion. It is the difference between highest and lowest
value in the series.
Range can be calculate as, Range = H-L
H−L
Coefficient of range =
H +L
Quartile deviation:
Q3−Q 1 Q3−Q 1
Quartile deviation, QD = and coefficient of QD =
2 Q 3+Q1
Mean deviation:
It is defines as the arithmetic mean deviations of all the values in series from their average.
The average selected may be mean median or mode.
Equation of mean deviation, MD =
∑ ¿ X− A∨¿ ¿ and mean deviation in discrete series, MD =
n
f
∑ ¿ X −A∨¿ ¿
N
MD
Coefficient of MD = , where A=mean, median or mode.
A
Standard deviation:
It may be defined as the square root of the arithmetic average of the squares of deviation taken
from the arithmetic average of the series.
It is denoted by the Greek letter σ
Equation of standard deviation, SD =
√ n √( )
∑ x2 - ∑ x
n
2
The coefficient of variation is useful for comparing the variability of two or more series.
standard deviation
CV = × 100
mean
Variance:
Variance of a series is defined as the mean of the squares of the deviation of all the values in the
series from their arithmetic mean.
Variance = ( SD )2
Module 3:
Random experiment:
A random experiment is an event whose outcome cannot be predicted with certainty beforehand.
Sample space:
The sample space of a random experiment is the set of all possible outcomes of that experiment.
The sample space of throwing 2 coins is {HH, HT, TT, TH}
Independent event:
In probability, two events are independent if the probability of one event occurring is not affected by
the occurrence of the other event.
Statistically, P(A|B) = P(A) and P(B|A) = P(B).
Proof:
To prove,
P(A ∩ B) = P(A) P(B)
If A is independent of B we have,
P(A|B) = P(A)
P(A) = P(A ∩ B) / P(B)
P(A ∩ B) = P(A)P(B) 1 (multiplication theorem)
To prove,
P(B|A) = P(B)
We take,
P(B|A) = P(A ∩ B) / P(A)
From eq 1,
P(B|A) = P(A)P(B) / P(A)
= P(B)
So B is also independent of A.
Axioms of probability:
Axiom 1: the probability of an event is always greater than or equal to 1.
Axiom 2: the probability of all possible outcome is equal to 1.
Axiom 3: the probability of union of a collection of mutually exclusive events is equal to the sum of
the individual probabilities.
Addition theorem:
If A and B are any two events then the probability of happening pf at least one of the events is
defined as,
P(AUB) = P(A) + P(B) – P(A∩B)
Bernoulli’s distribution:
A random variable which take 2 values, 0 and 1 with probabilities q and p respectively,
P(x=0) = q and P(x=1) = p
This variable is said to have Bernoulli’s distribution.
Mean of Bernoulli’s distribution is p and variance is pq.
Binomial distribution:
When random experiment is repeated a number of times, the event may or may not occur in each of
those experiments.
The occurrence of the events may be named successful and non-occurrence of the events a failure.
Therefore a random experiment has two outcomes, success and failure.
A discrete random variable x is said to follow binomial distribution with the parameters p and n, if its
probability function is,
q+p=1 (1 – p) = q
In the equation, x stands for number of success
P represents probability for success
q=1–p
n stands for number of trials.
In binomial distribution, the trials should be independent and repeated a finite number of times.