Vishal Vaid: 9829237778, 9829207778 Khandelwal Institute for Studies
Sampling:
Population or Universe. Population in statistics means the whole of the information which comes under the
purview of statistical investigation. It is the totality of all the observations of a statistical experiment or
enquiry. In other words, an aggregate of objects; animate or inanimate under study is the population. It is also
known ag the universe.
Thus, Population/ universe Aggregate of facts under consideration
Types of population:
1. Finite & Infinite.
2. Homogenous & Heterogeneous
3. Hypothetical & Actual
1. If it is possible to count the number of units practically, then it is finite & if it is not possible to count the number
of units, practically, then it is infinite.
2. If all the units are identical - Homogenous
If all the units are different - Heterogeneous
3. If assumed situation is taken into consideration then it is hypothetical.
Method of enumeration about population.
1. Census 2. Sampling
Census:
When each & every item of the universe is investigated Drawbacks:
1. More time & more money.
2. Unreliable results
3. Not applicable in infinite population.
4. Not required in homogenous population
5. Not applicable in case of perishable goods or items which are exhausted in data collection.
Sampling: When only a part of the universe is selected for the purpose of study, it is known as sampling method.
Sample: The part of universe which is selected for study.
Types of sample:
1. Random sample - When each & every item of the universe has the equal probability of being selected.
2. Non random sample - When each & every item of the universe don't equal probability of being selected.
Types of sampling:
I. Probability sampling - When each & every item of the universal has a pre assigned probability of being selected,
it is probability sampling. It includes simple random sampling, satisfied sampling, multi stage or multiphase
sampling.
i. Simple random sampling: When each & every item of universe has equal probability of being selected it is
known as simple random sampling. It can be done with replacement(SRSWR) or without
replacement(SRSWOR). All the major test of hypotheses are based on simple random sampling. The best method
of selecting random sample is use of random numbers.
It is mainly applied:
(i) when population is not very large,
(ii) sample is not very small &
(iii) population is not much heterogeneous.
If population is infinite, then simple random sampling will render same results with & without replacement.
ii. Stratified random sampling:
It is used when
(a) Population is large.
(b) Population is heterogeneous
(c) Prior information about the population is available.
In stratified sampling the population is divided into sub population or strata or stratum in such a way that the
variability among the stratas is maximum & variability within the strata is minimum.
After dividing the population into stratas, a simple random sample is drawn from each strata.
1|Page
Vishal Vaid: 9829237778, 9829207778 Khandelwal Institute for Studies
Advantages of stratified sampling:
(a) The presentation of all sub populations.
(b) Parameter about each sub population & overall parameter is available.
(c) Reduction in variability & thus increase in accuracy
The sample size depends upon the differences among the strata variances. If the difference between the strata variances is
low, then proportional or Bowley's allocation is applied. According to this allocation sample size is proportional to
universe size (i.e. n1 Ni ) . If there is much difference among strata variances, then Neyman's allocation is applied.
According to this allocation, sample size is jointly proportional to universe size and standard derivation ( ni Ni Si )
Disadvantages
(i) It may be difficult to divide the population into hetrogeneous groups.
(ii) There may be over-lapping of different strata of the population which will provide an unrepresentative
sample.
iii. Multi stage or Multi phase sampling.
In this sampling, sample of elementary unit is selected at first stage, then further a sample of elementary units is
selected among the above samples & so on. It is normally applied when population is very large. It adds flexibility
to the sampling process.
Advantages of random or probability sampling:
(i) It fulfills the object of the investigator & is unbiased.
(ii) The conclusions drawn about the parameter are reliable & accurate.
(iii) It is used for drawing statistical inferences.
(iv) The size of sample depends on demonstrable statistical method and therefore, it has a justification for the
expenditure involved.
(v) It provides a more accurate method of drawing conclusions about the characteristics of the population as
parameters.
(vi) The samples may be combined and evaluated, even though accomplished by different individuals.
(vii) The results obtained can be assessed in terms of probability, and the sample is accepted or rejected on a
consideration of the extent to which it can be considered representative.
II. Non probability sampling:
When sampling is made on the personal judgment or convenience of the investigator, then such sampling is
known as non probability sampling. It is also known as purposive sampling, judgment sampling or deliberate
sampling. It includes.
(1) Quota sampling: In quota sampling, quotas are fixed according to the basic parameters of the population.
Which are earlier determined & each investigation is assigned with quotas of elementary units to be
intervened.
(2) Convenience sampling: In convenience sampling a sample is obtained by selecting convenient population
elements from the population.
(3) Sequential sampling: In sequential sampling, a number of sample lots are drawn one after another
depending upon the results of earlier samples drawn from the same population. It has great use in
statistical quality control.
(4) Cluster sampling involves arrangement of elementary items in a population into heterogeneous sub
groups that are representative of overall population. Then, one of the sub groups is selected for sample study.
III. Mixed sampling or systematic sampling:
In this type of sampling both probability sampling & non probability sampling are used. In this type of sampling,
first unit is selected randomly (probabilistically) & then further units are selected after an equal interval according
to the judgement of the investigator which is non probabilistic in nature.
If universe size is a multiple of sample size, then it is known as linear systematic sampling (N= nk).
If universes size is not a multiple of sample size, the it is known as circular systematic sampling ( N = nk + p ).
Systematic sampling has a severe drawback if there is an unknown and undetected periodicity in the sampling
frame & the sampling interval is a multiple of that period, then we are going to get a most biased sample.
Parameter: The statistical measures of universe are known as parameter
Statistic: The statistical measures of sample are known as statistic.
Parameter Statistic
2|Page
Vishal Vaid: 9829237778, 9829207778 Khandelwal Institute for Studies
Mean X
Standard Deviation S
Proportion/Probability of success P p
No: of items N n
Errors in sample survey
Error or biasness is defined as the difference between the value of population parameter & the value of parameter
obtained/estimated from sample.
Types of errors:
1. Sampling error: Sampling errors are inherent in sampling method. Sampling depends on chance & due
to the existence of chance the sampling error occurs.
Types of sampling errors:
(i) Errors arising out due to defective sampling design.
(ii) Errors arising due to substitution.
(iii) Faulty demarcation of units.
(iv) Error owing to wrong choice of statistic
(v) Variability in the population
2. Non sampling errors: These are those may arise both in sampling method & census method. eg. Non
response, lapse of memory, preference to certain digit, psychological factors, ignorance, etc
(i) due to negligence and carelessness on the part of investigator;
(ii) due to faulty planning of sampling;
(iii) due to the faulty selection of sample units;
(iv) due to incomplete investigation and sample survey;
(v) due to framing of a wrong questionnaire;
(vi) due to negligence and non-response on the part of the respondents;
(vii) due to substitution of a selected unit by another,
(viii) due to error in compilation;
(ix) due to applying w r o n g statistical measure.
Basic principles / laws of sample survey
i. Law of statistical regularity.
ii. Law of inertia of large numbers.
iii. Principle of optimization.
iv. Principle of validity.
(i) Law of statistical regularity: This law states that if a sample is of fairly or moderately large size, then it would
possess the characteristics of that population on an average.
This law explains that if a reasonable large sample is selected at random without bias (i.e., probability
sampling), it is almost certain that on an average, the sample so chosen, shall have the same characteristics as those of
the parent population from where the units constituting the sample have been drawn, It is on the basis of this theory that
the law of statistical theory tells us that a random selection is very likely to give a representative sample.
(ii) Law of inertia of large numbers: This law states that as the sample size increases results are likely to be more
reliable & accurate.
(iii) Principle of optimisation: It states that by selecting a proper sampling design optimum level of efficiency can be
attained at minimum cost.
(iv) Principle of validity: It states that a sampling design is valid only when it is possible to obtain valid estimates &
valid tests about population parameter. Only probability sampling ensures this validity.
Objectives of sampling:
1. Hypothesis testing or significance testing.
2. Estimation of parameter.
Sampling distribution of a statistic:
When a large number of samples of equal size are drawn from a population & a particular statistical measure of each
sample is calculated then the frequency probability distribution of the above statistic is known as sampling distribution of
the above statistic.
3|Page
Vishal Vaid: 9829237778, 9829207778 Khandelwal Institute for Studies
n
Total number of possible samples with replacement = N
Total number of possible samples without replacement = N C
n
N = Universe n= Sample
N=3 n=2
( 1, 3, 5 )
With Replacement N n = 32 = 9
(1,1) 1 X f p
(1, 3) 2 1 1 1/9
(1, 5) 3 2 2 2/9
(3, 1) 2 3 3 3/9
(3, 3) 3 4 2 2/9
(3, 5) 4 5 1 1/9
(5, 1) 3 9
(5, 3) 4
(5, 5) 5
Without Replacement N C = 3C =3
n 2
X f p
(1, 3) 2 1 1/3
(1, 5) 3 1 1/3
(3, 5) 4 1 1/3
Standard Error of a statistic:
It is the standard duration of the sampling distribution. It is a measure which is used to measure the variability of
the values of a statistic computed from samples of the same size drawn from the population.
1. Standard error of mean:
(i) When population standard duration is given
Large sample Small sample
n 30 n < 30
n n
(ii) When population standard duration is not given
Large sample Small sample
n 30 n < 30
s s
n n −1
2. Standard error of difference between 2 sample means:
(i) When population standard duration is given.
Large sample Small sample
n 30 n < 30
1 1 1 1
2 + 2 +
n1 n2 n1 n2
(ii) When population standard duration is not given;
Large sample Small sample
n 30 n < 30
4|Page
Vishal Vaid: 9829237778, 9829207778 Khandelwal Institute for Studies
1 1 1 1
2 + Ŝ 2 +
n1 n2 n1 n2
n1s12 + n2 s22
Ŝ=
n1 + n2 − 2
3. Standard error for sample proportion
When population proportion is given.
PQ
Q=1-P
n
When population proportion is not given:
pq
q=1-p
n
4. Standard error for the difference between 2 sample proportions
When population proportion is given:
1 1
PQ +
n1 n2
When population proportion is not given
1. When samples are from same universe
1 1
P0Q0 +
n1 n2
n1 p1 + n2 p2
P0 = Q0 = 1 - P0
n1 + n2
ii. When samples are from different universe.
p1q1 p2 q2
+
n1 n2
5. Standard error for number of success:
i. When population proportion is given
nPQ
ii. When population proportion is not given.
npq
6. Standard error for standard duration.
(i) When population standard duration is known
2n
(ii) When population standard duration is not known
s
2n
7. Standard error for difference between 2 standard durations.
(i) When universe standard duration is known
5|Page
Vishal Vaid: 9829237778, 9829207778 Khandelwal Institute for Studies
2 1 1
+
2 n1 n2
(ii) When universe standard duration is not known
s2 1 1
+
2 n1 n2
Note: When population parameter is not known and sampling is done without replacement, then in all the above
standard errors, finite population correction (fpc) or finite population multiplier is multiplied.
N −n
fpc =
N −1
Normally fpc is multiplied when sample size is greater than 5% of universe size.
Significance Testing OR Hypothesis Testing
Steps:
1. Setting of Hypothesis:
Assumption / Statement about population
(i) Null Hypothesis - No significant different
(ii) Alternate Hypothesis - Significant different
Sampling design
2. Test Statistic Probability distribution
Sample size
3. Level of Significance ( )
% chance of type I error
Type I error - Rejecting the hypothesis (Null) which is true.
Type II error - Accepting the hypothesis which is false
4. Critical value / Critical Region
5. Decision.
If test statistic > Critical value
Null Hypothesis is rejected
If test statistic < Critical value
Null Hypothesis is accepted
"Z" Test:
It is applied in large samples ( n 30)
Condition or assumptions for "Z" test
1. The population from which sample is drawn is normally distributed.
2. The sample is random sample.
3. n 30
Steps:
1. Setting of hypothesis
Difference
2. Test Statistic (Z) =
S tan dard error
3. Level of significance ( )
10% 5% 4.55% 2% 1% .27%
.10 .05 .0455 .02 .01 .0027
4. Critical Value:
1.645 1.96 2 2.33 2.576 3
5. Decision
Estimation of parameter
6|Page
Vishal Vaid: 9829237778, 9829207778 Khandelwal Institute for Studies
I. Point estimate - When an exact value is calculated.
II. Interval estimate - Limits are taken and called
"Confidence Limits or Fiducial limits".
Ideal estimator
1. Minimum variance & unbiasedness
2. Consistency & efficiency
3. Sufficiency
X = MVUE
Minimum variable unbiased estimator
p = P MVUE
S → MVUE
Biased estimator
n
S→= S
n −1
II. Interval estimate
Confidence limits = Statistic Critical Value Standard error
7|Page