Sampling II
Sampling II
Chapter 4
SAMPLING DISTRIBUTIONS
Introduction
Statistical inference is the branch of statistics which deals with statistical techniques of making inferences
or drawing conclusions about population under study, on the basis of information available in a sample or
samples selected from the population.
The inferences drawn possess a certain amount of uncertainty. Therefore, the statistical methods are
applicable only when the sample used in making inferences about the population is random.
Sampling and sampling distributions are the fundamental basis of statistical inference. The two major areas
of statistical inference are theory of estimation and testing of hypothesis which will be discussed in later
chapters.
Sampling from a Distribution
1. Population
In Statistics, population is not only the human population, but also a group of items specified by certain
characteristics or defined under certain restrictions. So, a population is defined as any group or collection
of objects, animate or inanimate under study. It is also termed as a universe.
In other words, a population is an aggregate of all objects of a given type under consideration at a particular
point of time. The number of objects included in the population is called population size and it is denoted
by 'N'.
Population is the collection or aggregate of all possible objects or units similar to the area under study. In
statistics, population refers to the whole objects or thing under study. Population may be different in
different studies. The nature of elements to be included in a population depends upon the purposes of the
study.
A population may be of two types according to its size: Finite population and Infinite population. A finite
population consists of finite or limited number of objects is called a finite population. Similarly, an infinite
population contains infinitely many objects or elements.
2. Sample
A finite subset of a population selected from the population for the purpose of investigation is a sample. A
sample is selected for the purpose of getting a conclusion about the population on the basis of sample
observations. The number of units included in a sample is called sample size and denoted by (𝑛).
A finite part of a population selected from it with the objective of investigating its properties is called a
sample. The units selected from a population for the investigation is known as a sample. A sample refers to
a finite part of a population.
In many studies, data are to be collected from a population. A population may contain many or infinitely
many units; and their thorough study is not practically possible. Therefore, only a few units of the
population should be selected as a representative of the whole population. Such units form a set called a
sample. Thus, sample is a best representative of a population which estimates the population characteristics
with sufficient accuracy. The sample units are selected under two conditions. The conditions are:
DB THAPA (DMC) - 1
Sampling Distribution
Each and every unit of the population must have non-zero probabilities to be included in the sample.
The selection should be done according to the sampling technique.
While drawing a sample from a population, following criterion should be considered.
Purpose of investigation.
Types of sampling design to be used.
3. Parameter
A parameter is a fixed numerical value that summarizes a characteristic of a population. Since populations
can be large or even infinite, parameters are often unknown and must be estimated from sample data. Exa
mples of parameters include:
Instead of studying every member of the population (which is called a census), we study only a
representative part and draw conclusions about the entire group.
We define sampling as " It is the method of selecting a subset of individuals or observations from a
population to estimate characteristics of the whole population".
1. Saves Time – Studying the whole population may take too long.
2. Saves Cost – Census surveys are expensive.
3. Practicality – Sometimes it is impossible to study the whole population (e.g., testing blood
samples).
4. Accuracy – A carefully selected sample can give very reliable results.
Basic Terms
Advantages of Sampling
Economical
Faster results
Less manpower required
DB THAPA (DMC) - 2
Sampling Distribution
Limitations of Sampling
Objective of Sampling: The objective of sampling is to study about the population by selecting
representative units from the concerned population through minimum resources such as time, cost and
energy without losing accuracy. The major objectives of the sampling are:
a. To obtain maximum possible information about the population through the sample selected.
b. To obtain the limit of accuracy of the estimates of the population parameters.
c. To test the significance of population parameters on the basis of sample statistics.
Census Survey
A census is a complete enumeration (i.e. count) of the population units. In a census survey, each and every
unit of the population is studied. Therefore, there is less chance of error being committed in the survey. A
census survey requires more money, manpower and time. So it is not economic survey but it provides higher
accuracy. This method is free from sampling error.
Advantages of a sample survey over a census
a. It has a greater scope than the census survey.
b. It takes less time than the census survey.
c. It is more economic than the census survey.
d. It has higher quality of the survey work than the census.
Disadvantages of Census Survey
a. It is much more expensive survey.
b. It takes too much time and too much energy for its completion.
c. It is not suitable in certain problems having wide area.
Errors in Sampling: - The term error refers to the difference between the true value of the population
parameter and its estimated value obtained from sample statistics. These errors in statistics may arise due
to a number of factors such as;
a. Approximation in measurement,
b. Approximations in rounding of the figures to the nearest integer,
c. Biases due to faulty collection, presentation, analysis and interpretation of the results,
d. Personal biases of the investigator etc.
In statistical investigation, the discrepancies (errors) between the estimated and the actual values are the net
effect of a multiplicity of factors and can be broadly classified as sampling and non-sampling errors.
a. Sampling errors
DB THAPA (DMC) - 3
Sampling Distribution
When we use sampling instead of a census, the results may not exactly match the true population values.
The difference between the sample result and the true population value is called sampling error. These
occur because only a part of the population is selected, not the whole population.
Sampling error is the difference between a sample statistic and the corresponding population parameter.
Example
If the true average income of a population is Rs. 25,000 but a sample taken from it shows Rs. 24,500, then
the difference (Rs. 500) is sampling error.
b. Non-sampling error
Non sampling errors are men made error which arises at any stage of the survey such as planning and
execution of the survey, data collection, processing and analysis of the data. Non sampling errors are thus
present both in census surveys as well as sample surveys. The census survey, though free from sampling
error would still be subjected to the non-sampling error whereas the sampling survey would be subjected
to both sampling and non-sampling errors. Non sampling errors may occur at every stage of the planning
or execution of the survey.
These errors occur not because of sampling, but due to other mistakes in the process of data collection,
recording, or analysis.
DB THAPA (DMC) - 4
Sampling Distribution
Thus, in sampling,
Proper planning, good questionnaire design, trained investigators, and correct sampling methods help
minimize errors.
Types of Sampling
There are mainly two methods of selecting samples from population. They are probability sampling (or,
Random Sampling) and non-probability sampling (or, Non-random Sampling).
1. Probability/Random Sampling
Probability sampling is the scientific method of selecting samples in which samples are selected in such a
way that all the items of the population have equal chance of being selected in the sample. The probability
sampling is categorized as;
a. Simple random sampling
b. Stratified sampling
c. Systematic sampling
d. Cluster sampling
e. Multi-stage sampling
f. Sampling with Probability Proportion to Size (PPS)
The choice of the above methods will be determined by the purpose for which sampling is sought and the
nature of the population. However, the sample should be representative in nature.
a. Simple random sampling
In this method, samples are selected in such a way that each possible Sample has an equal opportunity of
being selected and each item in the population has an equal chance of being included in the sample. It is
the simplest and common method of sampling in which sample units are drawn one after another with equal
probability of selection for each unit at each draw. Simple random sampling is divided into two categories:
Simple Random sampling without replacement (SRSWOR): - It is the Simple Random Sampling in which
the unit selected in any draw is not replaced before the next draw. If 𝑋 , 𝑋 , … , 𝑋 denote population units
of size N and suppose that samples of size 𝑛 are drawn from it without replacement, then the mean of each
sample (𝑥̅ ) forms a distribution called distribution of sample mean.
DB THAPA (DMC) - 5
Sampling Distribution
a. Lottery Method: - In this method, numbers are assigned to each unit of the population on identical slips
of papers having same shape, size, color, thickness etc. That means each unit of the population is
identified with slips/tickets marked from 1 to N. Then the slips are folded and mixed together in a box.
A blindfold selection is then made to draw the slips one after another. Drawing of slips is continued
until a sample of desired size is obtained. The selection of sample units from the population may be
done either by replacement method or by without replacement method. For replacement method, each
selected ticket is noted and replaced back in the container and the next draw is made. For without
replacement method, each ticket selected is not replaced back in the container before the next draw.
This method is one of the most reliable methods of selecting a random sample.
Example: - Suppose we want to select 25 candidates out of 600. We assign the numbers from 1 to 600 on
the slips having identical shape, size, color etc. one each to a candidate. These slips are then folded,
take into a box and well shuffled. Then 25 slips are drawn one by one. The 25 candidates represented
by these slips will form a sample.
b. Random number table: - The lottery method is time consuming and bulky to use for large population.
The most practical and cheapest method of selecting a random sample is the use of random number
table. The random number table is so constructed that each of the digits 0, 1, 2, ………., 9 appear more
or less same number of times (i.e. frequency) and independent of each other. If the population size N
is of two digits, then the random number table should be formed consisting of two digit numbers from
00 to 99. If the population size N is of three digits, then the random number table should be formed
consisting of three digit numbers from 000 to 999 and so on. The drawing of random sample by random
number table involves the following steps.
1. All the units of the population are numbered from 1 to N.
2. Select any page of the random number table randomly and pick up the numbers in any row, column
or diagonal at random.
3. The population units corresponding to the numbers selected in step 2 form a sample.
There are different random number tables commonly used in practice, prepared by different people such as
Triplet random number table, Fishers and Yates random number table, Kendal and Smith random number
table, Millions Random number etc. A portion of Tippets random number table is presented below for
illustration and its use in the selection of random samples.
2952 6641 3992 9792 5911 3170
5624 4167 9524 1545 7203 5356
1300 2693 2370 7483 3408 3563
1089 6913 7691 0560 5246 1112
6008 8126 4233 8776 2754 9143
1405 7002 6111 8816 6446
c. Remainder Method:
Examples:
Q 1. Draw a random sample of size 100 without replacement from a population of size 4500, using
above random number table.
Solution: -
DB THAPA (DMC) - 6
Sampling Distribution
First of all, we have to identify all 4500 units of the population by serial numbers 0001,0002,0003,
………….,4500. Then we select first 100 numbers from the random table which has values up to 4500.
In the process, the numbers found over 4500 are discarded. If we go row-wise, then the units selected
will be 2952, 3992, 3170, 4167, 1545, 1396, 1300, 2693, 2370, 3408, 2762, 3563, 1089, 0560, 1112,
4233, 2754 and 1402 and so on up to 100 units. The items represented by these random numbers
generate a random sample of size 100.
Q 2. Draw a random sample of size 50 from a population of size 500 from the random number table
given above.
Solution: -
First of all we identify all units of the population by providing serial numbers from 001 to 500. Since
the random numbers are of 4 digits, only first three digits of first column are selected. After this, the
fourth digit of the first column and first two digits of the second column become three digits. Similarly,
the last two digits of second column and the first digit of the third column together form three digit
numbers and so on. Among these three digit numbers so formed, the numbers less than 500 are selected
as sample units and above 500 are discarded. The selection is continued till the desired sample size is
obtained. Thus, the selected sample would contain the numbers 292, 416, 237, 056, 275, 266, 074, 052,
491, 413, and so on up to 50 such numbers.
Q 3. If an investigator is given the following 2 digits' random numbers: 23, 70, 05, 60, 27, 54, 97, 92,
13, 95. How would he make a selection of 5 units from a population consisting of 50 units?
Solution: -
At first, all the units of the population are identified with serial numbers 01, 02, 03 ………, 50. But the
given list of random number has only four numbers less than 50. So, we can’t select a sample of size 5
from it. Therefore, it is required to manage 5 items from the random number table in an alternate way.
We go on selecting sample units by selecting the random numbers row-wise from the beginning. The
items selected above 50 will be divided by 50 and the remainder is taken as the sample unit. First we
take 23 as the first unit and then we take 70 which is greater than 50. So, divide it by 50, the remainder
will be 20. Take 20 as second unit of the sample and so on. Continuing in this way, the sample of size
5 would consist of the units 23, 20, 05, 10 and 27.
Q 4. Select a random sample of 5 students from a class of 45 students.
Solution:
a. Lottery method:
In this method, the names of 45 students or their serial numbers are written on 45 identical slips or
cards. These 45 identical slips are kept in a box and thoroughly shuffled. Then, 5 slips are selected one
by one from the box. If the slips once selected at any draw are replaced back before the next draw, there
is a chance of any slip to be selected more than once. But in sampling without replacement, there is no
chance of repetition of any slip in the sample. The 5 students corresponding to the names or serial
numbers written on the selected slips constitute the required random sample.
b. Using random numbers table:
Since the population size N = 45 is two digits' number, we select the two digit numbers from any page
or section of Fisher and Yates table of random numbers by starting from 1st row and 1st column and
move horizontally to the right side rejecting the random number 00 and the numbers greater than 45,
and 5 students corresponding to the selected random numbers
19, 21, 29, 02, 37
constitute the required random sample.
c. Remainder method:
Since the population size N = 45 is two digits' number, the two digit highest multiple of 45 is 90. If we
select 67, as 67 > 45 and 67 < 90, we divide 67 by 45 and select the remainder 22 as a unit of the random
sample. If we select 19, we accept this 19, as 19 < 45. Continuing this process using Fisher's and Yates
table of random numbers, the 5 students corresponding to the selected random numbers.
DB THAPA (DMC) - 7
Sampling Distribution
where 𝑡ᵢ stands for the value of statistic 𝑡 from the 𝑖ᵗʰ sample.
Similarly, the variance of the sampling distribution of the statistic 𝑡 is given by:
∑ ( ̅)
Var(𝑡) = .
DB THAPA (DMC) - 8
Sampling Distribution
replacement can therefore be considered the same as in sampling with replacement for all practical
purposes.
Sampling Distribution of Sample Mean
The frequency distribution or probability distribution formed from the values of means of the possible
random samples selected from a given population is called sampling distribution of the sample mean.
If the given finite population is normal distribution, then the sampling distribution so obtained is called
sampling distribution of mean of random samples drawn from normal population.
For example, if we select k = 20 random samples from a normal population and calculate sample means for
each sample and also the probability of selecting the samples, then we get a series of 20 means which would
form a probability distribution (or frequency distribution) of the sample mean. This distribution is known
as the sampling distribution of sample means of random samples selected from the normal population.
The sampling distribution of the sample mean describes how the average of a sample behaves when you
repeatedly draw samples from a population. It’s one of the most fundamental ideas in statistics because it
connects individual data to population-level inference.
In short, the sampling distribution of the sample mean describes how sample averages vary from sample
to sample, and why they cluster around the true population mean.
Two Integrals:
Proof:
2. The integral ∫ 𝑧𝑒 𝑑𝑧 = 0, because the function 𝑔(𝑥) = 𝑧 𝑒 is an odd function and for such odd
function,
∫ 𝑔(𝑥) 𝑑𝑥 = 0.
∫ 𝑔(𝑥) 𝑑𝑥 = 0.
Let 𝑥 , 𝑥 , … , 𝑥 be a sample of size 𝑛 drawn from a normal population with mean 𝜇 and variance 𝜎 .
∑
Then the sample mean is obtained as 𝑥̅ = .
DB THAPA (DMC) - 9
Sampling Distribution
Suppose, the probability density function of sample mean (𝑥̅ ) be given by:
where 𝑘 is a constant to be determined from the fact that the total probability is unity. Therefore,
∫ 𝑑𝐹(𝑥̅ ) = 1.
i.e. ∫ 𝑘𝑒 /√ 𝑑𝑥̅ = 1.
Or, 𝑘∫ 𝑒 /√ 𝑑𝑥̅ = 1.
̅
Put = 𝑧. Then, 𝑥̅ = 𝜇 + and 𝑑𝑥̅ = 𝑑𝑧.
√ √
√
Therefore, we get
𝑘∫ 𝑒 𝑑𝑧 = 1.
√
Or, 𝑘∫ 𝑒 𝑑𝑧 = 1.
√
Or, . = 1.
√ √
Or, 𝑘= .
√
√
√
∴ 𝑘= .
√
√
𝑓(𝑥̅ ) = 𝑒 /√ 𝑑𝑥̅ , −∞ < 𝑋 < ∞.
√
√
=∫ 𝑥̅ . 𝑒 √ 𝑑𝑥̅ .
√
√
= ∫ 𝑥̅ . 𝑒 √ 𝑑𝑥̅ . ................(i)
√
DB THAPA (DMC) - 10
Sampling Distribution
Suppose,
𝐼=∫ 𝑥̅ . 𝑒 √ 𝑑𝑥̅ .
̅
Put = 𝑧. Then 𝑥̅ = 𝜇 + and 𝑑𝑥̅ = 𝑑𝑧. Therefore, we have
√ √
√
𝐼=∫ 𝜇+ .𝑒 𝑑𝑧.
√ √
= 𝜇∫ 𝑒 𝑑𝑧 + ∫ 𝑧𝑒 𝑑𝑧 .
√ √
= 𝜎𝜇 .
√
𝐸(𝑥̅ ) = . 𝜎𝜇 = 𝜇.
√
Thus, the mean of the sampling distribution of mean is equal to the population mean. Hence, mean of the
sampling distribution of mean is an unbiased estimator of the population mean.
Similarly, variance of the distribution of sample mean is
𝑉(𝑥̅ ) = 𝐸(𝑥̅ ) − [𝐸(𝑥̅ )] . ............(ii)
Now,
= ∫ 𝑥̅ 𝑒 /√ 𝑑𝑥̅ .
√
√
= ∫ 𝜇+ 𝑒 . 𝑑𝑧.
√ √ √
√
= ∫ (𝜇 + 2 𝑧+ 𝑧 )𝑒 𝑑𝑧.
√ √
= 𝜇 ∫ 𝑒 𝑑𝑧 + 2 ∫ 𝑧𝑒 𝑑𝑧 + ∫ 𝑧 𝑒 𝑑𝑧 .
√ √
= 𝜇 . √2𝜋 + 0 + . √2𝜋 .
√
=𝜇 + .
DB THAPA (DMC) - 11
Sampling Distribution
𝑉(𝑥̅ ) = 𝜇 + − [𝜇] .
= .
Thus, the mean and variance of the sampling distribution of sample means are:
The above derivation proves that while the sample mean remains centered at the population mean 𝜇, its
spread (variance) decreases as the sample size 𝑛 increases.
Thus, the sampling distribution of mean also follows a normal distribution with mean 𝜇 and variance ,
i.e. 𝑥̅ ~𝑁 𝜇, .
The distribution of sample mean (𝑥̅ ) can also be derived by using moment generating function.
Let 𝑥 , 𝑥 , … , 𝑥 be a random sample of size 𝑛 drawn from a normal population 𝑁(𝜇, 𝜎 ). Then the
sample mean is given by
∑
𝑥̅ = .
𝑓(𝑥) = 𝑒 .
√
𝑀 (𝑡) = 𝐸(𝑒 ).
=∫ 𝑒 𝑓(𝑥) 𝑑𝑥.
=∫ 𝑒 . 𝑒 𝑑𝑥.
√
= ∫ 𝑒 𝑒 𝑑𝑥.
√
Therefore, we have
( )
𝑀 (𝑡) = ∫ 𝑒 𝑒 𝜎 𝑑𝑧.
√
DB THAPA (DMC) - 12
Sampling Distribution
( )
= ∫ 𝑒 𝑑𝑧.
√
= ∫ 𝑒 𝑑𝑧.
√
( )
= ∫ 𝑒 𝑑𝑧.
√
( )
= ∫ 𝑒 𝑑𝑧.
√
( )
= ∫ 𝑒 𝑑𝑧.
√
( )
= 𝑒 ∫ 𝑒 𝑑𝑧.
√
( )
=𝑒 ∫ 𝑒 𝑑𝑧.
√
( )
Since ∫ 𝑒 𝑑𝑧 is the integral of a normal density with mean (𝑡𝜎) and variance 1, it equals 1.
√
𝑀 (𝑡) = 𝑒 . ...................(i)
Similarly, the moment generating function of sample mean (𝑥̅ ) is obtained as:
𝑀 ̅ (𝑡) = 𝐸(𝑒 ̅ ),
∑( )
=𝐸 𝑒 ,
⋯
=𝐸 𝑒 ,
=𝐸 ∏ 𝑒 ,
=∏ 𝐸 𝑒 ,
=∏ 𝑀 ,
=∏ 𝑒 , [∵ From (i)]
DB THAPA (DMC) - 13
Sampling Distribution
=𝑒 .
( )
𝜇 = 𝐸(𝑥̅ ) = ,
= 𝜇+ 𝜎 𝑒 ,
= 𝜇.
Also,
( )
𝐸(𝑥̅ ) = = 𝜇+ 𝜎 𝑒
= 𝜇+ 𝜎 𝑒 + 𝑒 ,
=𝜇 + .
Then, Variance (𝑥̅ ) = 𝐸(𝑥̅ ) − [𝐸(𝑥̅ )] = . Thus, the sample mean 𝑥̅ follows normal distribution with
mean 𝜇 and variance .
Hence, we conclude that the sample mean (𝑥̅ ) is distributed as normal distribution with mean 𝜇 and variance
. That is, 𝑥̅ ~𝑁 𝜇, .
Some Results:
Proof: Here, each sample unit 𝑥 is a random draw from the population 𝑋 , 𝑋 , … , 𝑋 , 𝑥 has same
distribution as that of population. Also, the probability of selecting each unit of the population is . So, the
mathematical expectation 𝐸(𝑥 ) is
𝐸(𝑥 ) = ∑ 𝑋 . = 𝜇.
DB THAPA (DMC) - 14
Sampling Distribution
Proof: We know that, the sample units (𝑥 − 𝜇) obtains from the population units (𝑋 − 𝜇) with each of
probability . So, we have
∑ ( )
𝐸(𝑥 − 𝜇) = ∑ (𝑋 − 𝜇) . = =𝜇 .
For instance,
𝐸(𝑥 − 𝜇) = 𝜇 = 𝜎 .
𝝈𝟐
(iii) For SRSWR, 𝑬(𝒙 − 𝝁)𝟐 = 𝑽𝒂𝒓(𝒙)𝑾𝑹 = .
𝒏
Proof: We have,
𝐸(𝑥̅ − 𝜇) = 𝐸(𝑥̅ − 2 𝑥̅ 𝜇 + 𝜇 ) .
= 𝐸(𝑥̅ ) − 2𝜇 + 𝜇 .
= 𝐸(𝑥̅ ) − 𝜇 .
= 𝐸(𝑥̅ ) − [𝐸(𝑥̅ )] .
= 𝑉𝑎𝑟(𝑥̅ ) .
= .
Proof: We know,
= 𝐸(𝑥 − 𝜇) . [∵ 𝐸(𝑥 ) = 𝜇]
𝐸(𝑥 − 𝜇) = ∑ (𝑋 − 𝜇) . =𝜎 .
DB THAPA (DMC) - 15
Sampling Distribution
𝐸 (𝑥 − 𝜇) 𝑥 − 𝜇 = 𝐸(𝑥 − 𝜇) . 𝐸 𝑥 − 𝜇 = 𝜎 .𝜎 = 𝜎 .
Q. Let 𝒙𝟏 , 𝒙𝟐 , … , 𝒙𝒏 be a random sample of size 𝒏 drawn from a normal population 𝑵(𝝁, 𝝈𝟐 ). Then
prove that:
(i) 𝑬(𝒙) = 𝝁.
𝝈𝟐 𝑵 𝟏 𝑺𝟐
(ii) 𝑽𝒂𝒓(𝒙)𝑾𝑹 = = .
𝒏 𝑵 𝒏
∑
𝑥̅ = .
∑
𝐸(𝑥̅ ) = 𝐸 = ∑ 𝐸(𝑥 ).
Since the sample unit 𝑥 is to be selected from the population units 𝑋 , 𝑋 , … , 𝑋 , each of whose chance of
selection is , we have
∑
𝐸(𝑥 ) = ∑ (𝑋 . ) = = 𝜇,
we have,
𝐸(𝑥̅ ) = ∑ 𝜇 = × 𝑛𝜇 = 𝜇.
The mean of the sample mean is the population mean. This shows that the sample mean is an unbiased
estimator of the population mean
Since, sampling is done with replacement (i.e. population is considered to be infinite), the sample
observations 𝑥 , 𝑥 , … , 𝑥 are independent and identically distributed. So, we have
We know,
DB THAPA (DMC) - 16
Sampling Distribution
∑
𝑉(𝑥̅ ) = 𝑉 ,
= ∑ 𝜎 ,
= . 𝑛𝜎 ,
= .
Alternative Method:
= 𝐸(𝑥̅ − 𝜇) ,
⋯
=𝐸 −𝜇 ,
⋯
=𝐸 .
( ) ( ) ⋯ ( )
=𝐸 .
∑ ( )
=𝐸 .
= 𝐸[∑ (𝑥 − 𝜇)] .
= 𝐸∑ (𝑥 − 𝜇) + 2 ∑ (𝑥 − 𝜇) 𝑥 − 𝜇 .
𝐸(𝑥 − 𝜇) = ∑ (𝑥 − 𝜇) = ∑ (𝑥 − 𝜇) = 𝜎 .
∴ 𝑉(𝑥̅ ) = ∑ 𝜎 .
DB THAPA (DMC) - 17
Sampling Distribution
= .𝜎 ∑ (1).
= 𝑛.
= .
∑ ( )
𝑆 = ,
Then, we have
∑ (𝑥 − 𝜇) = 𝑁 𝜎 = (𝑁 − 1)𝑆 .
𝜎 = 𝑆 .
Hence, we have
Q. A sample of size 𝒏 is taken from from a population of size 𝑵 without replacement, then show that
𝑵 𝒏 𝝈𝟐
𝑽(𝒙)𝑾𝑶𝑹 = .
𝑵 𝟏 𝒏
𝑵 𝒏 𝑺𝟐 𝑺𝟐
𝑽(𝒙)𝑾𝑶𝑹 = = (𝟏 − 𝒇) ,
𝑵 𝒏 𝒏
𝒏
where 𝒇 = , called sampling fraction and
𝑵
𝟏 𝟏 𝑺𝟐
𝑽(𝒙)𝑾𝑶𝑹 = − 𝑺𝟐 → as 𝑵 → ∞.
𝒏 𝑵 𝒏
Solution: We know,
= 𝐸(𝑥̅ − 𝜇) ,
DB THAPA (DMC) - 18
Sampling Distribution
⋯
=𝐸 −𝜇 ,
⋯
=𝐸 .
( ) ( ) ⋯ ( )
=𝐸 .
∑ ( )
=𝐸 .
= 𝐸[∑ (𝑥 − 𝜇)] .
= ∑ 𝐸(𝑥 − 𝜇) + 2 ∑ ∑ 𝐸 (𝑥 − 𝜇) 𝑥 − 𝜇 . .....(a)
The sample value (𝑥 − 𝑥̅ ) is expected to come one after another without replacement from any one of the
population items (𝑋 − 𝜇) , (𝑋 − 𝜇) , … , (𝑋 − 𝜇) with each of probability and the sample value (𝑥 −
𝜇)(𝑥 − 𝜇) is expected to come from the population items (𝑋 − 𝜇)(𝑋 − 𝜇) with probabilities . , we
have,
𝐸(𝑥 − 𝜇) = ∑ (𝑋 − 𝜇) = 𝜎 .
Also, we know
∑ (𝑋 − 𝜇) = 0.
⇒ ∑ (𝑋 − 𝜇) = 0.
Or, ∑ (𝑋 − 𝜇) + 2 ∑ ∑ (𝑋 − 𝜇)(𝑋 − 𝜇) = 0.
∴ ∑ ∑ (𝑋 − 𝜇)(𝑋 − 𝜇) = − 𝑁𝜎 .
Then,
𝐸 (𝑥 − 𝜇) 𝑥 − 𝜇 = .( ∑ ∑ (𝑋 − 𝜇) 𝑋 − 𝜇 .
)
= − .( )
. 𝑁𝜎 .
DB THAPA (DMC) - 19
Sampling Distribution
=− . 𝜎 .
∴ 𝑉(𝑥̅ ) = ∑ 𝜎 + 2∑ ∑ − . .𝜎 .
= 𝜎 ∑ (1) − .𝜎 ∑ ∑ (1) .
( )
= 𝜎 𝑛− 𝜎 .
= 1− .
= .
= .
∴ 𝑉(𝑥̅ ) = .
The quantity is called finite population correction factor (FPC). This factor tends to unity when the
[population size is infinitely large. That means;
when 𝑁 → ∞, = = → 1.
The value of FPC factor is taken as unity when the value of sampling fraction ≤ 0.005 i.e. when the
sample includes 5% or less number of units in the sample.
The sampling distribution of the sample variance describes how the variance computed from random
samples fluctuates around the true population variance.
Definition:
∑ ( ̅)
𝑚 = .
DB THAPA (DMC) - 20
Sampling Distribution
Sampling Distribution of Sample Variance: The probability distribution of 𝑠 when repeated samples are
drawn from a population is called a sampling distribution of sample variance. It shows how much 𝑠 varies
from sample to sample. The key property of this distribution is that, the expected value of 𝑠 equals the
population variance 𝜎 . Thus, 𝑠 is an unbiased estimator of 𝜎 .
Case a. Mean and Variance of the Sampling Distribution of Sample Variance under SRSWR
(Sampling from an infinite Population)
Q. Find the mean and variance of the Sampling Distribution of Sample Variance when the sampling
is done with replacement from a population with mean 𝜇 and variance 𝜎 .
∑ ( ̅)
𝑚 = .
∑ ( ̅)
=𝐸 .
= 𝐸[∑ (𝑥 − 𝑥̅ ) ].
= 𝐸[∑ (𝑥 − 𝜇) − 𝑛(𝑥̅ − 𝜇) ].
= [∑ 𝐸(𝑥 − 𝜇) − 𝑛 𝐸(𝑥̅ − 𝜇) ].
= ∑ 𝐸(𝑥 − 𝜇) − 𝐸(𝑥̅ − 𝜇) .
DB THAPA (DMC) - 21
Sampling Distribution
= ∑ 𝐸(𝑥 − 𝜇) − 𝑉𝑎𝑟(𝑥̅ ) .
= ∑ 𝐸(𝑥 − 𝜇) − . .............(i)
Now, Since the sample units (𝑥 − 𝜇) , 𝑖 = 1,2, … , 𝑛 are drawn from the population units (𝑋 − 𝜇) , 𝑖 =
1,2, … , 𝑁 with each of probability , we have
𝐸(𝑥 − 𝜇) = ∑ (𝑋 − 𝜇) = ∑ (𝑋 − 𝜇) = 𝜎 .
𝐸(𝑚 ) = ∑ (𝜎 ) − .
= (𝑛𝜎 ) − .
= 1− 𝜎 .
= 𝜎 .
Therefore, mean of the sampling distribution of sample variance, samples being taken WR from a normal
population with mean 𝜇 and variance 𝜎 is 𝜎 .
= 𝐸(𝑚 ) − 𝜎 . ......................(ii)
Now,
∑ ( ̅)
𝐸(𝑚 ) = 𝐸 .
= 𝐸∑ (𝑥 − 𝑥̅ ) .
= 𝐸∑ 𝑥 − 2𝑥 𝑥̅ + 𝑥̅ .
= 𝐸∑ 𝑥 − 2 𝑥̅ ∑ 𝑥 + 𝑛𝑥̅ .
= 𝐸∑ 𝑥 − 2𝑛 𝑥̅ + 𝑛 𝑥̅ .
DB THAPA (DMC) - 22
Sampling Distribution
= 𝐸∑ 𝑥 − 𝑛 𝑥̅ .
=𝐸 ∑ 𝑥 − 𝑥̅ .
=𝐸 ∑ 𝑥 −2 ∑ 𝑥 . 𝑥̅ + (𝑥̅ ) .
=𝐸 ∑ 𝑥 − 2 𝑥̅ ∑ 𝑥 . +(𝑥̅ ) .
=𝐸 ∑ 𝑥 − (∑ 𝑥) ∑ 𝑥 + {(∑ 𝑥) } .
=𝐸 ∑ 𝑥 + 2∑ 𝑥 𝑥 − ∑ 𝑥 + 2∑ 𝑥𝑥 ∑ 𝑥 +
∑ 𝑥 + 2∑ 𝑥𝑥 .
=𝐸 ∑ 𝑥 + 2∑ 𝑥 𝑥 − ∑ 𝑥 + 2∑ 𝑥 𝑥 + 2∑ 𝑥 𝑥 +
∑ 𝑥 𝑥𝑥 + ∑ 𝑥 + 6∑ 𝑥 𝑥 .
= ∑ 𝐸 𝑥 + 2∑ 𝐸 𝑥 𝑥 − ∑ 𝐸 𝑥 + 2∑ 𝐸 𝑥 𝑥 +
2∑ 𝐸 𝑥 𝑥 +∑ 𝐸 𝑥 𝑥𝑥 + ∑ 𝐸 𝑥 + 6∑ 𝐸 𝑥 𝑥 .
( ) ( ) ( )
= 𝜇 + 𝜇 − 𝜇 − 𝜇 + 𝜇 + 𝜇 .
( ) ( ) ( )
= − + 𝜇 + − + 𝜇 .
( )
= 𝜇 + 1− + 𝜇 .
( ) ( )
= 𝜇 + 1− + 𝜇 + 𝜇 .
( ) ( ( )
= 𝜇 + . 𝜇 + 𝜇 .
𝑉𝑎𝑟(𝑚 ) = 𝐸(𝑚 ) − 𝜎 .
( ) ( ) ( ) ( )
= 𝜇 + . 𝜇 + 𝜇 − 𝜇 .
DB THAPA (DMC) - 23
Sampling Distribution
( ) ( ) ( ) ( )
= 𝜇 + 𝜇 + 𝜇 − 𝜇 .
( ) ( ) ( ) ( )
= 𝜇 − 𝜇 + 𝜇 + 𝜇 .
( ) ( )
= (𝜇 − 𝑛𝜇 + (𝑛 − 1)𝜇 ) + 𝜇 .
( ) ( )
= (𝜇 − 𝜇 ) + 𝜇 .
( )
𝑉𝑎𝑟(𝑚 ) = (𝜇 − 𝜇 ).
( )
𝑉𝑎𝑟(𝑚 ) = (3𝜎 − 𝜎 ).
( )
= . 2𝜎 .
𝑉𝑎𝑟(𝑚 ) = 𝜎 .
Q. Show that the modified sampling variance (𝒔𝟐 ) is unbiased estimate of the population variance
for SRSWR. That means, 𝑬 𝒔𝟐 = 𝝈𝟐 . Also, find the variance of modified sample varince.
Proof: Let 𝑥 , 𝑥 , … , 𝑥 be a sample of size 𝑛 taken from a normal population 𝑋 , 𝑋 , , … , 𝑋 of size 𝑁 with
mean 𝜇 and variance 𝜎 . Then, the modified sample variance denoted by 𝑠 is given by
∑ ( ̅)
𝑠 = .
We want to show that 𝐸(𝑠 ) = 𝜎 and need to find the variance of modified sample variance (𝑠 ).
We know,
∑ ( ̅)
𝐸(𝑠 ) = 𝐸 .
= 𝐸[∑ (𝑥 − 𝑥̅ ) ].
DB THAPA (DMC) - 24
Sampling Distribution
= 𝐸[∑ (𝑥 − 𝜇) − 𝑛(𝑥̅ − 𝜇) ].
= ∑( ) 𝐸(𝑥 − 𝜇) − 𝑛𝐸(𝑥̅ − 𝜇) .
= (𝑛𝜎 − 𝜎 ).
=𝜎 .
Therefore, 𝐸(𝑠 ) = 𝜎 . Hence, the modified sampling variance (𝑠 ) is unbiased estimate of the
population variance.
Also, the variance of the distribution of modified sample variance (𝑠 ) is obtained as follows:
∑ ( ̅) ∑ ( ̅)
We know, 𝑚 = and 𝑠 = . Therefore, we have a relation
𝑚 𝑛 = 𝑠 (𝑛 − 1).
∴𝑠 = 𝑚 .
We know,
( ) ( )
𝑉𝑎𝑟(𝑚 ) = (𝜇 − 𝜇 ) + 𝜇 .
Then,
𝑉𝑎𝑟(𝑠 ) = 𝑉𝑎𝑟 𝑚 .
= 𝑉𝑎𝑟(𝑚 ).
( ) ( )
=( )
× (𝜇 − 𝜇 ) + 𝜇 .
= + 𝜇 .
( )
DB THAPA (DMC) - 25
Sampling Distribution
= +2 − 𝜇 .
=( − )+ 𝜇 .
= + 𝜇 .
𝑉𝑎𝑟(𝑠 ) = 𝜎 . [∵ 𝜇 = 𝜎 ].
Case b. Mean and Variance of the Sampling Distribution of Sample Variance under SRSWOR
(Sampling from a finite Population)
Q. Find the mean and variance of the sampling distribution of sample variance when the samples of
size 𝑛 are drawn without replacement from a normal population of size 𝑁 with mean 𝜇 and variance
𝜎 .
∑ ( ̅)
𝑚 = .
∑ ( ̅)
=𝐸 .
= 𝐸[∑ (𝑥 − 𝑥̅ ) ].
DB THAPA (DMC) - 26
Sampling Distribution
= 𝐸[∑ (𝑥 − 𝜇) − 𝑛(𝑥̅ − 𝜇) ].
= [∑ 𝐸(𝑥 − 𝜇) − 𝑛 𝐸(𝑥̅ − 𝜇) ].
= ∑ 𝐸(𝑥 − 𝜇) − 𝐸(𝑥̅ − 𝜇) .
= ∑ 𝜎 − 𝑉𝑎𝑟(𝑥̅ ) .
= × 𝑛𝜎 − .
=𝜎 − .
= 1− 𝜎 .
( )
( )
= 𝜎 .
( )
( )
Thus, 𝐸(𝑚 ) = , which is the mean of the distribution of sample variance. That is;
( )
( )
𝐸(𝑚 ) = 𝜎 .
( )
∑ ( )
𝑆 = ,
then, (𝑁 − 1)𝑆 = 𝑁𝜎 .
So, we have
( ) ( )
𝐸(𝑚 ) = 𝜎 = × 𝑆 = 𝑆 .
( ) ( )
∴ 𝐸(𝑚 ) = 𝑆 .
Case b. Mean of the modified sample variance (Modified sample variance is unbiased estimate of
population mean square).
∑ ( ̅)
𝑠 = .
DB THAPA (DMC) - 27
Sampling Distribution
∑ ( ̅)
𝑚 = .
Therefore, we have
(𝑛 − 1)𝑠 = 𝑛𝑚 .
∴ 𝑠 = 𝑚 .
⇒ 𝐸(𝑠 ) = 𝐸 𝑚 .
= 𝐸(𝑚 ).
= × 𝑆 .
=𝑆 .
∴ 𝐸(𝑠 ) = 𝑆 .
Hence, modified sample variance 𝑠 is an unbiased estimator of population mean square 𝑆 in SRSWOR.
Case c. For the sufficiently large population, the mean of the sample variance in SRSWOR is 𝜎 ,
and modified sample variance is an unbiased estimator of the population variance.
( )
𝐸(𝑚 ) = .
𝐸(𝑠 ) = 𝐸 𝑚 = 𝐸(𝑚 ) = × 𝜎 =𝜎 .
Thus, modified sample variance is an unbiased estimator of the population variance in SRSWR.
Next, we find the variance of the sampling distribution of sample variance in SRSWOR. We know,
DB THAPA (DMC) - 28
Sampling Distribution
Note: Although the sample mean 𝑥̅ is an unbiased estimator of the population mean 𝜇, the sample variance
∑ ( ̅)
𝑚 = is not an unbiased estimator of the population variance 𝜎 . For this reason, when 𝜎 is
unknown, 𝑚 can't be used for practical purposes. In such case, modified sample variance 𝑠 , which is an
unbiased estimator of population variation 𝜎 can be used for practical purposes. Therefore, 𝑠 plays very
vital role in sampling theory. Moreover, the sample standard deviation 𝑠 is not an unbiased estimator of the
population standard deviation 𝜎.
Q. Define sample and population proportion. Find the sample mean, population mean, sample
variance, population variance, population mean square and modified sample variance in-terms of
sample proportion and population proportion. Also, find mean and variance of the distribution of
sample proportion.
Solution: The sampling distribution of the sample proportion is the probability distribution of all possible
values of the sample proportion (𝑝) obtained from all possible samples of a fixed size (𝑛) drawn from a
population. In simple words, if we repeatedly take many samples of size 𝑛 and calculate the proportion of
successes each time, then the distribution of those values is called the sampling distribution of 𝑝.
Again, suppose 𝑋 be the ith unit of the population and 𝑥 be the ith unit of the sample. For the statistical
analyses, we code the population and sample units in such a way that
1, 𝑖𝑓 𝑖𝑡 𝑖𝑠 𝑎 𝑠𝑢𝑐𝑐𝑒𝑠𝑠
𝑋 = ,
0, 𝑖𝑓 𝑖𝑡 𝑖𝑠 𝑎 𝑓𝑎𝑖𝑙𝑢𝑟𝑒
and,
1, 𝑖𝑓 𝑖𝑡 𝑖𝑠 𝑎 𝑠𝑢𝑐𝑐𝑒𝑠𝑠
𝑥 = .
0, 𝑖𝑓 𝑖𝑡 𝑖𝑠 𝑎 𝑓𝑎𝑖𝑙𝑢𝑟𝑒
∑ 𝑋 =𝑁 =∑ 𝑋 ,
and, ∑ 𝑥 =𝑛 =∑ 𝑥 .
DB THAPA (DMC) - 29
Sampling Distribution
∑
Population mean (𝜇) = = = 𝑃, and
∑
Sample Mean (𝑥̅ ) = = = 𝑝.
Thus, the population mean and sample mean in-terms of proportions are 𝜇 = 𝑃 and 𝑥̅ = 𝑝.
∑ ( )
Population Variance (𝜎 ) =
∑
= .
= ∑ 𝑋 − 2𝜇 ∑ 𝑋 + 𝑛𝜇 .
= ∑ 𝑋 − 2𝑁𝜇 + 𝑁𝜇 .
= ∑ 𝑋 − 𝑁𝜇 .
= (𝑁 − 𝑁𝜇 ).
= −𝜇 .
=𝑃−𝑃 .
= 𝑃(1 − 𝑃).
= 𝑃𝑄.
∑ ( ̅)
Sample variance (𝑚 ) = .
= ∑ 𝑥 − 2 𝑥̅ ∑ 𝑥 + 𝑛𝑥̅ .
= ∑ 𝑥 − 𝑛 𝑥̅ .
= (𝑛 − 𝑛𝑝 ).
=𝑝−𝑝 .
= 𝑝(1 − 𝑝).
DB THAPA (DMC) - 30
Sampling Distribution
= 𝑝𝑞.
Thus, the population variance and sample variance in-terms of proportions are 𝜎 = 𝑃𝑄 and 𝑚 = 𝑝𝑞.
∑ ( )
Population mean square (𝑆 ) = .
∑
= .
= ∑ 𝑋 − 𝑁𝜇 .
= [𝑁 − 𝑁𝑃 ].
= (𝑁𝑃 − 𝑁𝑃 ).
( )
= .
= .
∑ ( ̅)
Modified Sample Variance (𝑠 ) = .
= ∑ 𝑥 − 𝑛𝑥̅ .
= [𝑛 − 𝑛𝑥̅ ].
= − 𝑥̅ .
= (𝑝 − 𝑝 ).
= 𝑝(1 − 𝑝).
= 𝑝𝑞.
Thus, population mean square and modified sample variance in-terms of proportions are
𝑆 = and 𝑠 = 𝑝𝑞.
Now, mean and variance of the sampling distribution of sample proportion are derived as follows:
DB THAPA (DMC) - 31
Sampling Distribution
∑
=𝐸 .
= ∑ 𝐸(𝑥 ).
= ∑( ) ∑ 𝑋 . .
= ∑ 𝑃.
= 𝑃.
Therefore, 𝐸(𝑝) = 𝑃.
= 𝑉𝑎𝑟(𝑛 ).
= 𝑉𝑎𝑟(∑ 𝑥 ).
= ∑ 𝑉𝑎𝑟(𝑥 ).
= ∑ 𝜎 .
= ∑ 𝑃𝑄.
= × 𝑛𝑃𝑄.
= .
Therefore, 𝑉𝑎𝑟(𝑝) = .
Hence, mean and variance of the sampling distribution of sample proportion are 𝑃 and . Thus, the
sampling proportion 𝑝 follows binomial distribution with mean 𝑃 and variance .
Q. Find moment generating function of sample proportion and hence find mean and variance of the
sampling distribution of sample proportion.
DB THAPA (DMC) - 32
Sampling Distribution
Solution: Let a Bernoulli experiment results into 𝑥 , 𝑥 , … , 𝑥 outcomes with 𝑝 as the probability of success
and (1 − 𝑝) as the probability of failure. That is 𝑥 , 𝑥 , … , 𝑥 ~Bernoulli (𝑝). Let us assign
𝑥 = 0 (failure) with probability (𝑞). If, out of 𝑛 trials, 𝑛 outcomes are success
and 𝑛 outcomes are failures, then 𝑛 + 𝑛 = 𝑛. Also,
Now,
∑
proportion of success (𝑝) = = = 𝑥̅ , and
Therefore, 𝑝 + 𝑞 = 1.
∑
𝑝= .
𝑀 (𝑡) = 𝐸(𝑒 ).
∑
.
=𝐸 𝑒 .
.∑
=𝐸 𝑒 .
= 𝑀∑ .
=∏ 𝑀 .
=∑ (1 − 𝑃)𝑒 + 𝑃𝑒 .
DB THAPA (DMC) - 33
Sampling Distribution
𝑀 (𝑡) = 𝑄 + 𝑃𝑒 .
Now,
( )
𝐸(𝑝) = .
= 𝑒 .
= 𝑃+ 𝑒 .
= 𝑃.
And,
( )
𝐸(𝑝 ) = .
= 𝑃+ 𝑒 .
= + (𝑃 + 𝑒 .
= +𝑃 .
Thus, the sampling proportion 𝑝 follows binomial distribution with mean 𝑃 and variance .
Attribute/Categorical variables (Def): The variables which are not measured quantitatively are called
categorical variables. For example, sex, habit, honesty, intelligence, beauty etc. are categorical data.
DB THAPA (DMC) - 34
Sampling Distribution
Case a. Mean and Variance of the Sampling Distribution of Sample Proportion under SRSWR
(Sampling from an infinite Population)
Q. Find the mean and variance of the Sampling Distribution of Sample Proportion under the
sampling with replacement and without replacement from a population containing 𝑁 units whose
mean is 𝜇 and variance 𝜎 .
Solution: Let us consider a population consisting of 𝑁 units 𝑋 , 𝑋 , , … , 𝑋 with mean 𝜇 and variance 𝜎 .
Out of these, suppose 𝑁 outcomes are successes and 𝑁 outcomes are failures. Then, the population
proportion of success (𝑃) is 𝑃 = , and population proportion of failure is 𝑄 = . Suppose a sample
𝑥 , 𝑥 , … , 𝑥 of size 𝑛 is selected from the population and found that 𝑛 outcomes are successes and 𝑛 are
failures. Then, sample proportion of success is 𝑝 = and sample proportion of failure is 𝑞 = .
∑
𝐸(𝑝) = 𝐸 .
= ∑ 𝐸(𝑥 ).
= ∑ 𝑃.
= × 𝑛𝑃.
= 𝑃.
∴ 𝐸(𝑝) = 𝑃. This shows that sample proportion is an unbiased estimator of the population proportion.
= × .
= .
DB THAPA (DMC) - 35
Sampling Distribution
Again, for the sampling with replacement (i.e. for SRSWR), the sampling is considered to be done from an
infinite population. When the population size 𝑁 is sufficiently large, the ratio , called finite population
correction (FPC) would be 1,
Q. What do you mean by standard error of statistics? Discuss its importance in Statistics.
Solution: The standard error (SE) of a statistic is defined as the standard deviation of its sampling
distribution. In simple words, it measures how much a statistic (like sample mean or sample variance,
sample proportion) varies from sample to sample.
Standard error plays very important role in statistics. It measures the reliability of sampling due to chance.
It gives index of the reliability or precision of the estimate of the parameter. Greater is the SE, greater is the
deviation of the actual values from the expected ones. Hence, smaller the value of the SE, greater is the
reliability or precision of the estimate (sample statistics). Thus, the measure of reliability is
Reliability coefficient = .
( )
𝝈
Q. Prove that standard error of sample mean 𝑺𝑬(𝒙) = . Also, express standard error of sample
√𝒏
mean in terms of modified sample sd (𝒔) and sample sd (𝒎𝟐 ).
Solution: Let a random sample 𝑥 , 𝑥 , … , 𝑥 of size 𝑛 be drawn from a normal population with mean 𝜇 and
variance 𝜎 . Then, sample mean (𝑥̅ ) is given by
∑
𝑥̅ = ,
𝑉𝑎𝑟(𝑥̅ ) = .
Therefore, SE(𝑥̅ ) = 𝑉𝑎𝑟(𝑥̅ ) = . Thus, SE of sample mean 𝑥̅ is inversely proportional to the square
√
root of the sample size.
If the population variance (𝜎 ) is unknown, we use modified sample variance (𝑠 ) because it is an unbiased
estimator of the population variance, that is;
DB THAPA (DMC) - 36
Sampling Distribution
𝐸(𝑠 ) = 𝜎 .
SE(𝑥̅ ) = .
√
Further, we know,
𝑛𝑚 = (𝑛 − 1)𝑠 .
∴ 𝑠= 𝑚 .
Hence, we have
SE(𝑥̅ ) = 𝑚 .
√
𝑷𝑸
Q. Prove that standard error of sample proportion is 𝑺𝑬(𝒑) = .
𝒏
Solution: Let a sample of 𝑛 trials consists of 𝑥 number of successes. If the sample is taken from a population
with 𝑋 number of successes and (𝑁 − 𝑋) failures, then, the probability distribution of 𝑋 follows Binomial
Distribution with mean 𝜇 = 𝑃 and variance 𝜎 = 𝑃𝑄. We know,
𝑉𝑎𝑟(𝑝) = .
∴ SE(𝑝) = 𝑉𝑎𝑟(𝑝) = .
If the population proportion (𝑃) of success is unknown, we use sample proportion (𝑝) because it is an
unbiased estimator of the population proportion (𝑃), that is;
𝐸(𝑝) = 𝑃.
So, using sample proportion (𝑝) in place of population proportions 𝑃 and 𝑄, SE(𝑝) becomes
DB THAPA (DMC) - 37
Sampling Distribution
b. i.e., Precision of t =
. .( )
c. The larger the S.E., the less precise (efficient) is the estimate.
d. It is used to test whether the sample statistic differs significantly from the corresponding
hypothetical value in the population.
e. It is used to test the significance of the difference between two independent sample estimates of the
same population parameter.
f. It is used for point estimation of the population parameter.
g. It is used for the interval estimation of the population parameter.
Example: - A population consists of five numbers 1, 3, 5, 7 and 9. Enumerate all possible samples of size
2 drawn from the population without replacement. Find the mean and variance of the population. Find the
mean of the sampling distribution of means and show that it is equal to the population mean. Also, find the
variance of sampling distribution of means and hence verify it.
Solution: -
Here, N = 5 and n = 2
So, possible number of samples of size 2 drawn without replacement = C (N, n) = C(5,2) = = 10
Then the possible samples are: - (1,3), (1,5), (1,7), (1,9), (3,5), (3,7), (3,9), (5,7), (5,9), (7,9)
Calculation of population mean and variance:
X X–𝜇 (𝑋 − 𝜇)
1 -4 16
3 -2 4
5 0 0
7 2 4
DB THAPA (DMC) - 38
Sampling Distribution
9 4 16
∑ 𝑋 =25 ∑(𝑋 − 𝜇) =
40
∑
Here, Population mean (𝜇) = = =5
∑( )
And population variance (𝜎 ) = = =8
Calculation of mean and variance of sampling distribution of means:
Sample Sample Sample 𝑥̅ -𝑥̿ (𝑥̅ -
No. mean(𝑥̅ ) 𝑥̿ )2
1 (1,3) 2 -3 9
2 (1,5) 3 -2 4
3 (1,7) 4 -1 1
4 (1,9) 5 0 0
5 (3,5) 4 -1 1
6 (3,7) 5 0 0
7 (3,9) 6 1 1
8 (5,7) 6 1 1
9 (5,9) 7 2 4
10 (7,9) 8 3 9
Total = 50 =30
∑ ̅
Now, mean of sample means (𝑥̿ ) = = =5
( , )
Since the population mean is also found to be 5, the mean of sampling distribution of means is equal to the
population mean.
And, variance of sample means by using definition is,
∑( ̅ ̿)
Var(𝑥̅ ) = = =3
( , )
Also, the variance of sample mean by using formula is,
Var(𝑥̅ ) = . = . =3
Hence the formula for computing variance of sample mean is verified.
Standard error of means, s. e.(𝑥̅ ) = 𝑉𝑎𝑟(𝑥̅ ) = √3 = 1.732
Example: - A population consists of four numbers 2, 5, 8 and 1. Enumerate all possible samples of size two
which can be drawn from this population with replacement. Find the mean and variance of the population.
Find mean of the sampling distribution of means and show that it is equal to the population mean. Find the
variance of the sampling distribution of means and verify it. Find the standard error of mean.
Solution: Here,
No. of possible samples of size two drawn with replacement from the population = 𝑁 = 4 = 16
And the possible samples are: (1,2), (1,5), (1,8), (2,1), (2,5), (2,8), (5,1), (5,2), (5,8), (8,1), (8,2), (8,5) and
(8,5).
∑
Population mean (𝜇) = = =4
∑( ) ( ) ( ) ( ) ( )
Population variance (𝜎 ) = = = 7.5.
Calculation of mean and variance of the sampling distribution of means:
DB THAPA (DMC) - 39
Sampling Distribution
3 (1,5) 3 -1 1
4 (1,8) 4.5 o.5 0.25
5 (2,1) 1.5 -2.5 6.25
6 (2,2) 2 -2 4
7 (2,5) 3.5 -0.5 0.25
8 (2,8) 5 1 1
9 (5,1) 3 -1 1
10 (5,2) 3.5 -0.5 0.25
11 (5,5) 5 1 1
12 (5,8) 6.5 2.5 6.25
13 (8,1) 4.5 0.5 0.25
14 (8,2) 5 1 1
15 (8,5) 6.5 2.5 6.25
16 (8,8) 8 4 16
=64 (𝑥̅ − 𝑥̿ ) = 60
∑ ̅
Now, mean of sample means (𝑥̿ ) = = = 4. This is equal to population mean. Hence, mean of the
distribution of sample mean is equal to the population mean.
Also, variance of sample mean by the definition is,
∑( ̅ ̿)
Var(𝑥̅ ) = = = 3.75
And variance of the sample mean by using formula is,
.
var(𝑥̅ ) = = = 3.75
Hence the formula for variance of the distribution of sample mean is verified.
Standard error of mean is,
s.e. (𝑥̅ ) = 𝑉𝑎𝑟(𝑥̅ ) = √3.75 = 1.94
Exercise for practice
1. A sample of size 25 is drawn from a population consisting of 150 units. If the population [Link] 10,
find the standard error of sample mean when the sample is drawn (i) without replacement (ii) with
replacement.
2. A simple random sample of size 20 is drawn without replacement from a finite population of 75
units. If the number of defective units in the population is 12, find the standard error of the sample
proportion.
3. Consider a population of four units 3,6,2,1. a) Write down all possible samples of size 2 that can
be drawn with replacement from the sample. b) Find mean and variance of the population. c) Find
mean of the sampling distribution of means and show that it is equal to the population mean. d)
Find the variance of the sampling distribution of means and verify that it agrees with the formula.
e) Find the standard error of mean.
4. How does sampling with replacement differ from that without replacement? Which of them gives
lower value of S.D. of the sample mean? Explain by considering sample of size 2 from a population
consisting five members 2,3,6,8 and 11. Verify that sample mean is an unbiased estimator of the
population means and that its variance is given by (1-f), where the symbols have their usual
meanings.
DB THAPA (DMC) - 40