0% found this document useful (0 votes)
2 views40 pages

Sampling II

Uploaded by

db thapa
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
2 views40 pages

Sampling II

Uploaded by

db thapa
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Sampling Distribution

Chapter 4

SAMPLING DISTRIBUTIONS

Introduction

Statistical inference is the branch of statistics which deals with statistical techniques of making inferences
or drawing conclusions about population under study, on the basis of information available in a sample or
samples selected from the population.
The inferences drawn possess a certain amount of uncertainty. Therefore, the statistical methods are
applicable only when the sample used in making inferences about the population is random.
Sampling and sampling distributions are the fundamental basis of statistical inference. The two major areas
of statistical inference are theory of estimation and testing of hypothesis which will be discussed in later
chapters.
Sampling from a Distribution
1. Population
In Statistics, population is not only the human population, but also a group of items specified by certain
characteristics or defined under certain restrictions. So, a population is defined as any group or collection
of objects, animate or inanimate under study. It is also termed as a universe.
In other words, a population is an aggregate of all objects of a given type under consideration at a particular
point of time. The number of objects included in the population is called population size and it is denoted
by 'N'.
Population is the collection or aggregate of all possible objects or units similar to the area under study. In
statistics, population refers to the whole objects or thing under study. Population may be different in
different studies. The nature of elements to be included in a population depends upon the purposes of the
study.

A population may be of two types according to its size: Finite population and Infinite population. A finite
population consists of finite or limited number of objects is called a finite population. Similarly, an infinite
population contains infinitely many objects or elements.
2. Sample
A finite subset of a population selected from the population for the purpose of investigation is a sample. A
sample is selected for the purpose of getting a conclusion about the population on the basis of sample
observations. The number of units included in a sample is called sample size and denoted by (𝑛).
A finite part of a population selected from it with the objective of investigating its properties is called a
sample. The units selected from a population for the investigation is known as a sample. A sample refers to
a finite part of a population.
In many studies, data are to be collected from a population. A population may contain many or infinitely
many units; and their thorough study is not practically possible. Therefore, only a few units of the
population should be selected as a representative of the whole population. Such units form a set called a
sample. Thus, sample is a best representative of a population which estimates the population characteristics
with sufficient accuracy. The sample units are selected under two conditions. The conditions are:

DB THAPA (DMC) - 1
Sampling Distribution

 Each and every unit of the population must have non-zero probabilities to be included in the sample.
 The selection should be done according to the sampling technique.
 While drawing a sample from a population, following criterion should be considered.
 Purpose of investigation.
 Types of sampling design to be used.
3. Parameter
A parameter is a fixed numerical value that summarizes a characteristic of a population. Since populations
can be large or even infinite, parameters are often unknown and must be estimated from sample data. Exa
mples of parameters include:

 The average height of all adults in a country (population mean).


 The proportion of voters who support a specific candidate (population proportion).
4. Statistics
Statistics is a numerical value calculated from a sample data. It is used to estimate the corresponding
population parameters. Since, statistics are derived from samples, they can vary from one sample to another.
5. Sampling
Sampling is a statistical process of selecting a small number of units from a larger group to obtain
information about the whole population. The process of selecting a sample from a population is called
sampling. Sampling is a tool which helps us to draw conclusions about the characteristics of the population
after studying those units that are included in the sample.

Instead of studying every member of the population (which is called a census), we study only a
representative part and draw conclusions about the entire group.

We define sampling as " It is the method of selecting a subset of individuals or observations from a
population to estimate characteristics of the whole population".

Why Do We Use Sampling?

1. Saves Time – Studying the whole population may take too long.
2. Saves Cost – Census surveys are expensive.
3. Practicality – Sometimes it is impossible to study the whole population (e.g., testing blood
samples).
4. Accuracy – A carefully selected sample can give very reliable results.

Basic Terms

 Sampling Unit – Individual member of the population.


 Sampling Frame – List of all population units.

Advantages of Sampling

 Economical
 Faster results
 Less manpower required

DB THAPA (DMC) - 2
Sampling Distribution

 Often more manageable and efficient

Limitations of Sampling

 Sampling error may occur


 Results depend on sample representativeness
 Risk of bias if selection is improper

Sampling Technique (Steps in Sampling)


The following are some basic steps in sampling.
1. Defining population.
2. Defining sample units.
3. Listing of population units.
4. Deciding sample size.
5. Deciding sampling procedure to be used.
6. Testing of reliability of the sample.

Objective of Sampling: The objective of sampling is to study about the population by selecting
representative units from the concerned population through minimum resources such as time, cost and
energy without losing accuracy. The major objectives of the sampling are:
a. To obtain maximum possible information about the population through the sample selected.
b. To obtain the limit of accuracy of the estimates of the population parameters.
c. To test the significance of population parameters on the basis of sample statistics.

Census Survey
A census is a complete enumeration (i.e. count) of the population units. In a census survey, each and every
unit of the population is studied. Therefore, there is less chance of error being committed in the survey. A
census survey requires more money, manpower and time. So it is not economic survey but it provides higher
accuracy. This method is free from sampling error.
Advantages of a sample survey over a census
a. It has a greater scope than the census survey.
b. It takes less time than the census survey.
c. It is more economic than the census survey.
d. It has higher quality of the survey work than the census.
Disadvantages of Census Survey
a. It is much more expensive survey.
b. It takes too much time and too much energy for its completion.
c. It is not suitable in certain problems having wide area.

Errors in Sampling: - The term error refers to the difference between the true value of the population
parameter and its estimated value obtained from sample statistics. These errors in statistics may arise due
to a number of factors such as;
a. Approximation in measurement,
b. Approximations in rounding of the figures to the nearest integer,
c. Biases due to faulty collection, presentation, analysis and interpretation of the results,
d. Personal biases of the investigator etc.
In statistical investigation, the discrepancies (errors) between the estimated and the actual values are the net
effect of a multiplicity of factors and can be broadly classified as sampling and non-sampling errors.
a. Sampling errors

DB THAPA (DMC) - 3
Sampling Distribution

When we use sampling instead of a census, the results may not exactly match the true population values.
The difference between the sample result and the true population value is called sampling error. These
occur because only a part of the population is selected, not the whole population.

Sampling error is the difference between a sample statistic and the corresponding population parameter.

Causes of Sampling Errors

 Sample size is too small


 Sample is not properly selected
 Natural variation between samples

Controlling of Sampling Errors

Sampling error can be reduced by:

 Increasing sample size


 Using proper random sampling methods

Example

If the true average income of a population is Rs. 25,000 but a sample taken from it shows Rs. 24,500, then
the difference (Rs. 500) is sampling error.

b. Non-sampling error
Non sampling errors are men made error which arises at any stage of the survey such as planning and
execution of the survey, data collection, processing and analysis of the data. Non sampling errors are thus
present both in census surveys as well as sample surveys. The census survey, though free from sampling
error would still be subjected to the non-sampling error whereas the sampling survey would be subjected
to both sampling and non-sampling errors. Non sampling errors may occur at every stage of the planning
or execution of the survey.

These errors occur not because of sampling, but due to other mistakes in the process of data collection,
recording, or analysis.

Causes of Non-Sampling Errors

 Wrong data entry


 Faulty questionnaire
 Biased interviewer
 Non-response from respondents
 Measurement mistakes

Important Points about non-sampling errors:

 It can occur in both sampling and census


 They are often more serious than sampling errors
 It cannot be reduced by increasing sample size

DB THAPA (DMC) - 4
Sampling Distribution

Difference Between Sampling and Non-Sampling Errors

Basis Sampling Error Non-Sampling Error


Meaning Due to studying part of population Due to other mistakes
Occurs in Only sampling Both sampling & census
Control Reduced by larger sample Controlled by careful planning
Nature Random Random or systematic

Thus, in sampling,

Total error = Sampling Error + Non-Sampling Error

Controlling of Non-Sampling Errors

Proper planning, good questionnaire design, trained investigators, and correct sampling methods help
minimize errors.

Types of Sampling
There are mainly two methods of selecting samples from population. They are probability sampling (or,
Random Sampling) and non-probability sampling (or, Non-random Sampling).
1. Probability/Random Sampling
Probability sampling is the scientific method of selecting samples in which samples are selected in such a
way that all the items of the population have equal chance of being selected in the sample. The probability
sampling is categorized as;
a. Simple random sampling
b. Stratified sampling
c. Systematic sampling
d. Cluster sampling
e. Multi-stage sampling
f. Sampling with Probability Proportion to Size (PPS)

The choice of the above methods will be determined by the purpose for which sampling is sought and the
nature of the population. However, the sample should be representative in nature.
a. Simple random sampling
In this method, samples are selected in such a way that each possible Sample has an equal opportunity of
being selected and each item in the population has an equal chance of being included in the sample. It is
the simplest and common method of sampling in which sample units are drawn one after another with equal
probability of selection for each unit at each draw. Simple random sampling is divided into two categories:
Simple Random sampling without replacement (SRSWOR): - It is the Simple Random Sampling in which
the unit selected in any draw is not replaced before the next draw. If 𝑋 , 𝑋 , … , 𝑋 denote population units
of size N and suppose that samples of size 𝑛 are drawn from it without replacement, then the mean of each
sample (𝑥̅ ) forms a distribution called distribution of sample mean.

Simple random sampling with replacement (SRSWR)


It is the simple random sampling in which the unit selected in any draw is replaced before the next draw.
If 𝑋 , 𝑋 , … , 𝑋 denote population units of size N and suppose samples of size n are drawn from it with
replacement then the mean of each sample (𝑥̅ ) forms a distribution called distribution of sample mean.

DB THAPA (DMC) - 5
Sampling Distribution

Randomness of a sample (Method of Drawing a Random Sample):-


The theory of sampling is based on the on the assumption that the sample drawn from a population is
random in nature. But it is difficult to ensure true randomness in the selection of items in a sample. In order
to ensure the randomness, we use (i) Lottery Method and (ii) Random number table.

a. Lottery Method: - In this method, numbers are assigned to each unit of the population on identical slips
of papers having same shape, size, color, thickness etc. That means each unit of the population is
identified with slips/tickets marked from 1 to N. Then the slips are folded and mixed together in a box.
A blindfold selection is then made to draw the slips one after another. Drawing of slips is continued
until a sample of desired size is obtained. The selection of sample units from the population may be
done either by replacement method or by without replacement method. For replacement method, each
selected ticket is noted and replaced back in the container and the next draw is made. For without
replacement method, each ticket selected is not replaced back in the container before the next draw.
This method is one of the most reliable methods of selecting a random sample.

Example: - Suppose we want to select 25 candidates out of 600. We assign the numbers from 1 to 600 on
the slips having identical shape, size, color etc. one each to a candidate. These slips are then folded,
take into a box and well shuffled. Then 25 slips are drawn one by one. The 25 candidates represented
by these slips will form a sample.

b. Random number table: - The lottery method is time consuming and bulky to use for large population.
The most practical and cheapest method of selecting a random sample is the use of random number
table. The random number table is so constructed that each of the digits 0, 1, 2, ………., 9 appear more
or less same number of times (i.e. frequency) and independent of each other. If the population size N
is of two digits, then the random number table should be formed consisting of two digit numbers from
00 to 99. If the population size N is of three digits, then the random number table should be formed
consisting of three digit numbers from 000 to 999 and so on. The drawing of random sample by random
number table involves the following steps.
1. All the units of the population are numbered from 1 to N.
2. Select any page of the random number table randomly and pick up the numbers in any row, column
or diagonal at random.
3. The population units corresponding to the numbers selected in step 2 form a sample.

There are different random number tables commonly used in practice, prepared by different people such as
Triplet random number table, Fishers and Yates random number table, Kendal and Smith random number
table, Millions Random number etc. A portion of Tippets random number table is presented below for
illustration and its use in the selection of random samples.
2952 6641 3992 9792 5911 3170
5624 4167 9524 1545 7203 5356
1300 2693 2370 7483 3408 3563
1089 6913 7691 0560 5246 1112
6008 8126 4233 8776 2754 9143
1405 7002 6111 8816 6446
c. Remainder Method:

Examples:
Q 1. Draw a random sample of size 100 without replacement from a population of size 4500, using
above random number table.
Solution: -

DB THAPA (DMC) - 6
Sampling Distribution

First of all, we have to identify all 4500 units of the population by serial numbers 0001,0002,0003,
………….,4500. Then we select first 100 numbers from the random table which has values up to 4500.
In the process, the numbers found over 4500 are discarded. If we go row-wise, then the units selected
will be 2952, 3992, 3170, 4167, 1545, 1396, 1300, 2693, 2370, 3408, 2762, 3563, 1089, 0560, 1112,
4233, 2754 and 1402 and so on up to 100 units. The items represented by these random numbers
generate a random sample of size 100.
Q 2. Draw a random sample of size 50 from a population of size 500 from the random number table
given above.
Solution: -
First of all we identify all units of the population by providing serial numbers from 001 to 500. Since
the random numbers are of 4 digits, only first three digits of first column are selected. After this, the
fourth digit of the first column and first two digits of the second column become three digits. Similarly,
the last two digits of second column and the first digit of the third column together form three digit
numbers and so on. Among these three digit numbers so formed, the numbers less than 500 are selected
as sample units and above 500 are discarded. The selection is continued till the desired sample size is
obtained. Thus, the selected sample would contain the numbers 292, 416, 237, 056, 275, 266, 074, 052,
491, 413, and so on up to 50 such numbers.
Q 3. If an investigator is given the following 2 digits' random numbers: 23, 70, 05, 60, 27, 54, 97, 92,
13, 95. How would he make a selection of 5 units from a population consisting of 50 units?
Solution: -
At first, all the units of the population are identified with serial numbers 01, 02, 03 ………, 50. But the
given list of random number has only four numbers less than 50. So, we can’t select a sample of size 5
from it. Therefore, it is required to manage 5 items from the random number table in an alternate way.
We go on selecting sample units by selecting the random numbers row-wise from the beginning. The
items selected above 50 will be divided by 50 and the remainder is taken as the sample unit. First we
take 23 as the first unit and then we take 70 which is greater than 50. So, divide it by 50, the remainder
will be 20. Take 20 as second unit of the sample and so on. Continuing in this way, the sample of size
5 would consist of the units 23, 20, 05, 10 and 27.
Q 4. Select a random sample of 5 students from a class of 45 students.
Solution:
a. Lottery method:
In this method, the names of 45 students or their serial numbers are written on 45 identical slips or
cards. These 45 identical slips are kept in a box and thoroughly shuffled. Then, 5 slips are selected one
by one from the box. If the slips once selected at any draw are replaced back before the next draw, there
is a chance of any slip to be selected more than once. But in sampling without replacement, there is no
chance of repetition of any slip in the sample. The 5 students corresponding to the names or serial
numbers written on the selected slips constitute the required random sample.
b. Using random numbers table:
Since the population size N = 45 is two digits' number, we select the two digit numbers from any page
or section of Fisher and Yates table of random numbers by starting from 1st row and 1st column and
move horizontally to the right side rejecting the random number 00 and the numbers greater than 45,
and 5 students corresponding to the selected random numbers
19, 21, 29, 02, 37
constitute the required random sample.
c. Remainder method:
Since the population size N = 45 is two digits' number, the two digit highest multiple of 45 is 90. If we
select 67, as 67 > 45 and 67 < 90, we divide 67 by 45 and select the remainder 22 as a unit of the random
sample. If we select 19, we accept this 19, as 19 < 45. Continuing this process using Fisher's and Yates
table of random numbers, the 5 students corresponding to the selected random numbers.

DB THAPA (DMC) - 7
Sampling Distribution

22, 19, 26, 29, 15


constitute the required random sample.

Sampling Distribution of a Statistic


If we select a number of independent random samples 𝑥 , 𝑥 , … , 𝑥 of definite size (say n), from a given
parent population of size N and calculate some statistic say 𝑡 = 𝑡(𝑥 , 𝑥 , . . . , 𝑥 ) from each sample of the
values 𝑥 , 𝑥 , . . . , 𝑥 , then we shall get a series of values of the statistic. All these values of the statistic can
be assembled together with their relative frequency or the probability with which they occur. This frequency
distribution or probability distribution so formed from the different values of the statistic ′𝑡′ for all the
possible samples is called sampling distribution of that statistic ′𝑡′. In particular, the statistic may be either
sample mean (𝑥̄ ) or sample variance (𝑠 ) or similar statistical measures. As the values of the statistic 𝑡
vary from sample to sample, the differences in the values of 𝑡 are called sampling fluctuations.
Several sampling distributions of statistics can be generated from the given population. For example,
sampling distribution of sample mean (𝑥̄ ), sampling distribution of sample variance (𝑠 ), etc. However,
for practical purposes, it is sufficient to deal with the sampling distribution of only two important statistics
namely, sample mean (𝑥̄ ) and sample variance (𝑠 ).

Mean and Variance of a Sampling Distribution:


The mean and variance of the sampling distribution of a statistic ′𝑡′ (say) can be determined as follows:
Let a sample of size 𝑛 be selected from a given finite population of size 𝑁 with replacement. Then the total
𝑁
number of possible random samples is = 𝑘 (say). We compute some statistic 𝑡 = 𝑡(𝑥₁, 𝑥₂, … , 𝑥ₙ) for
𝑛
each of these 𝑘 samples.
Then the mean of the sampling distribution of the statistic 𝑡 is given by:

𝑡̅ = .

where 𝑡ᵢ stands for the value of statistic 𝑡 from the 𝑖ᵗʰ sample.
Similarly, the variance of the sampling distribution of the statistic 𝑡 is given by:
∑ ( ̅)
Var(𝑡) = .

Note that the sampling distribution tends to normal distribution, if


(i) the size of the sample is large and/or
(ii) number of samples is large.
Also, the positive square root of Var(t) is called standard error of the statistic 't'.
In practice, if the population size 𝑁 is very large in comparison to the sample size 𝑛 drawn from the
population, then there will be a very large number of possible samples of the same size 𝑛. In such case, we
do not get a true or theoretical sampling distribution of a statistic ′𝑡′ but only an experimental sampling
distribution. When the number of samples is large, there should be close agreement between the two
sampling distributions. Therefore, in such case the expected mean and standard deviation of the
experimental sampling distribution would be close to those of the theoretical sampling distribution. Also,
the expected mean and standard deviation of sampling distribution generated from sampling without

DB THAPA (DMC) - 8
Sampling Distribution

replacement can therefore be considered the same as in sampling with replacement for all practical
purposes.
Sampling Distribution of Sample Mean
The frequency distribution or probability distribution formed from the values of means of the possible
random samples selected from a given population is called sampling distribution of the sample mean.
If the given finite population is normal distribution, then the sampling distribution so obtained is called
sampling distribution of mean of random samples drawn from normal population.

For example, if we select k = 20 random samples from a normal population and calculate sample means for
each sample and also the probability of selecting the samples, then we get a series of 20 means which would
form a probability distribution (or frequency distribution) of the sample mean. This distribution is known
as the sampling distribution of sample means of random samples selected from the normal population.

The sampling distribution of the sample mean describes how the average of a sample behaves when you
repeatedly draw samples from a population. It’s one of the most fundamental ideas in statistics because it
connects individual data to population-level inference.

In short, the sampling distribution of the sample mean describes how sample averages vary from sample
to sample, and why they cluster around the true population mean.

Two Integrals:

1. The integral ∫ 𝑒 𝑑𝑧 = √2𝜋 has a great importance in this unit.

Proof:

2. The integral ∫ 𝑧𝑒 𝑑𝑧 = 0, because the function 𝑔(𝑥) = 𝑧 𝑒 is an odd function and for such odd
function,

∫ 𝑔(𝑥) 𝑑𝑥 = 0.

But the limits (−∞ 𝑡𝑜 ∞) are symmetric, we have

∫ 𝑔(𝑥) 𝑑𝑥 = 0.

Derivation of Sampling Distribution of Sample Mean (𝒙).

i. Using Distribution Function:

Let 𝑥 , 𝑥 , … , 𝑥 be a sample of size 𝑛 drawn from a normal population with mean 𝜇 and variance 𝜎 .

Then the sample mean is obtained as 𝑥̅ = .

We know that 𝑋~𝑁(𝜇, 𝜎 ) and the probability density function of 𝑋 is:

DB THAPA (DMC) - 9
Sampling Distribution

𝑓(𝑥) = 𝑒 , −∞ < 𝑋 < ∞.


Suppose, the probability density function of sample mean (𝑥̅ ) be given by:

𝑓(𝑥̅ ) = 𝑘 𝑒 /√ 𝑑𝑥̅ , −∞ < 𝑋 < ∞,

where 𝑘 is a constant to be determined from the fact that the total probability is unity. Therefore,

∫ 𝑑𝐹(𝑥̅ ) = 1.

i.e. ∫ 𝑘𝑒 /√ 𝑑𝑥̅ = 1.

Or, 𝑘∫ 𝑒 /√ 𝑑𝑥̅ = 1.

̅
Put = 𝑧. Then, 𝑥̅ = 𝜇 + and 𝑑𝑥̅ = 𝑑𝑧.
√ √

Therefore, we get

𝑘∫ 𝑒 𝑑𝑧 = 1.

Or, 𝑘∫ 𝑒 𝑑𝑧 = 1.

Or, . = 1.
√ √

Or, 𝑘= .


∴ 𝑘= .

Therefore, the probability density function of the sample mean (𝑥̅ ) is


𝑓(𝑥̅ ) = 𝑒 /√ 𝑑𝑥̅ , −∞ < 𝑋 < ∞.

Then, mean of the distribution of sample mean is

𝐸(𝑥̅ ) = ∫ 𝑥̅ 𝑓(𝑥̅ ) 𝑑𝑥̅ .


=∫ 𝑥̅ . 𝑒 √ 𝑑𝑥̅ .


= ∫ 𝑥̅ . 𝑒 √ 𝑑𝑥̅ . ................(i)

DB THAPA (DMC) - 10
Sampling Distribution

Suppose,

𝐼=∫ 𝑥̅ . 𝑒 √ 𝑑𝑥̅ .
̅
Put = 𝑧. Then 𝑥̅ = 𝜇 + and 𝑑𝑥̅ = 𝑑𝑧. Therefore, we have
√ √

𝐼=∫ 𝜇+ .𝑒 𝑑𝑧.
√ √

= 𝜇∫ 𝑒 𝑑𝑧 + ∫ 𝑧𝑒 𝑑𝑧 .
√ √

= 𝜇√2𝜋 + 0 . [∵ 𝑔(𝑧) = 𝑧𝑒 is odd function]


= 𝜎𝜇 .

Therefore, from equation (i),


𝐸(𝑥̅ ) = . 𝜎𝜇 = 𝜇.

Thus, the mean of the sampling distribution of mean is equal to the population mean. Hence, mean of the
sampling distribution of mean is an unbiased estimator of the population mean.
Similarly, variance of the distribution of sample mean is
𝑉(𝑥̅ ) = 𝐸(𝑥̅ ) − [𝐸(𝑥̅ )] . ............(ii)
Now,

𝐸(𝑥̅ ) = ∫ 𝑥̅ 𝑓 ̅(𝑥̅ ) 𝑑𝑥̅ .

= ∫ 𝑥̅ 𝑒 /√ 𝑑𝑥̅ .

= ∫ 𝜇+ 𝑒 . 𝑑𝑧.
√ √ √

= ∫ (𝜇 + 2 𝑧+ 𝑧 )𝑒 𝑑𝑧.
√ √

= 𝜇 ∫ 𝑒 𝑑𝑧 + 2 ∫ 𝑧𝑒 𝑑𝑧 + ∫ 𝑧 𝑒 𝑑𝑧 .
√ √

= 𝜇 . √2𝜋 + 0 + . √2𝜋 .

=𝜇 + .

Then, from equation (ii),

DB THAPA (DMC) - 11
Sampling Distribution

𝑉(𝑥̅ ) = 𝜇 + − [𝜇] .

= .

Thus, the mean and variance of the sampling distribution of sample means are:

Mean = 𝜇 and Variance = .

The above derivation proves that while the sample mean remains centered at the population mean 𝜇, its
spread (variance) decreases as the sample size 𝑛 increases.
Thus, the sampling distribution of mean also follows a normal distribution with mean 𝜇 and variance ,
i.e. 𝑥̅ ~𝑁 𝜇, .

ii. Using Moment Generating Function:

The distribution of sample mean (𝑥̅ ) can also be derived by using moment generating function.

Let 𝑥 , 𝑥 , … , 𝑥 be a random sample of size 𝑛 drawn from a normal population 𝑁(𝜇, 𝜎 ). Then the
sample mean is given by


𝑥̅ = .

We know, the probability density function of 𝑋~𝑁(𝜇, 𝜎 ) is given by

𝑓(𝑥) = 𝑒 .

Then, the moment generating function of 𝑋 is

𝑀 (𝑡) = 𝐸(𝑒 ).

=∫ 𝑒 𝑓(𝑥) 𝑑𝑥.

=∫ 𝑒 . 𝑒 𝑑𝑥.

= ∫ 𝑒 𝑒 𝑑𝑥.

Put = 𝑧. Then, 𝑥 = 𝜇 + 𝜎𝑧 and 𝑑𝑥 = 𝜎. 𝑑𝑧

Therefore, we have

( )
𝑀 (𝑡) = ∫ 𝑒 𝑒 𝜎 𝑑𝑧.

DB THAPA (DMC) - 12
Sampling Distribution

( )
= ∫ 𝑒 𝑑𝑧.

= ∫ 𝑒 𝑑𝑧.

( )
= ∫ 𝑒 𝑑𝑧.

( )
= ∫ 𝑒 𝑑𝑧.

( )
= ∫ 𝑒 𝑑𝑧.

( )
= 𝑒 ∫ 𝑒 𝑑𝑧.

( )
=𝑒 ∫ 𝑒 𝑑𝑧.

( )
Since ∫ 𝑒 𝑑𝑧 is the integral of a normal density with mean (𝑡𝜎) and variance 1, it equals 1.

Therefore, the moment generating function of 𝑋 is

𝑀 (𝑡) = 𝑒 . ...................(i)

Similarly, the moment generating function of sample mean (𝑥̅ ) is obtained as:

𝑀 ̅ (𝑡) = 𝐸(𝑒 ̅ ),

∑( )
=𝐸 𝑒 ,


=𝐸 𝑒 ,

=𝐸 ∏ 𝑒 ,

=∏ 𝐸 𝑒 ,

=∏ 𝑀 ,

=∏ 𝑒 , [∵ From (i)]

DB THAPA (DMC) - 13
Sampling Distribution

= 𝑒 , [∵ There are n same terms in product]

=𝑒 .

Now, the mean of the sampling distribution of mean is given by

( )
𝜇 = 𝐸(𝑥̅ ) = ,

= 𝜇+ 𝜎 𝑒 ,

= 𝜇.

Also,

( )
𝐸(𝑥̅ ) = = 𝜇+ 𝜎 𝑒

= 𝜇+ 𝜎 𝑒 + 𝑒 ,

=𝜇 + .

Then, Variance (𝑥̅ ) = 𝐸(𝑥̅ ) − [𝐸(𝑥̅ )] = . Thus, the sample mean 𝑥̅ follows normal distribution with
mean 𝜇 and variance .

Hence, we conclude that the sample mean (𝑥̅ ) is distributed as normal distribution with mean 𝜇 and variance
. That is, 𝑥̅ ~𝑁 𝜇, .

Some Results:

Let 𝑥 , 𝑥 , … , 𝑥 be a random sample of size 𝑛 selected from a population 𝑋 , 𝑋 , … , 𝑋 of size 𝑁, with


mean 𝜇 and variance 𝜎 . Then, the following results hold true.

(i) 𝑬(𝒙𝒊 ) = 𝝁. (𝒙𝒊 has same distribution as that of population).

Proof: Here, each sample unit 𝑥 is a random draw from the population 𝑋 , 𝑋 , … , 𝑋 , 𝑥 has same
distribution as that of population. Also, the probability of selecting each unit of the population is . So, the
mathematical expectation 𝐸(𝑥 ) is

𝐸(𝑥 ) = ∑ 𝑋 . = 𝜇.

DB THAPA (DMC) - 14
Sampling Distribution

(ii) For SRSWR, 𝑬(𝒙𝒊 − 𝝁)𝒓 = 𝝁𝒓 (rth central moment).

Proof: We know that, the sample units (𝑥 − 𝜇) obtains from the population units (𝑋 − 𝜇) with each of
probability . So, we have

∑ ( )
𝐸(𝑥 − 𝜇) = ∑ (𝑋 − 𝜇) . = =𝜇 .

For instance,

𝐸(𝑥 − 𝜇) = 𝜇 = 𝜎 .

𝝈𝟐
(iii) For SRSWR, 𝑬(𝒙 − 𝝁)𝟐 = 𝑽𝒂𝒓(𝒙)𝑾𝑹 = .
𝒏

Proof: We have,

𝐸(𝑥̅ − 𝜇) = 𝐸(𝑥̅ − 2 𝑥̅ 𝜇 + 𝜇 ) .

= 𝐸(𝑥̅ ) − 2𝜇𝐸(𝑥̅ ) + 𝐸(𝜇 ).

= 𝐸(𝑥̅ ) − 2𝜇 + 𝜇 .

= 𝐸(𝑥̅ ) − 𝜇 .

= 𝐸(𝑥̅ ) − [𝐸(𝑥̅ )] .

= 𝑉𝑎𝑟(𝑥̅ ) .

= .

Similarly, For SRSWOR, we can prove 𝐸(𝑥̅ − 𝜇) = 𝑉𝑎𝑟(𝑥̅ ) .

(iv) For SRSWR, 𝑽𝒂𝒓(𝒙𝒊 ) = 𝝈𝟐 .

Proof: We know,

𝑉(𝑥 ) = 𝐸[𝑥 − 𝐸(𝑥 )] .

= 𝐸(𝑥 − 𝜇) . [∵ 𝐸(𝑥 ) = 𝜇]

Since, each sample unit (𝑥 − 𝜇) is taken from the population units (𝑋 − 𝜇) , (𝑋 − 𝜇) , … , (𝑥 − 𝜇) ,


each of whose selection being , we have

𝐸(𝑥 − 𝜇) = ∑ (𝑋 − 𝜇) . =𝜎 .

Therefore, 𝑉(𝑥 ) = 𝜎 . Hence proved.

DB THAPA (DMC) - 15
Sampling Distribution

Also, we note the following results:

i. For independent variables 𝑥 and 𝑦, 𝐸(𝑥𝑦) = 𝐸(𝑥). 𝐸(𝑦).

ii. If the sampling is done with replacement (independent), then

𝐸 (𝑥 − 𝜇) 𝑥 − 𝜇 = 𝐸(𝑥 − 𝜇) . 𝐸 𝑥 − 𝜇 = 𝜎 .𝜎 = 𝜎 .

Q. Let 𝒙𝟏 , 𝒙𝟐 , … , 𝒙𝒏 be a random sample of size 𝒏 drawn from a normal population 𝑵(𝝁, 𝝈𝟐 ). Then
prove that:

(i) 𝑬(𝒙) = 𝝁.

𝝈𝟐 𝑵 𝟏 𝑺𝟐
(ii) 𝑽𝒂𝒓(𝒙)𝑾𝑹 = = .
𝒏 𝑵 𝒏

Solution: Let 𝑥 , 𝑥 , … , 𝑥 be a random sample of size 𝑛 selected from a population 𝑋 , 𝑋 , … , 𝑋 of size


𝑁, with mean 𝜇 and variance 𝜎 . Then,

(i) Mean of the sample mean 𝑬(𝒙):

The sample mean 𝑥̅ is given by


𝑥̅ = .

The mean of the sample mean is:


𝐸(𝑥̅ ) = 𝐸 = ∑ 𝐸(𝑥 ).

Since the sample unit 𝑥 is to be selected from the population units 𝑋 , 𝑋 , … , 𝑋 , each of whose chance of
selection is , we have

𝐸(𝑥 ) = ∑ (𝑋 . ) = = 𝜇,
we have,
𝐸(𝑥̅ ) = ∑ 𝜇 = × 𝑛𝜇 = 𝜇.
The mean of the sample mean is the population mean. This shows that the sample mean is an unbiased
estimator of the population mean

(ii) Variance of the Sample Mean 𝑽𝒂𝒓(𝒙)𝑾𝑹:

Since, sampling is done with replacement (i.e. population is considered to be infinite), the sample
observations 𝑥 , 𝑥 , … , 𝑥 are independent and identically distributed. So, we have

𝐸(𝑥 ) = 𝜇 and 𝑉(𝑥 ) = 𝜎 .

We know,

DB THAPA (DMC) - 16
Sampling Distribution


𝑉(𝑥̅ ) = 𝑉 ,

= 𝑉(∑ 𝑥 ), [∵ 𝑉(𝑎𝑥) = 𝑎 𝑉(𝑥)]

= ∑ 𝑉(𝑥 ), [∵ Events are independent, V(x + y) = V(x) + V(y)]

= ∑ 𝜎 ,

= . 𝑛𝜎 ,

= .

Alternative Method:

𝑉(𝑥̅ ) = 𝐸[𝑥̅ − 𝐸(𝑥̅ )]

= 𝐸(𝑥̅ − 𝜇) ,


=𝐸 −𝜇 ,


=𝐸 .

( ) ( ) ⋯ ( )
=𝐸 .

∑ ( )
=𝐸 .

= 𝐸[∑ (𝑥 − 𝜇)] .

= 𝐸∑ (𝑥 − 𝜇) + 2 ∑ (𝑥 − 𝜇) 𝑥 − 𝜇 .

= ∑ 𝐸(𝑥 − 𝜇) . [∵ Events are independent]

Since, the events (𝑥 − 𝜇) , 𝑖 = 1,2, … , 𝑛 occur from the population units (𝑋 − 𝜇) , (𝑋 − 𝜇) , … , (𝑋 −


𝜇) , each of whose probability of selection is , we have

𝐸(𝑥 − 𝜇) = ∑ (𝑥 − 𝜇) = ∑ (𝑥 − 𝜇) = 𝜎 .

∴ 𝑉(𝑥̅ ) = ∑ 𝜎 .

DB THAPA (DMC) - 17
Sampling Distribution

= .𝜎 ∑ (1).

= 𝑛.

= .

Therefore, 𝑉(𝑥̅ ) = . Hence proved.

Next, the population mean square is given by

∑ ( )
𝑆 = ,

Then, we have

∑ (𝑥 − 𝜇) = 𝑁 𝜎 = (𝑁 − 1)𝑆 .

This gives that

𝜎 = 𝑆 .

Hence, we have

𝑉(𝑥̅ ) = . Hence proved.

Q. A sample of size 𝒏 is taken from from a population of size 𝑵 without replacement, then show that

𝑵 𝒏 𝝈𝟐
𝑽(𝒙)𝑾𝑶𝑹 = .
𝑵 𝟏 𝒏

Also, show that

𝑵 𝒏 𝑺𝟐 𝑺𝟐
𝑽(𝒙)𝑾𝑶𝑹 = = (𝟏 − 𝒇) ,
𝑵 𝒏 𝒏

𝒏
where 𝒇 = , called sampling fraction and
𝑵

𝟏 𝟏 𝑺𝟐
𝑽(𝒙)𝑾𝑶𝑹 = − 𝑺𝟐 → as 𝑵 → ∞.
𝒏 𝑵 𝒏

Solution: We know,

𝑉(𝑥̅ ) = 𝐸[𝑥̅ − 𝐸(𝑥̅ )]

= 𝐸(𝑥̅ − 𝜇) ,

DB THAPA (DMC) - 18
Sampling Distribution


=𝐸 −𝜇 ,


=𝐸 .

( ) ( ) ⋯ ( )
=𝐸 .

∑ ( )
=𝐸 .

= 𝐸[∑ (𝑥 − 𝜇)] .

= 𝐸[(𝑥 − 𝜇) + ⋯ + (𝑥 − 𝜇) + 2(𝑥 − 𝜇)(𝑥 − 𝜇) + ⋯ ].

= [𝐸(𝑥 − 𝜇) + 𝐸(𝑥 − 𝜇) + ⋯ + 2 𝐸(𝑥 − 𝜇)(𝑥 − 𝜇) + ⋯ ].

= ∑ 𝐸(𝑥 − 𝜇) + 2 ∑ ∑ 𝐸 (𝑥 − 𝜇) 𝑥 − 𝜇 . .....(a)

The sample value (𝑥 − 𝑥̅ ) is expected to come one after another without replacement from any one of the
population items (𝑋 − 𝜇) , (𝑋 − 𝜇) , … , (𝑋 − 𝜇) with each of probability and the sample value (𝑥 −
𝜇)(𝑥 − 𝜇) is expected to come from the population items (𝑋 − 𝜇)(𝑋 − 𝜇) with probabilities . , we
have,

𝐸(𝑥 − 𝜇) = ∑ (𝑋 − 𝜇) = 𝜎 .

Also, we know

∑ (𝑋 − 𝜇) = 0.

⇒ ∑ (𝑋 − 𝜇) = 0.

Or, ∑ (𝑋 − 𝜇) + 2 ∑ ∑ (𝑋 − 𝜇)(𝑋 − 𝜇) = 0.

Or, 2∑ ∑ (𝑋 − 𝜇)(𝑋 − 𝜇) = − ∑ (𝑋 − 𝜇) = −𝑁𝜎 .

∴ ∑ ∑ (𝑋 − 𝜇)(𝑋 − 𝜇) = − 𝑁𝜎 .

Then,

𝐸 (𝑥 − 𝜇) 𝑥 − 𝜇 = .( ∑ ∑ (𝑋 − 𝜇) 𝑋 − 𝜇 .
)

= − .( )
. 𝑁𝜎 .

DB THAPA (DMC) - 19
Sampling Distribution

=− . 𝜎 .

Hence, from equation (a), we get

∴ 𝑉(𝑥̅ ) = ∑ 𝜎 + 2∑ ∑ − . .𝜎 .

= 𝜎 ∑ (1) − .𝜎 ∑ ∑ (1) .

( )
= 𝜎 𝑛− 𝜎 .

= 1− .

= .

= .

∴ 𝑉(𝑥̅ ) = .

The quantity is called finite population correction factor (FPC). This factor tends to unity when the
[population size is infinitely large. That means;

when 𝑁 → ∞, = = → 1.

Therefore, 𝑉(𝑥̅ ) = when 𝑁 → ∞.

The value of FPC factor is taken as unity when the value of sampling fraction ≤ 0.005 i.e. when the
sample includes 5% or less number of units in the sample.

Sampling Distribution of Sample Variances

The sampling distribution of the sample variance describes how the variance computed from random
samples fluctuates around the true population variance.

Definition:

Sample Variance: Let 𝑥 , 𝑥 , … , 𝑥 be a sample of size 𝑛 taken from a population 𝑋 , 𝑋 , … , 𝑋 of size 𝑁.


Then, the sample variance denoted by 𝑚 is defined as

∑ ( ̅)
𝑚 = .

DB THAPA (DMC) - 20
Sampling Distribution

Sampling Distribution of Sample Variance: The probability distribution of 𝑠 when repeated samples are
drawn from a population is called a sampling distribution of sample variance. It shows how much 𝑠 varies
from sample to sample. The key property of this distribution is that, the expected value of 𝑠 equals the
population variance 𝜎 . Thus, 𝑠 is an unbiased estimator of 𝜎 .

Case a. Mean and Variance of the Sampling Distribution of Sample Variance under SRSWR
(Sampling from an infinite Population)

Q. Find the mean and variance of the Sampling Distribution of Sample Variance when the sampling
is done with replacement from a population with mean 𝜇 and variance 𝜎 .

Solution: Let 𝑥 , 𝑥 , … , 𝑥 be a sample of size 𝑛 taken from a normal population 𝑋 , 𝑋 , , … , 𝑋 of size 𝑁


with mean 𝜇 and variance 𝜎 . Then, the sample variance denoted by 𝑚 is given by

∑ ( ̅)
𝑚 = .

We know, Mean of 𝑚 = 𝐸(𝑚 ) and 𝑉𝑎𝑟(𝑚 ) = 𝐸(𝑚 ) − [𝐸(𝑚 )] .

Now, mean = 𝐸(𝑚 ).

∑ ( ̅)
=𝐸 .

= 𝐸[∑ (𝑥 − 𝑥̅ ) ].

= 𝐸[∑ [(𝑥 − 𝜇) − (𝑥̅ − 𝜇) ].

= 𝐸[∑ {(𝑥 − 𝜇) − 2(𝑥 − 𝜇)(𝑥̅ − 𝜇) + (𝑥̅ − 𝜇) }].

= 𝐸[∑ (𝑥 − 𝜇) − 2(𝑥̅ − 𝜇) ∑ (𝑥 − 𝜇) + (𝑥̅ − 𝜇) ∑ (1)].

= 𝐸[∑ (𝑥 − 𝜇) − 2(𝑥̅ − 𝜇)(∑ 𝑥 − 𝑛𝜇) + 𝑛(𝑥̅ − 𝜇) ].

= 𝐸[∑ (𝑥 − 𝜇) − 2(𝑥̅ − 𝜇)(𝑛𝑥̅ − 𝑛𝜇) + 𝑛(𝑥̅ − 𝜇) ].

= 𝐸[∑ (𝑥 − 𝜇) − 2𝑛(𝑥̅ − 𝜇) + 𝑛(𝑥̅ − 𝜇) ].

= 𝐸[∑ (𝑥 − 𝜇) − 𝑛(𝑥̅ − 𝜇) ].

= [∑ 𝐸(𝑥 − 𝜇) − 𝑛 𝐸(𝑥̅ − 𝜇) ].

= ∑ 𝐸(𝑥 − 𝜇) − 𝐸(𝑥̅ − 𝜇) .

DB THAPA (DMC) - 21
Sampling Distribution

= ∑ 𝐸(𝑥 − 𝜇) − 𝑉𝑎𝑟(𝑥̅ ) .

= ∑ 𝐸(𝑥 − 𝜇) − . .............(i)

Now, Since the sample units (𝑥 − 𝜇) , 𝑖 = 1,2, … , 𝑛 are drawn from the population units (𝑋 − 𝜇) , 𝑖 =
1,2, … , 𝑁 with each of probability , we have

𝐸(𝑥 − 𝜇) = ∑ (𝑋 − 𝜇) = ∑ (𝑋 − 𝜇) = 𝜎 .

Therefore, from equation (i), we have

𝐸(𝑚 ) = ∑ (𝜎 ) − .

= (𝑛𝜎 ) − .

= 1− 𝜎 .

= 𝜎 .

Therefore, mean of the sampling distribution of sample variance, samples being taken WR from a normal
population with mean 𝜇 and variance 𝜎 is 𝜎 .

Also, the variance of the sampling distribution of sample variance is given by

𝑉𝑎𝑟(𝑚 ) = 𝐸(𝑚 ) − [𝐸(𝑚 )] .

= 𝐸(𝑚 ) − 𝜎 . ......................(ii)

Now,

∑ ( ̅)
𝐸(𝑚 ) = 𝐸 .

= 𝐸∑ (𝑥 − 𝑥̅ ) .

= 𝐸∑ 𝑥 − 2𝑥 𝑥̅ + 𝑥̅ .

= 𝐸∑ 𝑥 − 2 𝑥̅ ∑ 𝑥 + 𝑛𝑥̅ .

= 𝐸∑ 𝑥 − 2𝑛 𝑥̅ + 𝑛 𝑥̅ .

DB THAPA (DMC) - 22
Sampling Distribution

= 𝐸∑ 𝑥 − 𝑛 𝑥̅ .

=𝐸 ∑ 𝑥 − 𝑥̅ .

=𝐸 ∑ 𝑥 −2 ∑ 𝑥 . 𝑥̅ + (𝑥̅ ) .

=𝐸 ∑ 𝑥 − 2 𝑥̅ ∑ 𝑥 . +(𝑥̅ ) .

=𝐸 ∑ 𝑥 − (∑ 𝑥) ∑ 𝑥 + {(∑ 𝑥) } .

=𝐸 ∑ 𝑥 + 2∑ 𝑥 𝑥 − ∑ 𝑥 + 2∑ 𝑥𝑥 ∑ 𝑥 +
∑ 𝑥 + 2∑ 𝑥𝑥 .

=𝐸 ∑ 𝑥 + 2∑ 𝑥 𝑥 − ∑ 𝑥 + 2∑ 𝑥 𝑥 + 2∑ 𝑥 𝑥 +
∑ 𝑥 𝑥𝑥 + ∑ 𝑥 + 6∑ 𝑥 𝑥 .

= ∑ 𝐸 𝑥 + 2∑ 𝐸 𝑥 𝑥 − ∑ 𝐸 𝑥 + 2∑ 𝐸 𝑥 𝑥 +
2∑ 𝐸 𝑥 𝑥 +∑ 𝐸 𝑥 𝑥𝑥 + ∑ 𝐸 𝑥 + 6∑ 𝐸 𝑥 𝑥 .

= [𝑛𝜇 + 𝑛(𝑛 − 1)𝜇 ] − [𝑛𝜇 + 𝑛(𝑛 − 1)𝜇 ] + [𝑛𝜇 + 3𝑛(𝑛 − 1)𝜇 ].

( ) ( ) ( )
= 𝜇 + 𝜇 − 𝜇 − 𝜇 + 𝜇 + 𝜇 .

( ) ( ) ( )
= − + 𝜇 + − + 𝜇 .

( )
= 𝜇 + 1− + 𝜇 .

( ) ( )
= 𝜇 + 1− + 𝜇 + 𝜇 .

( ) ( ( )
= 𝜇 + . 𝜇 + 𝜇 .

Therefore, from equation (ii), we get

𝑉𝑎𝑟(𝑚 ) = 𝐸(𝑚 ) − 𝜎 .

( ) ( ) ( ) ( )
= 𝜇 + . 𝜇 + 𝜇 − 𝜇 .

DB THAPA (DMC) - 23
Sampling Distribution

( ) ( ) ( ) ( )
= 𝜇 + 𝜇 + 𝜇 − 𝜇 .

( ) ( ) ( ) ( )
= 𝜇 − 𝜇 + 𝜇 + 𝜇 .

( ) ( )
= (𝜇 − 𝑛𝜇 + (𝑛 − 1)𝜇 ) + 𝜇 .

( ) ( )
= (𝜇 − 𝜇 ) + 𝜇 .

∴ The variance of 𝑚 of order is

( )
𝑉𝑎𝑟(𝑚 ) = (𝜇 − 𝜇 ).

For the normal population, 𝜇 = 3𝜎 . Hence, for the normal population,

( )
𝑉𝑎𝑟(𝑚 ) = (3𝜎 − 𝜎 ).

( )
= . 2𝜎 .

Thus, the variance of sample variance 𝑚 of the order is

𝑉𝑎𝑟(𝑚 ) = 𝜎 .

Q. Show that the modified sampling variance (𝒔𝟐 ) is unbiased estimate of the population variance
for SRSWR. That means, 𝑬 𝒔𝟐 = 𝝈𝟐 . Also, find the variance of modified sample varince.

Proof: Let 𝑥 , 𝑥 , … , 𝑥 be a sample of size 𝑛 taken from a normal population 𝑋 , 𝑋 , , … , 𝑋 of size 𝑁 with
mean 𝜇 and variance 𝜎 . Then, the modified sample variance denoted by 𝑠 is given by

∑ ( ̅)
𝑠 = .

We want to show that 𝐸(𝑠 ) = 𝜎 and need to find the variance of modified sample variance (𝑠 ).

We know,

∑ ( ̅)
𝐸(𝑠 ) = 𝐸 .

= 𝐸[∑ (𝑥 − 𝑥̅ ) ].

= 𝐸[∑ {(𝑥 − 𝜇) − (𝑥̅ − 𝜇)}] .

DB THAPA (DMC) - 24
Sampling Distribution

= 𝐸[∑ {(𝑥 − 𝜇) − 2(𝑥 − 𝜇)(𝑥̅ − 𝜇) + (𝑥̅ − 𝜇) }].

= 𝐸[∑ (𝑥 − 𝜇) − 2(𝑥̅ − 𝜇) ∑ (𝑥 − 𝜇) + 𝑛(𝑥̅ − 𝜇) ].

= 𝐸[∑ (𝑥 − 𝜇) − 2(𝑥̅ − 𝜇)(∑ 𝑥 − 𝑛𝜇) + 𝑛(𝑥̅ − 𝜇) ].

= 𝐸[∑ (𝑥 − 𝜇) − 2𝑛(𝑥̅ − 𝜇) + 𝑛(𝑥̅ − 𝜇) ].

= 𝐸[∑ (𝑥 − 𝜇) − 𝑛(𝑥̅ − 𝜇) ].

= ∑( ) 𝐸(𝑥 − 𝜇) − 𝑛𝐸(𝑥̅ − 𝜇) .

= (𝑛𝜎 − 𝜎 ).

=𝜎 .

Therefore, 𝐸(𝑠 ) = 𝜎 . Hence, the modified sampling variance (𝑠 ) is unbiased estimate of the
population variance.

Also, the variance of the distribution of modified sample variance (𝑠 ) is obtained as follows:

∑ ( ̅) ∑ ( ̅)
We know, 𝑚 = and 𝑠 = . Therefore, we have a relation

𝑚 𝑛 = 𝑠 (𝑛 − 1).

∴𝑠 = 𝑚 .

We know,

( ) ( )
𝑉𝑎𝑟(𝑚 ) = (𝜇 − 𝜇 ) + 𝜇 .

Then,

𝑉𝑎𝑟(𝑠 ) = 𝑉𝑎𝑟 𝑚 .

= 𝑉𝑎𝑟(𝑚 ).

( ) ( )
=( )
× (𝜇 − 𝜇 ) + 𝜇 .

= + 𝜇 .
( )

DB THAPA (DMC) - 25
Sampling Distribution

= +2 − 𝜇 .

=( − )+ 𝜇 .

= + 𝜇 .

For normal population, 𝜇 = 3𝜎 . So, we have

𝑉𝑎𝑟(𝑠 ) = 𝜎 . [∵ 𝜇 = 𝜎 ].

Case b. Mean and Variance of the Sampling Distribution of Sample Variance under SRSWOR
(Sampling from a finite Population)

Q. Find the mean and variance of the sampling distribution of sample variance when the samples of
size 𝑛 are drawn without replacement from a normal population of size 𝑁 with mean 𝜇 and variance
𝜎 .

Solution: Let 𝑥 , 𝑥 , … , 𝑥 be a sample of size 𝑛 taken from a normal population 𝑋 , 𝑋 , , … , 𝑋 of size 𝑁


with mean 𝜇 and variance 𝜎 . Then, the sample variance denoted by 𝑚 is given by

∑ ( ̅)
𝑚 = .

We know, Mean of 𝑚 = 𝐸(𝑚 ) and 𝑉𝑎𝑟(𝑚 ) = 𝐸(𝑚 ) − [𝐸(𝑚 )] .

Now, Mean = 𝐸(𝑚 ).

∑ ( ̅)
=𝐸 .

= 𝐸[∑ (𝑥 − 𝑥̅ ) ].

= 𝐸[∑ [(𝑥 − 𝜇) − (𝑥̅ − 𝜇) ].

= 𝐸[∑ {(𝑥 − 𝜇) − 2(𝑥 − 𝜇)(𝑥̅ − 𝜇) + (𝑥̅ − 𝜇) }].

= 𝐸[∑ (𝑥 − 𝜇) − 2(𝑥̅ − 𝜇) ∑ (𝑥 − 𝜇) + (𝑥̅ − 𝜇) ∑ (1)].

= 𝐸[∑ (𝑥 − 𝜇) − 2(𝑥̅ − 𝜇)(∑ 𝑥 − 𝑛𝜇) + 𝑛(𝑥̅ − 𝜇) ].

= 𝐸[∑ (𝑥 − 𝜇) − 2(𝑥̅ − 𝜇)(𝑛𝑥̅ − 𝑛𝜇) + 𝑛(𝑥̅ − 𝜇) ].

= 𝐸[∑ (𝑥 − 𝜇) − 2𝑛(𝑥̅ − 𝜇) + 𝑛(𝑥̅ − 𝜇) ].

DB THAPA (DMC) - 26
Sampling Distribution

= 𝐸[∑ (𝑥 − 𝜇) − 𝑛(𝑥̅ − 𝜇) ].

= [∑ 𝐸(𝑥 − 𝜇) − 𝑛 𝐸(𝑥̅ − 𝜇) ].

= ∑ 𝐸(𝑥 − 𝜇) − 𝐸(𝑥̅ − 𝜇) .

= ∑ 𝜎 − 𝑉𝑎𝑟(𝑥̅ ) .

= × 𝑛𝜎 − .

=𝜎 − .

= 1− 𝜎 .
( )

( )
= 𝜎 .
( )

( )
Thus, 𝐸(𝑚 ) = , which is the mean of the distribution of sample variance. That is;
( )

( )
𝐸(𝑚 ) = 𝜎 .
( )

Case a. Mean of the sample variance in terms of population mean square:

The population mean square denoted by 𝑆 is defined as

∑ ( )
𝑆 = ,

then, (𝑁 − 1)𝑆 = 𝑁𝜎 .

So, we have

( ) ( )
𝐸(𝑚 ) = 𝜎 = × 𝑆 = 𝑆 .
( ) ( )

∴ 𝐸(𝑚 ) = 𝑆 .

Case b. Mean of the modified sample variance (Modified sample variance is unbiased estimate of
population mean square).

Proof: The modified sample variance is given by

∑ ( ̅)
𝑠 = .

DB THAPA (DMC) - 27
Sampling Distribution

The sample variance is given by

∑ ( ̅)
𝑚 = .

Therefore, we have

(𝑛 − 1)𝑠 = 𝑛𝑚 .

∴ 𝑠 = 𝑚 .

⇒ 𝐸(𝑠 ) = 𝐸 𝑚 .

= 𝐸(𝑚 ).

= × 𝑆 .

=𝑆 .

∴ 𝐸(𝑠 ) = 𝑆 .

Hence, modified sample variance 𝑠 is an unbiased estimator of population mean square 𝑆 in SRSWOR.

Case c. For the sufficiently large population, the mean of the sample variance in SRSWOR is 𝜎 ,
and modified sample variance is an unbiased estimator of the population variance.

Proof: We know, the mean of sample variance in SRSWOR is

( )
𝐸(𝑚 ) = .

If 𝑁 → ∞, then = = → 1. Therefore, we have

𝐸(𝑚 ) = 𝜎 , which is the mean of the sample variance in SRSWR.

So, for SRSWR, we have

𝐸(𝑠 ) = 𝐸 𝑚 = 𝐸(𝑚 ) = × 𝜎 =𝜎 .

Thus, modified sample variance is an unbiased estimator of the population variance in SRSWR.

Next, we find the variance of the sampling distribution of sample variance in SRSWOR. We know,

𝑉𝑎𝑟(𝑠 ) = 𝐸[(𝑠 ) ] − [𝐸(𝑠 )] .

DB THAPA (DMC) - 28
Sampling Distribution

= 𝐸[(𝑠 ) ] − (𝑆 ) , for SRSWOR

= 𝐸[(𝑠 ) ] − (𝜇 ) , for SRSWR.

Note: Although the sample mean 𝑥̅ is an unbiased estimator of the population mean 𝜇, the sample variance
∑ ( ̅)
𝑚 = is not an unbiased estimator of the population variance 𝜎 . For this reason, when 𝜎 is
unknown, 𝑚 can't be used for practical purposes. In such case, modified sample variance 𝑠 , which is an
unbiased estimator of population variation 𝜎 can be used for practical purposes. Therefore, 𝑠 plays very
vital role in sampling theory. Moreover, the sample standard deviation 𝑠 is not an unbiased estimator of the
population standard deviation 𝜎.

Sampling Distribution of Sample proportion:

Q. Define sample and population proportion. Find the sample mean, population mean, sample
variance, population variance, population mean square and modified sample variance in-terms of
sample proportion and population proportion. Also, find mean and variance of the distribution of
sample proportion.

Solution: The sampling distribution of the sample proportion is the probability distribution of all possible
values of the sample proportion (𝑝) obtained from all possible samples of a fixed size (𝑛) drawn from a
population. In simple words, if we repeatedly take many samples of size 𝑛 and calculate the proportion of
successes each time, then the distribution of those values is called the sampling distribution of 𝑝.

Let us consider a population (sample space of an experiment) consisting of 𝑁 units 𝑋 , 𝑋 , , … , 𝑋 with


mean 𝜇 and variance 𝜎 . Out of these, suppose 𝑁 outcomes are successes and 𝑁 outcomes are failures.
Then, the population proportion of success (𝑃) is 𝑃 = , and population proportion of failure is 𝑄 = .
Suppose a sample 𝑥 , 𝑥 , … , 𝑥 of size 𝑛 is observed and found 𝑛 outcomes to be successes and 𝑛 as
failures. Then, sample proportion of success is 𝑝 = and sample proportion of failure is 𝑞 = . Then,
we have 𝑃 + 𝑄 = 1 and 𝑝 + 𝑞 = 1 also.

Again, suppose 𝑋 be the ith unit of the population and 𝑥 be the ith unit of the sample. For the statistical
analyses, we code the population and sample units in such a way that

1, 𝑖𝑓 𝑖𝑡 𝑖𝑠 𝑎 𝑠𝑢𝑐𝑐𝑒𝑠𝑠
𝑋 = ,
0, 𝑖𝑓 𝑖𝑡 𝑖𝑠 𝑎 𝑓𝑎𝑖𝑙𝑢𝑟𝑒

and,

1, 𝑖𝑓 𝑖𝑡 𝑖𝑠 𝑎 𝑠𝑢𝑐𝑐𝑒𝑠𝑠
𝑥 = .
0, 𝑖𝑓 𝑖𝑡 𝑖𝑠 𝑎 𝑓𝑎𝑖𝑙𝑢𝑟𝑒

Then it is clear that

∑ 𝑋 =𝑁 =∑ 𝑋 ,

and, ∑ 𝑥 =𝑛 =∑ 𝑥 .

DB THAPA (DMC) - 29
Sampling Distribution

Now, the population mean and the sample mean are:


Population mean (𝜇) = = = 𝑃, and


Sample Mean (𝑥̅ ) = = = 𝑝.

Thus, the population mean and sample mean in-terms of proportions are 𝜇 = 𝑃 and 𝑥̅ = 𝑝.

Also, population variance and sample variance are given by:

∑ ( )
Population Variance (𝜎 ) =


= .

= ∑ 𝑋 − 2𝜇 ∑ 𝑋 + 𝑛𝜇 .

= ∑ 𝑋 − 2𝑁𝜇 + 𝑁𝜇 .

= ∑ 𝑋 − 𝑁𝜇 .

= (𝑁 − 𝑁𝜇 ).

= −𝜇 .

=𝑃−𝑃 .

= 𝑃(1 − 𝑃).

= 𝑃𝑄.

∑ ( ̅)
Sample variance (𝑚 ) = .

= ∑ 𝑥 − 2 𝑥̅ ∑ 𝑥 + 𝑛𝑥̅ .

= ∑ 𝑥 − 𝑛 𝑥̅ .

= (𝑛 − 𝑛𝑝 ).

=𝑝−𝑝 .

= 𝑝(1 − 𝑝).

DB THAPA (DMC) - 30
Sampling Distribution

= 𝑝𝑞.

Thus, the population variance and sample variance in-terms of proportions are 𝜎 = 𝑃𝑄 and 𝑚 = 𝑝𝑞.

Similarly, the population mean square (𝑆 ) for the proportion is

∑ ( )
Population mean square (𝑆 ) = .


= .

= ∑ 𝑋 − 𝑁𝜇 .

= [𝑁 − 𝑁𝑃 ].

= (𝑁𝑃 − 𝑁𝑃 ).

( )
= .

= .

And, modified sample variance (𝑠 ) is given by:

∑ ( ̅)
Modified Sample Variance (𝑠 ) = .

= ∑ 𝑥 − 𝑛𝑥̅ .

= [𝑛 − 𝑛𝑥̅ ].

= − 𝑥̅ .

= (𝑝 − 𝑝 ).

= 𝑝(1 − 𝑝).

= 𝑝𝑞.

Thus, population mean square and modified sample variance in-terms of proportions are

𝑆 = and 𝑠 = 𝑝𝑞.

Now, mean and variance of the sampling distribution of sample proportion are derived as follows:

DB THAPA (DMC) - 31
Sampling Distribution

Mean of sample proportion 𝐸(𝑝) = 𝐸 .


=𝐸 .

= ∑ 𝐸(𝑥 ).

= ∑( ) ∑ 𝑋 . .

= ∑ 𝑃.

= 𝑃.

Therefore, 𝐸(𝑝) = 𝑃.

Also, Variance of sample proportion 𝑉𝑎𝑟(𝑝) = 𝑉𝑎𝑟 .

= 𝑉𝑎𝑟(𝑛 ).

= 𝑉𝑎𝑟(∑ 𝑥 ).

= ∑ 𝑉𝑎𝑟(𝑥 ).

= ∑ 𝜎 .

= ∑ 𝑃𝑄.

= × 𝑛𝑃𝑄.

= .

Therefore, 𝑉𝑎𝑟(𝑝) = .

Hence, mean and variance of the sampling distribution of sample proportion are 𝑃 and . Thus, the
sampling proportion 𝑝 follows binomial distribution with mean 𝑃 and variance .

Moment Generating Function of Sample Proportion:

Q. Find moment generating function of sample proportion and hence find mean and variance of the
sampling distribution of sample proportion.

DB THAPA (DMC) - 32
Sampling Distribution

Solution: Let a Bernoulli experiment results into 𝑥 , 𝑥 , … , 𝑥 outcomes with 𝑝 as the probability of success
and (1 − 𝑝) as the probability of failure. That is 𝑥 , 𝑥 , … , 𝑥 ~Bernoulli (𝑝). Let us assign

𝑥 = 1 (success) with probability 𝑝, and

𝑥 = 0 (failure) with probability (𝑞). If, out of 𝑛 trials, 𝑛 outcomes are success
and 𝑛 outcomes are failures, then 𝑛 + 𝑛 = 𝑛. Also,

∑ 𝑥 = 𝑛 if 𝑥 is a success with probability 𝑝 and

∑ 𝑥 = 𝑛 , if 𝑥 is a failure with probability 𝑞.

Now,


proportion of success (𝑝) = = = 𝑥̅ , and

proportion of failure (𝑞) = = =1− = 1 − 𝑝.

Therefore, 𝑝 + 𝑞 = 1.

We define the sample proportion 𝑃 as


𝑝= .

Now, moment generating function of 𝑝 is given by

𝑀 (𝑡) = 𝐸(𝑒 ).


.
=𝐸 𝑒 .

.∑
=𝐸 𝑒 .

= 𝑀∑ .

=∏ 𝑀 .

=∑ (1 − 𝑃)𝑒 + 𝑃𝑒 .

= 1 − 𝑃 + 𝑃𝑒 . [∵ Since variables are independent]

DB THAPA (DMC) - 33
Sampling Distribution

= 𝑄 + 𝑃𝑒 (For small sample)

=𝑒 , (For large Sample)

Therefore, the moment generating function of sample proportion 𝑝~𝑁 𝑃, is

𝑀 (𝑡) = 𝑄 + 𝑃𝑒 .

Now,

( )
𝐸(𝑝) = .

= 𝑒 .

= 𝑃+ 𝑒 .

= 𝑃.

And,

( )
𝐸(𝑝 ) = .

= 𝑃+ 𝑒 .

= + (𝑃 + 𝑒 .

= +𝑃 .

Then, 𝑉𝑎𝑟(𝑝) = [𝐸(𝑝 ) − [𝐸(𝑝)] = +𝑃 −𝑃 = .

Thus, the sampling proportion 𝑝 follows binomial distribution with mean 𝑃 and variance .

Attribute/Categorical variables (Def): The variables which are not measured quantitatively are called
categorical variables. For example, sex, habit, honesty, intelligence, beauty etc. are categorical data.

Dichotomous/Manifold Classification (Def): It is necessary to measure such attributes quantitatively for


statistical analysis. So, in order to measure the attributes, the concerned population should be classified into
different classes. If the population is divided into two mutually disjoint classes, it is called a dichotomous
classification and if it is classified into more than two disjoint classes, it is called manifold classification.

DB THAPA (DMC) - 34
Sampling Distribution

Simple Random Sampling with and without replacement:

Case a. Mean and Variance of the Sampling Distribution of Sample Proportion under SRSWR
(Sampling from an infinite Population)

Q. Find the mean and variance of the Sampling Distribution of Sample Proportion under the
sampling with replacement and without replacement from a population containing 𝑁 units whose
mean is 𝜇 and variance 𝜎 .

Solution: Let us consider a population consisting of 𝑁 units 𝑋 , 𝑋 , , … , 𝑋 with mean 𝜇 and variance 𝜎 .
Out of these, suppose 𝑁 outcomes are successes and 𝑁 outcomes are failures. Then, the population
proportion of success (𝑃) is 𝑃 = , and population proportion of failure is 𝑄 = . Suppose a sample
𝑥 , 𝑥 , … , 𝑥 of size 𝑛 is selected from the population and found that 𝑛 outcomes are successes and 𝑛 are
failures. Then, sample proportion of success is 𝑝 = and sample proportion of failure is 𝑞 = .

Then, mean of the sample proportion is


𝐸(𝑝) = 𝐸 .

= ∑ 𝐸(𝑥 ).

= ∑ ∑ 𝑋. . ∵ 𝑥 𝑎𝑟𝑒 𝑡𝑎𝑘𝑒𝑛 𝑓𝑟𝑜𝑚 𝑋 𝑤𝑖𝑡ℎ 𝑝𝑟𝑜𝑏𝑎𝑏𝑖𝑙𝑖𝑡𝑦 .

= ∑ 𝑃.

= × 𝑛𝑃.

= 𝑃.

∴ 𝐸(𝑝) = 𝑃. This shows that sample proportion is an unbiased estimator of the population proportion.

Also, variance of sample proportion in SRSWOR 𝑉𝑎𝑟(𝑝) = 𝑆 .

= × .

= .

Thus, variance of 𝑝 in SRSWOR is 𝑉𝑎𝑟(𝑝) = .

DB THAPA (DMC) - 35
Sampling Distribution

Again, for the sampling with replacement (i.e. for SRSWR), the sampling is considered to be done from an
infinite population. When the population size 𝑁 is sufficiently large, the ratio , called finite population
correction (FPC) would be 1,

lim = lim = lim = 1.


→ → →

Thus, variance of 𝑝 in SRSWOR is, 𝑉𝑎𝑟(𝑝) = .

Standard Error of Statistics:

Q. What do you mean by standard error of statistics? Discuss its importance in Statistics.

Solution: The standard error (SE) of a statistic is defined as the standard deviation of its sampling
distribution. In simple words, it measures how much a statistic (like sample mean or sample variance,
sample proportion) varies from sample to sample.

𝑆𝐸(𝑠𝑡𝑎𝑡𝑖𝑠𝑡𝑖𝑐) = 𝑉𝑎𝑟𝑖𝑎𝑛𝑐𝑒 𝑜𝑓 𝑡ℎ𝑒 𝑠𝑎𝑚𝑝𝑙𝑖𝑛𝑔 𝑑𝑖𝑠𝑡𝑟𝑖𝑏𝑢𝑡𝑖𝑜𝑛.

Standard error plays very important role in statistics. It measures the reliability of sampling due to chance.
It gives index of the reliability or precision of the estimate of the parameter. Greater is the SE, greater is the
deviation of the actual values from the expected ones. Hence, smaller the value of the SE, greater is the
reliability or precision of the estimate (sample statistics). Thus, the measure of reliability is

Reliability coefficient = .
( )

𝝈
Q. Prove that standard error of sample mean 𝑺𝑬(𝒙) = . Also, express standard error of sample
√𝒏
mean in terms of modified sample sd (𝒔) and sample sd (𝒎𝟐 ).

Solution: Let a random sample 𝑥 , 𝑥 , … , 𝑥 of size 𝑛 be drawn from a normal population with mean 𝜇 and
variance 𝜎 . Then, sample mean (𝑥̅ ) is given by


𝑥̅ = ,

and, variance of the sample mean 𝑉𝑎𝑟(𝑥̅ ) is

𝑉𝑎𝑟(𝑥̅ ) = .

Therefore, SE(𝑥̅ ) = 𝑉𝑎𝑟(𝑥̅ ) = . Thus, SE of sample mean 𝑥̅ is inversely proportional to the square

root of the sample size.

If the population variance (𝜎 ) is unknown, we use modified sample variance (𝑠 ) because it is an unbiased
estimator of the population variance, that is;

DB THAPA (DMC) - 36
Sampling Distribution

𝐸(𝑠 ) = 𝜎 .

So, using 𝑠 in place of population variance 𝜎 , SE(𝑥̅ ) becomes

SE(𝑥̅ ) = .

Further, we know,

𝑛𝑚 = (𝑛 − 1)𝑠 .

∴ 𝑠= 𝑚 .

Hence, we have

SE(𝑥̅ ) = 𝑚 .

𝑷𝑸
Q. Prove that standard error of sample proportion is 𝑺𝑬(𝒑) = .
𝒏

Solution: Let a sample of 𝑛 trials consists of 𝑥 number of successes. If the sample is taken from a population
with 𝑋 number of successes and (𝑁 − 𝑋) failures, then, the probability distribution of 𝑋 follows Binomial
Distribution with mean 𝜇 = 𝑃 and variance 𝜎 = 𝑃𝑄. We know,

𝑉𝑎𝑟(𝑝) = .

∴ SE(𝑝) = 𝑉𝑎𝑟(𝑝) = .

Therefore, variance of sample proportion is .

If the population proportion (𝑃) of success is unknown, we use sample proportion (𝑝) because it is an
unbiased estimator of the population proportion (𝑃), that is;

𝐸(𝑝) = 𝑃.

So, using sample proportion (𝑝) in place of population proportions 𝑃 and 𝑄, SE(𝑝) becomes

SE(𝑝) = . [∵ the denominator is (𝑛 − 1) because 1 d.f. is lost in estimating


the unknown population proportion (𝑃).

Utility of Standard Error in Testing of hypothesis: -


a. It is used to determine the precision of the sample estimate of population parameter. The precision
of sampling distribution of the estimate is reciprocal of the standard error of the estimate.

DB THAPA (DMC) - 37
Sampling Distribution

b. i.e., Precision of t =
. .( )
c. The larger the S.E., the less precise (efficient) is the estimate.
d. It is used to test whether the sample statistic differs significantly from the corresponding
hypothetical value in the population.
e. It is used to test the significance of the difference between two independent sample estimates of the
same population parameter.
f. It is used for point estimation of the population parameter.
g. It is used for the interval estimation of the population parameter.

Merits of simple random sampling:


 It is scientific method since there is little possibility of personal bias affecting the results because
the item in the sample depends entirely on chance.
 As the size of the sample increases, it becomes more and more representative of the population.
 This method provides the maximum number of most reliable information at the least possible cost
and thus, saves time, money and labor.
 As the standard error of sampling distribution can easily be calculated, the efficiency of estimates
is easy to declare.

Demerits of Simple Random Sampling:


 Simple random sampling requires up to date list of population which is not often available in
practice.
 If the sample size is not sufficiently large, the sample may not be true representative of the
population.
 If the population is infinitely large, selecting samples by lottery method is uneconomic.
 If the population is heterogeneous in nature, the results obtained from simple random sampling may
not be accurate. For this, stratified random sampling is preferred.

Disadvantages of random number table:


 This method needs a complete update list of units of the population which may not be available in
practice.
 The sample will not be true representative of the universe if its size is small.
 The numbering of population units and the preparation of slips is quite time consuming and a bit
expensive if the population is too large.

Example: - A population consists of five numbers 1, 3, 5, 7 and 9. Enumerate all possible samples of size
2 drawn from the population without replacement. Find the mean and variance of the population. Find the
mean of the sampling distribution of means and show that it is equal to the population mean. Also, find the
variance of sampling distribution of means and hence verify it.
Solution: -
Here, N = 5 and n = 2
So, possible number of samples of size 2 drawn without replacement = C (N, n) = C(5,2) = = 10
Then the possible samples are: - (1,3), (1,5), (1,7), (1,9), (3,5), (3,7), (3,9), (5,7), (5,9), (7,9)
Calculation of population mean and variance:

X X–𝜇 (𝑋 − 𝜇)
1 -4 16
3 -2 4
5 0 0
7 2 4

DB THAPA (DMC) - 38
Sampling Distribution

9 4 16
∑ 𝑋 =25 ∑(𝑋 − 𝜇) =
40


Here, Population mean (𝜇) = = =5
∑( )
And population variance (𝜎 ) = = =8
Calculation of mean and variance of sampling distribution of means:
Sample Sample Sample 𝑥̅ -𝑥̿ (𝑥̅ -
No. mean(𝑥̅ ) 𝑥̿ )2
1 (1,3) 2 -3 9
2 (1,5) 3 -2 4
3 (1,7) 4 -1 1
4 (1,9) 5 0 0
5 (3,5) 4 -1 1
6 (3,7) 5 0 0
7 (3,9) 6 1 1
8 (5,7) 6 1 1
9 (5,9) 7 2 4
10 (7,9) 8 3 9
Total = 50 =30
∑ ̅
Now, mean of sample means (𝑥̿ ) = = =5
( , )
Since the population mean is also found to be 5, the mean of sampling distribution of means is equal to the
population mean.
And, variance of sample means by using definition is,
∑( ̅ ̿)
Var(𝑥̅ ) = = =3
( , )
Also, the variance of sample mean by using formula is,
Var(𝑥̅ ) = . = . =3
Hence the formula for computing variance of sample mean is verified.
Standard error of means, s. e.(𝑥̅ ) = 𝑉𝑎𝑟(𝑥̅ ) = √3 = 1.732
Example: - A population consists of four numbers 2, 5, 8 and 1. Enumerate all possible samples of size two
which can be drawn from this population with replacement. Find the mean and variance of the population.
Find mean of the sampling distribution of means and show that it is equal to the population mean. Find the
variance of the sampling distribution of means and verify it. Find the standard error of mean.
Solution: Here,
No. of possible samples of size two drawn with replacement from the population = 𝑁 = 4 = 16
And the possible samples are: (1,2), (1,5), (1,8), (2,1), (2,5), (2,8), (5,1), (5,2), (5,8), (8,1), (8,2), (8,5) and
(8,5).

Population mean (𝜇) = = =4
∑( ) ( ) ( ) ( ) ( )
Population variance (𝜎 ) = = = 7.5.
Calculation of mean and variance of the sampling distribution of means:

Sample Sample Sample mean(𝑥̅ ) (𝑥̅ − 𝑥̿ ) (𝑥̅ − 𝑥̿ )


No.
1 (1,1) 1 -3 9
2 (1,2) 1.5 -2.5 6.25

DB THAPA (DMC) - 39
Sampling Distribution

3 (1,5) 3 -1 1
4 (1,8) 4.5 o.5 0.25
5 (2,1) 1.5 -2.5 6.25
6 (2,2) 2 -2 4
7 (2,5) 3.5 -0.5 0.25
8 (2,8) 5 1 1
9 (5,1) 3 -1 1
10 (5,2) 3.5 -0.5 0.25
11 (5,5) 5 1 1
12 (5,8) 6.5 2.5 6.25
13 (8,1) 4.5 0.5 0.25
14 (8,2) 5 1 1
15 (8,5) 6.5 2.5 6.25
16 (8,8) 8 4 16
=64 (𝑥̅ − 𝑥̿ ) = 60

∑ ̅
Now, mean of sample means (𝑥̿ ) = = = 4. This is equal to population mean. Hence, mean of the
distribution of sample mean is equal to the population mean.
Also, variance of sample mean by the definition is,
∑( ̅ ̿)
Var(𝑥̅ ) = = = 3.75
And variance of the sample mean by using formula is,
.
var(𝑥̅ ) = = = 3.75
Hence the formula for variance of the distribution of sample mean is verified.
Standard error of mean is,
s.e. (𝑥̅ ) = 𝑉𝑎𝑟(𝑥̅ ) = √3.75 = 1.94
Exercise for practice
1. A sample of size 25 is drawn from a population consisting of 150 units. If the population [Link] 10,
find the standard error of sample mean when the sample is drawn (i) without replacement (ii) with
replacement.
2. A simple random sample of size 20 is drawn without replacement from a finite population of 75
units. If the number of defective units in the population is 12, find the standard error of the sample
proportion.
3. Consider a population of four units 3,6,2,1. a) Write down all possible samples of size 2 that can
be drawn with replacement from the sample. b) Find mean and variance of the population. c) Find
mean of the sampling distribution of means and show that it is equal to the population mean. d)
Find the variance of the sampling distribution of means and verify that it agrees with the formula.
e) Find the standard error of mean.
4. How does sampling with replacement differ from that without replacement? Which of them gives
lower value of S.D. of the sample mean? Explain by considering sample of size 2 from a population
consisting five members 2,3,6,8 and 11. Verify that sample mean is an unbiased estimator of the
population means and that its variance is given by (1-f), where the symbols have their usual
meanings.

DB THAPA (DMC) - 40

You might also like