Understanding Sampling Distributions
Understanding Sampling Distributions
9.1 Introduction
9.2 Objectives
9.3 Basics of Sampling
9.4 Sampling Distribution
9.4.1 Standard Error
9.4.2 Central Limit Theorem
9.1 INTRODUCTION
In general, extracting information from all the elements or items of a large
population may be time consuming and expensive if the population size is
infinitely large. But even then sometimes there are so many problems
attached to large size populations where it becomes necessary to draw
inference about population [Link] example, one may wish to
estimate the average height of all the two thousand students in a college, a
businessman may be interested to estimate the proportion of defective items
in a production line, a manufacturer of car tyres may want to estimate the
variations took place in the diameter of produced tyres, a pharmacist may
want to estimate the difference of effect of two types of drugs, [Link] all such
cases, there is always an unknown population involved whose characteristics
are described through some parameters.
Due to large sizes of populations, in which we may be interested, for
drawing inferences about the population parameters, generally we draw a
224 sample and determine a function of sample values, which is called statistic.
Selection of a sample, that is, only a part of the population saves a lot of time, Sampling Distributions
money and labour and results drawn on sample values are quicklyavailable
forinterpretation and as good sometimes as obtained on the basis of the entire
population. The process of generalising sample results to the population is
called Statistical Inference. Sincethere might be a large number of samples of
same size drawn from the population, the values of statistic generally vary
from sample to sample and is associated with the probability of selection of
the particular sample. Therefore, the sample statistic is a random variable
following any probability distribution. Therefore, it may be a matter of
interest for a statistician to know what distribution the statistic follows if the
samples are assumed to be selected from a theoretical distribution like, a
normal distribution with given mean and variance; a binomial distribution; a
gamma distribution; a Poisson distribution and so on. In contrast to
theoretical distributions, probability distribution of a statistic in popularly
called a sampling distribution. In this unit we shall discuss the sampling
distribution of sample mean; of sample median; of sample proportion; of
difference between two sample means and sample proportions.
Due to this curiosity, Prof. R.A. Fisher, Prof. G. Snedecor and some other
statisticians worked in this area and obtained exact sampling distributions
which are followed by some of the important statistics. In present unit of this
block, we shall also discuss some important sampling distributions such as 2
(read as chi-square), t and F. Generally, these sampling distributions are
named on the name of the originator, for instance, Fisher’s F-distribution is
named like this after the name of its inventor Prof. R.A. Fisher.
9.2 OBJECTIVES
After studying this unit, you should be able to:
explain the concept of sampling and sampling distribution;
explain the concept of standard error and Central Limit Theorem;
define the sampling distribution; and
describe the sampling distribution of sample mean and difference of
two sample means.
As we have discussed in Section 9.1, in many practical and real situations the
population under consideration is either infinitely large in size or it is
unbounded in the sense that its boundaries are not well-defined or it is of
destructive nature. As for example, in ordert to determine average life of two
hundred produced electric bulbs, it would be necessary to light these bulbs
unless they all get fused, resulting destruction of the entire lot of production.
Similarly, in order to estimate the proportion of smokers in a large city,
information have to be gathered from each and every person dwelling in the
city which would be a very difficult task in terms of manpower, money and
time to be required. Sometimes, the population size is not known if its
boundaries are not well-defined, like number of fish in a pond or lake. In all
such cases, it is not feasible to gather information on all the units, rather it
becomes almost impossible. Keeping in view the difficulty in contacting each
and every unit of the population due to these reasons and also to save the
time, money and manpower to be required for this, generally a part of the
population, popularly known as a “sample” is selected in some pre-asigned
manner which in turn is used for drawing the inferences on the population
itself or on the parameters of the population. The results obtained from the
sample are projected in such a way that they are valid for concluding about
the entire population. Therefore, the sample works like a “Vehicle” to reach
(drawing) at valid conclusions about the population. In fact, a sample helps
us to reach to the “whole”(population) from a “part” (sample) in all types of
statistical studies. Thus, statistical inference can be defined as:
“It is a process of concluding (projecting or infering) something desired about
a given population on the basis of sample results”.
Parameter and Statistic
A parameter is a function of population values which is used to represent
certain characteristics of the [Link] example, population total,
population mean, population variance, population coefficient of variation,
population proportion, population correlation coefficient, etc., are all
parameters since their calculation involve all the population values.
A statistic is a function of sample values only and does not contain any
unknown population parameter in it. For example, if X1 ,X2 , ..., Xn represent
values of the variable X in a random sample of size n taken from a population
1 n
then sample mean X X i is a statistic.
n i 1
of the statistic ̂ k and the probability of its appearance due to the kth
sample; k = 1, 2,…, nCN.
Generally, in practice onlya single random sample is taken from a given
population and its mean X is considered to be representative of the
population mean µ. This sample mean may or may not represent the
population mean. Since we cannot determine the proximity of sample mean
and population mean on the basis of single random sample, we can use the
concept of sampling distribution to bring the value of sample mean close to
that of Population mean. Being a random variable, the statistic must possess a
probability distribution which may be used to answer some questions
regarding the nature of the statistic, for example, what is the probability that
its value as obtained from the given sample is differing from the parameter
value by amargin of 10? The probability distribution of a statistic is called
sampling distribution of the statistic. Let us illustrate how the sampling
distribution of a statistic can be obtained given a population. For instance, let
we wish to estimate the population mean using the sample mean as an
estimator.
Suppose that a baby-sitter has 5 children under her supervision. The ages of
children are 2, 4, 6, 8 and 10 years. Supposing this group of children as a
population of size 5, we get the population mean as
1 N 2 4 6 8 10
X
N i 1
Xi
5
6
228
Therefore, the variance of this population is given by: Sampling Distributions
2 2 2
1 N
X i X
2 2 2 6 4 6 ....... 10 6 8
N i 1 5
Now, let us take all the possible simple random samples of size 2 without
replacement from this population. There are 5C2 = 10 such possible samples
which are listed below along with their respective means:
Sample
Sample Sample Mean
No.
1 2, 4 3
2 2, 6 4
3 2, 8 5
4 2, 10 6
5 4, 6 5
6 4, 8 6
7 4, 10 7
8 6, 8 7
9 6, 10 8
10 8, 10 9
Now, suppose that our selected sample is either (2, 4) with mean 3 or it is the
sample (8, 10) with mean 9. In both cases, the sample is not a good
representative of the population since sample mean is far from the population
mean 6. But if the selected samples are coincidently either (2, 10) or (4, 8),
then these are good representatives of the population in the sense that both
have sample means exactly equal to population mean. Thus, this example
illustrates that a single random sample may or may not be representative for
the decision maker to reach meaningful conclusion. However, the grand
mean of the distribution of these ten sample means, is calculated and
observed to be equal to the population mean as follows:
Hence, the mean of the sample means can be considered to represent the
population mean for analysis and decition making purposes.
Now let us put sample means along with their probabilities of occurrence as
follows:
229
Statistical Analysis Table 9.2: Probability Distribution
Probability
Sample
Frequency of
Mean
Occurrence
3 1 0.1
4 1 0.1
5 2 0.2
6 2 0.2
7 2 0.2
8 1 0.1
9 1 0.1
Total 10 1.0
This distribution which shows the distribution of probabilities over all the
possible values of the sample mean is refered to as sampling distribution of
the sample mean. Symbolically, it can be denoted as {( x , p( x )}. As
explained above, the sampling distribution of a statistic can be defined as:
“The probability distribution of all possible values of a statistic that would be
obtainedby drawing all possible samples of the same size from the population
is called sampling distribution of that statistic.”
The mean, variance and other measures of a sampling distribution can be
obtained in a similar way as computed in a frequency distribution, taking
probabilities as frequencies. Thus, mean of the above sampling distribution
will be
7
x x i p x i 3 x 0 . 1 4 x 0 . 1 5 x 0 . 2 6 x 0 . 2 7 x 0 . 1 8 x 0 . 1 9 x 0 . 1 6
i 1
This value is same as the population meanµ. The variance of the distribution
is obtained as:
2 2 2
1 N
x i x
2 2 3 6 4 6 ....... 9 6 28 4
x
7 i 1 7 7
SE X ,
n
σ12 σ 22
SE X Y =
n1 n 2
where, σ12 and σ 22 are the population variances of two different populations
and n1 and n2 are the sample sizes of two independent samples selected from
the two populations respectively.
P1Q1 PQ
SE( p1 p 2 ) = 2 2
n1 n2
From all the above formulae, we can understand that standard error is
inversely proportional to the sample size. Therefore, as sample size increases
the standard error decreases.
The standard error is used to express the accuracy or precision of the estimate
of population parameter because the reciprocal of the standard error is the
measure of reliability or precision of the [Link] error also
determines the probable limits or confidence limitswithin which the
231
Statistical Analysis population parameter may be expected to lie with certain level of
[Link] error is also applicable in testing of hypothesis.
9.4.2 Central Limit Theorem
The central limit theorem is the most important theorem of Statistics. It was
first introduced by De Movers in the early eighteenth century. The theorem
states that regardless of the nature of the distribution of the population, the
distribution of the sample mean approaches the Normal Probability
distribution as the sample size increases. In general, the larger the sample
size, the closer the proximity of the distribution of sample mean to the
Normal distribution. However, in practice, sample sizes of 30 or larger are
considered adequate for this purpose. It should be noted, however, that the
sampling distribution of sample mean would be always normally distributed
if the original population is normally distributed.
2
X ~ N ,
n
X
Z ~ N 0,1
/ n
follows normal distribution with mean 0 and variance unity, that is, the
variate Z follows standard normal distribution.
X X2 ... X n
Mean of X E X E 1 By defination of X
n
1
[E(X1 ) E(X 2 ) ... E(X n )]
n
Since E(Xi) = mean of Xi for all i=1, 2,…,n; therefore, we have E(Xi) = µ.
Similarly, Var(Xi) = σ2 for all i. We, therefore, have
1
E X ...... (n times) n
n n
and variance
1
Var(X) Var (X1 X 2 ... X n )
n
1
Var(X1 ) Var(X 2 ) ... Var(X n )
n2
1 2 n 2 2
n
2 2 2
2 ...... (n times) 2
n n
We, therefore, conclude that If X i ~ N , 2 then
2
X ~ N ,
n
and
SE X SD X Var X
n
Let us illustrate these results using some examples.
Example 1: Diameter of a steel ball bearing produced on a semi-automatic
machine is known to be distributed normally with mean 12 cm and standard
deviation 0.1 cm. If we take a random sample of 10 ball bearings then find
mean and variance of sampling distribution of mean.
Solution: Here, we are given that
= 12, σ = 0.1, n = 10
Since the sample is taken from the population of ball bearings in which
diameter follows normal distribution N(12, 0.01), we have 233
Statistical Analysis E X 12
2 (0.1) 2
Var X 0.001 .
n 10
2
54 602 56 602 ....... 66 60 2
104
20.08
5 5
Now, all possible samples (without replacement) are 10 which are shown in
the Table given below:
Table 9.3: Possible Samples
1 54, 56 55
2 54, 60 57
3 54, 64 59
4 54, 66 60
5 56, 60 58
6 56, 64 60
7 56, 66 61
8 60, 64 62
9 60, 66 63
10 64, 66 65
234
Table 9.4: Probability Distribution Sampling Distributions
Probability
Sample
Frequency of
Mean
Occurrence
55 1 0.1
57 1 0.1
58 1 0.1
59 1 0.1
60 2 0.2
61 1 0.1
62 1 0.1
63 1 0.1
65 1 0.1
Total 10 1.0
This value is same as the population meanµ. The variance of the distribution
is obtained as:
2 2 2
1 N
x i x
2 2 55 60 57 60 ....... 65 60 78 8.66
x
9 i 1 9 9
Note: i) Check your answers with those given at the end of the unit.
E X Y E X E Y =µ1− µ2
236
and variance Sampling Distributions
12 22
Var X Y Var X Var Y
n1 n 2
12 22
SE X Y Var X Y .
n1 n 2
and variance
2 2
12 22 200 100
Var(X Y) 400
n1 n 2 125 125
SE X Y Var X Y 400 20 .
If population variances 12 and 22 are unknown then we estimate 12 and 22
by the values of the sample variances of the samples taken from the first and
second population respectively. For large sample sizes n1 and n2 30 , the
sampling distribution of (X Y) is very closely normally distributed with
s2 s2
mean (µ1− µ2) and variance 1 2 .
n1 n 2 237
Statistical Analysis If population variances 12 and 22 are unknown and 12 22 2 then σ2 is
estimated by pooled sample variance s 2p where,
1
s 2p
n1 n 2 2
n 1s12 n 2 s 22
and variate
t
X Y 2
1
~ t ( n1 n 2 2 )
1 1
sp
n1 n 2
Note: i) Check your answers with those given at the end of the unit.
Carton Number of
Defectives Bulbs
A 2
B 4
C 1
D 3
From the above table, we can see that value of the sample proportion is
varying from sample to sample. So we consider all possible sample
proportions and calculate their probability of occurrence. Since there are 6
possible samples therefore the probability of selecting a sample is 1/8. Then
we arrange the possible sample proportions with their respective probability
in Table 2.7:
239
Statistical Analysis Table 9.7: Sampling Distribution of Sample Proportion
1 3 4 5 6 7 30 1 1
1 1 2 1 ... 1
6 40 40 40 40 40 40 6 8
Thus, we have seen that mean of sample proportion is equal to the population
proportion.
If a population whose elements are divided into two mutually exclusive
groups− one containing the elements which possess a certain attribute and
other containing elements which do not possess the attribute, then number of
successes (elements possess a certain attribute) follows a binomial
distribution with mean
E(X) nP
and variance
Var(X) nPQ where Q 1 P
where, P is the probability or proportion of success in the population.
Now, we can easily find the mean and variance of the sampling distribution
of sample proportion by using the above expression as
X 1 1
E(p) E E(X) nP P
n n n
and variance
X 1 Var aX a 2Var X
Var(p) Var 2 Var(X)
n n
1 PQ
nPQ . Var X nPQ
n 2
n
240
Also standard error of sample proportion can be obtained as Sampling Distributions
PQ
SEp Var( p)
n
If the sampling is done without replacement from a finite population then the
mean and variance of sample proportion is given by
E p P
and variance
N n PQ
Var p
N 1 n
where, N is the population size and the factor (N-n) / (N-1) is called finite
population correction.
If sample size is sufficiently large, such that np > 5 and nq > 5 then by central
limit theorem, the sampling distribution of sample proportion p is
approximately normally distributed with mean P and variance PQ/n where,
Q= 1 P.
Let us see an application of the sampling distribution proportion with the help
of an example.
Example 4: A machine produces a large number of items of which 15% are
found to be defective. If a random sample of 200 items is taken from the
population and sample proportion is calculated then find mean and standard
error of sampling distribution of proportion.
Solution: Here, we are given that
15
P= = 0.15, n = 200
100
We know that when sample size is sufficiently large, such that np > 5 and nq
> 5 then sample proportion p is approximately normally distributed with
mean P and variance PQ/n where, Q = 1– P. But here the sample proportion
is not given so we assume that the conditions of normality hold, that is, np >
5 and nq > 5. So mean of sampling distribution of sample proportion is given
by
E ( p ) P 0.15
and variance
PQ 0.15 0.85
Var(p) 0.0006
n 200
Therefore, the standard error is given by
241
Statistical Analysis CHECK YOUR PROGRESS 3
Note: i) Check your answers with those given at the end of the unit.
PQ PQ
p1 ~ N P1, 1 1 and p2 ~ N P2 , 2 2
n1 n2
and variance
P1Q1 P2Q 2
Var(p1-p2) = Var(p1)+Var(p2)
n1 n2
That is,
PQ P Q
p1 p 2 ~ N P1 P2 , 1 1 2 2
n1 n2
242
Sampling Distributions
Thus, standard error is given by
P1Q1 P2Q2
SE p1 p2 Var p1 p2
n1 n2
and variance
P1Q1 P2 Q 2 0.30 0.70 0.20 0.80
Varp1 p 2 0.0019
n1 n2 200 200
Note: i) Check your answers with those given at the end of the unit.
243
Statistical Analysis
9.6 EXACT SAMPLING DISTRIBUTION
As we have discussed in Section 9.1, some of the well known statisticians,
like Prof. R. A. Fisher, Prof. G. Snedecor, etc. worked on finding some of the
statistic and determined the exact sampling distributions their properties and
applications in differenct areas. These sampling distributions are named on
the nmame of its originator for example, F- distribution is named as Fisher’s
F-distribution and t-distribution as student’s t-distribution on the name of
Prof. W.S. Gosset. Before describing the Exact Sampling distribution first we
will discuss the term “Degree of Freedom” a very useful concept which is
necessarily to be understood before discussing and understanding the
concepts of Exact sampling distributions. The exact sampling distributions
are described with the help of degrees of freedom.
Degrees of Freedom (df)
The term degree of freedom (df) is related to the independency of sample
observations. In general, the number of degree of freedom is the total number
of observations minus the number of independent constraints or restrictions
imposed on the observations. For example, let x1, x2,…, xn be n independent
observations in a sample. Unless some condition is imposed on these x
values, it would have n df. Now, let one condition x1+x2+…+ xn = 100 be
imposed on this set, then it looses 1 df, that is, now df will be n – 1 since the
last value xn or any other value xi will be dependent on all other remaing
values and therefore, the number of independent values will be n-1. Further,
let x12+x22+…+ xn2 = 4000 be the another condition imposed, then now the df
will be n – 2, etc.
For a sample of n observations, if there are k restrictions among observations
(k < n), then the degrees of freedom will be (n–k).
9.6.1 Chi-square Distribution
The chi-square distribution was first discovered by Helmert in 1876 and later
independently explained by Karl- Pearson in 1900. The chi-square
distribution was discovered mainly as a measure of goodness of fit of any
model on the given frequency distribution or probability distribution. .
If a random sample X1, X2,…, Xn of size n is drawn from a normal
population having mean and variance σ2 then the sample variance can be
defined as
1 n n
s2 ( x i x ) 2 or (x i x ) 2 (n 1) s 2 s 2
n 1 i1 i 1
s 2
Then, the variate 2 which is the ratio of sample variance multiplied
2
by its degrees of freedom and the population variance follows the 2-
244 distribution with ν degrees of freedom.
The probability density functionof 2-distribution with ν df is given by Sampling Distributions
1 2 / 2 1
f 2 e /2
2
; 0 2 … (1)
2 / 2
2
where, ν = n −1.
2. Chi-square distribution has only one parameter n, that is, the degrees of
freedom.
3. Chi-square probability curve is highly positive skewed for smaller values
of n but becomes a symmetrical curve for larger values of n.
4. Chi-square-distribution is a uni-modal distribution, that is, it has single
mode.
5. The mean and variance of chi-square distribution with n df are n and 2n
respectively.
Note: i) Check your answers with those given at the end of the unit.
In general, the standard deviation σ is not known and in such a situation the
only alternative left is to estimate it from a sample. The value of sample
variance (S2) is used to estimate it where,
2 1 n1
s (x i x) 2
n 1 i1
X
But then in this case the variate is not normally distributed whereas it
S/ n
follows t-distribution with (n−1) df, that is,
X
t ~ t ( n 1) … (2)
s n
The t-variate is a widely used variable and its distribution is called student’s
t-distribution on the pseudonym name ‘Student’ of W.S. Gosset. The
probability density function of variable t with (n-1) = ν degrees of freedom is
given by
1
f t 1 / 2 ; t … (3)
2
1 t
B , 1
2 2
1
where, B , is known as beta function.
2 2
The probability curve of t-distribution is bell shaped and symmetric about t =
0 line. The probability curves of t-distribution is shown in Fig. 9.2 at two
different values of degrees of freedom n = 4 and 12.
Note: i) Check your answers with those given at the end of the unit.
9.6.3 F-Distribution
As we have mentioned in previous unit, F-distribution was introduced by
Prof. R. A. Fisher and defined as the ratio of two independent chi-square
variates when divided by their respective degrees of freedom. If we draw a
random sample X1 ,X 2 ,..., X n1 of size n1 from a normal population with mean
1 and variance σ 12 and another independent random sample Y1 , Y2 ,..., Yn2 of
size n2 from another normal population with mean 2 and variance 22
respectively then 1s12 / 12 is distributed as chi-square variate with ν1 df, that
1s12
is, 12 2
~ (21 ) … (1)
1
1 n1 1 n1
where, 1 n 1 1, X X i and s12 (X i X ) 2
n1 i1 n1 1 i1
2s 22
22 2
~ (2 2 ) … (2)
1
n2
1 1 n2
where, 2 n 2 1, Y
n2
X i and s 22
i 1
(Yi Y) 2
n 2 1 i1
Now, if we take the ratio of the above chi-square variates given in equations
(1) and (2), then we get
12 1s12 / 12 s12 / 12 12 / 1
~ F( 1 , 2 ) … (3)
22 2s 22 / 22 s 22 / 22 22 / 2
In the above expression F stands for the F-distribution. In the suffix, υ1 and υ2
are called the degrees of freedom of the F-distribution.
Now, if variances of both the populations are equal, that is, σ12 σ 22 , then F-
variate is written in the form of ratio of two sample variances which is as
follows:
249
Statistical Analysis s12
F ~ F( 1 , 2 ) … (4)
s 22
1 5 , 2 20
1 5 , 2 5
1 20 , 2 5
Fig. 9.3: Probability curves of F-distribution for (5, 5), (5, 20) and (20, 5)
degrees of freedom.
As it appears from the figure, F-distribution is uni-modal curve. It can be
seen from the figure that by increasing the first degrees of freedom from
1 5 to 1 20 the mean of the distribution (shown by vertical line) does
not change but probability curve shifs from the tail to the centre of the
distribution whereas increasing the second degrees of freedom from
2 5 to 2 20 the mean of the distribution (shown by vertical line)
decrease and the probability curve shifts from the tail to the centre of the
distribution. One can also get an idea about the skewness of the F-
distribution. We observe from the probability curve that it is positively
skewed curve and it becomes very highly positive skewed if ν 2 becomes
small. Now, we shall discuss some of the important properties of F-
distribution.
The F-distribution has the following important properties:
250
1. The probability curve of F-distribution is positively skewed curve. The Sampling Distributions
curve becomes highly positive skewed when ν2 is smaller than ν1.
2. F-distribution curve extends on abscissa from 0 to .
3. F-distribution is a uni-modal distribution, that is, it has single mode.
4. The square of t-variate with ν df follows F-distribution with 1 and ν
degrees of freedom.
2
5. The mean of F-distribution with (ν1,ν2) df is for 2 2.
2 2
222 1 2 2
2
for 2 4.
1 2 2 2 4
Women 40 45
Men 80 75
s12 / 12
F
s 22 / 22
F
45 / 40
1.27
1.93
2 2
75 / 80 0.88
For the above calculation, the degrees of freedom ν1 for women’s data
are7−1= 6 and the degrees of freedom ν2 for men’s data are 12 −1 =11.
Note: i) Check your answers with those given at the end of the unit.
13) For the purpose of a survey 15 students are selected randomly from
class A and 10 students are selected randomly from class B. At the
stage of the analysis of the sample data, the following information is
available:
15) What are the mean and variance of F-distribution with 1 5 and
2 12 degrees of freedom?
16) Write four applications of F-distribution.
We now end this unit by giving a summary of what we have covered in it.
253
Statistical Analysis Table 9.8: Calculation of Sample Mean
Sample Sample Sample
Number Observatio Mean ( X )
n
1 4, 6 5
2 4, 8 6
3 4,10 7
4 4, 12 8
5 6, 8 7
6 6, 10 8
7 6, 12 9
8 8, 10 9
9 8, 12 10
10 10, 12 11
Since the arrangement of all possible values of sample mean with their
corresponding probabilities is called the sampling distribution of mean,
thus,we arrange every possible value of sample mean with their respective
probabilities in the following Table 9.9 given below:
Table 9.9: Sampling distribution of sample means
= 200, σ = 4, n = 10
distribution is given by
E X 200
254
and variance Sampling Distributions
2
2 4 16
Var X 1.6
n 10 10
= 2550, n = 100, s = 54
First of all, we find the sampling distribution of sample mean. Since sample
size is large (n = 100 > 30) therefore, by the central limit theorem, the
sampling distribution of sample mean follows normal distribution. Therefore,
the mean of this distribution is given by
E X 2550
and variance
s 2 (54) 2 2916
Var ( X ) 29.16
n 100 100
1 = 68, σ1 = 2.3, n1 = 35
2 = 65, σ2 = 2.5, n2 = 50
To find the mean and standard error, first of all we find the sampling
distribution of difference of two sample means. Let X and Y denote the
mean height of male and female workers of hospital, respectively. Since n1
and n2 are large (n1, n2 > 30) therefore, by the central limit theorem, the
sampling distribution of (X Y) follows normal distribution with mean
E X Y 1 2 68 65 3
and variance
2 2
12 22 2.3 2.5
Var X Y
n1 n 2 35 50
0.2761 0.525
X 7000
N =10000, X = 7000 P 0.70 & n = 100
N 10000 255
Statistical Analysis First of all, we find the sampling distribution of sample proportion. Here, the
sample proportion is not given and n is large so we can assume that the
conditions of normality hold. So the sampling distribution is approximately
normally distributed with mean
E (p) P 0.70
and variance
PQ 0.70 0.30
Var(p) 0.0021 Q 1 P
n 100
25 20
P1 0.25, P2 0.20, n 1 250, n 2 200
100 100
Let p1 and p2be the sample proportions of alcohol drinkers in two cities A and
B respectively. Here the sample proportions are not given and n1 and n2 are
large n1 , n 2 30 so we can assume that conditions of normality hold. So the
sampling distribution of difference of proportions is approximately normally
distributed with mean
and variance
1 2 / 2 2 3
f 2 e ; 0 2
96
256
We have the probability density function of 2 distribution as: Sampling Distributions
1 2
f ( 2 ) e /2
( 2 ) ( / 2)1 ; 0 2
2 / 2
2
Where v = n = 1
by comparision we have
1 3 4 8
2 2
Thus, 8 n 1 8 n 9
Mean = n = 9 and Variance = 2n = 18.
9) Refer Sub-Section 9.6.1.
10) Here, we are given that
n
Mean = 0 and Variance ; n2
n 2
In our case, n = 8, therefore,
8 8
Mean = 0 and Variance 0.8
(80 2) 10
s12 / 12
F
s 22 / 22
257
Statistical Analysis Therefore, we have
2 2
F
60 / 65
0.85
0.69
2 2
50 / 45 1.23
2
Mean for 2 2
2 2
and
2 22 1 2 2
Variance 2
for 2 4.
1 2 2 2 4
2 12
Mean 1 .2
2 2 10
2(12) 2 (5 12 2) 30 144
Variance 10.8
5(12 2) 2 (12 4) 40 100
258