0% found this document useful (0 votes)
5 views35 pages

Understanding Sampling Distributions

Unit 9 of the Statistical Analysis document focuses on sampling distributions, covering the basics of sampling, the concept of standard error, and the Central Limit Theorem. It explains the necessity of sampling in statistical inference, the differences between population and sample, and various types of sampling methods. The unit also discusses exact sampling distributions such as chi-square, t-distribution, and F-distribution, providing a foundation for understanding how sample statistics relate to population parameters.

Uploaded by

katakaml8272
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
5 views35 pages

Understanding Sampling Distributions

Unit 9 of the Statistical Analysis document focuses on sampling distributions, covering the basics of sampling, the concept of standard error, and the Central Limit Theorem. It explains the necessity of sampling in statistical inference, the differences between population and sample, and various types of sampling methods. The unit also discusses exact sampling distributions such as chi-square, t-distribution, and F-distribution, providing a foundation for understanding how sample statistics relate to population parameters.

Uploaded by

katakaml8272
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Statistical Analysis

UNIT 9 SAMPLING DISTRIBUTIONS


Structure

9.1 Introduction
9.2 Objectives
9.3 Basics of Sampling
9.4 Sampling Distribution
9.4.1 Standard Error
9.4.2 Central Limit Theorem

9.5 Sampling Distribution of Statistics


9.5.1 Sampling Distribution of Mean
9.5.2 Sampling Distribution of Diference in Two Means
9.5.3 Sampling Distribution of Proportion
9.5.4 Sampling Distribution of Diference in Two Proportions

9.6 Exact Sampling Distributions


9.6.1 Chi-square Distribution
9.6.2 Student’s t-Distribution
9.6.3 F-Distribution

9.7 Let Us Sum Up


9.8 Key Words
9.9 Suggested Further Reading/References
9.10 Answers to Check Your Progress

9.1 INTRODUCTION
In general, extracting information from all the elements or items of a large
population may be time consuming and expensive if the population size is
infinitely large. But even then sometimes there are so many problems
attached to large size populations where it becomes necessary to draw
inference about population [Link] example, one may wish to
estimate the average height of all the two thousand students in a college, a
businessman may be interested to estimate the proportion of defective items
in a production line, a manufacturer of car tyres may want to estimate the
variations took place in the diameter of produced tyres, a pharmacist may
want to estimate the difference of effect of two types of drugs, [Link] all such
cases, there is always an unknown population involved whose characteristics
are described through some parameters.
Due to large sizes of populations, in which we may be interested, for
drawing inferences about the population parameters, generally we draw a
224 sample and determine a function of sample values, which is called statistic.
Selection of a sample, that is, only a part of the population saves a lot of time, Sampling Distributions
money and labour and results drawn on sample values are quicklyavailable
forinterpretation and as good sometimes as obtained on the basis of the entire
population. The process of generalising sample results to the population is
called Statistical Inference. Sincethere might be a large number of samples of
same size drawn from the population, the values of statistic generally vary
from sample to sample and is associated with the probability of selection of
the particular sample. Therefore, the sample statistic is a random variable
following any probability distribution. Therefore, it may be a matter of
interest for a statistician to know what distribution the statistic follows if the
samples are assumed to be selected from a theoretical distribution like, a
normal distribution with given mean and variance; a binomial distribution; a
gamma distribution; a Poisson distribution and so on. In contrast to
theoretical distributions, probability distribution of a statistic in popularly
called a sampling distribution. In this unit we shall discuss the sampling
distribution of sample mean; of sample median; of sample proportion; of
difference between two sample means and sample proportions.
Due to this curiosity, Prof. R.A. Fisher, Prof. G. Snedecor and some other
statisticians worked in this area and obtained exact sampling distributions
which are followed by some of the important statistics. In present unit of this
block, we shall also discuss some important sampling distributions such as 2
(read as chi-square), t and F. Generally, these sampling distributions are
named on the name of the originator, for instance, Fisher’s F-distribution is
named like this after the name of its inventor Prof. R.A. Fisher.

9.2 OBJECTIVES
After studying this unit, you should be able to:
 explain the concept of sampling and sampling distribution;
 explain the concept of standard error and Central Limit Theorem;
 define the sampling distribution; and
 describe the sampling distribution of sample mean and difference of
two sample means.

9.3 BASICS OF SAMPLING


Before discussing sampling distributions of different kinds of statistic, in this
section we shall be discussing basic concepts and definitions of some of the
important terms which areverymuch helpful in understanding the
fundamentals of statistical inference and are frequently used in this unit.
Population, Sample and Sampling
Literally, “population” means a well-defined group (collection or bunch) of
some objects (units or elements). Examples of populations and its units are : a
city with some clear-cut territory having its dwellers as units (elements) or
225
Statistical Analysis having houses as its units; a hospital with indoor patients as units or its
doctors as units; a library with its employees as units or books as its units; a
river with its fishes as units; a school with enrolled students as units; a
banyan tree with leaves as units; sky with stars as units and so on. Thus,
population is the collection or group of individuals /items /units/ observations
under study. The total number of elements / items / units / observations in a
population is known as its size and generally denoted by N.
A finite subset of units of a population is called a “sample” and the number
of units belonging to the sample is called the sample size. If n denotes the
sample size, then necessarily we have the condition n < N and then the
sample is said to be a proper sample. However, in some situations, we have n
= N, that is, all the units of the population are selected as sample. As
discussed above, in almost all kinds of statistical studies, a sample is selected
for inferential purposes because of the reasons stated above. However, even
if the population is not large enough or it has a well-defined territory or it is
not of destructive type; a sample is useful since besides it is always less
costly, less time consuming, it helps us to get a handy result in very short
time.

A sample of size n is not selected simultaneously, rather n units are selected


one by one in n draws in order to constitute a sample. The process through
which units of the population are selected in different draws to constitute a
sample of specified size, is known as “sampling”. Thus, sampling is a draw-
to-draw mechanism of selecting some units from a population. Sampling is
mainly of two types, namely,

(a) Probability Sampling or Random Sampling, and


(b) Non-Probability Sampling.
The technique of random sampling is of fundamental importance in the
application of statistics. The Estimation theory is based on the assumption of
random sampling. It is a scientific method of selecting samples accordingly
to some laws of chance in which each unit in the population has some
definite pre assigned probability of being selected in the [Link]
Random Sampling, Stratified Sampling, Systematic Random sampling, etc.,
are the examples of random sampling.
In non-probability sampling, the sample is selected with definite purpose
and, hence, the choice of sampling units depends entirely on the discretion
and judgment of the investigator. It is, therefore, not a scientific method and
attracts severe criticisms in many circumstances. Purposive sampling,
Judgement Sampling, etc., are the examples of non-probability sampling.
We know that if n < N, it is a sampling procedure. When n = N the process is
known as “complete enumeration or census”. An example of complete
enumeration is the decennial census of India.
226
Need for Sampling in Statistical Inference Sampling Distributions

As we have discussed in Section 9.1, in many practical and real situations the
population under consideration is either infinitely large in size or it is
unbounded in the sense that its boundaries are not well-defined or it is of
destructive nature. As for example, in ordert to determine average life of two
hundred produced electric bulbs, it would be necessary to light these bulbs
unless they all get fused, resulting destruction of the entire lot of production.
Similarly, in order to estimate the proportion of smokers in a large city,
information have to be gathered from each and every person dwelling in the
city which would be a very difficult task in terms of manpower, money and
time to be required. Sometimes, the population size is not known if its
boundaries are not well-defined, like number of fish in a pond or lake. In all
such cases, it is not feasible to gather information on all the units, rather it
becomes almost impossible. Keeping in view the difficulty in contacting each
and every unit of the population due to these reasons and also to save the
time, money and manpower to be required for this, generally a part of the
population, popularly known as a “sample” is selected in some pre-asigned
manner which in turn is used for drawing the inferences on the population
itself or on the parameters of the population. The results obtained from the
sample are projected in such a way that they are valid for concluding about
the entire population. Therefore, the sample works like a “Vehicle” to reach
(drawing) at valid conclusions about the population. In fact, a sample helps
us to reach to the “whole”(population) from a “part” (sample) in all types of
statistical studies. Thus, statistical inference can be defined as:
“It is a process of concluding (projecting or infering) something desired about
a given population on the basis of sample results”.
Parameter and Statistic
A parameter is a function of population values which is used to represent
certain characteristics of the [Link] example, population total,
population mean, population variance, population coefficient of variation,
population proportion, population correlation coefficient, etc., are all
parameters since their calculation involve all the population values.
A statistic is a function of sample values only and does not contain any
unknown population parameter in it. For example, if X1 ,X2 , ..., Xn represent
values of the variable X in a random sample of size n taken from a population
1 n
then sample mean X   X i is a statistic.
n i 1

Estimator and Estimate


A statistic which is used to guessor to estimate ( infer) an unknown
population parameter then it is known as estimator and the value of the
estimator based on observed value of the sample is known as estimate of
227
Statistical Analysis parameter. For example, suppose the parameter λ of the Poisson
1 n
populationf(x, λ) is unknown. If we use sample mean X   Xi ,
n i 1
calculated on the basis of sample values X1 ,X2 , ..., Xn to estimate λ then X
is an estimator and any particular value of x is called estimateof parameter λ.

9.4 SAMPLING DISTRIBUTION


As discussed under Section 9.1, a statistic is calculated on the basis of a
sample of specific size, selected from the given population for drawing the
inference about the population. But due to the fact that sample size is always
smaller than the population size; (n < N), a number of samples of same size
can theoretically be drawn from the population (theoretically nCN possible
samples). Thus, in fact, we have in all nCN values of the statistic. Moreover,
since in probability sampling, each sample is selected with some pre-defined
probability of selection, it means that all these statistic values have some
probability to appear. In this sense, we can say that a statistic ̂ is a random
 k k 
variable defined as ˆ , p(ˆ ) , where ̂ and p(ˆ ) are respectively the value
k k

of the statistic ̂ k and the probability of its appearance due to the kth
sample; k = 1, 2,…, nCN.
Generally, in practice onlya single random sample is taken from a given
population and its mean X is considered to be representative of the
population mean µ. This sample mean may or may not represent the
population mean. Since we cannot determine the proximity of sample mean
and population mean on the basis of single random sample, we can use the
concept of sampling distribution to bring the value of sample mean close to
that of Population mean. Being a random variable, the statistic must possess a
probability distribution which may be used to answer some questions
regarding the nature of the statistic, for example, what is the probability that
its value as obtained from the given sample is differing from the parameter
value by amargin of 10? The probability distribution of a statistic is called
sampling distribution of the statistic. Let us illustrate how the sampling
distribution of a statistic can be obtained given a population. For instance, let
we wish to estimate the population mean using the sample mean as an
estimator.
Suppose that a baby-sitter has 5 children under her supervision. The ages of
children are 2, 4, 6, 8 and 10 years. Supposing this group of children as a
population of size 5, we get the population mean as

1 N 2  4  6  8  10
X 
N i 1
Xi 
5
6

228
Therefore, the variance of this population is given by: Sampling Distributions

2 2 2
1 N
   X i  X  
2 2 2  6  4  6  .......  10  6  8
N i 1 5

Now, let us take all the possible simple random samples of size 2 without
replacement from this population. There are 5C2 = 10 such possible samples
which are listed below along with their respective means:

Table 9.1: Possible samples

Sample
Sample Sample Mean
No.
1 2, 4 3
2 2, 6 4
3 2, 8 5
4 2, 10 6
5 4, 6 5
6 4, 8 6
7 4, 10 7
8 6, 8 7
9 6, 10 8
10 8, 10 9

Now, suppose that our selected sample is either (2, 4) with mean 3 or it is the
sample (8, 10) with mean 9. In both cases, the sample is not a good
representative of the population since sample mean is far from the population
mean 6. But if the selected samples are coincidently either (2, 10) or (4, 8),
then these are good representatives of the population in the sense that both
have sample means exactly equal to population mean. Thus, this example
illustrates that a single random sample may or may not be representative for
the decision maker to reach meaningful conclusion. However, the grand
mean of the distribution of these ten sample means, is calculated and
observed to be equal to the population mean as follows:

1 10 3 45 65 6 7 89


x 
10 i 1
xi 
10
 6.

Hence, the mean of the sample means can be considered to represent the
population mean for analysis and decition making purposes.
Now let us put sample means along with their probabilities of occurrence as
follows:

229
Statistical Analysis Table 9.2: Probability Distribution
Probability
Sample
Frequency of
Mean
Occurrence

3 1 0.1

4 1 0.1

5 2 0.2

6 2 0.2

7 2 0.2

8 1 0.1

9 1 0.1

Total 10 1.0

This distribution which shows the distribution of probabilities over all the
possible values of the sample mean is refered to as sampling distribution of
the sample mean. Symbolically, it can be denoted as {( x , p( x )}. As
explained above, the sampling distribution of a statistic can be defined as:
“The probability distribution of all possible values of a statistic that would be
obtainedby drawing all possible samples of the same size from the population
is called sampling distribution of that statistic.”
The mean, variance and other measures of a sampling distribution can be
obtained in a similar way as computed in a frequency distribution, taking
probabilities as frequencies. Thus, mean of the above sampling distribution
will be
7
x   x i p x i   3 x 0 . 1  4 x 0 . 1  5 x 0 . 2  6 x 0 . 2  7 x 0 . 1  8 x 0 . 1  9 x 0 . 1  6
i 1

This value is same as the population meanµ. The variance of the distribution
is obtained as:
2 2 2
1 N
   x i  x  
2 2 3  6  4  6  .......  9  6  28  4
x
7 i 1 7 7

9.4.1 Standard Error

The expression obtained above, obviously measures the variability of the


sample mean around the actual mean, that is, how much samplestatistic may
vary from sample to sample.
This is equivalent to population variance which measures the deviation of the
population values around the population mean. If we consider the positive
230 square root of this , it would be equivalent to standard deviation. In order to
differentiate with population standard deviation, it is called as “standard Sampling Distributions
error” (SE) of that statistic. Thus, standard error ofa statistic can be defined
as:
“The standard deviation of a sampling distribution is known as standard
error”.
The computation of the standard error is a tedious process. However, the
formula for calculating standard error of sample mean on the basis of a
sample size n is seen always to be


SE  X   ,
n

where σstands for the standard deviation of the population.


Standard errors of some otherwell known statistics are given below:
1. The SE of sample proportion (p)is given by
PQ
SE(p) =
n

where, P is population proportion, n is sample size and Q = 1-P.

2. The SE of difference of two sample means is given by

σ12 σ 22
SE  X  Y  = 
n1 n 2

where, σ12 and σ 22 are the population variances of two different populations
and n1 and n2 are the sample sizes of two independent samples selected from
the two populations respectively.

3. The SE of difference between two sample proportions is given by

P1Q1 PQ
SE( p1  p 2 ) =  2 2
n1 n2

where, P1 and P2 are population proportions in two different populations from


which samples of sizes n1 and n2 are taken respectively; Q1= 1-P1 and Q2 = 1-
P2.

From all the above formulae, we can understand that standard error is
inversely proportional to the sample size. Therefore, as sample size increases
the standard error decreases.

The standard error is used to express the accuracy or precision of the estimate
of population parameter because the reciprocal of the standard error is the
measure of reliability or precision of the [Link] error also
determines the probable limits or confidence limitswithin which the
231
Statistical Analysis population parameter may be expected to lie with certain level of
[Link] error is also applicable in testing of hypothesis.
9.4.2 Central Limit Theorem

The central limit theorem is the most important theorem of Statistics. It was
first introduced by De Movers in the early eighteenth century. The theorem
states that regardless of the nature of the distribution of the population, the
distribution of the sample mean approaches the Normal Probability
distribution as the sample size increases. In general, the larger the sample
size, the closer the proximity of the distribution of sample mean to the
Normal distribution. However, in practice, sample sizes of 30 or larger are
considered adequate for this purpose. It should be noted, however, that the
sampling distribution of sample mean would be always normally distributed
if the original population is normally distributed.

Therefore, according to the central limit theorem, if X1 ,X2 , ..., Xn is a


random sample of size n taken from a population with mean µ and variance
σ2 thenthe sampling distribution of the sample mean tends to normal
distribution with mean µ and variance σ2/n as sample size tends to large
(n  30) whatever be the form of parent population, that is,

 2 
X ~ N  , 
 n 

and the variate

X 
Z ~ N  0,1
/ n

follows normal distribution with mean 0 and variance unity, that is, the
variate Z follows standard normal distribution.

9.5 SAMPLING DISTRIBUTION OF STATISTICS


Sample mean is one of the most commonly used statistic in any of the
statistical studies in order to study the nature and charactristics of a given
population. Examples are average income in a locality in a social survey,
average life of produced items in a manufacturing process, average
temperature in a day in meterological survey, etc. Oweing to these reasons, it
is important to obtain the sampling distribution of sample mean.

9.5.1 Sampling Distribution of Sample Mean


In the previous section we have elaborated the meaning of sampling
distribution of the sample mean and presented an example how it can be
obtained. On the basis of these, the sampling distribution of sample mean can
be defined as follows: “The probability distribution of sample mean that
232 would be obtainedby drawing all possible samples of same size from the
population is called sampling distribution of sample mean or simply of Sampling Distributions
mean.”
If X1 , X 2 , ..., X n is an independent random sample of size n taken from a
normalpopulation with mean µ and variance σ2 then it has been established
that sampling distribution of sample mean X is also normal. The mean and
variance of sampling distribution of X can be obtained as

 X  X2  ...  X n 
Mean of X  E  X   E  1   By defination of X 
 n
1
 [E(X1 )  E(X 2 )  ...  E(X n )]
n
Since E(Xi) = mean of Xi for all i=1, 2,…,n; therefore, we have E(Xi) = µ.
Similarly, Var(Xi) = σ2 for all i. We, therefore, have
1
E X        ......  (n times)  n  
n n
and variance

1 
Var(X)  Var  (X1  X 2  ...  X n ) 
n 

1
  Var(X1 )  Var(X 2 )  ...  Var(X n )
n2

1 2 n 2  2
n
 2 2 2
 2       ......   (n times)  2 
n n

We, therefore, conclude that If X i ~ N ,  2  then

 2 
X ~ N  , 
 n 

and


SE  X   SD  X   Var  X  
n
Let us illustrate these results using some examples.
Example 1: Diameter of a steel ball bearing produced on a semi-automatic
machine is known to be distributed normally with mean 12 cm and standard
deviation 0.1 cm. If we take a random sample of 10 ball bearings then find
mean and variance of sampling distribution of mean.
Solution: Here, we are given that
 = 12, σ = 0.1, n = 10
Since the sample is taken from the population of ball bearings in which
diameter follows normal distribution N(12, 0.01), we have 233
Statistical Analysis E  X     12

2 (0.1) 2
Var  X     0.001 .
n 10

Thus, X ~N(12, 0.001).

Example 2: If ages of 5 employees of Account Section of a Company are 58,


66, 64, 62 and 50 years then construct the sampling distribution of average
age of employees by taking all samples of size 2 with replacement.
Solution:Here, we are given that N= 5,n = 2
Hence the population mean:
54  56  60  64  66
X  60
5
Therefore, the variance of this population is given by:

2 
54  602  56  602  .......  66  60 2 
104
 20.08
5 5
Now, all possible samples (without replacement) are 10 which are shown in
the Table given below:
Table 9.3: Possible Samples

Sample Sample Sample


Number Observation Mean

1 54, 56 55
2 54, 60 57
3 54, 64 59
4 54, 66 60
5 56, 60 58
6 56, 64 60
7 56, 66 61
8 60, 64 62
9 60, 66 63
10 64, 66 65

We, therefore, have the grand mean of the distribution as


1 10 55  57  58  ...............  62  63  65
x 
10 i 1
xi 
10
 60

which is same as the population mean. Now, let us construct sampling


distribution of the sample mean as given under:

234
Table 9.4: Probability Distribution Sampling Distributions

Probability
Sample
Frequency of
Mean
Occurrence

55 1 0.1

57 1 0.1

58 1 0.1

59 1 0.1

60 2 0.2

61 1 0.1

62 1 0.1

63 1 0.1

65 1 0.1

Total 10 1.0

Therefore, the mean of sample means can be obtained by the formula:


7
x   x i px i   55x 0.1  (57 x 0.1)  58x 0.1  59 x 0.1  60 x 0.2
i 1

 61x 0.1  62 x 0.1  63x 0.1  65x 0.1  60

This value is same as the population meanµ. The variance of the distribution
is obtained as:
2 2 2
1 N
   x i  x  
2 2 55  60  57  60  .......  65  60  78  8.66
x
9 i 1 9 9

CHECK YOUR PROGRESS 1

Note: i) Check your answers with those given at the end of the unit.

1) If lives of 5 televisions of certain company are 4, 6, 8, 10 and 12 years


then construct the sampling distribution of average life of televisions by
taking all samples of size 2.
2) The weight of certain type of a truck tyre is known to be distributed
normally with mean 200 pounds and standard deviation 4 pounds. A
random sample of 10 tyres is selected. What is the sampling distribution
of sample mean? Also obtain the mean and variance of this distribution.
235
Statistical Analysis 3) The mean life of CFL blubs produced by a company is 2550 hours. A
random sample of 100 CFL bulbs is selected and the standard deviation
is found to be 54 hours. Find the mean and variance of the sampling
distribution of mean.

9.5.2 Sampling Distribution of Difference Between Two


Sample Means

Instead of sample mean of a single population, sometimes one may be


interested into two populations and, hence, in finding the sampling
distribution of the difference of sample means which are obtained on the
basis of the samples drawn from two populations. Such cases arise when we
wish to compare average lives of elecrtric bulbs of two kinds or to compare
the efficiency of two different drugs of same disease or to compare the
average marks of two different sections in a school, etc.
Let the same characteristic measured on two populations, say population-I
and population-II be represented by X and Y variables. Suppose population-I
have mean 1 and variance 12 whereas population-II have mean  2 and
variance  22 . Then,let X be the sample mean based on a sample of size n1
selected from population-I and Y be the sample mean based on a sample of
size n2 selected from the population-II. As before, it is obvious that, if
necessary, one may select all possible samples of sizes n1 and n2 respectively
from the two populations. Then considering all possible differences of means
and sampling distribution of difference of population means can be obtained.
The sampling distribution of difference of sample means can be defined as:

“The probability distribution of all values of the difference of two sample


means would be obtained by drawing all possible samples from both the
populations. Such a distribution is called sampling distribution of difference
of two sample means.”

If both the parent populations are normal, that is,


  
X ~ N 1 , 12 and Y ~ N  2 ,  22 
then as we discussed in previous section
 2   2 
X ~ N  1 , 1  and Y ~ N   2 , 2 
 n1   n2 
If two independent random variables X and Y are normally distributed then
the difference (X  Y) also normally distributed. Therefore, the sampling
distribution of difference of two sample means (X  Y) also follows normal
distribution with mean

E  X  Y   E  X   E  Y  =µ1− µ2
236
and variance Sampling Distributions

12  22
Var  X  Y   Var  X   Var  Y   
n1 n 2

Therefore, standard error of difference of two sample means is given by

12 22
SE  X  Y   Var  X  Y    .
n1 n 2

Let us see an application of the sampling distribution of difference of two


sample means with the help of an example.
Example 3: LED Bulbs manufactured by company A have mean lifetime of
2400 hours with standard deviation 200 hours, while LED Bulbs
manufactured by company B have mean lifetime of 2200 hours with standard
deviation of 100 hours. If random samples of 125 LED Bulbs of each
company are tested, find. The mean and standard error of the sampling
distribution of the difference of mean lifetime of LED Bulbs.
Solution: Here, we are given that
1 = 2400, σ1 = 200, 2 = 2200, σ2 = 100 and n1 = n2 = 125
Let X and Y denote the mean lifetime of CFLs taken from companies A and
B respectively. Since n1 and n2 are large (n1, n2 > 30) therefore, by central
limit theorem, the sampling distribution of (X  Y) follows normal
distribution with mean

E(X  Y)  1   2  2400  2200  200

and variance
2 2
12 22 200 100
Var(X  Y)      400
n1 n 2 125 125

Therefore, the standard error is given by

SE  X  Y   Var  X  Y   400  20 .

Now, continuing our discussion, the sampling distribution of difference of


two means, we consider another situation.

If population variances 12 and  22 are unknown then we estimate 12 and  22
by the values of the sample variances of the samples taken from the first and
second population respectively. For large sample sizes n1 and n2   30  , the
sampling distribution of (X  Y) is very closely normally distributed with

 s2 s2 
mean (µ1− µ2) and variance  1  2  .
 n1 n 2  237
Statistical Analysis If population variances 12 and  22 are unknown and 12   22   2 then σ2 is
estimated by pooled sample variance s 2p where,

1
s 2p 
n1  n 2  2

n 1s12  n 2 s 22 
and variate

t
X  Y     2 
1
~ t ( n1  n 2  2 )
1 1
sp 
n1 n 2

follows t-distribution with (n1 + n2 − 2) degrees of freedom. Similar to


sampling distribution of mean, for large sample sizes n1 and n2   30  the
sampling distribution of (X  Y) is very closely distributed as normal with
1 1 
mean (µ1− µ2) and variance s 2p    .
 n1 n 2 

CHECK YOUR PROGRESS 2

Note: i) Check your answers with those given at the end of the unit.

4) The average height of Male workers in a hospital is found to be 68


inches with a standard deviation of 2.3 inches whereas the average
height of Female workers in a hospital is found to be 65 inches with a
standard deviation of 2.5 inches. If a sample of 35 Male and 50 Female
mean and standard error of the sampling distribution of the difference
between workers are selected at random, find is the the sample means
of height of Male workers and female workers.

9.5.3 Sampling Distribution of Sample Proportion


In Section 9.5.1, we have discussed the sampling distribution of sample mean
if some quantitative characteristic is taken into consideration..But in many
real word situations,when some qualitatrive characteristic is under
consideration, sample proportion, instead of sample mean is computed on the
basis of a sample and hence, sampling distribution of sample proportion is
needed. Examples of cases in which sample proportions are needed for
analysis are proportion of male births; proportion of defective items in
manufacturing process; proportion of cancer cases in population, etc.
For sampling distribution of sample proportion, we need sample proportion p.
Let a sample of size n is taken from the population, then p is given by
X
p 1
n
where X is the number of observations /individuals / items / units in the
sample which have the particular characteristic under study. For better
238
understanding of the process, we consider the following example in which Sampling Distributions
size of the population is very small:
Suppose, there is a lot of 4 cartons A, B C and D of electric Tubes and each
carton contains 20 Electric Tubes. The number of defective bulbs in each
carton is given below:
Table 9.4: Number of Defective Bulbs per Carton

Carton Number of
Defectives Bulbs

A 2
B 4
C 1
D 3

The population proportion of defective tubes can be obtained as


2  4  1  3 10 1
P  
20  20  20 80 8
Now, let us assume that we do not know the population proportion of
defective tubes. So we decide to estimate population proportion of defective
tubes on the basis of samples of size n = 2. There are N C n  4 C 2  6 possible
samples of size 2 with replacement. All possible samples and their respective
proportion defectives are given in the following table:

Table 9.6: Calculation of Sample Proportion

Sample Sample Sample Sample


Carton Proportion(p)
Observation

1 (A, B) (2, 4) 6/40


2 (A, C) (2, 1) 3/40
3 (A,D) (2, 3) 5/40
4 (B, C) (4, 1) 5/40
5 (B, D) (4, 3) 7/40
6 (C, D) (1, 3) 4/40

From the above table, we can see that value of the sample proportion is
varying from sample to sample. So we consider all possible sample
proportions and calculate their probability of occurrence. Since there are 6
possible samples therefore the probability of selecting a sample is 1/8. Then
we arrange the possible sample proportions with their respective probability
in Table 2.7:

239
Statistical Analysis Table 9.7: Sampling Distribution of Sample Proportion

[Link]. Sample Frequency Probability


Proportion(p)
1 3/40 1 1/6
2 4/40 1 2/6
3 5/40 2 2/6
4 6/40 1 1/6
5 7/40 1 1/6
Total 1.00

This distribution is called the sampling distribution of sample proportion.


The mean of sampling distribution of sample proportion can be obtained as
1 k k
p  ii
K i 1
p f where, K  
i 1
fi

1 3 4 5 6 7  30 1 1
 1   1   2   1  ...   1   
6  40 40 40 40 40  40 6 8
Thus, we have seen that mean of sample proportion is equal to the population
proportion.
If a population whose elements are divided into two mutually exclusive
groups− one containing the elements which possess a certain attribute and
other containing elements which do not possess the attribute, then number of
successes (elements possess a certain attribute) follows a binomial
distribution with mean
E(X)  nP

and variance
Var(X)  nPQ where Q  1  P
where, P is the probability or proportion of success in the population.
Now, we can easily find the mean and variance of the sampling distribution
of sample proportion by using the above expression as

X 1 1
E(p)  E    E(X)  nP  P
n n n
and variance
X 1  Var  aX   a 2Var  X 
Var(p)  Var    2 Var(X)
n n
1 PQ
 nPQ  .  Var  X   nPQ 
n 2
n 
240
Also standard error of sample proportion can be obtained as Sampling Distributions

PQ
SEp  Var( p) 
n
If the sampling is done without replacement from a finite population then the
mean and variance of sample proportion is given by

E p  P

and variance
N  n PQ
Var  p  
N 1 n
where, N is the population size and the factor (N-n) / (N-1) is called finite
population correction.
If sample size is sufficiently large, such that np > 5 and nq > 5 then by central
limit theorem, the sampling distribution of sample proportion p is
approximately normally distributed with mean P and variance PQ/n where,
Q= 1 P.
Let us see an application of the sampling distribution proportion with the help
of an example.
Example 4: A machine produces a large number of items of which 15% are
found to be defective. If a random sample of 200 items is taken from the
population and sample proportion is calculated then find mean and standard
error of sampling distribution of proportion.
Solution: Here, we are given that
15
P= = 0.15, n = 200
100
We know that when sample size is sufficiently large, such that np > 5 and nq
> 5 then sample proportion p is approximately normally distributed with
mean P and variance PQ/n where, Q = 1– P. But here the sample proportion
is not given so we assume that the conditions of normality hold, that is, np >
5 and nq > 5. So mean of sampling distribution of sample proportion is given
by
E ( p )  P  0.15

and variance
PQ 0.15  0.85
Var(p)    0.0006
n 200
Therefore, the standard error is given by

SE  p   Var  p   0.0006  0.025

241
Statistical Analysis CHECK YOUR PROGRESS 3

Note: i) Check your answers with those given at the end of the unit.

5) A state introduced a policy to give loan to unemployed doctors to start


own clinic. Out of 10000 unemployed doctors 7000 accept the policy
and got the loan. A sample of 100 unemployed doctors is taken at the
time of allotment of loan. Find the mean and standard error of the
sampling distribution of proportion of doctors who accepted the policy
and got the loan.

9.5.4 Sampling Distribution of Difference of Two Sample


Proportions
Just like the sampling distribution of difference of two population means,
which has been obtained in sub-Section 9.5.2, the sampling distribution of
difference of two population proportions can be obtained which is required
many times in statistical inference. We shall show how it can be obtained.
Suppose, there are two populations, say, population-I and population-II under
study and the proportions of some attribute in populations I and II are
respectively P1 and P2. For finding sampling distribution of difference of two
sample proportions, let samples of sizes n1 and n2 be selected respectively
from populations I and II and the sample proportions of the attribute in these
samples are p1 and p2 respectively. Since sample proportions obviously vary
from sample to sample, their difference will be a random variable having
some probability distribution, which would be termed as the sampling
distribution of difference of two sample proportions. On the basis of
arguments made for one sampling proportion in the previous section, for the
distribution of sampling proportion p due to central limit theorem, here also
we can assume that if n1p1  5, n1q1  5, n2p2  5 and n2q2  5 then

 PQ   PQ 
p1 ~ N  P1, 1 1  and p2 ~ N  P2 , 2 2 
 n1   n2 

where, Q1 = 1 P1 and Q2 = 1 P2.


Also, by the property of normal distribution, the sampling distribution of the
difference of sample proportions follows normal distribution with mean

E(p1-p2) = E(p1)-E(p2) = P1-P2

and variance

P1Q1 P2Q 2
Var(p1-p2) = Var(p1)+Var(p2)  
n1 n2

That is,

 PQ P Q 
p1  p 2 ~ N P1  P2 , 1 1  2 2 
 n1 n2 
242
Sampling Distributions
Thus, standard error is given by

P1Q1 P2Q2
SE  p1  p2   Var  p1  p2   
n1 n2

Thus, p1-p2 follows a normal distribution with above mentioned mean,


variance and standard error. Let us see an application of the sampling
distribution of difference of two sample proportions with the help of an
example.
Example 5: In one population, 30% persons had hair colour Black and in
second population 20% had the same hair colour. A random sample of 200
persons is taken from each population independently and calculate the
sample proportion for both samples, then find the mean and variance of the
sampling distribution of the difference between two sample proportions.
Solution: Here, we are given that
P1 = 0.30, P2 = 0.20, n1= n2 = 200
Let p1 and p2be the sample proportions of blue-eye persons in the samples
taken from both the populations respectively. We know that when n1 and n2
are sufficiently large, such that n1p1  5, n1q1  5, n2p2  5 and n2q2  5 then
sampling distribution of the difference between two sample proportions is
approximately normally distributed. But here the sample proportions are not
given so we assume that the conditions of normality hold. So mean of
sampling distribution of the difference between two sample proportions is
given by

E  p1  p 2   P1  P2  0.30  0.20  0.10

and variance
P1Q1 P2 Q 2 0.30  0.70 0.20  0.80
Varp1  p 2       0.0019
n1 n2 200 200

Thus, standard error

Standard Error = (p1  p 2 )  Var(p1  p 2 )  0.0019  0.04

CHECK YOUR PROGRESS 4

Note: i) Check your answers with those given at the end of the unit.

6) In city A, 25% persons were found to be smokers and in another city B,


20% persons were found smokers. If 250 persons of city A and 200
persons of city B are selected randomly, then find the mean and
standard deviation of sampling distribution of the difference in sample
proportions.

243
Statistical Analysis
9.6 EXACT SAMPLING DISTRIBUTION
As we have discussed in Section 9.1, some of the well known statisticians,
like Prof. R. A. Fisher, Prof. G. Snedecor, etc. worked on finding some of the
statistic and determined the exact sampling distributions their properties and
applications in differenct areas. These sampling distributions are named on
the nmame of its originator for example, F- distribution is named as Fisher’s
F-distribution and t-distribution as student’s t-distribution on the name of
Prof. W.S. Gosset. Before describing the Exact Sampling distribution first we
will discuss the term “Degree of Freedom” a very useful concept which is
necessarily to be understood before discussing and understanding the
concepts of Exact sampling distributions. The exact sampling distributions
are described with the help of degrees of freedom.
Degrees of Freedom (df)
The term degree of freedom (df) is related to the independency of sample
observations. In general, the number of degree of freedom is the total number
of observations minus the number of independent constraints or restrictions
imposed on the observations. For example, let x1, x2,…, xn be n independent
observations in a sample. Unless some condition is imposed on these x
values, it would have n df. Now, let one condition x1+x2+…+ xn = 100 be
imposed on this set, then it looses 1 df, that is, now df will be n – 1 since the
last value xn or any other value xi will be dependent on all other remaing
values and therefore, the number of independent values will be n-1. Further,
let x12+x22+…+ xn2 = 4000 be the another condition imposed, then now the df
will be n – 2, etc.
For a sample of n observations, if there are k restrictions among observations
(k < n), then the degrees of freedom will be (n–k).
9.6.1 Chi-square Distribution
The chi-square distribution was first discovered by Helmert in 1876 and later
independently explained by Karl- Pearson in 1900. The chi-square
distribution was discovered mainly as a measure of goodness of fit of any
model on the given frequency distribution or probability distribution. .
If a random sample X1, X2,…, Xn of size n is drawn from a normal
population having mean  and variance σ2 then the sample variance can be
defined as
1 n n
s2   ( x i  x ) 2 or  (x i  x ) 2  (n  1) s 2  s 2
n  1 i1 i 1

where, ν = n −1; the symbol ν read as ‘nu’.

s 2
Then, the variate  2  which is the ratio of sample variance multiplied
2
by its degrees of freedom and the population variance follows the 2-
244 distribution with ν degrees of freedom.
The probability density functionof 2-distribution with ν df is given by Sampling Distributions

1 2   / 2 1
f  2   e /2
 
2
; 0  2   … (1)

2 / 2
2
where, ν = n −1.

It can be seen that 2-distribution is a sampling distribution of a statistic


s 2
2  , the shape of probability distribution of which is shown by
2
probability curve shown in Fig. 3.1 for n = 1, 4, 10 and 22.

Fig. 9.1: Chi-square probability curves for n = 1, 4, 10 and 22

After looking the probability curves of 2-distribution for n = 1, 4, 10 and 22,


one can understand that probability curve takes shape of inverse J for n = 1.
The probability curve of chi-square distribution gets skewed more and more
to the right as n becomes smaller and smaller. It becomes more and more
symmetrical as n increases because as n tends to ∞, 2-distribution converges
to a normal distribution. It is apparent from the figure that even when n= 22,
it is tending to a symmetrical probability curve which is an essential property
of a normal curve.
Example 8: What are the mean and variance of chi-square distribution with 5
degrees of freedom?
Solution: The mean and variance of chi-square distribution with n degrees of
freedom are given by
Mean = n and Variance = 2n
In our case, n = 5, therefore,
Mean = 5 and Variance = 10.

Some of the salient features of 2-distribution are mentioned here without


giving proofs of each:
245
Statistical Analysis 1. The probability curve of the chi-square distribution lies in the first
quadrant because the range of  2 -variate is from 0 to ∞.

2. Chi-square distribution has only one parameter n, that is, the degrees of
freedom.
3. Chi-square probability curve is highly positive skewed for smaller values
of n but becomes a symmetrical curve for larger values of n.
4. Chi-square-distribution is a uni-modal distribution, that is, it has single
mode.
5. The mean and variance of chi-square distribution with n df are n and 2n
respectively.

The applications of chi-square distribution are very wide in Statistics. The


chi-square distribution is used (i) to test the hypothesis that whether the
population variance is same as the specified value or not in parametric test
procedures, (ii) to test the goodness of fit, that is, to judge whether there is a
discrepancy between theoretical and experimental observations or not and
(iii) to test the independence of two attributes. The applications listed above
shall be discussed in detail in Unit 4 of this block.

CHECK YOUR PROGRESS 5

Note: i) Check your answers with those given at the end of the unit.

7) What are the mean and variance of chi-square distribution with 10


degrees of freedom?
8) What are the mean and variance of chi-square distribution with pdf
given below
1 2 / 2 2 3
f  2   e   ; 0  2  
96

9) List the applications of chi-square distribution.

9.6.2 Student’s t-Distribution


The t-distribution was discovered by W.S. Gosset in 1908. He was better
known by the pseudonym ‘Student’ and hence t-distribution is called
‘Student’s t-distribution’.
If a random sample X1, X2,…, Xn of size n is drawn from a normal
population having mean  and variance σ2 then we know that the sample
mean X is distributed normally with mean  and variance 2 / n , that is, if
Xi ~ N  , 2  then X ~ N , 2 / n  . Then as mentioned under sub-section
2.3.2, the variate
X 
Z
246 / n
is distributed normally with mean 0 and variance 1, that is, Z ~ N 0, 1 . Sampling Distributions

In general, the standard deviation σ is not known and in such a situation the
only alternative left is to estimate it from a sample. The value of sample
variance (S2) is used to estimate it where,

2 1 n1
s   (x i  x) 2
n  1 i1

X 
But then in this case the variate is not normally distributed whereas it
S/ n
follows t-distribution with (n−1) df, that is,
X 
t ~ t ( n 1) … (2)
s n

The t-variate is a widely used variable and its distribution is called student’s
t-distribution on the pseudonym name ‘Student’ of W.S. Gosset. The
probability density function of variable t with (n-1) = ν degrees of freedom is
given by
1
f  t   1 / 2 ;  t  … (3)
2
 1   t 
 B  ,  1  
 2 2  

1 
where, B ,  is known as beta function.
2 2
The probability curve of t-distribution is bell shaped and symmetric about t =
0 line. The probability curves of t-distribution is shown in Fig. 9.2 at two
different values of degrees of freedom n = 4 and 12.

Fig. 9.2: Probability curves for t-distribution at n = 4, 12 along with


standard normal curve
In the figure given above, we have drawn the probability curves of t-
distribution at two different values of degrees of freedom along with
probability curve of standard normal distribution. By looking at the figure,
one can easily understand that the probability curve of t-distribution is similar 247
Statistical Analysis in shape to that of normal distribution and asymptotic to the horizontal-axis
whereas it is flatter than standard normal curve. The probability curve of the
t-distribution is tending to the normal curve as the value of n increases.
Therefore, for sufficiently large value of sample size n, practically for(> 30),
the t-distribution tends to the normal distribution.
The mean and variance of the t-distribution with n df are given by
n
Mean = 0 and Variance = n>2
n2

Now, we shall discuss some of the important properties of the t-


[Link] t-distribution has the following properties:
1. The t-distribution is a uni-modal distribution, that is, t-distribution has
single mode.
n
2. The mean and variance of the t-distribution with n df are zero and
n2
respectively. The variance of the distribution exists only if n > 2.
3. The probability curve of t-distribution is similar in shape to the standard
normal distribution and is symmetric about t = 0 line but flatter than
normal curve.
4. The probability curve is bell shaped and asymptotic to the horizontal
axis.
Above we have discussed some important properties of t-distribution without
mentioning their proof. You may be interested to know the applications of the
t-distribution. The t-distribution has wide number of applications in Statistics.
The t-distribution is used (i) to test the hypothesis about the significance of a
population mean, (ii) the the hypothesis of equality of two population means
of two normal populations and (iii) to test the hypothesis whether the
population correlation coefficient is zero or not.
Example 9: The life of light bulbs manufactured by the company A is known
to be normally distributed. The CEO of the company claims that an average
life time of the light bulbs is 300 days. A researcher randomly selects 25
bulbs for testing the life time and he observed the average life time of the
sampled bulbs is 290 days with standard deviation of 50 days. Calculate
value of t-variate.
Solution: Here, we are given that
  300, N  25, X  290 and s  50
The value of t-variate can be calculated by the formula
X 
t
s/ n
Therefore, we have
290  300 10
t   1
248 50/ 25 10
CHECK YOUR PROGRESS 6 Sampling Distributions

Note: i) Check your answers with those given at the end of the unit.

10) The scores on an IQ test of the students of a class are assumed to be


normally distributed with a mean of 60. From the class, 15 students are
randomly selected and an IQ test of the similar level is conducted. The
average test score and standard deviation of test scores in the sample
group are found to be 65 and 12 respectively. Calculate the value of t-
statistic.
11) What are the mean and variance of t-distribution with 8 degrees of
freedom?
12) Write any three applications of t-distribution.

9.6.3 F-Distribution
As we have mentioned in previous unit, F-distribution was introduced by
Prof. R. A. Fisher and defined as the ratio of two independent chi-square
variates when divided by their respective degrees of freedom. If we draw a
random sample X1 ,X 2 ,..., X n1 of size n1 from a normal population with mean
1 and variance σ 12 and another independent random sample Y1 , Y2 ,..., Yn2 of
size n2 from another normal population with mean 2 and variance 22
respectively then 1s12 / 12 is distributed as chi-square variate with ν1 df, that
1s12
is, 12  2
~  (21 ) … (1)
1

1 n1 1 n1
where, 1  n 1  1, X   X i and s12   (X i  X ) 2
n1 i1 n1  1 i1

Similarly,  2s 22 /  22 is distributed as chi-square variate with ν 2 df, that is,

 2s 22
 22  2
~  (2 2 ) … (2)
1
n2
1 1 n2
where,  2  n 2  1, Y 
n2
 X i and s 22 
i 1
 (Yi  Y) 2
n 2  1 i1
Now, if we take the ratio of the above chi-square variates given in equations
(1) and (2), then we get
12 1s12 / 12 s12 / 12 12 / 1
   ~ F( 1 , 2 ) … (3)
 22  2s 22 /  22 s 22 /  22  22 /  2

In the above expression F stands for the F-distribution. In the suffix, υ1 and υ2
are called the degrees of freedom of the F-distribution.
Now, if variances of both the populations are equal, that is, σ12  σ 22 , then F-
variate is written in the form of ratio of two sample variances which is as
follows:
249
Statistical Analysis s12
F ~ F( 1 , 2 ) … (4)
s 22

The probability density function of F-variate is given as


ν1 / 2
F
ν1 / 2  1
f(F) 
 ν1/ν 2  ; 0F … (5)
 ν1  ν2  / 2
ν ν 
B  1 , 2   1  ν1 F 
 2 2  ν 2 

As shown in (5), F-variate varies from 0 to , therefore, it is always positive
so probability curve of F-distribution wholly lies in the first quadrant. ν1 and
ν 2 are the degrees of freedom and are called the parameters of F-distribution.
Hence, the shape of probability curve of F-distribution depends on ν1 and ν 2 .
Probability curves of F-distribution taking ( 1 , 2 ) as (5, 5), (5, 20) and (20,
5) are shown in Fig.2.3 below:

 1  5 ,  2  20 
 1  5 ,  2  5 
 1  20 ,  2  5 

Fig. 9.3: Probability curves of F-distribution for (5, 5), (5, 20) and (20, 5)
degrees of freedom.
As it appears from the figure, F-distribution is uni-modal curve. It can be
seen from the figure that by increasing the first degrees of freedom from
1  5 to 1  20 the mean of the distribution (shown by vertical line) does
not change but probability curve shifs from the tail to the centre of the
distribution whereas increasing the second degrees of freedom from
 2  5 to  2  20 the mean of the distribution (shown by vertical line)
decrease and the probability curve shifts from the tail to the centre of the
distribution. One can also get an idea about the skewness of the F-
distribution. We observe from the probability curve that it is positively
skewed curve and it becomes very highly positive skewed if ν 2 becomes
small. Now, we shall discuss some of the important properties of F-
distribution.
The F-distribution has the following important properties:
250
1. The probability curve of F-distribution is positively skewed curve. The Sampling Distributions
curve becomes highly positive skewed when ν2 is smaller than ν1.
2. F-distribution curve extends on abscissa from 0 to .
3. F-distribution is a uni-modal distribution, that is, it has single mode.
4. The square of t-variate with ν df follows F-distribution with 1 and ν
degrees of freedom.
2
5. The mean of F-distribution with (ν1,ν2) df is for  2  2.
2  2

6. The variance of F-distribution with (ν1,ν2) df is

222  1  2  2 
2
for  2  4.
1  2  2    2  4 

7. If we interchange the degrees of freedom ν1 and ν2 then there exists a


very useful relation as
1
F 1 , 2 ,1 
F 2 , 1 , 

F-distribution has a lot of applications in Statistics. It is used to test the


hypothesis of equality of variances of two normal populations, for the
significance of multiple correlation coefficients, correlation ratio in the
population. It is also used in one-way and two-way analysis of variance for
testing the equality of several means at a time.
Example 10: A statistician selects 7 women randomly from the population of
women, and 12 men from a population of men. The table given below shows
the standard deviation of each sample and population:
Population Sample
Population Standard Standard
Deviation Deviation

Women 40 45
Men 80 75

Compute the value of F-variate.


Solution: Here, we are given that
n 1  7, n 2  12, 1  40,  2  80, s1  45, s 2  75
The value of F-variate can be calculated by the formula given below

s12 / 12
F
s 22 /  22

where, s12 & s 22 are the sample variances.


251
Statistical Analysis Therefore, we have
2 2

F
 45 /  40 
1.27
 1.93
2 2
 75 /  80 0.88

For the above calculation, the degrees of freedom ν1 for women’s data
are7−1= 6 and the degrees of freedom ν2 for men’s data are 12 −1 =11.

CHECK YOUR PROGRESS 7

Note: i) Check your answers with those given at the end of the unit.

13) For the purpose of a survey 15 students are selected randomly from
class A and 10 students are selected randomly from class B. At the
stage of the analysis of the sample data, the following information is
available:

Class Population Sample


Standard Standard
Deviation Deviation
Class A 65 60
Class B 45 50

Calculate the value of F-variate.


14) Write any five properties of F-distribution.

15) What are the mean and variance of F-distribution with 1  5 and
 2  12 degrees of freedom?
16) Write four applications of F-distribution.

We now end this unit by giving a summary of what we have covered in it.

9.7 LET US SUM UP


In this unit, we have covered the following points:
1. The statistical procedure which is used for drawing conclusions about
the population parameter on the basis of the sample data is called
“statistical inference”.
2. A group of units or items under study is known as “Population”
whereas a part or a fraction of it is known as “sample”.
3. A “parameter” is a function of population values which is used to
represent certain characteristics of the population and any quantity
calculated from sample values does not contain any unknown population
parameter is known as “statistic”.
4. Any statistic used to estimate an unknown population parameter is
known as “estimator” and the particular value of the estimator is known
as “estimate”.
252
5. The probability distribution of any statistic is called “sampling Sampling Distributions
distribution”of that statistic.
6. The standard deviation of the sampling distribution of a statistic is
known as “standard error”.
7. The most important theorem of Statistics is “central limit theorem”
which states that the sampling distribution of the sample mean tends to
normal distribution as sample size n tends to large (generally when n >
30).
8. The sampling distribution of sample mean and difference between two
sample means.
9. The sampling distribution of sample proportion and difference of two
sample proportions.
10. The properties, probability curve and applications of χ2, t and F
distributions, and

11. Mean and variance of χ2, t and F distributions.

9.8 KEY WORDS


Standard Error of the : A rough measure of the average amount by
Mean which sample means deviate from the
population mean.

Sampling Distribution : The probability distribution of means for all


of the Mean possible random samples of a given size from
some population.

9.9 SUGGESTED FURTHER READING/


REFERENCES
Witte, R., & Witte, J. (2017). Statistics. Hoboken, NJ: John Wiley & Sons.

9.10 ANSWERS TO CHECK YOUR PROGRESS


1) Here, we are given that
N=5, n=2
Since we have to estimate the average life of Televisions on the basis of
samples of size n = 2 therefore, all possible samples (with replacement) are
N
C N  5 C 2  10 and for each sample we calculate the sample meanas shown
in Table 2.8 given below:

253
Statistical Analysis Table 9.8: Calculation of Sample Mean
Sample Sample Sample
Number Observatio Mean ( X )
n

1 4, 6 5
2 4, 8 6
3 4,10 7
4 4, 12 8
5 6, 8 7
6 6, 10 8
7 6, 12 9
8 8, 10 9
9 8, 12 10
10 10, 12 11

Since the arrangement of all possible values of sample mean with their
corresponding probabilities is called the sampling distribution of mean,
thus,we arrange every possible value of sample mean with their respective
probabilities in the following Table 9.9 given below:
Table 9.9: Sampling distribution of sample means

[Link] Sample Frequency Probability


. Mean ( X )
1 5 1 1/10
2 6 1 1/10
3 7 2 2/10
4 8 2 2/10
5 9 2 2/10
6 10 1 1/10
7 11 1 1/10

Here, we are given that

 = 200, σ = 4, n = 10

2) Since parent population is normal so sampling distribution of sample


means is also normal. Therefore, the mean of this

distribution is given by

E  X     200

254
and variance Sampling Distributions

2
2  4  16
Var  X      1.6
n 10 10

3) Here, we are given that

 = 2550, n = 100, s = 54

First of all, we find the sampling distribution of sample mean. Since sample
size is large (n = 100 > 30) therefore, by the central limit theorem, the
sampling distribution of sample mean follows normal distribution. Therefore,
the mean of this distribution is given by

E  X     2550

and variance

s 2 (54) 2 2916
Var ( X )     29.16
n 100 100

4) Here, we are given that

1 = 68, σ1 = 2.3, n1 = 35

2 = 65, σ2 = 2.5, n2 = 50

To find the mean and standard error, first of all we find the sampling
distribution of difference of two sample means. Let X and Y denote the
mean height of male and female workers of hospital, respectively. Since n1
and n2 are large (n1, n2 > 30) therefore, by the central limit theorem, the
sampling distribution of (X  Y) follows normal distribution with mean

E  X  Y   1   2  68  65  3

and variance
2 2
12 22 2.3 2.5
Var  X  Y     
n1 n 2 35 50

 0.1511  0.1250  0.2761

Thus standard Error = Var ( X  Y )

 0.2761  0.525

5) Here, we are given that

X 7000
N =10000, X = 7000  P    0.70 & n = 100
N 10000 255
Statistical Analysis First of all, we find the sampling distribution of sample proportion. Here, the
sample proportion is not given and n is large so we can assume that the
conditions of normality hold. So the sampling distribution is approximately
normally distributed with mean

E (p)  P  0.70

and variance

PQ 0.70  0.30
Var(p)    0.0021  Q  1  P
n 100

Thus, standard error is

standard error = (p)  Var(p)  0.0021  0.0458

6) Here, we are given that

25 20
P1   0.25, P2   0.20, n 1  250, n 2  200
100 100

Let p1 and p2be the sample proportions of alcohol drinkers in two cities A and
B respectively. Here the sample proportions are not given and n1 and n2 are
large  n1 , n 2  30  so we can assume that conditions of normality hold. So the
sampling distribution of difference of proportions is approximately normally
distributed with mean

E(p1  p 2 )  P1  P2  0.25  0.20  0.05

and variance

P1Q1 P2 Q 2 0.25  0.75 0.20  0.80


Varp1  p 2     
n1 n2 250 200

0.75 0.80 1.55


    0.00155
1000 1000 1000

Hence, Standard Error (p1-p2) =0.0394


7) We know that the mean and variance of chi-square distribution with n
degrees of freedomare

Mean = n and Variance = 2n


In our case, n = 10, therefore,
Mean = 10 and Variance = 20
8) Here, we are given that

1 2 / 2 2 3
f  2   e   ; 0  2  
96
256
We have the probability density function of  2  distribution as: Sampling Distributions

1 2
f ( 2 )  e  /2
( 2 ) (  / 2)1 ; 0  2  

2 / 2
2

Where v = n = 1
by comparision we have

 
  1  3   4    8
2 2

Thus,   8  n  1  8  n  9
Mean = n = 9 and Variance = 2n = 18.
9) Refer Sub-Section 9.6.1.
10) Here, we are given that

  60, n  15, X  65 and s  12


We know that the t-variate is
X 
t
s/ n
Therefore, we have
65  60 5
t   1.14
12 / 15 4.39
11) We know that the mean and variance of t-distribution with n degrees
of freedomare

n
Mean = 0 and Variance  ; n2
 n  2
In our case, n = 8, therefore,

8 8
Mean = 0 and Variance    0.8
(80  2) 10

12) Refer Sub-Section 9.6.2

13) Here, we are given that

n 1  15, n 2  10, 1  65,  2  45, s1  60, s 2  50


Thevalue of F-variate can be calculated as follows:

s12 / 12
F
s 22 /  22

where, s12 & s 22 are the values of sample variances.

257
Statistical Analysis Therefore, we have
2 2

F
 60 /  65 
0.85
 0.69
2 2
50 /  45 1.23

14) Refer Sub-Section 9.6.3.

15) We know that the mean and variance of F-distributionwith 1 and 2


degrees of freedom are

2
Mean  for  2  2
2  2

and

2 22  1   2  2 
Variance  2
for 2  4.
1  2  2    2  4 

In our case, 1  5 and  2  12, therefore,

2 12
Mean    1 .2
 2  2 10

2(12) 2 (5  12  2) 30 144
Variance    10.8
5(12  2) 2 (12  4) 40  100

16) Same as Sub-section 9.6.3.

258

You might also like