Statistics Module 2
Statistics Module 2
Learning Objectives
When you have completed this chapter, you will be able to:
Define sample, sampling and sampling distribution.
Understand and identify the different types of sampling
techniques in statistics.
Understand and calculate the sampling distribution of the
different types of sample Statistics.
Use and Understand the central Limit Theorem.
5.1 Introduction
In statistics, Sampling plays vital role. Because most of the time we are occupied
with a class of problems that involve an attempt to say something about the
properties of a large group of objects, given information on a relatively small subset
of them. This is done, especially in Least Developed Countries (LDCs), because
there is no resource to undertake census or complete enumeration there. So the
major motivation for examining a sample rather than the whole population is that
the collection of complete information on the latter would typically be prohibitively
expensive. Even in circumstances where sufficient resources are apparently
available to contact the whole population, it may be preferable to devote these
resources to just a subset of the population in the hope that such a concentration
of effort will produce more accurate measurements. But if we take a sample from a
population, the eventual aim is to make statements that have some validity for the
population at large. Therefore, it is important that the sample be representative of
the whole population.
Generally, the overall purpose of this chapter is to introduction and equip students
with the concepts of sample, sampling and sampling distribution.
Check yourself
1. Why sampling?
2. What is the difference and similarity between Sampling
And Census (complete enumeration).
If we are convinced that the sample statistics are accurate estimate of the
population characteristics, we could use sample statistics to estimate the
population parameter without measuring the entirety of the items under study.
In order to be consistent, Statisticians use lower case Roman letters to denote
sample statistics, and Greek or capital letters to denote population parameters.
Population Sample
Definition Collection of all items being dealt in Part (sub-set) of the
a study population
Characteristic Population Parameter Sample Statistics
Symbol • Population size = N • Sample size = n
• Population mean = µ • Sample mean = x
• Population variance = σ2 • Sample variance = S2
• Population standard • Sample standard
deviation = σ deviation = S
Example:
Suppose that a restaurant has four branches (N,S, E and W) and that it wants to
select samples of two branches at a time in order to evaluate the operation of the
branches. Using simple random sampling, six different samples of size 2 that can be
drawn from the population (i.e., the four branches). These six samples are (NS);
(NE); (NW); (SE); (SW); and (EW).
The probability of each sample is 1/6 to be selected from the population and the
probability of an element in the sample is ½.
Some Definitions:
Finite population: means that the population has limited size,
that is to say, there is a whole number (N) that tells us how many items are there
in the population.
N.B: The numbers of different possible samples of size n that can be drown
without replacement from a population of N elements equals:
Illustration:- Suppose that there are 1000 resident or households in one Kebelle
with different income levels. If the Statistician /researcher has the list of all
households randomly listed and wants to study the income disparity in that kebelle
by taking 50 samples. Since there are 1000 households the sampling can be
accomplished by taking every 20th household on the list. To determine which of
the first 20 element to begin with the statician/researcher can randomly chose a
number from 1 to 20 Once this number is chosen (let’s say 3), then the statician
selects the 3rd, 23rd, 33rd, 43rd, … households from the list. Such kind of sampling is
systematic sampling.
Illustration: Still taking the study of the income disparity condition in Mekelle. In
this case, the Mekelle city will be classified by locality (i.e., in to Northern, southern
part of Mekelle, etc). Once the city is classified in to various clusters, randomly
some of the clusters (i.e., locality in our case) will be chosen and the researcher can
take all elements in the cluster or randomly selects elements from the chosen
cluster. This depends on the cost and other considerations.
Most of the time students face difficulty in differentiating stratified and cluster
sampling. The main distinguishing criteria of stratified from cluster sampling is that
in the case of stratified sampling the population is divided in to well-defined groups,
where each group has homogeneity with in itself but wider heterogeneity (or
variation) among the groups. In the case of cluster sampling, the situation is the
reveres for stratified sampling (i.e., the different clusters are homogeneous but
elements in each cluster are heterogeneous).
o In statistical inference the assumption is that the samples are selected
using simple random sampling (other probability sampling technique
attempt to approximate the simple random sampling).
o Very few so-called random samples are truly random. Why?
5. Multi-Stage Sampling:- The four methods we have covered so far viz. simple,
stratified, systematic and cluster, are the simplest random sampling strategies. In
most real applied social research, we would use sampling methods that are
considerably more complex than these simple variations. The most important
principle here is that we can combine the simple methods described earlier in a
variety of useful ways that help us address our sampling needs in the most efficient
and effective manner possible. When we combine sampling methods, we call this
multi-stage sampling.
For example, consider the idea of sampling Amhara region residents for face-to-
face interviews. Clearly we would want to do some type of cluster sampling as the
first stage of the process. We might sample townships or census tracts throughout
So, we might set up a stratified sampling process within the clusters. In this case,
we would have a two-stage sampling process with stratified samples within cluster
samples. Or, consider the problem of sampling students in grade schools. We might
begin with a national sample of school districts stratified by economic status and
educational level. Within selected districts, we might do a simple random sample of
schools.
All of the methods that follow can be considered subcategories of purposive sampling methods.
We might sample for specific groups or types of people as in modal instance, expert, or quota
sampling. We might sample for diversity as in heterogeneity sampling. Or, we might capitalize on
informal social networks to identify specific respondents who are hard to locate otherwise, as in
snowball sampling. In all of these methods we know what we want -- we are sampling with a
purpose.
3. Modal Instance Sampling: - In statistics, the mode is the most frequently occurring value in
a distribution. In sampling, when we do a modal instance sample, we are sampling the most
frequent case, or the "typical" case. In a lot of informal public opinion polls, for instance, they
interview a "typical" voter.
There are a number of problems with this sampling approach. First, how do we know what the
"typical" or "modal" case is? We could say that the modal voter is a person who is of average age,
educational level, and income in the population. But, it's not clear that using the averages of these
is the fairest (consider the skewed distribution of income, for instance). And, how do you know
that those three variables -- age, education, income -- are the only or event the most relevant for
classifying the typical voter? What if religion or ethnicity is an important discriminator? Clearly,
modal instance sampling is only sensible for informal sampling contexts.
4. Expert Sampling: - Expert sampling involves the assembling of a sample of persons with
known or demonstrable experience and expertise in some area. Often, we convene such a
sample under the support of a "panel of experts." There are actually two reasons you might do
For instance, let's say you do modal instance sampling and are concerned that the criteria you
used for defining the modal instance are subject to criticism. You might convene an expert panel
consisting of persons with acknowledged experience and insight into that field or topic and ask
them to examine your modal definitions and comment on their appropriateness and validity. The
advantage of doing this is that you aren't out on your own trying to defend your decisions i.e. you
have some acknowledged experts to back you. The disadvantage is that even the experts can be,
and often are, wrong.
5 .Quota sampling:-In quota sampling, you select samples non-randomly according to some
fixed quota. There are two types of quota sampling: proportional and non proportional. In
proportional quota sampling you want to represent the major characteristics of the population by
sampling a proportional amount of each. For instance, if you know the population has 40%
women and 60% men, and that you want a total sample size of 100, you will continue sampling
until you get those percentages and then you will stop.
So, if you have already got the 40 women for your sample, but not the sixty men, you will
continue to sample men but even if legitimate women respondents come along, you will not
sample them because you have already "met your quota." The problem here (as in much
purposive sampling) is that you have to decide the specific characteristics on which you will base
the quota will it be by gender, age, education race, religion, etc.
Non-proportional quota sampling is a bit less restrictive. In this method, you specify the minimum
number of sampled units you want in each category. Here, you're not concerned with having
numbers that match the proportions in the population. Another term for this is sampling for
diversity. In many brainstorming or nominal group processes (including concept mapping), we
would use some form of heterogeneity sampling because our primary interest is in getting broad
spectrum of ideas, not identifying the "average" or "modal instance" ones.
In effect, what we would like to be sampling is not people, but ideas. We imagine that there is a
universe of all possible ideas relevant to some topic and that we want to sample this population,
not the population of people who have the ideas. Clearly, in order to get all of the ideas, and
especially the "outlier" or unusual ones, we have to include a broad and diverse range of
participants. Heterogeneity sampling is, in this sense, almost the opposite of modal instance
sampling.
6. Snow Ball Sampling:-In snowball sampling, you begin by identifying someone who meets the
criteria for inclusion in your study. You then ask them to recommend others who they may know
who also meet the criteria. Although this method would hardly lead to representative samples,
there are times when it may be the best method available. Snowball sampling is especially useful
when you are trying to reach populations that are inaccessible or hard to find. For instance, if you
are studying the homeless, you are not likely to be able to find good lists of homeless people
3, 6, 15 8 µ=
∑x =
45
= 9
N 5
3, 9, 12 8
3, 9, 15 9 The population mean value ( µ ) varies from
3, 12, 15 10 some of the sample means ( x i ). This
6, 9, 12 9 leads us in to concept of sampling
6, 9, 15 10 distribution.
6, 12, 15 11
9, 12, 15 12
∑x= 90
The above table is sampling distribution of the mean. In this section our objective is to describe
the characteristics of the sampling distribution of the mean and the shape of the sampling
distribution.
The concept of sampling distribution is helpful in letting us to make probability statements about
the error involved when the sample mean ( x ) is used to estimate the population mean (µ). That
is, the practical value of the sampling distribution of the mean can be used to provide probability
information about the sampling ERROR
δ
δx = - - - - - - - - - For infinite population.
n
δ N −n
δx = -------- For finite population.
n N −1
With Simple random Sampling the value of Standard deviation of the mean depends on
whether the population is finite or infinite.
3. The sampling distribution of the mean is normally distributed regardless of the population from
which it is drawn
The aforementioned are some of the basic common properties of the sampling distribution the
mean. Next one examines the shape of the sampling distribution of the mean, which is already
stated as the third characteristics of the sampling distribution of the mean, when the population
from which the samples drawn is normal or non-normal.
Table 5.2 Foreign Direct Investment (FDI) of East Asian Countries in millions of USD
Country Philippines Indonesia Taiwan [Link] Singapore
FDI in millions 3 3 7 9 14
of USD
Population Mean FDI in millions of USD ( µ) = 36/5 =7.2
Considering the data we agree that the population may not be normal, because there are only 5
elements involved in the population hence too small to be approximated by a normal distribution.
Let us draw samples of size 3, compute the sample means ( x ) and list them; and calculate the
mean of sampling distribution ( µ x ) . This is done and put in table below.
µx =
∑x i
=
72
= 7 .2 = µ
NCn 10
From the table we recognize that even in a case in which the population is not normal, the mean
of the sampling distribution ( µ x ) is still equal to the population mean.
Probability distribution figures.
Examples:
1. A population of 100 elements has a mean of 19.2 and standard deviation of [Link] is the
mean and standard deviation of the sampling distribution of the mean for samples of size
25?
Solution:
µ x = µ ⇒ µ x = µ = 19.2
δ N −n n 25
δx = , Since > 0.05 ⇒ > 0.05
n N −1 N 100
1 100 − 25 1 75
δx = . = = 0.174
25 100 − 1 5 99
Check Your Self
Interpret δ x = 0.174 of the above result.
Solution:
Given of µ = 320 δ = 75
Required: P (300 < x < 340)?
This is the case of sampling distribution of x . Since the population is normal, the
sampling distribution of x is also normal. So, using the normal probability distribution
way of computing probability, P (300 < x < 340) should be converted in to standard
normal probability distribution.
Given
δ 75 75
x 1 = 300 , x 2 = 340 , and δ x = = = = 13.6
n 30 5.5
3. The distribution of annual earnings of all economics graduates with zero year experience
is skewed negatively, as shown in figure (a) below. This distribution has a mean of 19,000
Birr and a standard deviation of 2,000 Birr. If we draw a random sample of 30 fresh
economic graduates, what is the probability that their earnings will average more than
19,750 Birr annually?
Solution
Given
µ = 19,000 σ = 2,000
Required: P( x > 19,750)?
X
µ = 19,000
(a
.
In order to answer this question, first let’s calculate the standard error of the mean ( x )
δ 2,000 2000
δx = = = 365.16 Birr
n 30 5.477
Next, let’s convert the random variable in to standard normal probability
value (Z)
Z= x - µ x = 19750
19,750 − 19000 750
δx Z= = = 2.05
365.16 365.16
Then, Z = 2.05 corresponds to area equal to 0.4798; we are interested to area above Z = 2.05.
This is obtained by taking the difference between the area to the right of Z = 0, which is 0.5 and
the area between Z = 0 and Z = 2.05 that is 0.4798. Thus, the area to the right of Z = 2.05 is
(0.5 – 0.4798) = 0.0202. Thus, P( x > 19,750) = 0.0202. This is the area to the right of Z= 2.05 of
the normal distribution graph shown below.
In this problem’s case we assumed that the sampling distribution of the mean is normal, using the
central limit the area since n = 30.
4. If the number of miles per gallon achieved by all cars of a particular Model has mean of 25 and
standard deviation of 2, what is the probability that, for a random sample of 20 such cars, average
miles per gallon will be less than 24? Assume that the population distribution is normal.
Solution:
Let x denote the sample mean. Then we need to find
24 − µ X
P( x < 24) = p[ Z < ]
σX
Sampling distribution of the proportion is the probability distribution of all possible values of the
sample proportion ( P ).
If necessitates the understanding of the properties of sampling distribution of the proportion
( P ): the mean value of ( P ), standard deviation of ( P ) and the shape or form of the sampling
distribution of ( P ).
Properties of the sampling distribution of the proportion ( P ).
1. The expected value of the sample proportion E( P ) is equal to the population
proportion, P.
Symbolically: E ( P ) = P
− 0.03
Z1 = = − 0.857
0.035
-0.857 0 0.86
∑
i = 1
( Xi − X ) 2
= ≤
2
S x
n − 1
, for n 30
or
∑
i = 1
( Xi − X ) 2
= >
2
S x
n
, for n 30
Definition
Let X1, X2, …, Xn be a random sample from a population. The quantity
n
n
∑ ( Xi
i =1
− X )2
= n ≤ 30
2
S x
n −1
, for
or
∑ ( Xi
i =1
− X )2
= n > 30
2
S x
n
, for
2 2
S is called the variance. [Again we distinguish between the random variable S x and
x
specific values it can take. Thus, if the actual sample observed is X1, X2, …, Xn, then
2
the realization of S x
is by using the above formula.
At the first, the use of (n – 1) rather than n as the divisor in our definition of the sample variance
may be rather surprising. The motivation for this formulation is that, if the sample variance is
defined in this way1, it can be shown that the mean of its sampling distribution is the true
population variance; that is,
E (Sx) = σ
2 2
x
∑ ( X i − X ) 2 = ∑ [( X i − µ X ) − ( X − µ X )] 2
i =1 i =1
n
= ∑ [( X i − µ X ) 2 − 2( X − µ X )( X i − µ X ) + ( X − µ X ) 2 ]
i =1
n n
= ∑ ( X i − µ X ) − 2( X − µ X ) ∑ ( X i − µ X ) + ∑ ( X − µ X ) 2
2
i =1 i =1
n
= ∑ ( X i − µ X ) 2 − 2n ( X − µ X ) 2 + n( X − µ X ) 2
i =1
n
= ∑ ( X i − µ X ) 2 − n( X − µ X )
i =1
Taking expectations then gives
n n
E[∑ ( X − X ) 2 ] = E[∑ ( X i − µ X ) 2 ] − nE[( X − µ X ) 2 ]
i =1 i =1
n
= ∑ E[( X i − µ X ) 2 − nE[( X − µ X ) 2 ]
i =1
Now, the expectation of each ( ( X i − µ X ) 2 is the population variance σ 2 X , and the expectation
of ( ( X − µ X ) 2 is the variance of the sample mean, i.e. σ 2 X n .Hence, we have
n
nσ 2 X
E[ [∑ ( X i − X ) 2 ] = nσ 2 X −
= (n − 1)σ 2 X
i =1 n
Finally, for the expected value of the sample variance, we have
1 n
∑
n − 1 i =1
(X i − X )2 ]
n
1
[ E [
E ( S x ) = E [ n − 1 i =1
2 ∑ (X i − X )2 ]
1
= .(n − 1)σ 2 X
n −1
=σ x
2
The conclusion that the expected value of the sample variance is the population variance is quite
general. However, in order to characterize further the sampling distribution we need to know
more about the underlying population distribution. In many practical applications the assumption
that the population distribution is normal is not unreasonable. In this case it can be shown that
the random variable
(n − 1) S 2 ∑ (X
i =1
i − X )2
= S x2 =
X
σ σ
2 2
X
x
has a distribution known as the χ distribution (Chi-Square distribution) with (n- 1) degrees of
2
freedom.(N.B: The chi-square distribution with v degrees of freedom is the distribution of the sum of
squares of v independent standard normal random variables.)
The chi-square family of distributions is frequently employed in statistical analysis. The
distributions are defined only for positive values of a random variable, which is appropriate in the
present context since a sample variance can not be negative. The density function, which is
illustrated in the figure below which is asymmetric. A specific member of the chi-square family is
characterized by a single parameter, referred to as the number of degrees of freedom, for
which the symbol v is typically used. If a random variable has a χ 2 distribution with v degrees of
χ
2
freedom, it will be denoted . The mean and variance of this distribution are equal to the
v
number of degrees of freedom and twice the number of degrees of freedom, respectively; that is,
Figure 5.2 probability functions of the chi-square distribution with v= 4, 6, and 8 degree
of freedom
χ χ
2 2
E( ) = vi Var ( ) = 2v
v v
χ
2
In the present context, the random variable (n – 1) S2x/ σ 2
x has (n-1) distribution, so that
v
(n − 1) S 2 x
its mean is E ( ) = n −1
σ 2X
So that
E ( S x2 ) = σ x2
2σ X
2
Var ( S x2 ) = = 2(n-1)
(n − 1)
2σ x2
So that Var ( S x2 ) =
(n − 1)
The properties of the X2 distribution can therefore be used to find the variance of the sampling
distribution of the sample variance. (It should be repeated that this result holds only when the parent
population is normal)
χ
2
The parameter v of the distribution is known as the number of degrees of freedom. To
v
understand this terminology, let us look at the sample variance. It involves the sum of the squares
of the quantities: ( X 1 − X ), ( X 2 − X ) K ( X n − X )
Thus, these n pieces of information are employed to calculate the sample variance. However,
they are not independent pieces of information, since they must sum to 0 (as follows from the
definition of X). Hence, if we know any (n – 1) of the ( X i − X ) , we can calculate the other one
from the first (n-1). For example, since
∑ (X
i =1
i − X) = 0
it follows that:
n −1
Xn - X = ∑ (X
i =1
i − X)
The n quantities (Xi – X ) are equivalent to a set of (n-1) independent pieces of information. The
situation can be thought of as follows: we want to make inference about the unknown σ 2x. If the
population mean µx were known, this inference could be based on the sum of squares of
(X1 – µx); (X2 – µx); …. ; (Xn – µx)
These quantities are independent of one another, and we would say that we have n degrees of
freedom for the estimation of σ 2x. However, since the unknown population mean must be
replaced in practice by its estimate X , one of these degrees of freedom is used up and we are left
with the equivalent of (n -1) independent observations for use in making inference about the
population variance. It is then said that (n- 1) degrees of freedom are available.
Or, equivalently,
P ( χ10 >K) = 0.10
2
The distribution function of the chi-square random variable is available in the Appendix of many
statistical books and from these tables these cutoff points can be read directly. For the χ10
2
random variable, it can be seen from a table that the P( χ10 >K) = 0.10, then K= 15.99. This
2
probability is shown as an area under the density function of the random variable in the following
Figure.
Figure: probability (0.9) that a chi-square random variable with 10 degrees of freedom is less than
15.99
Sampling distribution of the sample variance
2
Let S x denote the sample variance for a random sample of n observations from a
population with variance σ 2x. Then
has mean σ 2x; that is,
2
i. The sampling distribution of S x
E ( S x ) = σ 2x
2
ii. The variance of the sampling distribution of S2x depends on the underlying
population distribution. If that distribution is normal, then
2 2σ 4 X
Var ( S x ) =
n −1
iii. If the population distribution is normal, then
(n − 1) S 2 x
is distributed as χ 2 ( n −1)
σ 2
X
Example:
A manufacturer of canned peas is concerned that the mean weight of the product be close to the
advertised weight. In addition, he does not want too much variability in the weights of the cans of
peas; otherwise, a large proportion will differ markedly from the advertised weight. Assume that
the population distribution of weights is normal. If a random sample of twenty cans is checked,
find the numbers K1, and K2 such that
S 2x S 2x
P( 2 < K1) = 0.05; P ( 2 > K1) = 0.05;
σ X σ X
We have
S 2x (n − 1) S x2
0.05 = P ( 2 < K1) , P < (n − 1) K 1
σ X σx
2
= P[ χ ( n −1) < (n-1) K1]
2
Where n = 20 is the sample size and X2 (n-1) is a chi-square random variable with
(n – 1) = 19 degrees of freedom. Then
0.05 = P( χ 192 < 19K1) or 0.95 = P ( χ 192 > 19K1)
From χ 2 table, we therefore have
19K1 = 10.12
So that K1 = 0.533
The conclusion then is that the probability is 0.05 that the sample variance will be less than 53.3%
of the population variance.
We also require the number K2 such that
S 2x
0.533 = P 2 > K2)
σ X
Equivalently, we have
(n − 1) S x2
0.05 = P > (n − 1) K 2
σx
2
= P[ χ ( n −1) > (n-1) K2]
2
Figure: probability is 0.05 that a chi- square random variable with 19 degrees of freedom is less
than 10.12, and also that this random variable is bigger than 30.14
One of the principal objectives of statistical investigation is to make reasonable estimates. The
over all objective of this chapter is to describe how statisticians or professionals that use
statistical tools go about doing statistical estimates.
Introduction
Every one makes estimates. When you are ready to cross a street, you see a car approaching to
you on the street; you estimate the speed of any car that is approaching, the distance between
you and that car, and your own speed. Having made these quick estimates, you decide whether
to wait, walk, or run.
So far we have covered discussions on the concepts of probability theory and sampling
distribution that forms foundation for statistical inference. In this chapter we begin to explore the
possibility of making inferential statements about a population, based on the information
contained in a random sample.
Estimation: - is a method that enables us to estimate, with reasonable accuracy, the population
parameter. In reality, to calculate the population parameter is extremely difficult or an impossible
goal, hence we need to make estimates and some times make a statement about the error that
will likely accompany the estimate.
The procedure of marking estimation is to have random sample of size n from the know
probability distribution, compute sample statistics and use it as an estimate of the population
parameters. This procedure of estimation can be categorized in to two: Point Estimation and
Interval Estimation.
∑X
n =1
i
X= Where ( X ) is sample mean
n
Σ is summation
Xi values of random variables
n is sample size
The sample variance and standard deviation
∑ (X
n =1
i − X )2
S2 = ------- Sample variance formula
n −1
n
∑ (X
n =1
i − X )2
S= ------ Sample standard deviation formula
n −1
The Sample Proportion P
It is calculated by taking elements in the sample that have the same characteristics; we can use
P this estimator of P
Example
Price- earnings ratios for a random sample of ten stocks traded on the Addis Ababa stock
Exchange on December 27 2006 were:
10; 16; 5; 10; 12; 8; 4; 6; 5; 4
Find point estimates of the population mean, variance and standard deviation and of the
proportion of stocks in the population for which the price – earning ratio exceeded 8.5.
Solution
To find the first three of these sample quantities, we show the calculations in tabular form;
i Xi X 2i
1 10 100
2 16 256
3 5 25
4 10 100
5 12 144
6 8 64
7 4 16
8 6 36
9 5 25
10 4 16
Sum 80 782
We then have
n =10; ∑ X i = 80 ; ∑ X i = 782
2
n −1
Interval Estimation
An Interval Estimate is a range of values used to estimate a population parameter. It indicates the
error in two ways: By the extent of its range and the probability of the true population
parameter lying with in that range. Instead of relying on the point estimate alone, we may
construct an interval around the point estimator, say with in two or three standard error of the
mean on either side of the point estimator such that this interval has, for instance 0.95 probability
of including the true parameter value.
Assume that we want to find out how “close” is and estimator X to µ. For this purpose, we try
to find out two positive values U and X such that the probability that the random variable
(X − U and X + U ) contains the true µ is 1 - α. i.e.
P( X − U ≤ µ ≤ X + U ) = 1 − α
0 ≤ α ≤ 1 and 0 ≤ 1 − α ≤ 1
Where: α is the level of significance
1 – α: is confidence coefficient
Such an interval is known as confidence Interval. The random variables ( X − U ) and
( X + U ) are the lower and the upper confidence limit (or critical values) respectively.
Where Z α is the value of standard normal variable that is exceeded with a probability of α /2
2
or Z α is Z value providing an area of α/2 in the upper tail of the standard normal probability
2
distribution.
Example 1:
A normal infinite population has a standard deviation of 10. A random sample of size 25 has a
mean of 50. Construct a 95% confidence interval of the population mean?
Solution:
Given δ = 10 n = 25
χ = 50 95% confidence interval.
Required: X − Z α δ X ≤ µ X
≤ X + Zα δ X ?
2 2
1st find the standard error of the mean i.e.,
δ 10 10
δx = = =2
n 5 5
Since the population is infinite.
2 We have that α = level of significance = 0.05
nd
Confidence α α /2 Z α /2
Level
Solution:
Given χ = 24,000 Birr n = 250
δ = 5,000 Birr
a) α = 0.05
σ 5000
1 − α = 0.95, σ X = =
n 250
Z α = Z 0.025 = 1.96
2
S =
∑ ( Xi − χ ) 2
n −1
Thus, the standard deviation of the sampling distribution of the sample means,
δ X ------- is given by:
S
δX = -------- For infinite population
n
In this case, the construction of confidence interval estimate depends up on whether the sample
size is larger or small:
Case 1. When the sample size is large and unknown δ
(A sample size is large when n ≥ 30)
Confidence interval estimate for population mean (µ) is given by:
S
X ± Zx/2
n
Case 2: when the sample size is small and unknown δ (A sample size is small when
n < 30).
The confidence interval estimate of µ is:
S S S
χ ± Zx/2 , χ − tx/2 ≤ µ χ + tx/2 .
n n n
Degrees of freedom: - are the number of values we can choose freely. Assume that we are
dealing with two sample values a and b, and we know that they have a mean of 18, symbolically,
the situation is
A+ B
= 18
2
How can we find what values a and b take on this situation? The answer is that a and b can be
any two values whose sum is 36, because 36 ÷ 2 = 18. Suppose we learn that a has a value of 10.
Then, b will have the value of 26, because if a = 10, then 10 + b = 18 10 + b = 36, ∴ b = 26
This example shows that when there are two elements in a sample and we know the sample
mean of these two elements, we are free to specify only one of the elements, because the other
element will be determined by the fact that the two elements sum to twice the sample mean.
Staticians say, “We have one degree of freedom.”
In our context, the degree of freedom is related to the sample standard deviation. There are n-
values of X i − X , involved in computing ∑ ( X i − X ) 2 , which are X 1 − X , X 2 − X , ……
X n − X , and we know that ∑ ( X i − X ) = 0 for any data set.
Therefore, if we known n – 1 values of, the remaining value can be determined exactly by using
the condition that the sum of X − X value must be zero. Thus, there are n – 1 values degrees
of freedom or n – 1 values that can independently determine. This n – 1 is associated with
∑ (X i − X )2.
Example 1:- In the testing of a new production method, 18 employees were selected randomly
and asked to try the new method. The sample mean production rate for the 18 employees was
80 parts per hour and the sample standard deviation was 10 parts per hour. Provide 90% and
95% confidence intervals for the population mean production rate for the new method, assuming
the population has a normal probability distribution.
Solution: - Given n = 10 χ = 80
S = 10 construct CI for 90% and 95% CI?
S 10
δX = = = 2.36
n 18
At 10% level of significance
t x / 2 = t 0.05 = 6.314
µ = X ± tx/ 2 δ X
µ = 80 ± 6.314( 2.36)
65.1 ≤ µ ≤ 94.9
n = x/2
E
Where E is the maximum sampling error at some level of precision (1 – α ).
Example: The CSA of Ethiopia has past data that indicate the interview time for a consumer
opinion study has standard deviation of 6 minutes.
(a) How large a sample should be taken if the authority desires a 98% level of precision
that the mean interval time to be with in 2 minutes or less?
(b) Assume that the sample size recommended in (a) above is taken and that the mean
interview time for the sample is 32 minutes. What is the 98% confidence interval
estimate for the mean interview time for the population?
Solution: Given: δ = 6 confidence level = 0.98 = 1 - α
Level of significance = α = 0.02
6 6
32 − 2.32 ≤ µ ≤ 32 + 2.32 ⋅
49 49
32 - 2 ≤ µ ≤ 32 + 2
30 ≤ µ ≤ 34
P ± Z α /2 δ P , δ P = S P
P (1 − P )
P - Z α /2 S P ≤ P ≤ P + Zx/2 S P
n
P (1 − P ) P (1 − P )
P - Z α /2 ≤ P ≤ P + Z α /2
n n
Examples
1. A survey conducted by Ethiopia Economic Policy Research Institute (EEPRI) shows that 47% of
all investors included in the sample invested in agricultural sector.
Compute a 95% confidence interval estimate for the proportion of the population investors who
invested in agricultural sector assuming that a sample of size 250 investors was used in the study?
0.4081 ≤ P ≤ 0.5319
Solution:
Given: P = 0.47 n = 250
.47(0.53) .47(0.53)
Thus, 0.47 – 1.96 ⋅ ≤ P ≤ 0.47 + 1.96 ⋅
250 250
0.47 – (1.96 x 0.0316) ≤ P ≤ 0.47 + (1.96 x 0.0316)
Interpretation: of all the investors taken as a sample, the proportion of the investors that
invest in agricultural sectors varies between 40% and 53%.
2. A random sample of 400 members of the labour force in four regional states of Ethiopia
showed 32 week unemployed. Construct a 95% confidence interval for the proportion of
unemployed labor force in there four regions.
Solution:
32
Given: P = = 0.08
400
Interpretation
The unemployed labour force with a confidence interval of 95% varies between 5.3% and 10.67%.
The values Z α S P in confidence interval for the population proportion is the maximum sampling
2
error involved with , α ,a specified level of precision, and let represent it by E.
E = Zα SP
2
P (1 − P
E=
n
P (1 − P )
E = Zα
2
n
P (1 − P )
E 2 = Z 2α
2 n
P (1 − P )
n = Z 2α
2
E2
Note: - In most cases, the desired maximum sampling error (E) is given as 0.01 or less.
In some statistics text books, the sample size is determined by the following way:
Example 1. Suppose we want to estimate a population proportion with ± 0.04 and a confidence
coefficient of 90%. What sample size should be taken when P = 0.5.
Solution:
Z α = Z 0.05 = 1.64
Given E = 0.04 2
1 − P = 0 .5
P = 0.5
p (1 − P ) (1.64) 2 (0.5)(0.5)
n = Z 2α . = = 420.25 = 421
2 n (0.04) 2
Check yourself
How large a sample of investors should be used in example 1 above if it is desired to be at 95%
confidence level that the sampling error is 5% or less?
As another example, a farmer may consider the use of the two alternative fertilizers, his interest
being in the difference between the resulting mean crop yields per acre. In order to compare
population mean, a random sample is drawn from the two populations, and inference a bout the
difference between the population means is based on the population results. The appropriate
method for analyzing this information depend on the procedure used in selecting the samples.
We will consider the following two very common sampling schemes:
(a) MATCHED PAIRS: in this scheme, the sample members are chosen in pairs, one from each
population. The idea is that, a part from the factor under study, the members of these
pairs should resemble one another as closely as possible, so that the comparison of
interest can be directly made. For instance, suppose we want to measure the effectiveness
of the speed-reading course. One possible approach would be to record the number of
words per minute read by sample of student before taking the course ,and compare with
the result for the same students after completing the course, and compare with results
X-cars Y- Differences
cars
I xi yi di d 2i
nd
n − 1 i =1 i
1
[10.52 − (8)(0.775)2]0.816
=
7
t n −1.a Sd t n −1 , a S d
d+ 2 2
< µ X − µY < d +
n n
Where tn-1,α/2= is that number for which
a
P(tn-1>tn,a/2)= 2
and the random variable tn-1 has a student’s distribution with (n-1) degrees of freedom.
t n −1.a Sd t n −1 , a S d
d+ 2
< µ X − µY d + 2
n n
We therefore find that, based on the data of the above table, a 99% confidence interval for the
difference in population mean miles per gallon for these two types of automobile ranges from -
0.342 to [Link] the interval includes 0, the sample evidence against the conjecture that the
population means are the same in not very strong.
As a first step, we examine the situation where the two population distributions are normal with
known variances. Since the object of interest is the difference between the two population
means, it is natural to base inference on the difference between the corresponding sample means.
This random variable has mean
E ( X - Y ) = E ( X ) – E ( Y ) = µ X + µY
and, since the samples are independent, variance
σ 2X σ 2Y
+
Var( X - Y ) = Var( X ) +Var ( Y ) = n X nY
Furthermore, it can be shown that its distribution is normal. It therefore follows that the
random variable
( X + Y ) − (µ X − µ Y )
Z=
σ 2X σ 2Y
+
nX nY
Has a standard normal distribution. An argument parallel to that of confidence interval
computation for the mean of a normal population can then be used to obtain confidence
intervals for the difference between the population means. Since this interval requires
knowledge of true population variance, it is rarely of much direct use.
However, as was the case in confidence interval computation for the mean of a normal
population its range of applicability is greatly extended when the sample sizes are large, as
indicated in the box.
observed sample means are X and Y , then a 100 (1- α ) %confidence interval for (x-y) is
given by
σ 2X σ 2Y σ 2X σ2
( X − Y ) − Za + < µ X − µY < ( X − Y ) + Z a +
2 nX nY 2 nX nY
Za
Where is that number for which
2
a
Za
P(Z> 2 ) = 2 and the random variable Z has a standard normal distribution.
If the sample sizes nX and nY are large, then to a good approximation, a 100(1- α )%
confidence interval for ( µ X - µY ) is obtained by replacing the population variances in the
previous expression by the corresponding observed sample variances S2x and S2y. For large
sample sizes, this approximation will typically remain adequate even if the
population distributions are not normal.
Example1. A professor, teaching two sections of Statistics for Economists course, organized
quizzes differently in the two sections. In one section quizzes, based on pre-assigned reading,
were given on the first day of discussion of a topic. In the other section quizzes were given after
each chapter was completed. A common examination was set for the two sections, which each
contained 40
students. For the first group, the mean score was 143.7, and the standard deviation was 21.2,
while for the second the mean and standard deviation were 131.7 and 20.9.
If we can regard these students as independent random samples from the populations of all
students who might be exposed to these two approaches to quizzes, find a 99% confidence
interval for the difference between the population mean scores.
Solution:
For the pre-quiz group we have
X = 143.7; nx = 40; S2x = (21.2)2 = 449.44
and for the post-quiz group
Y = 131.7; ny = 40; S2y = (20.9)2 = 436.81
Since the sample sizes are quite large, we can use the sample variances in place of the
population variances, in the formula given above, to find confidence intervals for the difference
between the two population means. These intervals then take the form
2
( X − Y ) − Za S 2 X nX + S Y < µ X − µY < ( X − Y ) + Z a S 2 X nX + S 2Y nY
2 nY 2
-0.1< ( µ X - µY ) <24.1
This 99% confidence interval for the difference in population mean scores just includes zero,
suggesting that the evidence in the data against the conjecture that the two population means are
the same is not overwhelmingly strong.
2. Researchers interested in verbal responses to survey questions have found that not only the
responses themselves, but also subjects’ judgment time in responding to the question, contain
useful information. Typically the data analyzed are the reciprocals of judgment time, and are called
the certainty values. Thus, the larger the time taken to produce a response, the smaller the
certainty value. A random sample13 of 143 people were contacted by telephone and asked:
“Assuming that it’s convenient for you, do you expect to get a swine flu shot?” One hundred of
these subjects answered “yes.” Their mean certainty value was 0.76, and the standard deviation
was 0.50. For the 43 subjects answering “no,” the mean certainty value was 1.41 and the standard
deviation was [Link] by µ X the population mean of those answering “yes” and by µY
the population mean for the “no” group, find a 95% confidence interval for ( µ X - µY ).
Solution:
Again, since the sample sizes are large, we can use the sample variances in place of the population
variances and obtain intervals from
σ 2X σ 2Y σ 2X σ2
( X − Y ) − Za + < µ X − µY < ( X − Y ) + Z a +
2 nX nY 2 nX nY
Where: X = 0.76 nX = 100 Sx= 0.50
Y = 1.41 nY = 43 Sy =0.72 and for a 95% confidence interval,
Z α /2= Z0.025 =1.96
The interval is then
(0.50) 2 (0.720 2 (0.50) 2 (0.72) 2
(0.76 − 1.141) − (1.96) + < µ X − µY < (0.76 − 1.41) + (1.96) +
100 43 100 43 or
-0.89< ( µ X - µY ) >-0.41
This interval includes only negative values, indicating that those who say they will get a shot are
less certain of their answers, on the average, than those who say they will not, in the population
at large. The figure shows this confidence interval, together with 80%, 90% confidence intervals
for the difference in the population means.
We now have to consider the case where the sample sizes are not large, and a confidence
interval is needed for the difference between the means of two normal populations based on
independent random samples from the two populations.
In the above figure 80%, 90%, 95% and 99% confidence intervals for difference in population
means based on data of Example 2 that the two population variance are equal, a fairly
straightforward method is available.
Suppose again that we have independent random samples of nx and ny observations from normal
populations with mean µ X -and µY , and that the populations have a common (unknown)
variance σ 2 . Inference about the population means is, as before, based on the difference ( X - Y )
between the two sample means. This random variable has a normal distribution with mean ( X -
Y ) and variance
Var ( X - Y ) = Var ( X ) + Var( Y )
σ2 σ2
= +
nX nY
1 1
= σ 2( + )
n X nY
n X + nY
= σ 2( )
n X nY
It therefore follows that the random variable
( X − Y ) − (µ X + µY )
Z =
n + nY (a)
σ 2( X )
n X nY
Has a standard normal distribution. However, this result cannot be used as it stands because the
unknown population variance is involved. Since this variance is common to the two populations,
the two sets of sample information can be pooled together to estimate it. The estimator used is
(n X + nY − 2)
2 2
Where S x and S y are the two sample variances.
Reporting the unknown σ 2 by its estimator s2X in equation (a) gives the random variable
( X − Y ) − ( µ X − µY )
S2 =
n + nY
S( ( X )
nX nY
It can be shown that this random variable obeys the student’s t distribution with (nx+ny-2) degree
of freedom. Given this result, confidence intervals for the difference between the population
means can be obtained through an argument similar to that used in confidence interval
computation of a normal population.
(n X + nY − 2)
t
And n x + nY − 2,α / 2 is that number for which
α
P (t n X + n Y − 2 , > (t n X + n Y − 2 , α / 2
)= 2
Where the random variable tnt +ny-2 has a student’s t distribution with
(nx +ny -2) degree of freedom.
Example:
In a study of the effect of planning on the financial performance of banks, a random sample of six
“partial formal planners” showed mean annual percentage increase in net income of 9.972 and a
standard deviation of 7.470. An independent random sample of nine banks with no formal
planning system had a mean annual percentage increase in net income of 2.098 and a standard
deviation of 10.834. Assuming the two population distributions are normal with the same
variance, find a 90% confidence interval for the difference between their means.
Solution:
We have, with x referring to the “partial formal planners,” and y to those with no formal
planning.
(n X + nY − 2)
= (5) (7.470)2 + (8)(10.834)2 = 93.693
13
So that S = 93.693 = 9.680
The interval required is of the form
n X + nY n + nX
( X − Y ) − t n x + nY − 2,α / 2 S < µ X − µ Y < ( X + Y ) + t n X + nY − 2,α / 2 S X
n X nY n X nY
Where, for a 90% confidence interval, α = 0.10, so that
tn x + nY − 2,α / 2 = t13, 0.05 = 1.771 from a table.
Hence, the 90% confidence interval for the difference between the population mean percentage
increases in net incomes is
6+9 6+9
(9.972 − 2.098) − (1.771)(9.680( < µ X − µY < (9.972 − 2.098) + (1.771)(9.680) or -
54 54
1.161 < µ X - µY <16.909
Our 90% confidence interval for the difference between population mean annual percentage
increases in net income for these two groups of banks includes 0. This suggests that the evidence
in the data against the conjecture that the two population means are the same is not strong.
PX (1 − PX ) PY (1 − PY )
= +
nX nY
Furthermore, if the sample sizes are large, the distribution of this random variable is
approximately normal, so that subtracting its mean and dividing by its standard deviation gives a
standard normal random variable. Moreover, for large sample sizes, this approximation remains
good when the unknown population proportions in equation (b) are replaced by the
corresponding sample quantities. Thus, to a good approximation, the random variable
( PX − PY ) − ( PX − PY )
Z=
P X (1 − P X ) P Y (1 − P Y )
[ +
nX nY
has a standard normal distribution.
This result allows the derivation of confidence intervals for the difference between the two
population’s proportions when the sample sizes are large, as shown in the box.
Let P x denote the observed proportion of successes in a random sample of nx observations from
a population with proportion px successes, and P y the proportion of successes observed in an
independent random sample from a population with proportion py successes. Then, if the sample
sizes are large, a 100(1 - α ) % confidence interval for (px – py) is given by:
P (1 − P ) P (1 − P X ) P (1 − P ) P (1 − P X )
(P X − PY ) − Z a + < (px – py)< ( P X − P Y ) − Z a +
2
nX nY 2
nX nY
Where za/2 is that number for which
P (z>za/2) = α / 2
and the random variable Z has a standard normal distribution
(0.79)(0.21) (0.89)(0.11))
(0.79 – 0.89) – (1.645) +
200 200
(0.79)(0.21) (0.89)(0.11))
< Px – Py < +
200 200
or
-0.16 < Px – Py < -0.04
The figure below also shows 80%, 95% and 99% confidence intervals for the difference between
the two population proportions. Notice that none of these intervals contains the difference 0 in
population proportions. Thus, the data strongly suggest that accounts attracted by premiums are
less likely to be held for as long as 6 months than those not attracted by premiums.
2. Frequently populations are surveyed by mail questionnaires, and investigations are anxious to
obtain as high a response rate as possible. A questionnaire, printed on a single sheet, front and
back, was sent to a random sample of 220 households, of
This figure shows 80%, 90%, 95% and 99% confidence intervals for difference in population proportions,
using data of Example 1 above.
which 36% responded. The same questionnaire, printed on two sheets, front only, was sent to an
independent random sample of 220 households, and the achieved response rate was 30%. Find a
95% confidence interval for the difference between the two population proportions responding.
Solution:
The sample values are
nx = 220; Px = 0.36; ny = 220; Py = 0.30
For a 95% confidence interval, α = 0.05, and so
Za/2 = Z.025 = 1.96
X=
∑ X i = 1 ( x + X + X + ... + X )
1 2 3 n
n n
We see that X is a linear function of all observations in the sample.
When a sample is drawn from a population, the evidence obtained can be used to make
inferential statements about the characteristics of the population. As we have seen, one
possibility is to estimate the unknown population parameters through the calculation of point
estimates or confidence intervals. Alternatively, the sample information can be employed to
assess the validity of some conjecture, or hypothesis, that an investigator has formed about the
population. Examples of situations of this kind are as follows:
1. A manufacturer who produces boxes of cereal claims that, on average, the contents weigh at
least 20 ounces. In order to check this claim, the contents of random sample of boxes can be
weighed, and inference based on the sample results.
2. A company receiving a large shipment of parts may want to accept delivery only if no more
than 5% of the parts is defective. The decision on whether to take delivery might be based on
a check of a random sample of these parts.
3. An instructor is interested in the value of regularly administered quizzes in a statistics course.
She uses these quizzes in one section of the course, but not in another. At the end of the
course she compares the average performances of students in the two sections on the final
examination, in order to check here hypothesis that the quizzes raise average performance.
4. a political scientist wants to know if a tax reform proposal appeals equally to men and
women. In order to check whether this is so he obtains the options of randomly selected
samples of males and females.
The examples given here have a common theme. A hypothesis is formed about some population,
and conclusions about the merits of this hypothesis are to be formed on the basis of sample
information. In this section we introduce a general framework for approaching such problems.
Specific procedures are then developed in the following section.
To keep our discussion quite general, let us denote the population parameter of interest (for
example, the population mean, variance, or proportion) by =. Suppose that some hypothesis has
been formed about this parameter, and that this hypothesis will be believed unless sufficient
contrary evidence is produced. This can be thought of as a maintained hypothesis. In the language
of statistical hypothesis testing, it s called a null hypothesis.
For example, we weight, in the absence of evidence to dispute it, believe the manufacturer’s claim
that, on average, the contents of its boxes of cereal weigh at least 20 ounces. When sample
information is collected, this hypothesis is put in jeopardy, or tested. If the hypothesis is not true,
then some alternative must be true and, in carrying out a hypothesis test, the investigator
formulates an alternative hypothesis against which the null hypothesis is tested. For the cereal
manufacturer we could test the null hypothesis that the mean contents weight is at least 20
ounces against the alternative hypothesis that the mean weight is less than 20 ounces. The null
hypothesis will be denoted Ho and the alternative hypothesis H1.
On the other hand, a hypothesis could specify a range of values for the unknown population
parameter. Such a hypothesis is said to be composite. And will hold true for more than one
value of the population parameter. For instance, the null hypothesis that mean weight of boxes of
cereal is at least 20 ounces is composite. The hypothesis is true for any population mean weight
greater than or equal to 20 ounces.
In many applications, a simple null hypothesis, say Ho; µ 0 =o, is tested against a composite
alternative. In some cases, only alternatives on one side of the null hypothesis are of interest. For
example, we might want to test this null hypothesis against the alternative hypothesis that the
true value of µ 0 is bigger than X 0 , which we can write.
H1:> 0
Conversely, the alternative of interest might be
H1:< o
Such alternative hypotheses are called one-sided alternatives. Another possibility is that we
want to test this simple null hypothesis against the very general alternative that the true of = is
something other than =o, that is.
H1; ≠o
This is referred to as a two-sided alternative.
The specification of appropriate null and alternative hypotheses is problem specific To illustrate,
we return to our earlier examples:
1. Let = denote the population mean weight (in ounces) of cereal per box. The null
hypothesis is that this mean is at least 20 ounces. So we have the composite null
hypothesis
H0; ≥ 20
The obvious alternative is that the true mean weight is less than 20 ounces, that is,
H1; < 20
2. A company intends to accept delivery of parts unless it has evidence to suspect that more
than 5% are defective. Let P denote the population proportion of defectives. The null
hypothesis here is that this proportion is at most 0.05, that is,
H0= = < 0.05
On the basis of sample information, this hypothesis is tested against the alternative
H1= = > 0.05
The null hypothesis, then, is that the shipment of parts is of adequate quality overall, while
the alternative is that it is not.
If all that is available is a sample from population, then the population parameters will not be
precisely known. Accordingly, it cannot be known for sure adopted, there is some chance of
reaching an erroneous conclusion about the population parameter of interest. In fact, as indicated
in table 7.1, either of two possible kinds of error could be made. There are two possible states of
nature-either the null hypothesis is true or it is false. One error that could be made, called a
Type I error is the rejection of a true null hypothesis. If the decision rule is such that the
probability of rejecting the null hypothesis hen it is true is ά, then is ά said to be the significance
level of the test. Since the null hypothesis must either be accepted or rejected, it follows that the
probability of accepting the null hypothesis when it is true is (1- ά). The other possible error,
called a Type II error, arises when a false null hypothesis is accepted. Suppose that, for a
particular decision rule, the probability of making such an error when the null hypothesis is false
is denoted β. Then, the probability of rejecting a false null hypothesis is (1- β), which is called the
power of the test.
The null hypothesis is that, in the population, the proportion of men in favor of this proposal is
the same as the proportion of women. This null hypothesis is to be tested against the alternative
that the two population proportions differ. In order to test the null hypothesis, independent
Table 7.1 states of nature and decision on null hypothesis, with associated probabilities of making
the decisions, given the particular states of nature random samples of men and women are taken,
States of Nature
and the views of the sample members are solicited. It is natural to base inference about the null
hypothesis on the difference between the sample proportions of men and women in favor of the
proposal. If this difference is large, the null hypothesis of equality of the population proportions
would be rejected; otherwise, this null hypothesis would be accepted. Let P, denote the sample
proposal. If this difference is large, the null hypothesis of equality of the population proportions
would be rejected; otherwise, this null hypothesis would be accepted. Let P x denote the sample
proportion of men, and P y the sample proportion of women, in favor of the tax reform
proposal. Then a possible decision rule is
Reject H0 if ( p x − Py ) > 0.05 or ( Px − P y ) < -0.05
Now suppose that, in fact, the null hypothesis that the two population proportions favoring the
proposal are equal is true. It nevertheless could happen that the sample proportions differ by
more than 0.05 so that, according to our decision rule, the null hypothesis would be rejected. In
that case, a Type I error would have been made. The probability of this occurring (when the null
hypothesis is true) is the significance level ά. On the hand, suppose that the null hypothesis is
false and that, in fact, the population proportion of men and women in favor of the proposal are
not the same. It may still be the vase that the two sample proportions differ by less than 0.05.
Then according to our decision rule, the null hypothesis would be accepted, and a Type II error
would have been made. The probability of making such an error will depend on just how different
are the two population proportions.
We would be less likely, for given sample sizes, to accept the null hypothesis if 80% of men and
20% of women favored the proposal than if these percentage were 55% and 45%.
To illustrate this sequence, consider again the problem of testing based on a sample of thirty
observations, whether the true mean weight of boxes of cereal is at least 20 ounces given a
decision rule. We could determine the probabilities of Type.
Type I and type II errors associated with the test. However, in fact, we proceed by first fixing the
type I error probability. Suppose, for example, that we want to ensure that the probability of
rejecting the null hypothesis when it is true is at most 0.05. We can do this by choosing an
appropriate number, K, in the decision rule. “Reject the null hypothesis if the sample mean is less
that K ounce.” (We will discuss in the next section how this can be done.) Once the number K
is chosen, the type II error probabilities can be computed, using procedures to be discussed in
section 7.9.
We have seen that, since the decision rule is determined by the particular significance level
chosen, the concept of power plays no direct part in the decision as to whether to reject a null
hypothesis. However, calculations of power, stemming from particular significance level choices,
provide the investigator with valuable information about the properties of the decision rule.
Often an investigator has some flexibility in the choice of the number of sample observations to
take. For a given significance level, the bigger the sample size, the higher will be the power of the
test. In deciding how big the sample should be, the analyst must balance the benefits from
increased power against the costs acquiring additional sample information. Another important use
In sections 7.2 -7.8, we show how, for given significance levels, decision rules can be formulated
for some important classes of hypothesis testing problems, we will return in section7.9 to a
consideration of the power of test. For convenience, the new terminology introduced in this
section is summarized in the accompanying box.
The terminology “accept” and “reject” for the possible decisions about a null hypothesis is
commonly used in formal summaries of the outcomes of particular tests. However, these terms
do not adequately reflect the asymmetry of the status of the null and alternative hypotheses, or
the consequences of a procedure in which the significance level of fixed, and the probability of
Type II error is not controlled. As we have already noted, the null hypothesis has the status of a
maintained hypothesis a hypothesis that will be held true, unless the data contain sufficient contrary
evidence. Moreover, in fixing a significance level, generally at some small probability, we are
X − µ0
Figure7.2 the probability density of > Z α when the null hypothesis Ho: µ = µo is true, and the
σ
n
decision rule for testing Ho against the alternative H1: µ >µo at significance level ά is:
X − µ0
Reject Ho if > Zα
σ
n
Then the probability of rejecting Ho when it is true will be ά, so that ά is the significance level of
the test based on this decision rule. This situation is illustrated in figure 7.2, which shows the
sampling distribution of the random variable 7.2.1 when the null hypothesis is true, through a
graph of its probability density function. The figure shows the value Zά, which is such that the
probability of its being exceeded, when the null hypothesis is true, is significance level ά of the
test. It follows that the probability of a sample result in the corresponding rejection region,
shown as the shaded area in the figure, must be ά when the null hypothesis is correct.
Example:
1. When a process producing ball bearings is operating correctly, the weights of the ball bearings
have a normal distribution with mean 5 ounces and standard deviation 0.1 ounce. An adjustment
has been made to the process, and the plant manager suspects this has raised the mean weight of
ball bearing produced, leaving the standard deviation unchanged. A random sample of sixteen ball
bearings is taken, and their mean weight is found to be 5.038 ounces. Test at significance levels
0.05 and 0.10 (that is, at 5% and 10% levels) the null hypothesis that the population mean weight
is 5 ounces against the alternative that it is bigger.
Denoting by µ the population mean weight (in ounces), we want to test
Ho: µ = µo = 5
against H 1: µ>5
The decision rule is to reject Ho in favor of H1 if
X − µ0
> Zα
σ
n
From the statement of the example, we have
X = 5.038; µo = 5; σ = 0.1; n= 16
X − µ0 5.038 − 5
So that = = 1.52
σ 0 .1
n 16
Definition
The smallest significance level at which a null hypothesis can be rejected
is called the probability- value, or P-value, of the test.
X − µ0
In example 1 we found = 1.52
σ
n
Therefore, according to our decision rule, the null hypothesis is rejected for any significance level
ά for which Zά is less than 1.52. From a table we find that, when Zά is 1.52, ά is equal to 0.0643.
This, then, is the P– value of the test the implication is that the null hypothesis can be rejected at
all levels of significance
higher than 6.43% this is illustrated in figure 9.3, which shows the correspondence between the
significance level of the test, ά, and the corresponding value Zά, which enters the decision true.
Suppose that, in place of the simple null hypothesis, we had wanted to test the composite null
hypothesis
H 0: µ < µ o
against the alternative
H 1: µ > µo
at significance level ά. For the decision rule developed in the case of the simple null hypothesis,
we saw that if the population mean is precisely µo, then the probability of rejecting the null
hypothesis is ά. For this same decision rule, if the true population mean is anything less than µo
we would be even less likely to reject the null hypothesis. Hence, use of this decision rule in the
present context guarantees a probability of at most ά of rejecting the composite null hypothesis
when it is true.
A Test of the mean of a normal distribution (variance known):
composite Null and Alternative hypotheses
The appropriate procedure for testing, at significance level , the null hypothesis
H0: µ < µo
Against the alternative hypothesis
H 1: µ > µo
is precisely the same as when the null hypothesis is Ho: µ = µo
In this circumstance, doubt would be cast on the null hypothesis if the sample mean were a good
deal lower than the hypothesized population mean. Once again, if the null hypothesis were true,
the random variance 7.2.1 would follow a standard normal distribution. To achieve a test with
significance level, we only need to note that
P(Z< Z ά) = ά
If Z is a standard normal random variable: Hence, if X is the observed sample mean, the
appropriate decision rule is:
X − µ0
Reject Ho if < −Z α
σ
n
This is illustrated in figure 7.4, which should be compared with figure 7.2. Clearly the former is
simply the mirror image of the latter.
Using an analogous argument to that developed earlier, we can see that this decision rule
continues to be appropriate if, in place of the simple null hypothesis, we have the composite
hypothesis Ho:µ>µo
With the same alternative hypothesis
X − µ0
Reject Ho if < −Z α
σ
n
This is illustrated in figure7.5, from which we see that the region of sample outcomes for which
the null hypothesis is rejected is divided into two parts. The upper part of the region
corresponds to observed values of the sample mean greatly in excess of the hypothesized
population mean, and the lower part to values of the sample mean that are substantially below µo
X − µ0
Figure7.5 the probability density function of Z = when the null hypothesis Ho: µ= µo is true,
σ
n
and the decision rule for testing Ho against the alternative H1 : µ≠ µo at significance level ά
The reader has probably noticed the similarity between the developments of procedures for
determining confidence intervals and testing hypotheses. A review of the material in section 8.2
will clarify the relationship. The null hypothesis Ho: µ= µo is rejected against the two –sided
alternative H1: µ≠ µo at significance level ά if and only if the 100 (1- ά) % confidence interval for µ
does not contain µo.
Example:
A drill, as part of an assembly line operation, is used to drill holes in sheet metal. When the drill
is functioning properly, the diameters of these holes have a normal distribution with mean 2
inches and standard deviation 0.06 inch. Periodically, to check that the drill is functioning
properly, the diameters of a random sample of holes are measured. Assume that the standard
deviation does not vary. A random sample of nine measurements yield mean diameter 1.95
inches. Test the null hypothesis that the population mean is 2 inches against the alternative that it
is not. Use a 5% significance level and also find the P-value of the test.
Example:
It might be suspected that the firms most likely to attract take-over bids are those that have been
achieving relatively poor returns. One measure o such performance is through “abnormal
returns,” which average 0 over all firms. A random sample of 88 firms for which cash tender
offers had been made showed abnormal returns with a mean of -0.0029 and a standard deviation
of 0.0169 in the period from 24 months to 4 months prior to the take-over bids. Test the null
Figure 7.6 conclusion of the test in the above example; the null hypothesis Ho: µ = µo is rejected against
the alternative H0: µ < µo at significance levels greater than .0537
suppose we have a random sample of n observations from a normal population with mean µ. if
the observed sample mean and standard deviation are X and Sx, then the following tests have
significance level ά:
i) To test either null hypothesis
H0: µ= µo or H0: µ< µo
against the alternative
H1: µ>µo = 0
the decision rule is
Reject H0 if X − µ 0 > tn – 1,ά
sX
n
Figure 7.7 illustrates the setup of the test against a two-sided alternative. The probability density
function is now that of the student’s t- distribution and this figure is the analogue of figure 7.5,
which related to the case where the population variance was known. In an obvious way, tests
against one-sided alternative hypotheses can be viewed pictorially in a manner analogous to
figures 7.2 and 7.4.
Example:
A real chain knows that, on average, sales in its stores are 20% higher in December than in
November. For the current year, a random sample of six stores was selected. Their percentage
December sales increases were found to be:
19.2; 18.4; 19.8; 20.2; 20.4; 19.0
Assuming a normal population distribution, test the null hypothesis that the true mean
percentage sales increase is 20, against the two-sided alternative, at the 10% significance level.
Solution:
Letting µ denote the population mean percentage increase in sales in December, we want to test
the null hypothesis.
H0: µ= µo = 20
against the alternative
H1: µ ≠20
The sample mean and variance are obtained by using the computations in the accompanying table.
We have for the sample mean,
X=
∑ X i = 117 = 19.6
n 6
Xi Xi2
19.2 368.64
18.4 338.56
19.8 392.04
20.2 408.04
20.4 416.16
19.0 316.00
Sum=117 2,284.44
2
Sx2 =∑xi– n X = 2,284.44 – (6)(19.5) = 0.588
n -1 5
So that the sample standard deviation is
Sx= √0.588 = 0.767
We then have
X - µo = 19.5 – 20 = -1.597
Sx/√n 0.767/√6
Since a test of significance level ά = 0.10 is required, we have from a table,
tn – 1,ά/2 = t5.05 = 2.015
Thus, since –1.597 lies between - 2.015 and 2.015, the null hypothesis that the true mean
percentage increase is 20 is accepted at the 10% level. The evidence in the data against this
hypothesis is not terribly strong.
obeys a chi-square distribution with (n-1) degrees of freedom. Tests of hypotheses about the
variance of a normal population are then based on the sample value observed for the equation
7.4.1. If the alternative hypothesis is that the true variance exceeds σ 0 2, we would be suspicious
of the null hypothesis is the observed sample variance was much bigger than σ 0 2. Hence, the null
hypothesis would be rejected if a high value of 7.4.1 we observed. Conversely, if the alternative is
that the population variance is less than the value specified by the null hypothesis, the null
hypothesis would be rejected for low values of 7.4.1. Finally, for the two-sided alternative that
the population variance differs from σ 0 2, we would want to reject the null hypothesis on
observing either unusually high or unusually low values of 7.4.1.
The rationale for the development of appropriate tests now follows the same pattern as in
section 7.2
Suppose we have a random sample of n observations from a normal population with variance
σ 2. if the observed sample variance is Sx2, then the following tests have significance level ά:
i) to test either null hypothesis
Ho: σ 2 = σ 2o or Ho: σ 2 < σ 2o
against the alternative /
H1: σ 2 > σ 2o
the decision rule is
reject Ho if ( n-1) Sx2 >X2n-1,ά
σ 2o
ii) To test either null hypothesis
Ho: σ 2 = σ 02 or Ho: σ 2 > σ 2o
against the alternative /
H1: σ 2 < σ 2o
the decision rule is
reject Ho if ( n-1) Sx2 < X2n-1,1-ά
σ 02
iii) To test the null hypothesis
Ho: σ 2 = σ 2o
against the two-sided alternative
H 1: σ 2 ≠ σ 2o
Figure 7.9 the probability density function of X2n-1 = (n – 1) S2x/∂2o when the null hypothesis Ho: σ 2 = σ 2
0 is
true, and the decision rule for testing Ho against the alternative H1: σ 2 ≠ σ 2o at significance level ά
Example:
In order to meet established standards, it is important that the variance of the percentage
impurity levels in consignments of a chemical not exceed 4.0. A random sample of twenty
consignments had a sample variance of 5.62 in impurity level percentage. Test the null hypothesis
that the population variance is not more than 4.0.
Solution:
Let σ 2 denote the population variance of impurity concentrations. The null hypothesis
Ho: σ 2 < σ 20 = 4.0
Is to be tested against
H1: σ 2 > 4.0
Based on the assumption that the population distribution is normal, the decision rule, for a test of
significance level ά, is to reject Ho in favor of H1 if
(n- 1) S2x < X2n-1, ά
σ2
From the statement of the example we have
S2x = 5.62; n= 20; σ 2o = 4.0
Hence,
(n- 1) S2x = (19) (5.62) = 26.695
σ 20 4.0
For a 10% level test, ά = 0.10 and we see from table 5 in the appendix that the corresponding
cutoff point of the chi-square distribution with (n-1) = 19 degrees of freedom is
X219,.10 = 27.20
Therefore, since 26.695 is not bigger than 27.20, the null hypothesis can not be rejected at the
10% level. Hence, the data do not contain terribly strong evidence against the hypothesis that the
population variance in impurity level percentages is at most 4.0.
( PX − P0 )
Fig 7.10 The probability density function of Z = when the null hypothesis H0:
P0 (1 − P0 ) / n
P =P0 is true, and the decision rule for testing H0 against the alternative H1: P<P0 at a significance level ά
P − P0 P − P0
Reject Ho if > Z a or < −Z a
P0 (1 − P0 ) 2 P0 (1 − P0 ) 2
n n
Here, as previously, Zά is that number for which
P (Z > Zά) = ά
Where the random variable Z has a standard normal distribution.
The decision rule for the second of these tests is illustrated in figure 7.10.
Example:
Forecasts of corporate earnings per share are made on a regular basis by many financial analysts.
In a random sample of 600 forecasts, it was found that 382 of these forecasts exceed the actual
out come for earnings. Test against a two-sided alternative the null hypothesis that the population
proportion of forecasts that are higher than actual outcomes is 0.05. (This is the hypothesis we
would expect to be true if there were no overall tendency for financial analysts to be either
unduly optimistic or unduly pessimistic about earnings prospects.)
Solution:
Let P denote the population proportion of forecasts that are above actual out comes. We want
to test
Ho: P= Po = 0.50
Against H1: P ≠ 0.50
The decision rule is to reject. Ho in favor of the alternative if
P − P0 P − P0
> Z a or < −Z a
P0 (1 − P0) 2 P0 (1 − P0 ) 2
n n
We have
Po = 0.50; n = 600; p 0 = 382 = 0.637
600
P − P0 0.637 − 0.50
Then = = 6.71
P0 (1 − P0 ) (0.50)(0.50)
n 600
We now return to our example of brain activity of subjects watching television commercials;
from table 7.2, the sample mean of the differences is
n
d = ∑ di = 1 (210) = 21.0
i=1 10
The sample variance is
n
Sd = 1 (∑ di2 – nd2)
2
n – 1 i=1
= 1 [14,202 – (10)(21.0)2] = 1,088
9
So that the sample standard deviation is
Sd = √1,088 = 32.98
We want to test the null hypothesis
Ho: µx – µy = Do = 0
against the alternative
Ho: µx – µy > 0
d − D0 21.0
The test is based on or = =2.014
Sd 32.98 / 10
n
This quantity must be compared with tabulated values of the student’s t distribution with
(n- 1) = 9 degrees of freedom. From a table, we have for 5% level and 2.5% level tests,
t9, .05 = 1.833 and t9, .025 = 2.262
Hence, the null hypothesis of equality of the population means can be rejected at the 5% level,
but not at the 2.5% level of significance. We see then that the data of table 7.2 contain much
evidence suggesting that, on the average, brain activity were the same for these two is higher for
the high recall than for the low recall group. If, in fact, the mean brain activity as extreme or
more extreme than actually obtained would be between .025 and .05
− −
nX nY nX nY
Example:
1. The international banking crisis of 1974, involving the failure of the franking national bank in
New York, led the Federal Reserve System to guarantee the international as well as the domestic
deposits of the bank. It might be hypothesize that this “Franklin Message” would lead to a
decrease in the risk premium attached to large American Banks’ deposits. (Risk premium here is
taken to be measured by the excess of secondary market certificate of deposit rates over
Treasury bill yields.) For 48 months before the “Franklin Message,” the mean risk premium was
0.899, and the variance was 0.247. For 48 months after the message, the mean and variance were
0.703 and 0.320. If µx and µy denote respectively the means before and after the message, test the
null hypothesis Ho: µx – µy = 0
Against the alternative Ho: µx – µy > 0
Solution:
Assume that the data can be regarded as independent random samples from the two populations.
The decision rule is to reject Ho in favor of H1.
In this example,
X= 0.899; S2x= 0.247; nx = 48; Y = 0.703; S2y = 0.320; ny= 48
So that
X −Y 0.0899 − 0.703
= = 1.80 From a table, we find that the value of ά corresponding
2 2
S X S Y 0.247 0.320
+ + to Zά = 1.80 is 0.0359. Hence, the null hypothesis can
nx nY 48 48 be rejected at all levels of significance greater than
3.59%. Hence, were the null hypotheses of equality of population means true, the probability of
observing a sample result as extreme or more extreme than that found would be 0.0359. This
represents pretty strong evidence against the null hypothesis of equality of these means,
suggesting rather a decrease in the mean risk premium after the “Franklin Message.”
We will now treat the case where the sample sizes are not large. If it can be assumed that the
two population variances are equal, then tests can be based on the result our discussion of
confidence intervals for the difference between the means of two normal population that the
random variable.
( X − Y ) − ( µ X − µY )
t=
n + nY
S X
n X nY
has a student’s t distribution with (nx + ny -2) degrees of freedom, where
Tests For The Difference Between The Means Of Two Normal Populations:
Independent Samples, Population Variances Equal
Suppose we have independent random samples of nx and ny observations from normal
distributions with means µ X and µ Y and a common variance. If the observed sample variances
are S2x and S2y, an estimate of the common population variance is provided by
S2 = (nx – 1) S2x + (ny – 1)S2y
(nx + ny – 2)
Then, if the observed sample means are X and Y , the following tests have significance level ά:
i) To test either null hypothesis
Ho: µx – µy = D0 or Ho: µx – µy < D0
against the alternative H1: µx – µy > D0
the decision rule is
X − Y − D0
Reject Ho if > t n x + nY − 2, a
n x − nY
S
nx n y
ii) To test either null hypothesis
Ho: µx – µy = D0 or Ho: µx – µy > D0
against the alternative
H1: µx – µy < D0
the decision rule is
X − Y − D0
Reject Ho if < −t n x + nY − 2, a
n − nY
S x
nx n y
iii) To test the null hypothesis
Ho: µx – µy = D0
against the alternative
H1: µx – µy < D0
the decision rule is
X − Y − D0 X − Y − D0
Reject Ho if < −t n x + nY − 2, a or > t n x + nY − 2, a
n − nY n − nY
S x S x
nx n y nx n y
X −Y
< −t nx + nY − 2,a
n x − nY
S
nx n y
For these data, we have
X = 0.058; Sx= 0.055; nx = 23; Y = 0.146; Sy= 0.058; ny = 23
Hence
S2 = (nx -1) S2x + (ny – 1) S2y
nx + ny - 2
= (22) (0.055)2 + (22) (0.058)2 = 0.0031945
23 + 23 – 2
So that S = √0.0031945 = 0.0565
X −Y 0.058 − 0.146
Then = = −5.282
n x − nY 23 + 23
S 0.0565
nx n y 23 2
For a 0.5% level test, we have by interpolation from table 6, for the student’s t distribution with
(nx + ny – 2) = 44 degrees of freedom,
t44, .005 = 2.695
Then, since -5.282 is much less than -2.695, the null hypothesis is overwhelmingly rejected even
at this level of significance. The data cast considerable doubt on the hypothesis that the two
population means are equal. Rather, they suggest very strongly that the population mean return
on assets is lower for failed than for nonfailed retail firms.
The test just discussed and illustrated is based on an assumption that the two population
variances are equal. In fact, it is possible to develop tests that are valid when this assumption does
not hold. However, these will not be discussed further here.
7.7 Tests for the Difference between Two Population Proportions (Large Sample)
We turn now to the problem of comparing two population proportions. As we discussed in
chapter 6,Suppose that a random sample of nx observations from a population with proportion
px “successes” gives a sample proportion PX , and that an independent random sample of ny
observations from a population with proportion py ”successes” yields sample proportion PY .
( p X − PY ) − ( Px − PY )
Z =
PX (1 − PX ) PX (1 − PY )
+
nX nY
p X − PY
= (7.7.1)
n + nY
P0 (1 − P0 )( X )
n X nY
p X − PY p X − PY
<- Z α / 2 or > Zα / 2
n + nY n + nY
P0 (1 − P0 )( X ) P0 (1 − P0 )( X )
n X nY n X nY
For these data, we have
P x= 36 = 0.180; nx= 200; P y = 29 = 0.246; ny =118
200 118
Hence,
n X PX + nY PY
= = (200)(0.180)+(118)(0.246) /200+118 =0.204
n X + nY
Then
p X − PY 0.180 − 0.246
= = -1.41
n + nY 200 + 118
P 0 (1 − P 0 )( X ) (0.204)(0.796)[ ]
n X nY (200)(118)
The value of α /2 corresponding to z α /2= 1.41 is, from a table α /2=.0793,sothat α =.1586.
Hence, the null hypothesis can be rejected only at significance levels higher than 15.86% .The
evidence against the hypothesis that the population proportions viewing price as the most
important criterion are the same in these two age groups is very strong.
F= S X / σ X
2 2
(7.8.1)
S 2Y σ 2Y
Follows a distributions known as the F distribution ( Formally this distribution is defined as the
distribution followed by the ratio of two independent Chi-square variables , each divided by its associated
degree of freedom). This family of distribution is widely used in statistical analysis. A particular
member of the family is distinguished by two values-the degrees of freedom associated with the
numerator and with the denominator. In the preset context, recall that the degrees of freedom
associated with the sample variance s2xis (nx-1), and that with s2y is (ny-1). The random variable
7.8.1 then has an F distribution with numerator degrees of freedom (nx-1 and denominator
degrees of freedom (ny-1).
The F distribution has an asymmetric probability density function, defined only for nonnegative
values. This density function is illustrated in figure 7.11.
The F Distribution
Suppose that independent random samples of nx=and ny observations are taken from two
normal population with variances σ 2x and σ 2y .If the sample variances are s2xand s2y,then the
random variable
S X /σ 2 X
2
F= 2
S Y / σ 2Y
has an F distribution with numerator degrees of freedom (nx – 1) and denominator degrees of
freedom (ny – 1). An F distribution with numerator degrees of freedom V1 and denominator
degrees of freedom V2 will be denoted FV1, [Link] denoted by FV1, V2,ά that number for which
P(FV1 V2 > FV1, V2,ά) = ά
Figure 7.11 probability density function of the distribution with 6numerator degrees of freedom and 4
denominator degrees of freedom; the probability is ά that F6,4 exceeds F6,4,ά
S2y
Here, Fnx -1,ny- 1, ά is that number for which
P (Fnx -1,ny- 1, ά > Fnx -1,ny- 1, ά) = ά
where Fnx -1,ny- 1 has an F distribution with numerator degrees of freedom (nx -1)
denominator degrees of freedom (ny -1).
Example:
It is hypothesized that the market share of a corporation should vary more in an industry with
active price competition than in one with duopoly and tacit collusion. In a study of the steam
turbine generator industry, it was found that in 4 years of active price competition, the variance
of general electric’s market share was 114.0895. In the following 7 years, in which there was
duopoly and tacit collusion, this variance was 16.0780. If the two population variances are
denoted σ 2x and σ 2y test
Example:
In example 7.1 we tested the null hypothesis that the population mean weight (in ounces) of ball
bearings was
Ho: µ = µo =5
against the alternative
H1: µ >5
The population standard deviation was σ = 0.1, and the test was based on n =16 observation.
The test was carried out at significance level ά = 0.05, so that
Zά = Z.05=1.645
We now determine the probability that our decision rule will reject the null hypothesis when
the rule mean weight is µ1 = [Link] power is then
1 – β = p (Z> µo+ µ1 + Zά)
σ /√n
Figure 7.12 sampling distributions of sample mean in example 7.11 for population mean µ0 =5 µ1=5.02.
Figures show calculation of power,1-β,corresponding to significance level ά= .05 for testing H0 : µ =5
against H1: µ> 5;power is evaluated at µ=5.02
X- µ0 > 1.645
σ /√n or, with µ0=5, σ =0.1, and n =16,when
X − 5 >1.645
0.1/4
This is equivalent to requiring
X > 5.045
Thus, when the null hypothesis is true, the probability that the sample mean exceeds 5.041 is
0.05. This is show in part (a) of figure 7.12. Part (b) of the figure shows the density function of the
sampling distribution of the sample mean when the population mean is 5.02 . It differs from part
(a) of the figure in being shifted to the right by an amount 0.02 the difference between the means
5.02 and 5. The shaded area in this figure shows the probability that the sample mean exceeds
5.041 when the population mean is 5.02. This is the power, evaluated at that point, as calculated
previously.
In a similar manner, such probabilities can be calculated for any value of µ[Link] powers are
shown in the table and graphed in Figure 7.13.
Figure 7.13 power function for example 9.11; test of H0: µ=5.00 against H1: µ>5.00(ά =0.05, σ =0.1,
n=16)
Figure7.14 power functions for test of H0: µ=5.00 against H1: µ>5.00(ά=0.05, σ = 0.1), shown for
sample sizes 4,9,16
P − P0
> Z a ] + P[ PX − P0
P[ < −Z a 2 ]
P0 (1 − P0 ) n 2
P0 (1 − P0 ) n
Against, this probability is equal to the significance level ά when the null hypothesis is true.
Suppose now that the null hypothesis is false and that the population proportion of successes is
p1, which differs from p0 .In that case, provided the sample size is large, we know that to a good
approximation, the random variable
PX − P1
Z=
P1(1 − P1 ) n
has a standard normal distribution. The power of the test can then be found as
P − P0
> Za + PX − P0
Power = 1-β= p [ < −Z a 2 ]
P0 (1 − P0 ) n 2
P0 (1 − P0 ) n
PX − P1 P0 − P1 P0 (1 − P0 ) / n
=P[ < − Za 2
P1 (1 − p1 ) n P1 (1 − P1 ) / n P1 (1 − P1 ) / n
PX − P1 P0 − P1 P0 (1 − P0 ) / n
+ > + Za 2 ]
P1 (1 − p1 ) n P1 (1 − P1 ) / n P1 (1 − P1 ) / n
P0 − P1 P0 (1 − P0 ) / n
=P [Z< − Za 2 ]
P1 (1 − P1 ) / n P1 (1 − P1 ) / n
P0 − P1 P0 (1 − P0 ) / n
+ P[ Z > + Za 2 ]
P1 (1 − P1 ) / n P1 (1 − P1 ) / n
Where Z is standard normal random variable. This result allows the calculation of power a
function of p1 =and n.
Example:
In Example 7.6 we tested the null hypothesis that population of earnings forecasts by financial
analysts that exceed the actual outcome is
H0: P= P0 =0 .50
against the alternative
H1 : P ≠ 0 .50
The test was based on a random sample of n=600 observations. For a test at significance level
ά= 0.05, we have
Zά/2 =Z.205= 1.96
We now determine the probability that, for a 5% level test, the null hypothesis will be rejected
when the rule population proportion is P1=0.52.
Substitution into the formula gives the power
1-β = P{[Z < 0.05 -0.52 ] - [1.96 √(0.50)(0.50)/600 ] ] + p [z > 0.05 -0.52 ]
√(0.52)(0.48)/600 √(0.52)(0.48)/600 √(0.52)(0.48)/600
+ [1.96 √(0.50)(0.50)/600 ] }
√(0.52)(0.48)/600
= p (Z< 2.94)+p(Z>.98)
= 0.0016+0.1635=0.1651 , from a table.
So, the probability is 0.1651 that this decision rule will reject the null hypothesis when the
population proportion is 0.52.
Similarly this probability can be calculated for any population proportion p1.
Figure 7.15 shows the power function for this example. Because the alternative hypothesis is two
side, the power function differs in shape from that of figure 7.13
Mekelle University
Faculty of Business & Economics
Department of Economics
INSTRUCTIONS
⇒ Work out the following the following questions and show your steps clearly and neatly.
⇒ Legible & neat works credited more.
⇒ Make sure that the Exam contains four pages, two parts and ten questions.
⇒ Attempt all the questions.
3.A certain local investor has $10,000 two invest and two agriculture investment opportunities,
each requiring a minimum of $[Link] return per $100 from the first alternative can be
represented by a random variable, X, having the following probability function:
X -5 20
P(X) 0.4 0.6
The Return per $100 from the second is given by the random variable, y, whose probability
function is:
Y 0 25
P(Y) 0.6 0.4
The random variable X and Y are independent. And the investor has the following three possible
strategies:
I. $10,000 in the first investment opportunity
II. $10,000 in the second investment opportunity
III. $ 5000 in each investment.
Given this answer the following:
a. What will be the mean and variance of return from each strategy?
b. Give your opinion about which strategy is promising for the investor? Why? Would you
necessarily advise the investor to adopt this strategy?
4.A certain Brewery’s beer bottles are not always filled to its capacity. The brewery
advertises that its bottles contain, on average 12 ounces of beer with a standard
deviation of 0.4 ounces. If a sample of 100 bottles was taken from the production
line, what is the probability of observing a sample mean of 11.9 ounces or less?
[Link] years UNESCO report indicated that the per-capita income of LDCs, including
Ethiopia, distributed Normally with mean annual income of $ 120 and standard
deviation of $20. But an Economist from African Economic Commission (AEC)
challenged this report. According to the Economist, there is/was an absolute
poverty in LDCs, especially, in Ethiopia, where most of the citizens earn mean
annual income below Absolute Poverty Line, $100. To disproof the UNESCO report,
he collect the following 32 sample Ethiopians` annual incomes by using
Area (Cluster) sampling method. The following data shows the result of the
sampling.
[Link] studies indicated that the reason for most new products market failures is
associated with the setting of price higher than the consumers willingness to pay.
To build this hypothesis a random sample of 89 new product market failures were
taken and the Managers associated with this assessment indicated that the main
cause for Market failure for the sample is 18.2%. Find a 95% confidence interval
for the proportion of all market failures for which high price is held to be the main
cause.
7. From 5000 new extension package adopters 600 are observed not to benefit from
the extension package. And 50 interviewers are sent to 50 different villages to
collect data on 100 farmers each. After reaching on the site they use a random
number to sample the farmers. How many of the interviewers are expected to
find 5 to 8 farmers who do not benefit from the extension package. (Makes sure
that you explain why you use a given theoretical distribution).
8. A town has 50,000 population and small studies show that 30% of the population
expects inflation to be much higher next year but not the rest. But we want to
conform it by study which includes 1000 people. So we first get the list of
all individuals, we gave them an identification number and we pick them by lottery
with out replacement. What are the chances that our sample will include 400 to
500 individuals which does not expect prices to rise? (Makes sure that you
explain why you use a given theoretical distribution)
9. If the mean wheat price is 90 Birr and it has a standard deviation of 20 Birr and
the prices are highly skewed to the right. How large a sample size is need to be,
if we want with 90% confidence that sample average price to be in range of 80
to [Link] interpret the result and the confidence interval.
[Link] distribution of income among 100 individuals in a given area is normal and
20 of them are randomly selected and the mean and standard deviation is found
to be 200 Birr and 100 Birr, what is the 90% confidence interval for the
population mean income. Could you compare this result with the point estimate
of income?