0% found this document useful (0 votes)
9 views22 pages

Chapter Two

Chapter Two discusses statistical estimation, focusing on how sample means and proportions can estimate unknown population parameters. It differentiates between point estimates and interval estimates, emphasizing the importance of confidence intervals in providing a range of values for population parameters. The chapter also covers the construction of confidence intervals for population means and proportions, including scenarios where the population standard deviation is known or unknown.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
9 views22 pages

Chapter Two

Chapter Two discusses statistical estimation, focusing on how sample means and proportions can estimate unknown population parameters. It differentiates between point estimates and interval estimates, emphasizing the importance of confidence intervals in providing a range of values for population parameters. The chapter also covers the construction of confidence intervals for population means and proportions, including scenarios where the population standard deviation is known or unknown.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

CHAPTER TWO

STATISTICAL ESTIMATION
1.1. INTRODUCTION
The sampling distribution of the mean shows how far sample means could be from a known
population mean. Similarly, the sampling distribution of the proportion shows how far sample
proportions could be from a known population proportion. In estimation, our aim is to determine
how far an unknown population mean could be from the mean of a simple random sample
selected from that population; or how far an unknown population proportion could be from a
sample proportion. Those are the concerns of statistical inference, in which a statement about an
unknown population parameter is derived from information contained in a random sample
selected from the population.
Objectives of the Chapter
When you have completed this chapter you will be able to;
 Estimation
 Differentiate the types of estimation.
 Construct a confidence interval for the population mean when the population standard
deviation is known.
 Construct a confidence interval for the population mean when the population standard
deviation is unknown.
 Construct confidence interval for population proportion.
 Determine sample size for attribute and variable sampling.
2.2 BASIC CONCEPTS:
 Estimation: is the process of using statistics as estimates of parameters. It is any
procedure where sample information is used to estimate/ predict the numerical value of
some population measure (called a parameter).
 Estimator- refers to any sample statistic that is used to estimate a population parameter.

E.g. x for μ , p for p.


 Estimate- is a specific numerical value of our estimator. E.g. x= 9, 2, 5
x , p , s 2 , s ……………. Estimators
μ , p ,σ 2 , σ ………………… items being estimated
1, 0.5, 9, 3 …………………... Estimates
1.3. TYPES OF ESTIMATES:
We can make two types of estimates about a population: a point estimate and an interval
estimate.
 A point estimate: - is a single number that is used to estimate an unknown population
parameter. It is a single value that is measured from a sample and used as an estimate of

the corresponding population parameter.


 The most important point estimates (given that they are single values) are:

 Sample mean ( x ) for population mean ( μ ) ;

 Sample proportion ( p ) for population proportion( p ) ;

 Sample variance ( s ) for population variance ( σ 2 )


2
and

 Sample standard deviation ( s ) for population standard deviation ( σ )


 Interval estimation: - is a range of values used to estimate a population parameter. It
describes the range of values with in which a parameter might lie. Stated differently, an
interval estimate is a range of values with in which the analyst can declare with some
confidence that the population parameter will fall.
Point estimators of population parameters, while useful, do not convey as much information
as interval estimators.
Point estimation produces a single value as an estimate of the unknown population
parameter. The estimate may or may not be close to the parameter value; in other words, the
estimate may be incorrect.
Further, a measure of confidence in the interval estimator is provided; consequently, interval
estimates are also called confidence intervals. For these reasons, interval estimators are
considered more desirable than point estimators.
Example:
Suppose we have the sample 10,20,30,40 and 50 selected randomly from a population whose
mean μ is unknown.
∑ xi10+20+30+ 40+50
=30
The sample mean, x , n = 5 is a point estimate of μ .

On the other hand, if we state that the mean, μ , is between x±10 , the range of values from 20
(30-10) to 40 (30+10) is an interval estimate.
2.4. INTERVAL ESTIMATORS OF THE MEAN AND PROPORTION
Interval estimation for population means, μ
 As a result of the Central Limit Theorem (discussed in Chapter I) the following z formula
for sample means can be used when sample sizes are large, regardless of the shape of the
population distribution or for smaller sizes if the population is normally distributed.
X−μ
Z=
σ
n
Rearranging the formula:
σ
μ= X − Z
n
 Because the sample mean can be greater than or less than the population mean, z can be
positive or negative. Thus, the preceding expression takes the form:
σ
μ= X ± Z
n

 The value of the population mean, μ , lies somewhere within this range. Rewriting this
expression yields the confidence interval for population mean:
σ σ
X −Z ≤ μ ≤ X +Z
n n
The confidence interval for population mean is affected by:
1. The population distribution, i.e., whether the population is normally distributed or not
2. The standard deviation, i.e., whether σ is known or not.
3. The sample size, i.e., whether the sample size, n, is large or not.
Confidence internal estimate of μ - Normal population, σ known
 A confidence interval estimate for  is an interval estimate together with a statement of how
confident we are that the interval estimate is correct.
 When the population distribution is normal and at the same time σ is known, we can
estimate μ (regardless of the sample size) using the following formula.
σ
μ= X ± Z α / 2
n
Where:
X = sample mean
Z = value from the standard normal table reflecting confidence level
σ = population standard deviation
n = sample size
α = the proportion of incorrect statements (α = 1 – C)
 = unknown population mean
From the above formula we can learn that an interval estimate is constructed by adding and
subtracting the error term to and from the point estimate. That is, the point estimate is found at
the center of the confidence interval.
To find the interval estimate of population mean, μ we have the following steps.

1. Compute the standard error of the mean(


σ x)

2. Compute α /2 from the confidence coefficient.

3. Find the Z value for the α /2 from the table


4. Construct the confidence interval
5. Interpret the results
Example:
1. The vice president of operations for Ethiopian Tele Communication Corporation (ETC) is in the
process of developing a strategic management plan. He believes that the ability to estimate the
length of the average phone call on the system is important. He takes a random sample of 60
calls from the company records and finds that the mean sample length for a call is 4.26 minutes.
Past history for these types of calls has shown that the population standard deviation for call
length is about 1.1 minutes. Assuming that the population is normally distributed and he wants
to have a 95% confidence, help him in estimating the population mean.
Solution:

n= 60 calls X = 4.26 minutes σ = 1.1 minutes C= 0.95


σ 1 .1 σ
σ X= μ= X ± Z α / 2
i. √n = √60 = 0.142 iv. √n
ii. α = 1 – C = 1- 0.95 = 0.05 = 4.26 ± 1.96(0.142)
α /2 = 0.05/2 = 0.025 = 4.26 ± 0.28

iii.
Z α /2= Z 0.025 =1.96
3.98 ≤  ≤ 4.54
The vice-president of ETC can be 95% confident that the average length of a call for the
population is between 3.98 and 4.54 minutes.
2. A survey conducted by “Addis Zemen Gazetta” found that the sample mean age of men was 44
years and the sample mean age of women was 47 years. Altogether, 454 people from Addis
were included in the reader poll –340 women and 114 men. Assume that the population standard
deviation of age for both men and women is 8 years.
a. Develop a 95% confidence interval estimate for the mean age of the population men who
read the gazetta.
b. Develop a 95% confidence interval estimate for the mean age of the population women
who read the gazetta.
c. Compare the widths of the two interval estimates form part (a) & (b) which one has a
better precision? Why?
Solution:
a.

n= 114 men X = 44 years σ = 8 years C= 0.95


σ 8 σ
σ X= μ= X ± Z α / 2
i. √n = √114 = 0.75 iv. √n
ii. α = 1 – C = 1- 0.95 = 0.05 = 44 ± 1.96(0.75)
α /2 = 0.05/2 = 0.025 = 4.26 ± 1.47

iii.
Z α /2= Z 0.025 =1.96
42.53 ≤  ≤ 45.47
b.

n= 340 women X = 47 years σ = 8 years C= 0.95


σ 8 σ
σ X= μ= X ± Z α / 2
i. √n = √340 = 0.434 iv. √n
ii. α = 1 – C = 1- 0.95 = 0.05 = 47 ± 1.96(0.434)
α /2 = 0.05/2 = 0.025 = 47 ± 0.85

iii.
Z α /2= Z 0.025 =1.96
46.15 ≤  ≤ 47.85
c. Part b has a better precision because the sample size is larger as compared with part a.
3. Time magazine reports information on the time required for caffeine from products such as
coffee and soft drinks to leave the body after consumption. Assume that the 99% confidence
interval estimate of the population mean time for adults is 5.6 hrs to 6.4 hrs.
a. What is the point estimate of the mean time for caffeine to leave the body after
consumption?
b. If the population standard deviation is 2 hrs, how large a sample was used to provide the
interval estimate?
Solution:
C = 0.99 Confidence interval: 5.6 ≤  ≤6.4
5. 6+6. 4
=6 hours
a. point estimate = 2
Or;

{
+¿ 5.6=X−Z α /2
σ
√n
¿ ¿¿¿

12 = 2 X
X = 6 hours
b. 0.99 σ = 2 hours Confidence interval: 5.6 ≤  ≤6.4 n=?

α = 1- C = 1- 0.99 = 0.01 α/2 = 0.005


Z α /2= Z 0. 005 =2.58
σ
6 . 4= X + Z α /2
√n
2
6 . 4= 6+ 2 .58
√n
5 . 16
0 . 4=
√ n ; rearranging the expression
5 .16
√ n=
0.4
√ n=12 . 9 ; squaring both sides
n = 166 : We state with 99% confidence that the mean time required for caffeine to leave the
body after consumption lies between 5.6 and 6.4 hrs.
Confidence interval estimate of μ - Normal population, σ unknown, n large
 If we know that the population is normal, and we know the population standard deviation,
σ the confidence interval for μ should be constructed in the manner already shown, i.e.,
σ
μ= X ± Z α / 2
√n .
 If the population standard deviation is unknown, it has to be estimated from the sample; i.e.,

when σ is unknown, we use sample standard deviation:


S=
√ ∑ ( X i− X )2
n−1 .
σ X , is estimated by the sample standard error of the
 Then, the standard error of the mean,
S
SX =
mean: √n.
Therefore, the confidence interval to estimate μ when population standard deviation is unknown,
population normal and n is large is
S
μ= X ± Z α / 2
√n.
Example:
1. Suppose that a car rental firm in Addis wants to estimate the average number of miles
traveled by each of its cars rented. A random sample of 110 cars rented reveals that the
samples mean travel distance per day is 85.5 miles, with a sample standard deviation of
19.3 miles. Compute a 99% confidence interval to estimate μ .
Solution:

n= 110 rented cars X = 85.5 miles s = 19.3 miles C= 0.99


S 19 .3 s
SX = μ= X ± Z α / 2
i. √ n = √110 = 1.84 iv. √n
ii. α = 1 – C = 1- 0.99 = 0.01 = 85.5 ± 2.58(1.84)
α /2 = 0.01/2 = 0.005 = 85.5 ± 4.75

iii.
Z α /2= Z 0. 005 =2.58
80.75 ≤  ≤ 90.25
We state with 99% confidence that the average distance traveled by rented cars lies between
80.75 and 90.25 miles.
Example:
A study is being conducted in a company that has 800 engineers. A random sample of 50 of
these engineers reveals that the average sample age is 34.3 years, and the sample standard
deviation is 8 years. Assuming normality, construct a 98% confidence interval to estimate the
average age of all engineers in this company.
Confidence interval for μ− σ unknown, n-small, population normal
 If the sample size is small (n<30), we can develop an interval estimate of a population mean
only if the population has a normal probability distribution.
 If the sample standard deviation s is used as an estimator of the population standard
deviation σ and if the population has a normal distribution, interval estimation of the
population mean can be based up on a probability distribution known as t-distribution.
Characteristics of t-distribution
1. The t-distribution is symmetric about its mean (0) and ranges from - ∞ to ∞.
2. The t-distribution is bell-shaped (unimodal) and has approximately the same appearance as
the standard normal distribution (Z- distribution).
3. The t-distribution depends on a parameter ν (Greek Nu) 1, called the degrees of freedom of the
distribution. Ν = n -1, where n is sample size. The degree of freedom, ν, refers to the number
of values we can choose freely.
4. The variance of the t-distribution is ν/ (ν-2) for ν>2.
5. The variance of the t-distribution always exceeds 1.
6. As ν increases, the variance of the t-distribution approaches 1 and the shape approaches that
of the standard normal distribution.
7. Because the variance of the t-distribution exceeds 1.0 while the variance of the Z-distribution
equals 1, the t-distribution is slightly flatter in the middle than the Z-distribution and has
thicker tails.
8. The t-distribution is a family of distributions with a different density function corresponding
to each different value of the parameter ν. That is, there is a separate t-distribution for each
sample size. In proper statistical language, we would say, “There is a different t-distribution
for each of the possible degrees of freedom”.
9. The t formula for sample when σ is unknown, the sample size is small, and the population is
X−μ X −μ
t= =
SX s
normally distributed is: √n This formula is essentially the same as the z-
formula, but the distribution table values are not.
The confidence interval to estimate μ becomes:
s
μ= X ±t α / 2 , v
√n
Where: X = sample mean
α=1–C

1
What are degrees of freedom? We can define them as the number of values we can choose
freely. In general, the degrees of freedom for a t statistic are the degrees of freedom
associated with the sum of squares used to obtain an estimate of the variance. The variance
estimate depends on not only on the sample size but also on how many parameters must be
estimated with the sample:

Degrees of = Number of − Number of parameters that


freedom Observations must be estimated beforehand
Here we calculate sample variance by using n observations and estimating one parameter
(the mean). Thus, there are (n – 1) degrees of freedom.
ν = n – 1 (degrees of freedom)
s = sample standard deviation
n = sample size
 = unknown population mean
Steps:
1. Calculate degrees of freedom (v=n-1) and sample standard error of the mean.

2. Compute α /2

t
3. Look up α / 2, V
4. Construct the confidence interval
5. Interpret results

Example:

1. If a random sample of 27 items produces x= 128.4 and s = 20.6. What is the 98%
confidence interval for μ ? Assume that x is normally distributed for the population.
What is the point estimate?
Solution:
The point estimate of the population mean is the sample mean, in this case 128.4 is the point
estimate.

n= 27 X = 128.4 s = 20.6 C= 0.98


S 20 . 6
SX =
i. √ n = √27 = 3.96 ν = n – 1 = 27-1 = 26
ii. α = 1 – C = 1- 0.98 = 0.02
α /2 = 0.02/2 = 0.01

iii.
t α/2, v= t 0 .01,26 =2.479
s
μ= X ±t α / 2 , v
iv. √n
= 128.4 ± 2.479(3.96)
= 128.4 ± 9.82
118.56 ≤  ≤ 138.22
We state with 98% confidence that the population mean lies between 118.56 and 138.23.
2. A sample of 20 cab fares in Bahir Dar city shows a sample mean of Br 2.50 and a sample
standard deviation of Br. 0.50. Develop a 90% confidence interval estimate of the mean
cab fares in Bahir Dar city. Assume the population of cab fares has a normal distribution.

n= 20 X = Birr 2.50 s = Birr 0.50 C= 0.90


S 0 .5
SX =
i. √ n = √20 = 0.112 ν = n – 1 = 20-1 = 19
ii. α = 1 – C = 1- 0.90 = 0.10
α /2 = 0.10/2 = 0.05

iii.
t α/2, v= t 0 .05,19=1.729
s
μ= X ±t α / 2 , v
iv. √n
= 2.50 ± 1.729(0.112)
= 2.50 ± 0.194
2.31 ≤  ≤ 2.69
We state with 90% confidence that the mean of cab fares in Bahir Dar city lies between Birr 2.31
and 2.69.
Example: Thirty –six items are randomly selected from a population of 300 items. The sample
mean is 35 and the sample standard deviation [Link] a 95 percent confidence interval for
population mean.

Interval Estimation of the Population Proportion

 We know that a sample proportion, p , is an unbiased estimator of a population proportion


P and if the sample size is large then, the sampling distribution of p is normal with
P−P P−P
Z= =


σp Pq
n.
P−P
Z=

However, here p is unknown and we want to estimate p by p and hence z becomes √ pq


n.

σ
That is, p is substituted by
S p=
√ pq
n

Solving for P results in


P= p+ Z
√ pq
n and since Z can assume both positive and negative

values, it becomes
P= p±Z
√ pq
n.

Since Z represents the confidence level we write it as


P= p±Z α /2
= p±Z α /2 S p
√ pq
n

Where: p = sample proportion


q =1- p
α=1–C
n = sample size
P = unknown population proportion
Example:
1. Recently, a study of 87 randomly selected companies with telemarketing operation was
completed. The study revealed that 39% of the sampled companies had used
telemarketing to assist them in order processing. Using this information estimate the
population proportion of telemarketing companies who use their telemarketing operation
to assist them in order processing taking a 95% confidence level.
Solution:

n= 87 p = 0.39 q = 0.61 C = 0.95

i.
S p=
√ √
pq
n=
0 .61∗0 . 39
87 = 0.0523

ii. α = 1 – C = 1- 0.95 = 0.05


α /2 = 0.05/2 = 0.025
iii.
Z α /2= Z 0.025 =1.96

iv.
P= p±Z α / 2 S p
= 0.39 ± 1.96(0.0523)
= 0.39 ± 0.1025
0.2875 ≤ P ≤ 0.4925
We state with 955 confidence that the proportion of companies which use telemarketing to assist
order processing lies between 0.2875 and
2. A fast food restaurant took a random sample of 400 customers to determine the
proportion of customers who are female. A confidence interval of .73 to .87 was
reported.
a. Find the number of females and the sample proportion
b. Find the level of confidence of this interval
Solution:

a. n= 400 0.73 ≤ P ≤ 0.87 p =? Number of females=?


0 .73+0. 87
=0. 80
Point estimate = 2
Or;

+¿ {0.73=p−Zα/2 s p ¿ ¿¿¿
1.60 = 2 p

p = 0.8
Number of females (X) = n* p = 400*0.8 = 320

b.
P= p±Z α / 2 S p

0.87 = 0.8+
Zα /2 S p

0.07 =
Zα /2
√ 0 . 8∗0 . 2
400

0.07 =
Z α /2∗0. 02
3.50 =
Zα /2
(P/Z=3.5) = 0.49977
C = 0.49977*2
= 99.954%
1.4. INTERVAL ESTIMATION OF THE DIFFERENCE BETWEEN TWO
INDEPENDENT MEANS
 It is clear that the unbiased point estimate of the difference between the means of two

populations ( μ1 −μ2 ) is the difference between two sample means( x 1 −x 2 ) , where each
sample is a random sample taken from the respective target population. The confidence
interval is constructed by adding the relevant standard error value which is called standard
error of the difference between means and the confidence level desired.
 If the two parent populations are normal, then the sampling distribution of the difference
between two means will be normally distributed regardless of n (sample size). And we can

estimate
μ1 −μ2 (regardless of n1 ∧n 2 using the following formula; given that σ 1 &σ 2 are

known.


2 2
σ1 σ2
μ1 −μ2 =X 1 −X 2 ±Z α /2 σ X −X σX − X 2=
1
√σ 2
X1 +σ
2
X2 = +
n1 n 2
1 2

When
σ 1 and σ 2 are not known, the standard error between two sample means ( σ x 1 −x 2 ) is

estimated by the sample standard error of the difference between two sample means,

1 2

S X −X = S + S =
S21 S22
2
X1 +
2
X2

n1 n2 , and the interval estimation takes the following form:

μ1 −μ2 =X 1 −X 2 ±Z α /2 S X − X
1 2, given that the sample sizes are large.
Example:
1. In a sex discrimination case, an employee alleged that a large corporation paid men more
than women for comparable work. Let population 1 represent all male employees
performing certain jobs and population 2 represent all female employees performing

comparable jobs at the corporation. Independent samples are taken of


n1 =100 males and

n2 =100 females; the sample means are x 1=Birr 20 ,600 and x 2 =Birr 19 , 700 , and the
sample standard deviations are
s1 =Birr 3 , 000 and s2 =Birr 2, 500 . Construct a 95%

confidence interval for


μ1 −μ2 . What do you conclude from this?

Solution:
Male employees Female employees
n1 =100 males n2 =100 females C= 0.95
x 1=Birr 20 ,600 x 2 =Birr 19 , 700
s1 =Birr 3 , 000 s2 =Birr 2, 500
Steps:
i. Calculate the (sample) standard error of the difference between two means


S 21 S22

2 2
(3 ,000 ) (2 , 500)
S X −X = + = + = √142 , 500=390 . 51
1 2 n1 n2 100 100

ii. Compute α /2
α = 1-C = 1- 0.95 = 0.05
α/2 = 0.05/2 = 0.025

iii. Look up
Z α /2=Z 0. 025 =1. 96

iv. Construct the confidence interval


μ1 −μ2 =X 1 −X 2 ±Z α /2 S X − X
1 2

μ1 −μ2 =(20 , 600−19 ,700 )±1 . 96(390 .51 ) = 900 ± 765.40

134.60 ≤
μ1 −μ2 ≤ 1,665.40

We state with 95% confidence that the mean salary difference between the male and female
workers lies between Birr 134.60 and Birr 1665.40

Because this interval contains only positive values, we can be quite confident that ( μ1 −μ2 ) > 0.
Thus, it is reasonable to assume that the mean salary for males exceeds the mean salary for
females.
2. A farmer wants to determine if different types of feed can influence the mean member of
eggs that hens lay per month. In a random sample of 100 hens that ate feed 1, the average
member of eggs per month was
x 1=15. 2 with variance 4. In a random sample of 100 hens

that ate feed2, the average number of eggs per month was
x 2 =14 with variance 4. Construct

a 95% confidence interval for


μ1 −μ2 . What do you conclude?

Solution:
Feed 1 Feed 2
n1 =100 hens n2 =100 hens C= 0.95
x 1= 15. 2 eggs x 2 =14 eggs
s21 =4 eggs s22 =4 eggs
Steps:
i. Calculate the (sample) standard error of the difference between two means

S X −X =
1 2

S 21 S22
+ =
n1 n2
4
+

4
100 100
=√ 0. 08=0 .283

ii. Compute α /2
α = 1-C = 1- 0.95 = 0.05
α/2 = 0.05/2 = 0.025

iii. Look up
Z α /2=Z 0. 025 =1. 96
iv. Construct the confidence interval
μ1 −μ2 =X 1 −X 2 ±Z α /2 S X − X
1 2

μ1 −μ2 =( 15 . 2−14 )±1 . 96( 0 . 283)

= 1.2 ± 0.5547

0.6453 ≤
μ1 −μ2 ≤ 1.7547

We state with 95% confidence that the mean number of eggs laid by hens which ate the two type
of feeds lies between 0.6543 eggs and 1.7547 eggs.
Since the interval contains only positive values, then those hens which ate feed type 1 are more
productive than those hens that ate feed type 2.

Confidence interval for the difference between two population proportions (


p1 − p2 )
We know that the unbiased estimator of the difference between the proportions of two

populations ( p1 − p2 ) is the difference between two sample proportions ( p1 − p2 ) , where each


sample is a random sample taken from the respective target population. Moreover, based on

CLT, if
n1 p1 , n 1 q1 and n 2 p 2, n2 q 2 are greater than 5, the sampling distribution of p1 − p2 is

( P1−P 2 ) −( P1 −P2 )
Z=

normal with √ P1 q 1 p2 q2
n1
+
n2

However, here
p1 andp 2 are unknown, and we want to estimate p1 andp 2 by p1 and p2

respectively, and hence Z becomes:


( P1−P 2 ) −( P1 −P2 )
Z=

√ P1 q 1 p2 q2
n1
+
n2 . That is,
σ p −p
1 2 is substituted by
Sp − p
1 2

Solving for
p1 − p2 results in:

P1 −P2 =p 1− p2 +Z
√ p1 q 1
n1
+
p 2 q2
n 2 , and since Z can assume both positive and negative
values, it becomes:

P1 −P2 =p 1− p2 ±Z
√ p 1 q1
n1
+
p2 q 2
n2
Since z represents the confidence level we write it as

Where:
P1 −P2 =p 1− p2 ±Z α /2
√ p1 q 1 p 2 q2
n1
+
n2

p1 = the sample proportion of success in the first sample


p2 = the sample proportion of in the second sample
q 1 = 1- p1
q 2 = 1- p2
n1 = sample size drawn from the first population

n2 = sample size drawn from the second population

α=1-C

This formula holds true provided that


n1 p1 , n1 q1 ≥5 and n2 p 2 , n2 q 2 ≥5 .
Example:
1. A TV executive is interested in determining if the proportion of people who watch a late-
night talk show is higher with the regular host or a guest host. In a random sample of 400
people, 175 watch the show when the regular host is on. In an independent random sample
of 500 people, 185 watch the show a guest host is on. Calculate a 95% confidence interval

for
p1 − p2 . What do you conclude?

Solution:
Regular host Guest Host
n1 = 400 p1 = 0.4375
n2 = 500 p2 = 0.37

X1 = 175
q 1 = 0.5625 X2 = 185
q 2 = 0.63
C = 0.95
i. Calculate the sample standard error of the diff. between two proportions

Sp − p =
1 2
√ p1 q 1 p2 q2
n1
+
n2
=

0 . 4375∗0 .5625 0 .37∗0 . 63
400
+
500
=0 . 033

ii. Compute α /2
α = 1-C = 1- 0.95 = 0.05
α/2 = 0.05/2 = 0.025

iii. Look up
Z α /2=Z 0. 025 =1. 96
iv. Construct the confidence interval

P1 −P2 =p 1− p2 ±Z α /2
√ p1 q 1 p 2 q2
n1
+
n2
=( 0 . 4375−037 ) ±1. 96 (0. 033 )
= 0.0675 ± 0.065

0.0025 ≤
p1 − p2 ≤ 0.1325

We state with 95% confidence that the true difference between


p1 − p2 is between 0.0025 and

0.1325. Since this interval contains only positive value it is reasonable to say that the proportion
of people who watch TV when the regular host is on is greater than when the guest host is on.
1.6. DETERMINATION OF SAMPLE SIZE
The reason for taking a sample from a population is that it would be too costly to gather data for
the whole population. But collecting sample data also costs money; and the larger the sample,
the higher the cost. To hold cost down, we want to use as small a sample as possible. On the
other hand, we want a sample to be large enough to provide “good” approximation/estimates of
population parameters. Consequently, the question is “How large should the sample be?”
The answer depends on three factors:
1) How precise (narrow) do we want a confidence interval to be?
2) How confident do we want to be that the interval estimate is correct?
3) How variable is the population being sampled?
Sample size for estimating population mean, μ
σ
μ= X ± Z α / 2
The confidence interval for μ is √n .
σ
Zα /2
From the above expression √ n is called error of estimation (e). That is, the difference
between x and μ which results from the sampling process. So
σ
Zα /2
e= √n
δ2 Z2 σ 2
e 2 =Z 2α / 2 n= α / 22
n . Solving for n results in, e
Squaring both sides results in

( )
2
Zα /2 σ
nμ =
e
Example:
1. A gasoline service station shows a standard deviation of Birr 6.25 for the changes made
by the credit card customers. Assume that the station’s management would like to
estimate the population mean gasoline bill for its credit card customers to be with in ±
Birr 1.00. For a 95% confidence level, how large a sample would be necessary?
Solution:

e = Birr 1.00 σ = Birr 6.25 C = 0.95


Z α /2=Z 0. 025 =1. 96

( )
2
Zα /2 σ
nμ =
e

( )
2
1. 96∗6 . 25
nμ =
1
= 150. 06 ≈ 151
2. The National Travel and Tour Organization (NTO) would like to estimate the mean
amount of money spent by a tourist to be with in Birr 100 with 95% confidence. If the
amount of money spent by tourist is considered to be normally distributed with a standard
deviation of Br 200, what sample size would be necessary for the NTO to meet their
objective in estimating this mean amount?
Solution:

e = Birr 100 σ = Birr 200 C = 0.95


Z α /2=Z 0. 025 =1. 96

( )
2
Zα /2 σ
nμ =
e

( )
2
1. 96∗200
nμ =
100
= 15. 37 ≈ 16

Sample size for estimating population proportion, p.

The confidence interval for p is


P = p±Z α /2
√ pq
n . The expression √
Z α /2
pq
n is called the error
term (e). That is,
e=Z α /2
√ pq
n , squaring both sides
pq
e 2 =Z 2α /2
n , solving for n
Z 2α /2 p q
np =
e2

Since we are trying to determine n, we cannot have p and q . Instead, we should have p

( )
2
Zα /2
np = pq
and q. so it becomes e
Example
1. Suppose that a production facility purchases a particular component parts in large lots
from a supplier. The production manager wants to estimate the proportion of defective
parts received from this supplier. She believes that the proportion of defects is no more
than 0.2 and wants to be with in 0.02 of the true proportion of defects with a 90% level of
confidence. How large a sample should she take?
Solution:

e = 0.02 p = 0.2 q =0.8 C = 0.90


Z α /2=Z 0. 05=1 . 64

( )
2
Zα /2
np = pq
e

np = (
1 . 64 2
0. 02 )
0 . 2∗0 . 8
=1075 .84 ≈1076
2. What is the largest sample size that would be needed in estimating a population
proportion to within ± 0.02, with a confidence coefficient of 0.95?
Solution:

e = 0.02 C = 0.95
Z α /2=Z 0. 025 =1. 96

The largest sample size would be obtained when p = 0.5. So,


( )
2
Zα /2
np = pq
e

( ) 0 . 5∗0. 5
2
1 . 96
np =
0. 02
=2401
If p is unknown and there is no possibility of estimating it, use 0.5 as the value of p because it
will generate the greatest possible sample size as compared with other values.

You might also like