Ratio Estimation Techniques Explained
Ratio Estimation Techniques Explained
Introduction
Estimation of the population mean and totals in preceding chapters was based
on measuring a variable, y. By measuring a variable, y, and a subsidiary
variable, x, on each element in the sample, we obtain additional information for
estimating the population parameter of interest. When a strong positive
relationship exists between the variables x and y, such that the line pass through
the origin, the ratio estimation procedure usually provides more precise
estimators of the population mean and total. When the line does not pass
through the origin, the regression estimation procedure is used.
Ratio Estimation
In a sample survey, in addition to estimating the means, totals and proportions,
we may wish to estimate also the ratio of two characters. For instance, the
interest may be in the estimation of ratio of sales of yam to sales of cassava,
ratio of male to female in a workplace, ratio of farmer’s income to non-farmer’s
income, etc.
In addition, if it is known that the regression line of the variable of interest y on
auxiliary variable x is a linear and passes through the origin, then the population
ratio of y to x can be estimated and used in the estimation of the population
mean (total) of y.
𝑪𝑰 = 𝒓 ± 𝒁∝⁄𝟐 𝑺𝑬
B=𝒁𝝈⁄𝟐 𝑺𝑬 WHEN 𝝈 = 𝟓%, THEN 𝒁𝝈⁄𝟐 = 𝟏. 𝟗𝟔 ≈ 𝟐
Estimated variance of 𝒓:
𝑛 𝑛 2
∑ 𝑦𝑖 𝑁−𝑛 1 ∑ (𝑦 −𝑟𝑥 )
𝑉̂ (𝑟) = 𝑉̂ [∑𝑖=1
𝑛 ]=( ) ( 2 ) 𝑖=1 𝑖 𝑖 ……………….(4.2)
𝑖=1 𝑥𝑖 𝑛𝑁 𝜇𝑥 𝑛−1
Bond on error of estimation:
𝑁−𝑛 1 ∑𝑛
𝑖=1(𝑦𝑖 −𝑟𝑥𝑖 )
2
2√𝑉̂ (𝑦) = 2√( ) (𝜇2 ) ………………(4.3)
𝑛𝑁 𝑥 𝑛−1
Note
2
If the population mean for x, µ𝑥 , is unknown, one can use 𝑥 to approximate 𝜇𝑥2
in equations (4.2) and (4.3).
Example
In a survey to examine trends in a real estate, an investigator is interested in the
relative change over a two-year period in the assessed value of homes in a
particular community. A simple random sample of n = 20 homes is selected
from the N = 1000 homes in the community. From tax records, the investigator
obtains the assessed value for this year(y) and corresponding value for two
years ago(x) for each of the n = 20 homes included in the sample. He wishes to
estimate R, the relative change in assessed value for the N = 1000 homes, using
information contained in the sample.
Using the data for real estate survey presented in the table below to estimate R,
the relative change in real estate valuation over the given two-year period and
place a bound on the error of estimation.
Home 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20
Assessed
value 2-
year ago
(xi) 20.2 25.4 26.1 29.5 24.3 22.1 23.7 24.9 21.5 28.2 28.6 26.9 25.2 24.1 23.9 23.1 27.5 30.2 31.4 29.3
Current
value (yi) 24.2 29.9 31.8 36.0 28.7 26.0 28.9 30.3 25.2 33.3 34.2 32.0 30.3 29.4 28.2 28.1 33.2 35.6 38.6 34.3
Solution
GRAPH OF REAL
ESTATE VALUATION
6
0
0.000 0.100 0.200 0.300
Assessed
value Current
2-year value
Home ago (xi) (yi) xi 2 yi 2 xi y i
1 20.2 24.2 408.04 585.64 488.84
2 25.4 29.9 645.16 894.01 759.46
3 26.1 31.8 681.21 1011.24 829.98
4 29.5 36.0 870.25 1296.00 1062.00
5 24.3 28.7 590.49 823.69 697.41
6 22.1 26.0 488.41 676.00 574.60
7 23.7 28.9 561.69 835.21 684.93
8 24.9 30.3 620.01 918.09 754.47
9 21.5 25.2 462.25 635.04 541.80
10 28.2 33.3 795.24 1108.89 939.06
11 28.6 34.2 817.96 1169.64 978.12
12 26.9 32.0 723.61 1024.00 860.80
13 25.2 30.3 635.04 918.09 763.56
14 24.1 29.4 580.81 864.36 708.54
15 23.9 28.2 571.21 795.24 673.98
16 23.1 28.1 533.61 789.61 649.11
17 27.5 33.2 756.25 1102.24 913.00
18 30.2 35.6 912.04 1267.36 1075.12
19 31.4 38.6 985.96 1489.96 1212.04
20 29.3 34.3 858.49 1176.49 1004.99
516.1 618.2 13497.7 19380.8 16171.8
618.2
𝑟= = 1.19783 ≈ 1.20 𝑡𝑜 2 𝑑. 𝑝.
516.1
Hence, the real estate valuation has increased approximately 20% over a two-
year period in the area studied.
The bound on error of estimation can be obtain by using (4.3).
But ∑𝑛𝑖=1(𝑦𝑖 − 𝑟𝑥𝑖 )2 = ∑𝑛𝑖=1 𝑦𝑖2 + 𝑟 2 ∑𝑛𝑖=1 𝑥𝑖2 − 2𝑟 ∑𝑛𝑖=1 𝑥𝑖 𝑦𝑖
∑𝑛𝑖=1(𝑦𝑖 − 𝑟𝑥𝑖 )2 = 19380.80 + (1.19783)2 (13497.73) − 2(1.19783)(16171.81)
= 5.14024
Using (4.3)
𝑁−𝑛 1 ∑𝑛
𝑖=1(𝑦𝑖 −𝑟𝑥𝑖 )
2
2√𝑉̂ (𝑦) = 2√( ) (𝜇2 )
𝑛𝑁 𝑥 𝑛−1
1000−20 1 5.14024
2√𝑉̂ (𝑦) = 2√( ) ((25.805)2) ( )
20(1000) 20−1
= 0.00892
CI≈ 1.19783 ±0.00892 AT 95% C.I
Estimator of the population total 𝑻𝒚 :
𝑛
̂ 𝒚 = ∑𝑖=1
𝑻 𝑛
𝑦𝑖
(𝑇𝑥 ) = 𝑟𝑇𝑥 , ………….(4.4)
∑ 𝑖=1 𝑥𝑖
̂𝒚 :
Estimated variance of 𝑻
𝑉̂ (𝑻
̂ 𝒚 ) = 𝑉̂ (𝑟𝑇𝑥 )
𝑛 2
̂ 𝒚 ) = (𝑇𝑥 )2 𝑉̂(𝑟) = (𝑇𝑥 2 ) (𝑁−𝑛) ( 12 ) ∑𝑖=1(𝑦𝑖 −𝑟𝑥𝑖 )
𝑉̂ (𝑻 ……….(4.5)
𝑛𝑁 𝜇𝑥 𝑛−1
Where 𝜇𝑥 and 𝑇𝑥 are the population mean and total, respectively, for random
variable x.
Bond on error of estimation:
𝑁−𝑛 1 ∑𝑛
𝑖=1(𝑦𝑖 −𝑟𝑥𝑖 )
2
2√𝑉̂ (𝑦) = 2√(𝑇𝑥 2 ) ( ) (𝜇 2 ) ………(4.6)
𝑛𝑁 𝑥 𝑛−1
Note
It is not necessary to know N or µ𝑥 , we must know 𝑇𝑥 . N is frequently
unknown. Consequently, the investigator must decide under what conditions use
of the ratio estimator, 𝑇̂𝑦 = 𝑟𝑇𝑥 is better than the corresponding estimator 𝑇̂𝑦 =
𝑁𝑦, where both estimators are based on srs. Generally, 𝑟𝑇𝑥 possesses a smaller
variance than 𝑁𝑦 when there is a strong positive correlation coefficient between
x and y (say 𝜌 > 0.5). This is because of the additional information.
Example
To estimate the total number of sugar content of a truck load of oranges, a
random sample of n = 10 oranges was juiced and weighed (see table below).
The total weight of all the oranges, obtained by first weighing the truck loaded
and then unloaded, was found to be 1800 pounds. Estimate 𝑇𝑦 , the total sugar
content for the oranges, and place a bound on the error of estimation.
Oranges 1 2 3 4 5 6 7 8 9 10
Sugar 0.021 0.030 0.025 0.022 0.033 0.027 0.019 0.021 0.023 0.025
content
(in lbs)
Weight of 0.40 0.48 0.43 0.42 0.50 0.46 0.39 0.41 0.42 0.44
orange
(lbs)
Solution
sugar content (in lbs) Vs, weight
of orange (lbs)
0.60
0.40
0.20
0.00
0.000 0.010 0.020 0.030 0.040
sugar weight
content of
(in lbs) orange
yi (lbs) xi yi 2 xi 2 xi y i
0.021 0.40 0.000441 0.16 0.0084
0.030 0.48 0.000900 0.23 0.0144
0.025 0.43 0.000625 0.18 0.01075
0.022 0.42 0.000484 0.18 0.00924
0.033 0.50 0.001089 0.25 0.0165
0.027 0.46 0.000729 0.21 0.01242
0.019 0.39 0.000361 0.15 0.00741
0.021 0.41 0.000441 0.17 0.00861
0.023 0.42 0.000529 0.18 0.00966
0.025 0.44 0.000625 0.19 0.011
0.246 4.350 0.006 1.904 0.108
From (4.4)
∑𝑛𝑖=1 𝑦𝑖 . 246
̂ 𝒚 = 𝑟𝑇𝑥 = 𝑛
𝑻 (𝑇𝑥 ) = (1800) = 101.79 𝑝𝑜𝑢𝑛𝑑𝑠.
∑𝑖=1 𝑥𝑖 4.35
Using (4.6), the Bond on error of estimation:
𝑁−𝑛 1 ∑𝑛
𝑖=1(𝑦𝑖 −𝑟𝑥𝑖 )
2
B=2√𝑉̂ (𝑦) = 2√(𝑇𝑥 2 ) ( ) (𝜇2 )
𝑛𝑁 𝑥 𝑛−1
𝑁−𝑛 𝑁
Since N is unknown, (4.6) is modify by assuming fpc, (
𝑁
) = (𝑁 − 𝑁𝑛 ) =
(1 − 0) is ≈ 1. Also sample mean 𝑥 is to replace 𝜇𝑥 ,
But ∑𝑛𝑖=1(𝑦𝑖 − 𝑟𝑥𝑖 )2 = ∑𝑛𝑖=1 𝑦𝑖2 + 𝑟 2 ∑𝑛𝑖=1 𝑥𝑖2 − 2𝑟 ∑𝑛𝑖=1 𝑥𝑖 𝑦𝑖
∑𝑛𝑖=1(𝑦𝑖 − 𝑟𝑥𝑖 )2 = .006224 + (. 0566)2 (1.9035) − 2(.0566)(.10839)
= 0.000052285
𝑁−𝑛 1 ∑𝑛
𝑖=1(𝑦𝑖 −𝑟𝑥𝑖 )
2
2√𝑉̂ (𝑦) = 2√(𝑇𝑥 2 ) ( ) (𝜇 2 )
𝑛𝑁 𝑥 𝑛−1
1 1 0.000052285
= 2√(1800)2 ( ) ( )
10 (0.435)2 10−1
= 6.308
At 95% CI= 101.79 ± 6.308
Estimator of the population mean µ𝒚 :
∑𝑛 𝑦
𝑖
µ̂𝒚 = ∑𝑖=1
𝑛 (µ𝑥 ) = 𝑟µ𝑥 , ………….(4.7)
𝑖=1 𝑥𝑖
𝑉̂ (µ̂𝒚 ) = 𝑉̂ (𝑟(µ̂𝒙 ))
𝑛 2
𝑁−𝑛 1 ∑ (𝑦 −𝑟𝑥 )
𝑉̂ (µ̂𝒚 ) = (µ̂𝒙 )2 𝑉̂ (𝑟) = (𝜇𝑥2 ) ( ) ( 2 ) 𝑖=1 𝑖 𝑖 ……….(4.8)
𝑛𝑁 𝜇𝑥 𝑛−1
Note
It is not necessary to know N or𝑇𝑥 to estimate µ𝑦 using the ratio procedure;
however, we must know µ𝑥 .
Example
A company wishes to estimate the average amount of money, µ𝑦 , paid to
employees for medical expenses, during the first three months of the current
calendar year. Average quarterly figures are available in the fiscal reports of the
previous year. A random sample of 100 employee records is taken from the
population of 1000 employees. The sample results are summarised below. Use
the data to estimate µ𝑦 and place a bound on the error of estimation.
𝑛 = 100, 𝑁 = 1000
∑100
𝑖=1 𝑦𝑖 =1750
∑100
𝑖=1 𝑥𝑖 =1200
Solution
The estimate of µ𝑦 is
µ̂𝒚 = 𝑟µ𝑥
Where
𝑇𝑥 12500
µ𝑥 = = = 12.5
𝑁 1000
Thus,
∑100
𝑖=1 𝑦𝑖 1750
µ̂𝒚 = 100 (µ𝑥 ) = (12.5) = 18.23
∑𝑖=1 𝑥𝑖 1200
The bound on the error of estimation can be found using (4.9); however, we
must first calculate
∑𝑛𝑖=1(𝑦𝑖 − 𝑟𝑥𝑖 )2 = ∑𝑛𝑖=1 𝑦𝑖2 + 𝑟 2 ∑𝑛𝑖=1 𝑥𝑖2 − 2𝑟 ∑𝑛𝑖=1 𝑥𝑖 𝑦𝑖
∑𝑛𝑖=1(𝑦𝑖 − 𝑟𝑥𝑖 )2 = 31650 + (1.4583)2 (15620) − 2(1.4583)(2205835)
=441.68
Substituting into (4.9), the bound on the error of estimation is
𝑁−𝑛 ∑𝑛
𝑖=1(𝑦𝑖 −𝑟𝑥𝑖 )
2
2√𝑉̂ (µ̂𝒚 ) = 2√( ) 𝑛𝑁 𝑛−1
1000−100 441.68
= 2√(
100(1000)
) (100−1)
= 0.42
Selecting sample size
We shall consider the sample size required to estimate a population parameter
R, µ𝒚 𝒐𝒓 𝑻𝒚 to within B units for simple random sampling using ratio
estimators.
Note that the procedure for obtaining the sample size, n, is identical to that
presented when discussing simple random sampling. The number of
observations required to estimate R, a population ratio, with a bound on the
error of estimation of magnitude B is determined by setting two standard
deviations of the ratio, r, equal to B and solving this expression for n. That is,
we must solve
2√𝑉(𝑟)=B (4.10)
For n. Although we have not discussed the form of V(r), you will recall that
𝑉̂ (𝑟), the estimated variance of r, is given by the formula
𝑛 2
𝑁−𝑛 1 ∑ (𝑦 −𝑟𝑥 )
𝑉̂ (𝑟) = ( ) ( 2 ) 𝑖=1 𝑖 𝑖 (4.11)
𝑛𝑁 𝜇𝑥 𝑛−1
Let rewrite (4.11) as
𝑁−𝑛 1
𝑉̂ (𝑟) = ( ) ( 2 )s2 (4.12)
𝑛𝑁 𝜇𝑥
Where
∑𝑛
𝑖=1(𝑦𝑖 −𝑟𝑥𝑖 )
2
s2 =
𝑛−1
An approximate population variance, V(r), can be obtained from 𝑉̂ (𝑟)by
replacing s2 with the corresponding 𝝈2. Thus, the number of observations
required to estimate R with a bound, B, on the error of estimation is determined
by solving the following equation for n;
𝑁−𝑛 1
2√𝑉(𝑟) = 2√(
𝑛𝑁
) (𝜇2 ) 𝜎 2 = B (4.13)
𝑥
∑𝑛
𝑖=1(𝑦𝑖 −𝑟𝑥𝑖 )
2
𝜎̂ 2 =
𝑛−1
Then, we substitute this quantity for 𝜎 2 in (4.14) and we find an approximate
sample size. If 𝜇𝑥 is also unknown, it can be replaced by the sample mean, 𝑥,
calculated from the n/ preliminary observations.
Example
A manufacturing company wishes to estimate the ratio of change from last year
to this year in the number of man hour lost due to sickness. A preliminary study
of n/ = 10 employee records is made and the results are tabulated below:
The company records show that the total number of man hour lost because of
sickness for the previous year was Tx=16,300. Use the data to determine the
sample size required to determine R, the rate of change for the company, with a
bound on the error of estimation of magnitude B=0,01. Assume the company
has 1000 employees (N=1000).
Solution
∑𝑛 𝑦 𝑖 187
𝑟 = ∑𝑖=1
𝑛 = = 1.05
𝑖=1 𝑥𝑖 178
Next we estimate 𝜎 2 using the data from the preliminary study. Thus,
∑𝑛
𝑖=1(𝑦𝑖 −𝑟𝑥𝑖 )
2
𝜎̂ 2 = ,
𝑛−1
𝑁−𝑛 1
2√𝑉(𝑟) = 2√(
𝑛𝑁
) (𝜇2 ) 𝜎 2 = B
𝑥
2𝜇𝑥 √𝑉(𝑟)= B
𝑁−𝑛 1
2𝜇𝑥 √( 𝑛𝑁 ) (𝜇2) 𝜎2 = 𝐵
𝑥
𝑵𝜎 2
𝒏= (4.16)
𝑵𝑫+𝜎 2
Where
𝐵2
𝑫= .
𝟒
Note that we need not know the value of 𝜇𝑥 to determine n in (4.14): however,
we do need an estimate of 𝜎 2 , either from prior information if it is available or
from information obtained in a preliminary study.
Example
An investigator wishes to estimate the average number of trees, 𝜇𝑦 , per acre on
a N = 1000 acre plantation. He plans to sample n one-acre plots and count the
number of trees, y, on each plot. He also has aerial photographs of the
plantation from which he can estimate the number of trees, x, on each plot for
the entire plantation. Hence he knows 𝜇𝑥 . Therefore it seems appropriate to use
ratio estimator of 𝜇𝑦 . Determine the sample size needed to estimate 𝜇𝑦 with a
bound on the error of estimation of magnitude B = 1.0.
Solution
Plot 1 2 3 4 5 6 7 8 9 10
Aerial estate(x) 23 14 20 25 12 18 30 27 8 31
Actual number(y) 25 15 22 24 13 18 35 30 10 29
∑𝑛 𝑦 𝑖 221
𝑟 = ∑𝑖=1
𝑛 = = 1.06
𝑖=1 𝑥𝑖 208
Next we estimate 𝜎 2 using the data from the preliminary study. Thus,
2 ∑𝑛
𝑖=1(𝑦𝑖 −𝑟𝑥𝑖 )
2
𝜎̂ = ,
𝑛−1
∑𝑛𝑖=1 𝑦𝑖2 = (25)2 + (15)2 + . . . + (29)2 = 5469
𝐵 2 (1)2
𝐷= = = 0.25
4 4
2
𝑁𝜎
̂ 100(4.21)
Thus 𝑛 = 2 = 100(0.25)+4.21 = 16.56 ≈ 17
𝑁𝐷+𝜎
̂
H/W
Determine the value of n when B = 0.5, 0.6. 0.7, 0.8, 0.9, 1.0, 1.1, 1.2, 1.3, 1.4, 1.5 and plot
the graph of n against B.
The sample size required to estimate 𝑇𝑦 , with a bound on the error of estimation of
magnitude B can be found by solving the following equation for n:
2√𝑉(𝑇̂𝑦 ) = B (4.17)
Stated differently,
𝑁−𝑛 1
2√𝑉(𝑟) = 2√(
𝑛𝑁
) (𝜇2 ) 𝜎 2 = B
𝑥
2𝑇𝑥 √𝑉(𝑟)= B
𝑁−𝑛 1
2𝑇𝑥 √(
𝑛𝑁
) (𝜇2 ) 𝜎 2 = B (*)
𝑥
𝑵𝜎 2
𝒏= (4.18)
𝑵𝑫+𝜎 2
Where
𝐵2
𝑫= .
𝟒𝑁2
Example
An auditor wishes to compare the actual naira value of an inventory of a
hospital, 𝑇𝑦, with the recorded inventory, 𝑇𝑥 . The recorded inventory, 𝑇𝑥 , can be
summarized from computer-stored hospital records. The actual inventory, 𝑇𝑦 ,
could be determined by examining and counting all hospital supplies, but this
process would be time consuming and costly. Hence, the auditor plans to
estimate 𝑇𝑦 based on a sample of n/ different items randomly selected from the
hospital’s supplies.
Records in the computer list N = 2100 different item types and the number of
each particular item in the hospital’s inventory. Using these data, a total value
for each item, x. can be obtained by multiplying the total number of each
recorded item by the unit value per item. The total naira value of the inventory
obtained from the computer is given by
𝑇𝑥 = 𝑠𝑢𝑚 𝑜𝑓 𝑡ℎ𝑒 𝑛𝑎𝑖𝑟𝑎 𝑣𝑎𝑙𝑢𝑒 𝑓𝑜𝑟 𝑡ℎ𝑒 𝑁 = 2100 𝑖𝑡𝑒𝑚𝑠
= ∑2100
𝑖=1 𝑥𝑖
In this instance 𝑇𝑥 was found to be N950,000. Determine the sample size
(number of items) needed to estimate 𝑇𝑦 with a bound on the error of estimation
of magnitude B = N500.
Solution
Because there is no prior information available, a preliminary study must be
conducted to estimate 𝜎 2 . Two men can determine the actual naira value, y, for
each of 15 items I one day. For this example, we shall use the data from a single
day’s inventory (n/=15) as preliminary study to obtain a rough estimate of 𝜎 2
and consequently a rough approximation of the required sample size, n.
Actually, the investigator would probably take a preliminary study of two or
three days’ inventory to provide a good approximation to 𝜎 2 and hence, n;
however, to simplify computations, we will consider a preliminary study of
n/=15 items. These data are summarized below along with the corresponding
computer figures (entries in hundreds of naira).
Items 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15
Naira value from
computer(xi) 15.0 9.5 14.2 20.5 6.7 9.8 25.7 12.6 15.1 30.9 7.3 28.6 14.7 20.5 10.9
Actual naira value(yi) 14.0 9.0 12.5 22.0 6.3 8.4 28.5 10.0 14.4 28.2 15.5 26.3 13.1 20.5 9.8
∑𝑛𝑖=1 𝑦𝑖 = 237.5
∑𝑛𝑖=1 𝑥𝑖 = 242.0
∑𝑛 𝑦 𝑖 237.5
𝑟 = ∑𝑖=1
𝑛 = = 0.9814 ≈ 0.98
𝑖=1 𝑥𝑖 242.0
2 ∑𝑛
𝑖=1(𝑦𝑖 −𝑟𝑥𝑖 )
2 104.2218
𝜎̂ = = = 7.444
15−1 14
The required sample size can now be found using (4.18). Note
𝐵2 (500)2
𝐷= = = 0.01417
4𝑁 2 4(21000)2
2
𝑁𝜎
̂ 21000(7.444)
Thus 𝑛 = 2 = 21000(0.01417)+7.444 = 420.2326 ≈ 421
𝑁𝐷+𝜎
̂
CLUSTER SAMPLING
Estimation of a Population Mean and Total
Cluster sampling is simple random sampling with each sampling unit containing a number of
elements. Hence, the estimators of the population mean, µ, and total, T, are similar to those for
simple random sampling. In particular, the sample mean, 𝑦, is a good estimator of the
population mean, µ. An estimator of µ and two estimators of T will be discuss.
∑𝑛 𝑦
𝑦 = ∑𝑛𝑖=1𝑚𝑖
𝑖=1 𝑖
Thus, 𝑦, takes the form of a ratio estimator as developed in chapter 4, with mi, taking the place
of xi. Then, the estimated variance of 𝑦, has the form of the variance of a ratio estimator given
by (4.2).
∑𝑛 𝑦
𝑦 = ∑𝑛𝑖=1𝑚𝑖 ……. (5.1)
𝑖=1 𝑖
Estimated variance of 𝒚:
𝑛 2
𝑁−𝑛 ∑𝑖=1(𝑦𝑖 −𝑦𝑚𝑖 )
𝑉̂ (𝒚) = ( 2) ………. (5.2)
𝑁𝑛𝑀 𝑛−1
The estimated variance in (5.2) is biased and a good estimator of 𝑉(𝒚) on if n is large, say n ≥
20. The bias disappears if the cluster sizes 𝑚1 , 𝑚2 , . . . , 𝑚𝑁 , are equal.
Example 5.1
The data on income of adult females in a particular city from 25 out of 415 clusters are
presented in the table below. Use the data to estimate the average income per adult female in
the city and place a bound on error of estimation,
Cluster i 1 2 3 4 5 6 7 8 9
Number of Adult
female(mi) 8 12 4 5 6 6 7 5 8
Total Income per Cluster
(yi) 96,000 121,000 42,000 65,000 52,000 40,000 75,000 65,000 45,000
Cluster i 10 11 12 13 14 15 16 17 18
Number of Adult
female(mi) 3 2 6 5 10 9 3 6 5
Total Income per Cluster
(yi) 50,000 85,000 43,000 54,000 49,000 53,000 50,000 32,000 22,000
Cluster i 19 20 21 22 23 24 25
Number of Adult
female(mi) 5 4 6 8 7 3 8
Total Income per Cluster
(yi) 45,000 37,000 51,000 30,000 39,000 47,000 41,000
Solution
The estimate of the population mean is given in (5.1)
∑𝑛 𝑦
𝑦 = ∑𝑛𝑖=1𝑚𝑖
𝑖=1 𝑖
∑25 𝑛
𝑖=1 𝑚𝑖 = 151 , ∑𝑖=1 𝑦𝑖 = 𝑁1,329,000
𝑁1,329,000
𝑦= = 𝑁8.801.
151
𝑛 2
𝑁−𝑛 ∑𝑖=1(𝑦𝑖 −𝑦𝑚𝑖 )
𝑉̂ (𝒚) = ( 2)
𝑁𝑛𝑀 𝑛−1
But
2
∑𝑛𝑖=1(𝑦𝑖 − 𝑦𝑚𝑖 )2 = ∑𝑛𝑖=1 𝑦𝑖2 + 𝑦 ∑𝑛𝑖=1 𝑚𝑖2 − 2𝑦 ∑𝑛𝑖=1 𝑚𝑖 𝑦𝑖
Substituting
∑𝑛
𝑖=1 𝑚𝑖 151
𝑚= = = 6.04
𝑛 25
Given N=415,
415−25 (15,227,502,247)
𝑉̂ (𝒚) = ((415)(25)(6.04)2) = 653,785
25−1
𝒚 ± 2√𝑉̂ (𝒚),
𝟖, 𝟖𝟎𝟏 ± 2√653,785
The population total, T, is now Mµ because M denotes the total number of elements in the
population. Consequently, as in simple random sampling, M𝒚 provides an estimate of T.
∑𝑛 𝑦
𝑴𝑦 = 𝑴 ∑𝑛𝑖=1𝑚𝑖 ……. (5.4)
𝑖=1 𝑖
𝑛 2
𝑁−𝑛 ∑𝑖=1(𝑦𝑖 −𝑦𝑚𝑖 )
𝑉̂ (𝑴𝑦) = 𝑴2 𝑉̂ (𝑦) = 𝑴2 ( 2)
𝑛−1
𝑁𝑛𝑀
𝑁−𝑛 ∑𝑛
𝑖=1(𝑦𝑖 −𝑦𝑚𝑖 )
2 𝑁 2 𝑁−𝑛 ∑𝑛
𝑖=1(𝑦𝑖 −𝑦𝑚𝑖 )
2
= 𝑴2 ( 𝑀 2
) = 𝑴2 (𝑀) ( 𝑁𝑛 )
𝑁𝑛( ) 𝑛−1 𝑛−1
𝑁
𝑁−𝑛 ∑𝑛
𝑖=1(𝑦𝑖 −𝑦𝑚𝑖 )
2
𝑉̂ (𝑴𝑦) = 𝑵2 ( 𝑁𝑛 ) ………. (5.5)
𝑛−1
𝑁−𝑛 ∑𝑛
𝑖=1(𝑦𝑖 −𝑦𝑚𝑖 )
2
2√𝑉̂ (𝑴𝑦) = 2√𝑵2 ( 𝑁𝑛 ) ………… (5.6)
𝑛−1
Note that the estimator 𝑴𝑦 is useful only if the number of elements in the population, M, is
known.
Example 5.3
Use the data in Example 5.2 to estimate the total income of all adult females in the city, and
place a bound on the error of estimation. There are 2,500 adult females in the city.
Solution
𝑴𝑦 = 2500(8.801) {from Example 5.2, 𝑦 = 8.801}
= N22,002,500
This is large bound on error of estimation, and it could be reduced by increasing the sample
size.
Note often, M, the number of elements in the population is not known in which cluster sampling
is appropriate. This makes it impossible to use the estimator 𝑴𝑦, but another estimator can be
form which does not depend on M. The quantity 𝒚𝒕
∑𝑛
𝑖=1 𝑦𝑖
𝒚𝒕 = …………………… (5.7)
𝑛
is an unbiased estimator of the average of the N cluster totals in the population. By the same
reasoning as employed in simple random sampling, 𝑵𝒚𝒕 , is an unbiased estimator of the sum
of the cluster totals or, equivalently, of the population total, T.
∑𝑛
𝑖=1 𝑦𝑖
𝑵𝒚𝒕 = 𝑵( ) ……. (5.8)
𝑛
𝑛 2
𝑁−𝑛 ∑ (𝑦𝑖 −𝑌𝑡 )
𝑉̂ (𝑵𝒚𝒕 ) = 𝑁 2 𝑉̂ (𝒚𝒕 ) = 𝑵2 ( 𝑁𝑛 ) 𝑖=1𝑛−1 ………. (5.9)
𝑁−𝑛 ∑𝑛
𝑖=1(𝑦𝑖 −𝑌𝑡 )
2
2√𝑉̂ (𝑵𝒚𝒕 ) = 2√𝑵2 ( 𝑁𝑛 ) ………… (5.10)
𝑛−1
Note
If there is a large amount of variation among the cluster sizes and if cluster sizes are highly
correlated with cluster totals, the variance of 𝑵𝒚𝒕 ( equation (5.9)) is generally larger than the
variance of 𝑴𝑦 ( equation (5.5)). The estimator 𝑵𝒚𝒕 does not use the information provided by
the cluster sizes 𝑚1 , 𝑚2 , . . . . , 𝑚𝑛 , and, hence, may be less precise.
Example 5.4
Use the data in Example 5.2 to estimate the total income of all adult females in the city if M
is not known. Place a bound on the error of estimation.
Solution
Given N = 415 from Example 5.1 and the table in Example 5.2, the estimate of the total income,
T, is
∑𝑛𝑖=1 𝑦𝑖 1,329,000
𝑵𝒚𝒕 = 𝑵 ( ) = 𝟒𝟏𝟓 ( ) = 𝑁22,061,400
𝑛 25
The result is fairly close to the estimate given in Example 5.2.
To place a bound on the error of estimation
𝑁−𝑛 ∑𝑛
𝑖=1(𝑦𝑖 −𝑌𝑡 )
2
𝑵𝒚𝒕 ± 2√𝑉̂ (𝑵𝒚𝒕 )= 𝑵𝒚𝒕 ± 2√𝑵2 ( 𝑁𝑛 ) 𝑛−1
415−25 (11,389,000)
=22,061,400 ± 2√(𝟒𝟏𝟓)2 ((415)(25)) 25−1
=22,061,400 ± 3,505,920
Where
1
∑𝑛𝑖=1(𝑦𝑖 − 𝑌𝑡 )2 = ∑𝑛𝑖=1 𝑦𝑖2 − (∑𝑛𝑖=1 𝑦𝑖 )2
𝑛
1
= 82,039,000,000 − 25 (1,329,000)2
= 11,389,000.
The bound on the error of estimation is slightly smaller than the bound for the estimator 𝑀𝑦
(Example 5.3). This is partly because the cluster sizes are not highly correlated with the cluster
total in this example. The unbiased estimator 𝑵𝒚𝒕 appears to be better than the estimator 𝑴𝒚.
Note
When all cluster sizes are equal (that is 𝑚1 = 𝑚2 =. . . . = 𝑚𝑛 ), the estimators µ and T possess
special properties:
(i) The estimator 𝒚 given in (5.1), is an unbiased estimator of the population mean, µ.
(ii) 𝑉̂ (𝒚) given in (5.2), is an unbiased estimator of the population variance of 𝒚,
(iii) The two estimators 𝑴𝒚 and 𝑵𝒚𝒕 of the population total, T, are equivalent.
Example 5.5
The circulation manager of a newspaper wishes to estimate the average number of newspapers
purchased per household in a given community. Travel costs from household to household are
substantial. Therefore, the 4000 households in the community are listed in 400 geographical
clusters of 10 households each, and a simple random sample of 4 clusters is selected. Interviews
are conducted with the following results:
Estimate the average number of newspapers per household for the community, and place a
bound on the error of estimation.
Solution
∑𝑛 𝑦
𝑦 = ∑𝑛𝑖=1𝑚𝑖
𝑖=1 𝑖
When 𝑚1 = 𝑚2 =. . . . = 𝑚𝑛 = 𝑚
∑𝑛
𝑖=1 𝑦𝑖 19+20+16+20
𝑦= = = 1.875
𝑛𝑚 4(10)
Substituting
400−4 (10,75)
𝑉̂ (𝒚) = ((400)(4)(10)2 ) 4−1 = 0.0089
𝒚 ± 2√𝑉̂ (𝒚),
𝟏. 𝟖𝟕𝟓 ± 2√0,0089