Horvitz Thompson (HT) Estimator of Population Mean:
Define a random variable i (i 1, 2,.., N ) as
1 if Yi is included in a sample " s " of size n
i
0 otherwise.
nyi
Let zi , i 1...N assuming E ( i ) 0 for all i
NE ( i )
where E ( i ) 1.P (Yi s ) 0.P (Yi s ) i
is the probability of including the unit i in the sample and is
called as inclusion probability.
The HT estimator of Y based on a sample y1 , y2 ,..., yn is
1 n 1 N
zn YˆHT zi i zi .
n i 1 n i 1 2
Unbiasedness:
ˆ 1 N
E (YHT ) E ( zi i )
n i 1
1 N
zi E ( i )
n i 1
1 N nyi
E ( i )
n i 1 NE ( i )
1 N nyi
Y
n i 1 N
which shows that HT estimator is an unbiased estimator of
population mean.
3
Variance:
V (YˆHT ) V ( zn )
E ( z ) E ( zn )
2 2
n
E ( zn2) Y 2 .
Consider
2
1 N
E ( zn2 ) 2
E
n i 1
z
i i
1 N 2 2 N N
2 E i zi i j zi z j
n i 1 i ( j ) 1 j 1
1 N 2 N N
2 zi E ( i ) zi z j E ( i j ) .
2
n i 1 i ( j ) 1 j 1
4
Variance:
If S = {s} is the set of all possible samples and i is probability
of selection of ith unit in the sample s then
E ( i ) 1 P( yi s ) 0.P( yi s )
1. i 0.(1 i ) i
E ( i2 ) 12. P( yi s ) 02.P( yi s )
i.
So E (i ) E (i2 )
N
1 N N
E ( zn ) 2 zi i
2 2
ij zi z j
n i 1 i ( j ) 1 j 1
where ij is the probability of inclusion of ith and jth unit in the
sample. This is called as second order inclusion probability. 5
Variance:
Now
Y E ( zn )
2 2
2
1 N
2 E i zi
n i 1
1 N 2 N N
2 zi E ( i ) zi z j E ( i ) E ( j )
2
n i 1 i ( j ) 1 j 1
1 N 2 2 N N
2 zi i i j zi z j .
n i 1 i ( j ) 1 j 1
6
Variance:
Thus
ˆ 1 N N N 1 N 2 2 N N
Var (YHT ) 2 i zi ij zi z j 2 i zi i j zi z j
2
n i 1 i ( j ) 1 j 1 n i 1 i ( j ) 1 j 1
1 N N N
2 i (1 i ) zi ( ij i j ) zi z j
2
n i 1 i ( j ) 1 j 1
1 N n 2 yi2 N N n 2 yi y j
2 i (1 i ) 2 2 ( ij i j ) 2
n i 1 N i i ( j ) 1 j 1 N
i j
1 N 1 2 N N
2 yi y j .
ij i j
i
i
y
N i 1 i i ( j ) 1 j 1 i j
7
Estimate of Variance:
ˆ 1 n y 2 (1 ) n n yi y j
ˆ
V1 Var (YHT ) 2 .
i i ij i j
i j
2
N i 1 i i ( j ) 1 j 1 ij
This is an unbiased estimator of variance .
8
Drawback:
yi
It does not reduces to zero when all are same, i.e.,
i
when yi i .
Consequently, this may assume negative values for some
samples.
A more elegant expression for the variance of yˆ HT has
been obtained by Yates and Grundy.
9
Yates and Grundy Form of Variance:
Since there are exactly n values of i which are 1 and (N ‐ n)
values which are zero, so
N
i 1
i n.
N
Taking expectation on both sides E ( ) n.
i 1
i
Also
2
N
N N N
E i E ( i ) E ( i j )
2
i 1 i 1 i ( j ) 1 j 1
N N N
E n E ( i ) E (
2
i J ) (using E ( i ) E ( i2 ))
i 1 i ( j ) 1 j 1
10
Yates and Grundy Form of Variance:
Also
2
N
N N N
E i E ( i ) E ( i j )
2
i 1 i 1 i ( j ) 1 j 1
N N N
E n E ( i ) i J
2
E ( ) (using E ( i ) E ( 2
i ))
i 1 i ( j ) 1 j 1
N N
n n
2
E (
i ( j ) 1 j 1
i J )
N N
E (
i ( j ) 1 j 1
i J ) n(n 1).
Thus
E ( i j ) P( i 1, j 1)
P( i 1) P( j 1| i 1)
E ( i ) E ( j | i 1).
11
Yates and Grundy Form of Variance:
Therefore
N N
j ( i ) 1
E ( i j ) E ( i ) E ( j )
j ( i ) 1
E ( i ) E ( j | i 1) E ( i ) E ( j )
N
E ( i )
j ( i ) 1
E ( j | i 1) E ( j )
E ( i ) (n 1) (n E ( i ))
E ( i ) 1 E ( i ) i (1 i ). (1)
N
Similarly
i ( j ) 1
E ( i j ) E ( i ) E ( j ) j (1 j ). (2)
We had earlier derived the variance of HT estimator as
ˆ 1 N N N
Var (YHT ) 2 i (1 i ) zi ( ij i j ) zi z j .
2
n i 1 i ( j ) 1 j 1
Use 1 and 2 in this expression. 12
Yates and Grundy Form of Variance:
Using (1) and (2) in this expression, we get
ˆ 1 N N N N
Var (YHT ) 2 i (1 i ) zi j (1 j ) z j 2 ( i j ij ) z i z j
2 2
2n i 1 j 1 i j 1 j 1
1 N N 2
2 E ( i j ) E ( i ) E ( j ) zi
2n i 1 j ( i ) 1
N 2
E ( i j ) E ( i ) E ( j ) z j 2 E ( i ) E ( j ) E ( i j ) zi z j
N N n
j 1 i ( j ) 1 i ( j ) 1 j 1
1 N N N N N N
2 ( ij i j ) zi ( ij i j ) z j 2 ( ij i j ) zi z j
2 2
2n i ( j ) 1 j 1 i ( j ) 1 j 1 i ( j ) 1 j 1
1 N N
2 ( i j ij )( zi z j 2 zi z j ) .
2 2
2n i ( j ) 1 j 1
13
Yates and Grundy Form of Variance:
The expression for i and ij can be written for any given sample size.
For example, for n = 2, assume that at the second draw, the
probability of selecting a unit from the units available is proportional
to the probability of selecting it at the first draw.
Since
E(i Probability of selecting Yi in a sample of two
E ( i ) Pi1 Pi 2
where Pir is the probability of selecting Yi at rth draw (r = 1, 2).
14
Yates and Grundy Form of Variance:
If Pi is the probability of selecting the rth unit at the first draw then
we had earlier derived that
Pi1 Pi
yi is not selected yi is selected at the 2 nd draw|
Pi 2 P P
at the 1 draw
st
yi is not selected at the 1 draw
st
N Pj Pi
j ( i ) 1 1 Pj
N Pj Pi
Pi .
j 1 1 Pj 1 Pi
So
N Pj Pi
E ( i ) Pi .
j 1 1 Pj 1 Pi 15
Yates and Grundy Form of Variance:
Again
E ( i j ) Probability of including both yi and y j in a sample of size two
Pi1 Pj 2|i Pj1 Pi 2| j
Pj Pi
Pi Pj
1 Pi 1 Pj
1 1
PP
i j .
1 Pi 1 Pj
16
Estimate of Variance:
The estimate of variance is given by
ˆ 1 n n i j ij
Var (YHT ) 2
2n
i ( j ) j 1 ij
( z i z j ) 2
.
17
Midzuno System of Sampling:
Under this system of selection of probabilities, the unit in the
• first draw is selected with unequal probabilities of selection
(i.e., pps) and
• remaining all the units are selected with SRSWOR at all
subsequent draws.
18
Midzuno System of Sampling:
Under this system
E ( i ) i P (unit i (U i ) is included in the sample)
P (U i is included in 1st draw) + P (U i is included in any other draw )
Probability that U i is not selectedat the first draw and
Pi
is selected at anyof subsequent ( n -1) draws
n 1
Pi (1 Pi )
N 1
N n n 1
Pi .
N 1 N 1
19
Midzuno System of Sampling:
Similarly
E ( i j ) Probability that both the units U i and U j are in the sample
Probability that U i is selected at the first draw and
U is selected at any of the subsequent ( n 1) draws
j
Probability that U j is selected at the first draw and
U is selected at any of the subsequent ( n 1) draws
i
Probability that neither U i nor U j is selected at the first draw but
both of them are selected during the subsequent ( n 1) draws
n 1 n 1 (n 1)(n 2)
Pi Pj (1 Pi Pj )
N 1 N 1 ( N 1)( N 2)
(n 1) N n n2
( P P )
( N 1) N 2 N 2
i j
n 1 N n n2
ij ( P P ) .
N 1 N 2 N 2
i j 20
Midzuno System of Sampling:
Similarly,
E ( i j k ) ijk Probability of including U i , U j and U k in the sample
(n 1)(n 2) N n n3
( P P P ) .
( N 1)( N 2) N 3 N 3
i j k
By an extension of this argument, if Ui , Uj ,…, Ur are the r units
in the sample of size n(r < n) the probability of including these r
units in the sample is
E ( i j ... r ) ij ...r
( n 1)( n 2)...( n r 1) N n nr
( P P ... P ) .
( N 1)( N 2)...( N r 1) N r N r
i j r
21
Midzuno System of Sampling:
Similarly, if U1 , U2 ,…, Uq be the n units, the probability of
including these units in the sample is
( n 1)( n 2)...1
E (i j ... q ) ij ... q ( Pi Pj ... Pq )
( N 1)( N 2)...( N n 1)
1
( Pi Pj ... Pq )
N 1
n 1
which is obtained by substituting r = n.
22
Midzuno System of Sampling:
Thus if Pi ' s are proportional to some measure of size of units in
the population then the probability of selecting a specified
sample is proportional to the total measure of the size of units
included in the sample.
Substituting these i , ij , ijk etc. in the HT estimator, we can
obtain the estimator of population’s mean and variance. In
particular, an unbiased estimate of variance of HT estimator
given by
ˆ 1 n n i j ij
Var (YHT ) 2
2n
i j 1 j 1
( z i z j ) 2
ij
where N n n 1
i j ij ( N n ) PP (1 P Pj .
)
( N 1) 2 N 2
i j i
23
Midzuno System of Sampling:
The main advantage of this method of sampling is that it is
possible to compute a set of revised probabilities of selection
such that the inclusion probabilities resulting from the revised
probabilities are proportional to the initial probabilities of
selection.
It is desirable to do so since the initial probabilities can be
chosen proportional to some measure of size.
24