0% found this document useful (0 votes)
5 views5 pages

Chapter 5 Regression Estimation

Chapter 5 discusses difference and regression estimators, focusing on how to estimate changes between population characteristics over time or space. It introduces difference estimators as unbiased estimates of the difference between two population totals and explains the regression method of estimation, which improves conventional estimators by incorporating supplementary variables. The chapter also addresses bias and variance in regression estimators, noting their efficiency compared to ratio estimators under certain conditions.

Uploaded by

iceyxi002
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
5 views5 pages

Chapter 5 Regression Estimation

Chapter 5 discusses difference and regression estimators, focusing on how to estimate changes between population characteristics over time or space. It introduces difference estimators as unbiased estimates of the difference between two population totals and explains the regression method of estimation, which improves conventional estimators by incorporating supplementary variables. The chapter also addresses bias and variance in regression estimators, noting their efficiency compared to ratio estimators under certain conditions.

Uploaded by

iceyxi002
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Chapter 5.

Difference and Regression Estimators

I. Estimating a difference
Consider the question of estimation of changes or difference over time or space between
population characteristics or between populations for the same characteristic and of using the
estimator of certain types of differences to obtain better estimates for the population totals
and means of some characteristics.
Suppose Yˆ1 and Yˆ2 are unbiased estimators of the population totals Y1 and Y2 at the two end-
points of a specified period, then we get an unbiased estimator of the difference D=Y2-Y1 as
d = Yˆ2 − Yˆ1 .
This estimator may be termed a difference estimator. The sample variance of d is given by
( ) (
) ( )
V (d ) = V Yˆ1 − 2COV Yˆ1 , Yˆ2 + V Yˆ2

where the covariance term COV (Yˆ , Yˆ ) will be zero if the two samples selected at the two
1 2

points of time are independent or when the estimators are otherwise uncorrelated. Thus we
see that the variance of the difference estimator depends much on the correlation between Yˆ1

and Yˆ2 and that the variance decreases with increase in the correlation coefficient showing

that for efficient estimation of the differences, the correlation between the two estimators Yˆ1

and Yˆ2 should be positive and as large as possible.

II. Regression Method of Estimation


In ratio estimation, we considered the question of improving the conventional unbiased
X
estimator Yˆ by multiplying it with the factor , where X̂ is an unbiased estimator of the

total of a suitably chosen supplementary variable. Here we examine the possibility of
improving upon Yˆ by considering the estimator Xˆ − X , which is a zero function in the sense
that its expected value is zero. For instance, an unbiased estimator of Y can be taken as
(
Yˆ ′ = Yˆ + λ Xˆ − X )
where λ is a constant. The value of λ can be so fixed as to minimize the variance of Yˆ ′
which is given by
( ) () ( )
V Yˆ ′ = V Yˆ + 2λCOV Xˆ , Yˆ + λ2V Xˆ . ( )

1
Hence
( )
∂V Yˆ ′
( )
set
= 2COV Xˆ , Yˆ + 2λV Xˆ = 0 ( )
∂λ

=> λˆ = −
COV Xˆ , Yˆ( ) ∂ 2V Yˆ ′ ( ) > 0 => minimum
V Xˆ ( )
with
∂λ2 λˆ

Thus we see that the optimum value of λ if given by -β, where β is the regression coefficient
of Yˆ on X̂ and hence the estimator with the optimum value of λ is

(
Yˆr′ = Yˆ + β X − Xˆ . )
Its variance being
( ) [ ( )]
V Yˆr′ = V Yˆ + β X − Xˆ
() ( ) (
= V Yˆ + β 2V X − Xˆ + 2 βCOV X − Xˆ , Yˆ )
 COV (Xˆ , Yˆ )  COV (Xˆ , Yˆ )
2

= V (Yˆ ) +  ( ˆ ) + COV (− Xˆ , Yˆ )
V (X ) V (X )
 V X 2 
 ˆ  ˆ  

() [ (Xˆ , Yˆ )] V (Yˆ ) − 2 [COV (Xˆ , Yˆ )] V (Yˆ )


2 2
COV
= V Yˆ +
V (Xˆ ) V (Yˆ ) V (Xˆ ) V (Yˆ )
= V (Yˆ ) − [ρ (Xˆ , Yˆ )] V (Yˆ )
2

= V (Yˆ ){1 − [ρ (Xˆ , Yˆ )] }


2

where ρ (Xˆ , Yˆ ) is the correlation coefficient between X̂ and Yˆ . The estimator Yˆr′ is termed
the regression estimator and the procedure of estimation is known as regression method of
estimation.
N.B.:
(a) We see that if Yˆ and X̂ is highly correlated, then the estimator would be efficient.
(
(b) Yˆr′ is more efficient than Yˆ if ρ Xˆ , Yˆ is non-zero. )
(c) Comparing with the variance of the ratio estimator
( ) () ( ) ( )
V YˆR = V Yˆ − 2 R ⋅ COV Xˆ , Yˆ + R 2V Xˆ
−) V (Yˆ ′ ) = V (Yˆ ) − [ρ (Xˆ , Yˆ )] V (Yˆ )
2
r

V (Yˆ ) − V (Yˆ ′ ) = [ρ (Xˆ , Yˆ )] V (Yˆ ) − 2 R ⋅ COV (Xˆ , Yˆ ) + R V (Xˆ )


2 2
R r

= [ρσ (Yˆ ) − rσ (Xˆ )]


2

which shows that a regression estimator is more efficient than the corresponding ratio
estimator in general and that they are equally efficient when β=R, i.e. when the line of
regression passes through the origin. In actual practice the exact value of β may not be
known and it may have to be estimated on the basis of a sample. If β̂ is an estimator of β,

2
( )
we get Yˆr = Yˆ + βˆ X − Xˆ . This estimator is generally biased for Y and its bias and variance
are consider next.

Comment:
It may be mentioned that for this estimator to be efficient it is not necessary to get the exact
value or an estimate of β based on a current sample and that even an approximation to it
available from previous survey or census may be sufficient. However, the closer the
approximation, the higher will be the efficiency. So long as the value of β used is


uncorrelated with X̂ , the regression estimator remains unbiased. Noting that βˆ = , when

the line of regression of Yˆ on X̂ passes through the origin, we see that the regression

estimator reduces to the ratio estimator YˆR = X .

III. Bias and variance


The regression estimator
(
Yˆr = Yˆ + βˆ X − Xˆ )
is biased since
(i) The regression coefficient β is generally estimated by the ratio of an estimator of

( ) ( )
COV Xˆ , Yˆ to that of V Xˆ and

(ii) It involves the product of two estimators β̂X̂ .

Writing Yˆ = Y (1 + e ) , Xˆ = X (1 + e′) and βˆ = β (1 + e′′) and substitute in the estimator above,


we get
Yˆr = Y + (eY − e′βX ) − e′e′′βX
The bias of the regression estimator is given by
( ) ( ) ( )
B Yˆr = E Yˆr − Y = − E (e′e′′)βX = −COV Xˆ , βˆ since E(e’) = E(e”) = 0.

(
N.B. For n large, the bias is expected to be negligible, because usually COV X̂ , β̂ will )
decrease as the sample size increases.

1. If the sample is selected in the form of m independent sub-samples then the bias can be
unbiasedly estimated by

3
( )
ˆ
b Yr = −
1 m ˆ
∑ X i − Xˆ βˆ i − βˆ
m − 1 i =1
( )( )
1 m
where β̂ i and X̂ i are estimators of β and X based on the ith sub-sample and βˆ = ∑ βˆ i
m i =1

1 m
and Xˆ = ∑ Xˆ i , provided β̂ in Yˆr is also taken as the means of β̂ i ’s.
m i =1
2. In sampling n units with SRSWOR, an estimator of β, the coefficient of the regression of
y and x , which is the same as that of y on x is given by
n

∑ (x i − x )( y i − y )
β̂ = i =1
n
.
∑ (x − x)
2
i
i =1

Substituting in Yˆr = Yˆ + βˆ X − Xˆ , we get ( )


[
Yˆr = N y + β̂ (X − x ) . ]
The bias of the estimator is given by
( )
B Yˆr = − N ⋅ COV x , β̂ , ( )
with variance to the first approximation

( )
V Yˆr =
( )
N 2 1 − ρ 2 σ y2 N − n
.
n −1 N −1
An estimator of V Yˆr is given by ( )
( )
v Yˆr = N 2
1− f n

n(n − 1) i =1
[
( yi − y ) − βˆ (xi − x )
2
]
  
2
n
n ∑ ( xi − x )( y i − y ) 
2 1− f   i =1  
∑ ( y i − y ) −
2
=N 
n(n − 1)  i =1 n

∑ ( xi − x ) 2 
 
 i =1

Comments:
(a) The regression estimator is not commonly used in practice due to the fact that the
calculation of the estimate of the regression coefficient in large scale surveys becomes
cumbersome and time consuming.

4
(b) Further, since the regression line passes through the origin or close to the origin in
most of the cases usually not with, the ratio estimator is generally used instead of the
more complicated regression estimator.
(c) It is of interest to note that when the regression of y on x is perfectly linear, i.e. when
ρ =1, the variance becomes zero, and that if y and x are uncorrelated, the variance is
the same as in the case of the conventional unbiased estimator.

Comments:
When do ratio estimators offer substantial improvement over simple unbiased (mean per unit)
estimations:
(a) We must be able to observe simultaneously 2 variables x and y which appear to be
roughly proportional to each other (i.e. highly correlated).
(b) The auxiliary variable x must not have a substantially greater coefficient of variation than
y.
(c) The population mean X or total X must be known exactly.

You might also like