Econometric Methods
Ae
Rg
a
OT
Loe
i ile
abit alt
AUVea,Rei
faite
o,
ONIN
LPP de
RAte
NEr
danni
Johnston, J.(John),
O at
Te
[AO
| 2 as a
> | co
by Mm
Pp|< 7
mt :
V
jo e CN) eel =
ECONOMETRIC
METHODS
Third Edition
J. Johnston
University of California, Irvine
197964
IN MEMORY OF
B. and J.
ECONOMETRIC METHODS
ESBNeG-O?-03eb65-1
Preface Vil
iii
iv CONTENTS
Index 561
ae vachefate
eeie: Mt end ig Mi
ie “whe aah A
Oe
‘WAySee iér ne sil a gue at
«MC NA Mien ines
mi
ole
SS ’ To ; ais OThiiig Bes
ae t ond ’ eat dheott Cyee pee ae ;
aa ae
PREFACE
This edition has been completely rewritten. The main features of the new edition
are the following:
1. The mathematical, statistical threshold has been lowered in order to make the
material more accessible to students with only an elementary prior knowledge
of statistics. This has resulted in a somewhat larger proportion of words to
symbols in the early chapters than would otherwise have been the case. A
series of paragraphs on mathematical, statistical topics has also been provided
in Appendix A. These are keyed into the early chapters to ease the transition
into the heart of the book. For the same reason the chapter on matrix algebra
has been retained and, indeed, expanded to include a geometric as well as an
algebraic treatment of some topics.
2. All the inference procedures for the general linear model have been derived as
special cases of a single basic procedure, namely, the testing of a set of linear
restrictions on the parameters of the model (Chapter 5). This leads in turn to
an exhaustive treatment of tests for structural change (Chapter 6). Chapter 6
also contains extended treatments of the use of dummy variables and of
multicollinearity among the regressors.
3. Every effort has been made to cover both new and old topics on which
substantial work has been done in recent years and which are thought to be
significant and enduring rather than passing fancies. Such topics include the
estimation of sets of equations with special reference to transcendental
logarithmic approximations and applications in energy economics (Chapter 8),
autocorrelated error terms (Chapter 8), time series techniques (Chapter 9) and,
in A Smorgasbord of Further Topics (Chapter 10), the following “menu”:
recursive residuals, spline functions, pooling of time-series and cross-section
data, variable-parameter models, qualitative dependent variables, and errors
in variables. The author has also granted himself the indulgence. of some
personal comments on the present state of econometrics (Chapter 12).
vii
Vili PREFACE
4. The problem sets have been extended to become truly Anglo-American, with
offerings from the Royal Statistical Society plus the universities of Cambridge,
London, Manchester, and Oxford on one side of the pond, and Chicago,
Michigan, Yale, and Washington on the other side. Grateful acknowledgment
is made to various anonymous authorities for the first set, and to Arnold
Zellner, Jan Kmenta, Peter Phillips, and Charles Nelson for the second.
Appendix B also contains an extensive set of statistical, econometric tables,
and grateful acknowledgment is made to the appropriate sources for permis-
sion to publish them.
J. Johnston |
CHAPTER
ONE
THE NATURE OF ECONOMETRICS
Before asking the question, “What is econometrics?,” one must pose the prior
question, “What is economics?” The answer to the second question will indicate
the role that econometrics can play in the development of economics. Although
the focus of the exposition in this chapter will be on economic models, the
methods that have been developed in econometrics can and do play an important
role in other social sciences, where there is a concern with building and estimating
models of the interconnections between various sets of variables in a predomi-
nantly nonexperimental situation.
As an example of the model building process let us consider one of the simplest
forms of the national income model, which is used as a pedagogic device in most
elementary textbooks on economics. Such models begin with the national income
identity. For a closed economy with no foreign trade, this identity in any period is
Y =cChi ae (1-1)
where y= gross national product (GNP)
c=consumption expenditure
i= investment expenditure
g=government expenditure
all expenditure flows being measured in real terms. The construction of the model
proceeds with the formulation of hypotheses about the determinants of the
expenditure components of GNP.
Consumption expenditure might be hypothesized as dependent on disposable
income, net of tax, and the rate of interest. Thus we write}
c= f((1-7)y,r) (1-2)
where T= tax rate (assumed constant across the economy)
r=rate of interest
argument. The first assumption in Eq. (1-3) is that the marginal propensity to
consume out of disposable income is a positive fraction less than unity. The
second assumption is that a rise in the rate of interest will have a depressing effect
on consumption since it raises the return on savings, increases the cost. of
financing «consumer durables, and also reduces the nominal value of bonds, which
are a part of wealth, which in turn might appear as an argument of the
consumption function but has been omitted from Eq. (I--2) on grounds of
simplicity.
The investment function may be specified as
i= f(Ay,r) (1-4)
with
f a 0, h <0 (1-5)
The term A y indicates the change in GNP. Investment is positively influenced by
profit expectations, and the crude assumption here is that observed changes in
real GNP serve as a proxy for these profit expectations. The rate of interest is
again expected to be negatively related to this form of expenditure.
Collecting results, so far we have a three-equation model, namely,
Scale ge
Cm (Vy f)
Aya)
supplemented by the expected signs on derivatives expressed in Eqs. (1-3) and
(1-5). This model then constitutes a theory about the joint determination, or
“explanation,” of the three variables c, i, and y. Such an explanation is obviously
conditional on the values assumed for g, r, and r. The model builder now faces a
decision on how to treat these remaining variables. Should one formulate theories
to explain the determination of government expenditure, the rate of interest, and
the tax rate, thus expanding the system to one of six equations? If one does, the
new equations will almost certainly contain some explanatory variables on the
right-hand side that have not previously appeared in the system, and these, in
turn, raise the question of how they are to be treated. It might seem that economic
models must become infinitely large, but there is not, of course, an infinite
number of variables to be explained. In any case the behavior of model builders is
very pragmatic. Everything is relative: all depends on the problem at hand. For
some purposes a small model is sufficient and some variables, which in larger
models would have explanatory equations, may be left “unexplained.” In the
present instance we make no pretense at economic realism, but only require a
model for illustrative purposes, so we will restrict it to the three equations already
specified.
The model contains only two behavioral relations, one for consumption and
the other for investment. Economic theory has done two things. First, it has
specified the list of explanatory variables on the right-hand side of each equation,
and second, it has indicated the expected signs on the partial derivatives. This is
usually as far as theory per se can go, but it still leaves a series of important
questions unanswered.
4 ECONOMETRIC METHODS
Lag structure Somewhat allied with problems of data definition are problems of
lag structure. Should investment be specified as responding to the current interest
rate or to some set of previous interest rates in view of the inevitable time lags
involved in making and implementing investment decisions? Again, by the nature
of things, economic theory cannot be specific about appropriate lag structures.
Moreover, much of economic theorizing has necessarily been about equilibrium
positions, as, for example, the equilibrium rate of consumption corresponding to
some level of income, which has, in theory, remained constant long enough for
consumers to become fully adjusted to it. In practice, the world is always
THE NATURE OF ECONOMETRICS 5
Choice between theories So far, in discussing the previous four problems, we have
implicitly assumed that our theoretical model is “correct,” but how can we tell
whether a theory is sufficiently correct to be used as a valid tool of analysis?
Perhaps there are as many theories as there are theorists. There is, in practice, a
very important and very difficult problem involved in attempting to discriminate
between competing theories. Some theoretical models differ in degree but not in
kind. They might be regarded as variations on a theme. For example, another
theorist might accept the general form of our consumption and investment
functions but wish to add wealth as an additional explanatory variable to the first
equation and capital stock to the second. At the other end of the spectrum would
be a theorist who rejected the Keynesian flavor of our model and advanced
instead a supply-determined theory of output or a model in which the fundamen-
tal driving force was the money supply.
Econometrics tackles all five questions. Its basic task is to put empirical flesh and
blood on theoretical structures. This involves several crucial steps. First of all, the
‘theory or model must be specified in explicit functional form. The econometrician
does not have any special insights in this area that are denied to the economic
theorist, so one usually starts with the simplest functional forms that are con-
sistent with the a priori specifications. At the same time one makes an initial
specification of the lag structure. As an example we might specify the three-equa-
6 ECONOMETRIC METHODS
Ve Cp Bg, (1-8)
with a priori expectations
USar<st, a, < 0, B; > 0; B, <0
The subscripts on the variables refer to time periods. The unit time period can be
anything considered relevant by the econometrician, provided there exist ap-
propriate data in terms of that unit. However, it is typically a quarter or a year,
and the model is in discrete, not continuous, time.
The second task of the econometrician is to decide on the appropriate data
definitions and assemble the relevant data series for the variables which enter the
model. The third task is to perform a “marriage” of theory and data by means of
statistical methods. The “offspring” of the marriage are various sets of statistics,
which shed crucial light on the validity of the theoretical model that has been
specified. The most important set consists of the numerical estimates of the
parameters of the structural form. The Greek letters of Eqs. (1-6) and (1-7) are
now replaced by numbers. There are further statistics which enable one to assess
the reliability or precision with which these parameters have been estimated,
which in turn helps us to check whether the model conforms to the theoretical
expectations about signs of derivatives. There are still further statistics and
diagnostic tests that help one to assess the performance of the model and decide
whether or not to proceed sequentially by modifying the specification in certain
directions and testing out the new variant of the model against the data.
Most of this book will be concerned with the statistical methods used by
econometricians in estimating, testing, and evaluating economic models. Histori-
cally, econometrics started with the corpus of methods inherited from classical
Statistics. These methods, however, were mainly developed in the context of the
experimental sciences. Special problems of statistical inference arise in economics,
where the possibility of controlled experiments is the exception, not the rule, and
these will be described in the chapters to follow. All that remains to be done in
this introductory chapter is to indicate some of the possible applications of an
econometric model, once it has been estimated. This will again be done with the
simple model outlined above.
Equations (1-6) to (1-8) constitute the structural form of the model. The structural
form may be regarded as a theoretical explanation, or hypothesis, about the
determination of the three variables y,, c,, and i,, conditional on the values
currently assumed by g, and r, and also on the recent history of the system as
represented by y,_,, y,5, and r_,. This enables us to make the following
THE NATURE OF ECONOMETRICS 7
The important point about Eq. (1-9) is that only one current endogenous variable
appears in the equation, namely, y, on the left-hand side. The right-hand-side
variables are a mixture of current exogenous variables and lagged variables,
whether endogenous or exogenous. This collection of three sets of variables is
labeled the class of predetermined variables, since, from the viewpoint of the
model. in period ¢, their values either are determined -by_the past history of the
system or_are [Link] in the current period. The investment equation
already has nothing but predetermined variables on the right-hand side, so we
repeat it here:
i, = Bo + B\(Y-1 — Y-2) + Boni (1-10)
Finally, substituting Eq. (1-9) in the consumption function gives
The three Eqs. (1-9), (1-10), and (1-11) constitute the reduced form of the
model. Each equation of the reduced form expresses a current endogenous
variable as a function only of predetermined variables. The reduced form may be
8 ECONOMETRIC METHODS
Exogenous variables
current and lagged ~~_,
Predetermined Current endogenous
variables variables
Lagged endogenous oat
is variables
Figure 1-1
written compactly as
The 7’s of the reduced-form equations are economically very important parame-
ters. They measure the impact in the current period on each endogenous variable
of a unit change in any predetermined variable. Consider, for example, a unit
increase in the level of g,. From Eq. (1-8) of the structural form there would be a
simultaneous increase of one unit in GNP. But from the consumption function
(1-6), increases in GNP will induce increases in consumption, which in turn, from
Eq. (1-8), will induce further increases in GNP. The reduced-form coefficient
dy, 1
oe eras \
shows the end result of this process in period t. This is the national income
multiplier of simple Keynesian theory. For example, if + = 0.25 and a, = 0.8,
7, = 2.5, so that a unit increase in government expenditure, with tax rates and
all other parameters unchanged, would raise national income in the same period
by 2.5 units. Similarly, an inspection of
To = (a + By)
shows that a unit increase (upward shift) in the intercept of either the consump-
tion or the investment function would have equal multiplier effects on GNP. All
THE NATURE OF ECONOMETRICS 9
the 7’s are multipliers, and they are termed impact multipliers, because they show
the effect in the current period of changes in predetermined variables. Estimates
of the structural coefficients can yield estimates of the reduced-form coefficients,
and so these impact multipliers can be evaluated. Alternatively, the reduced-form
equations may be estimated directly. These topics will be discussed in the chapter
on simultaneous equation estimation later in the book.
The impact effects in period ¢ are not the end of the story. Let us write Eq.
(1-12) in first difference form,
Ay, = m Ag, + m2Ar, + m3An_) + MAAN + MsAY—2 (1-15)
where
Ay, See ile Vet aie:
Let us suppose that g and r have been held constant sufficiently long for y to settle
down at some constant equilibrium level. This involves the implicit assumption
that equilibrium values exist and that the system is stable, and we will return to
this point below. Equilibrium thus implies
Ag,_.=°°° =0
Ag, = Ag,_, =
= 3... = (0
= Ar,_,
Ar, Say ay
= hy ai
Ay = Aye
0 é
Weise (iy + m5)
7 two-period lag
08441
The estimated reduced form can be applied sequentially to trace out the dynamic
10 ECONOMETRIC METHODS
a + ja(a — 4)
UMguagar 2
If a < 4, the roots are complex and the economic structure is inherently cyclical.
The product of the roots is also a, and so if a < 1, the cycles are damped, but if
1 < a < 4, the cycles are explosive. We see that a depends on
Thus, once again, empirical estimates of these parameters shed crucial light on the
nature of the economic structure. The above model has been highly simplified for
expository purposes, but these methods of analysis can be and are applied to large
systems. We have attempted to illustrate the importance of econometric estima-
tion and testing by reference only to a simplified aggregate system. Other varied
illustrations of the power and range of econometrics will be given in the course of
the book.
CHAPTER
rwoO
THE TWO-VARIABLE LINEAR MODEL
The national income model of Chap. 1 has two complications that we do not wish
to tackle right away. First of all, it is a simultaneous equation model with three
equations to explain the determination of three endogenous variables. Second,
each behavioral equation contains more than two variables. We will, however,
begin our exposition of econometric methods by concentrating upon a single.
equation with just two variables. It is not claimed that a single two-variable
equation is an adequate model of any economic process, but starting with it has
the double advantage that certain fundamental ideas can be introduced in the
simplest of all settings and that the tools and concepts developed for the
two-variable model are essential building blocks for the more complicated cases
which are treated in the rest of the book.
Y =/(X) (2-1)
where Y indicates the dependent (explained) variable and X the independent
(explanatory) variable. We may have theoretical expectations about the sign of
f’(X) or about the range of. values in which it lies. In this chapter we will deal
only with linear specifications.
12
THE TWO-VARIABLE LINEAR MODEL 13
Y=a+ BX (2-2)
Y = aX? | (2-3)
and Yee exp a+ p>} (2-4)
are all linear specifications. The first is already linear in Y and X. The second, on
taking logarithms of both sides of the equation, may be written as
log Y = loga + Blog X (2-5)
which is linear in log Y and log X. The third is
logY=a+t Bo (2-6)
which is linear in the logarithm of Y and the reciprocal of X. The function
Y=a+ BX + yx?
is linear in Y, X, and X?, but it is not a two-variable linear function, and so its
treatment will be postponed to Chap. 3. The function
6
Y=a+
Ap
however, where a, 8, and 6 are unknown parameters, cannot be reduced to a
linear function of some transformations of Y and X, and so cannot be treated by
the methods of this chapter.
The first step in the econometric investigation of the relationship between Y
and X is to obtain a sample of n pairs of observations on the two variables. The
sample data are thus indicated by
X,, Y,ql b= lezen
Next we must make a choice between specifications such as Eqs. (2-2), (2-3), and
(2-4). At this stage the choice is made by plotting the raw data or various
transformations of them on two-dimensional scatter diagrams to see which, if
any, yields an approximately linear scatter. Examples of various typical shapes
and appropriate linearizing transformations will be given in Chap. 3. Here we will
assume that Y and X denote appropriately transformed data, and so we postulate
the linear relationship
Y=a+t+ BX
where a indicates the intercept made by the line on the vertical, Y, axis and B
indicates the slope of the line.
The econometrician now faces the task of using the sample data to obtain
numerical estimates of the unknown parameters a and B. If the postulated
relationship were really true, one would have no problems at all; one would need
just two sample points and a ruler to join them. Further sample points would lie
on the same straight line and would convey no additional information. However,
exact functional relationships such as Eq. (2-2) are inadequate descriptions of
economic behavior. Scatter diagrams do not yield points which all lie on a single
straight line. Thus the specification of the linear relationship is expanded to
Y=at+BX+u (2-7)
where u denotes a stochastic variable with some specified probability distribution.
The purpose of the u term is to characterize the discrepancies that emerge
between the actual, observed values of Y and the values that would be given by an
exact functional relationship. mi
To fix these ideas let us suppose that we have data from a budget survey with
X representing household disposable income and Y household consumption
expenditure. Clearly, household expenditure will depend on some crucial factors
in addition to income, such as household size and composition, so let us suppose
that we are looking at the relationship between Y and X within a subset of
households of a given size and composition. Nonetheless it would still be
unrealistic to expect all households with a given income X, to display exactly the
same expenditure a + £X;. First of all, even among households of the same size
and composition and with the same income, there will be variations in the precise
ages of the parents and children, in the number of years since marriage, in
whether the husband is a golfer, drinker, poker player, or bird-watcher, in
whether the wife is addicted to spring hats, Paris fashions, swimming pools, or
foreign sports cars, in whether the household income has been increasing or
decreasing, in whether the parents are themselves the children of thrifty, cautious
folks or carefree spendthrifts, and so forth. This list might be extended ad
infinitum. Many factors may not even be quantifiable, and even if they are, it is
not usually possible to obtain data on all of them. Even if it were, the number of
variables would almost certainly exceed the feasible number of observations, so
that no statistical means exist for estimating their influence. Moreover, many
variables may have very slight effects so that, even with substantial quantities of
data, the statistical estimation of their influence will be difficult and uncertain. We
thus let the ner effect of all these possible influences be represented by a single
stochastic variable uw.
A second reason for the addition of the stochastic term is that there may be a
basic and unpredictable element of randomness in human responses. For pur-
poses of practical statistics the distinction between these two reasons for variabil-
ity does not matter since, for reasons of both theory and data, we hardly ever
claim to have included all distinguishable and relevant factors in any relationship,
so that the insertion of a stochastic term is required on the first count, and the
second merely adds to its variance. Finally, we note that if there were measure-
ment errors in Y so that the recorded values did not accurately reflect the values
given by the theory, this would also be a component of the stochastic term and
add to its variance.
N(0,02)
The assumption that the values of u are drawn independently from a normal
distribution with zero mean and variance o? is written compactly as
u ~ NID(0, o2)
where the symbol ~ means “is distributed,” and NID stands for “normally and
independently distributed.”
The specified model is illustrated graphically in Fig. 2-1. For a household
with income X, the average or expected expenditure is given by a + BX;. The
actual expenditure will be a + BX; + u,, where u,; is a random drawing from
N(0, o2). The complete mathematical specification of the model is
Y,=a+ BX, + u; i= 1,2,...,n (2-82)
E(u;) = 6 for alli (2-8b)
\WN ~ uj= ie
E(u,u,)
0 ij, alli, j
parecalbi ei :
(2-8c)
Pu)
Figure 2-1
The reason for splitting them is that some of our subsequent derivations only
require assumptions (2-85) and (2-8c) and not the assumption of normality. The
first part of assumption (2-8c) states that all possible covariances of the u’s are
zero, and the second part states that the variances of the u distributions at each
point in Fig. 2-1 are the same.t
The three unknown parameters of the model are a, 8, and o2. We now turn
to methods by which these parameters may be estimated.
Figure 2-2
points will lie above the line and some below. We define the residuals from the
line by
ESV
eeYn a DX OT Sl en (2-10)
A
L L
Y=a+bx (2-11)
Equation (2-11) merely gives the condition that a and b should be chosen to make
the line go through the point ofmeans (X, Y). Thus we could pass a line with any
slope whatsoever through (X, Y), and it would make the algebraic sum of the
residuals zero. The criterion is thus inadequate to determine a specific line.
The least-squares criterion is stronger. If each residual is squared, negative
signs disappear, and the sum of squared residuals is a nonnegative quantity. The
least-squares principle is
+ Where there is no ambiguity about the range of summation, we will use Le;, or sometimes just
Le, instead of the more cumbersome L7_ e;, and similarly for other expressions.
+See App. A-3, Operations with Summation Signs.
18 ECONOMETRIC METHODS
a(ze?)
da
_ abe?)
0b
_ :
LY = na + bX
(2-13)
AY ay Ap
ae
These are termed the normal equations for the straight line, for reasons that will
become clear when we discuss the geometry of least squares in Chap. 4.
To estimate the line implied by Eq. (2-13) we first compute five quantities
from the sample data, namely,
sueda =I(Y
—aaa yma 2a
gives LY= na + bYX, and
2
Table 2-1
Y XY Xe Y e=Y-Y
2) 4 8 4 4.50 —0.50
3 7 21 9 6.25 0.75
1 3 3 1 DUIS 0.25
5 9 45 25 9.75 = 075
9 147 153 81 16.75 0.25
x 1. The regression line passes through the point of means X, Y (i.e., the sum of the
residuals is zero).
This follows directly from the first equation in Eqs. (2-13), which, on division
by n, gives
Y=a+bX
and it is also shown in the footnote on page 18.
+ 2. The residuals have zero covariance with the sample X values and also with the
predicted Y values.
20 ECONOMETRIC METHODS
This also follows from the footnote on page 18 where d(Le”)/db = 0 gives
Xe = 0. The sample covariance between X and e, by definition, is
] o :
cov( X,e) = Ax = X \(e-@)
= ee _ ee
n n
= “[Link] since Le = 0
il
Daya a
ey (2-2140
14 )
and
a= Y-bxX (2-14b)
where x and y denote derivations from sample means,
x=X-X, y=Y-Y
O i > X
>s| >
Figure 2-3
[b= bxy |
The total variation in Y may be expressed as the sum of just two components,
the variation “explained”’ by the linear regression and the variation “unexplained”
by the regression. From property 3 we have
Ve ae UX tec,
Squaring and summing over all n observations gives
Ly? = Ly? + Le? + Wye = b?Lx? + Le? + 2rYxe
or Ly? = Lp? + Le? = b?7Lx? + Le? (2-15)
22 ECONOMETRIC METHODS
since y and x each have zero covariance with e. The crucial quantities in Eq.
(2-15) are
Ly?=total sum of squares in the dependent variable, measured about its mean
(TSS)
Le*=residual or unexplained sum of squares (RSS)
~y*=explained sum of squares (ESS)
It follows from Eq. (2-15) that the explained sum of squares may be
expressed in several alternative ways,
2
ESS = Dj? = B°Ex? = bExy = ‘ey
xX
Example 2-2 The data of Table 2-1 may be expressed in deviation form, as in
Table 2-2. Thus
neylO
oa = 7 = 1-75
and a= Y —bX
= 8 — 1.75(4)
=1
as before. The explained sum of squares may be calculated as
ESS = bY xy = 1.75(70) = 122.5
and the residual sum of squares may be obtained by subtraction as
RSS = TSS — ESS = 124 — 122.5 = 1.5
The proportion of the Y variation explained by the linear regression is
ESS@ 21275
Tss 124 0-788
The u values underlying the sample data are unknown and unobservable, for
we could only measure them if we actually knew the true values « and 8. Thus the
variance of the disturbance distribution o7 cannot be estimated from a sample of
Table 2-2
5 y xy ee y?
peer SS SE Se SO FBS eee ee ee
—2 —4 8 4 16
—] —] 1 1 1
ao —5 15 9 DS
1 1 1 1 l
5 9 45 25 81
Sums 0 0 70 40 124
THE TWO-VARIABLE LINEAR MODEL 23
Both are in use, but for reasons to be explained later in this chapter we typically
use
so= (2-16)
The regression estimated from Eq. (2-13) fixes a line which passes through the
sample scatter of points in the X, Y space. The correlation coefficient indicates the
“closeness” of the scatter about the fitted regression line. A visual inspection of
the scatter cannot indicate the degree of closeness, since changes in the units of
measurement for X and Y can stretch or contract scatters to give very different
impressions of the relationship.
The correlation coefficient is defined as
Ee
nS,S,
(2-17)
where x and y denote deviations from sample means and s,, 5, are the sample
standard deviations,
2x? _
oes n y n
This is known as the Pearsonian (after the distinguished statistician Karl Pearson),
or product moment, coefficient of correlation. Its rationale may be explained as
follows. Referring to Fig. 2-3, the perpendiculars erected at X and Y divide the
diagram into four quadrants. We pay particular attention to the sign of the
product x;y, in each quadrant.
NE quadrant xy positive
NW quadrant xy negative
SW quadrant xy positive
SE quadrant xy negative
Thus if we have a positive relationship, with sample points lying mainly in the NE
and SW quadrants, Lxy tends to be a positive number. Conversely, a negative
24 ECONOMETRIC METHODS
relationship will generate points mainly in the NW and SE quadrants, with Lxy
tending to be a negative number. If there is little, if any, relationship between the
two variables, sample points will be scattered in all four quadrants, and Lxy will
tend to zero. Lxy, however, has two defects as a measure of association between X
and Y. The first is that its numerical value may be increased by simply adding
further observations. This is corrected by dividing by the sample size to give the
sample covariance
cov( X,Y) = ze
n
Second, the covariance depends on the units in which X and Y are measured.
Shifting from dollars to cents for each variable would increase the covariance by a
factor of ten thousand. The covariance is standardized by dividing each deviation
by the sample standard deviation of the variable in question. Defining
x
variance of X = var( X) = s? = ——
n
and so on, and by some algebraic manipulation, we have a variety of ways of
looking at and computing the correlation coefficient:
oN cov( X, Y) (2-18)
/var( X) yvar(Y)
s =o (2-185)
any (2-18¢)
se
nuXY —(ZX)(XY)
= 5 (2-18d)
\n=X? — (ZX) pay? — (cYy
Looking at Eq. (2-18c) and rearranging gives
ks(==) ae2 f
x y7
= b=Sy
from Eq. (2-14a)
Sy
or b=r—
i Pe
a (2-19)
-
x
which shows the relationship between the regression slope and the correlation
THE TWO-VARIABLE LINEAR MODEL 25
ae ee (Zxy)°
(Zx°)(Ly’)
-= bExy
Fy?
_ ESS
TSS
= ]— RSS
TSS
Bee 2
al (2-20)
Ly*/n
Thus r? measures the proportion of the total sum of squares explained by the
regression. The last two expressions in Eq. (2-20) show that the Jimits of r are +1.
The residual sum of squares is nonnegative. It is only equal to zero if each and
every residual e, is zero, that is, if all the scatter points lie exactly on a straight
line. A value of unity for r? thus corresponds to all points lying on the regression
line. The sign of r depends upon the sign of the regression slope, that is, on the
sign of the covariance term. The relationships in Eq. (2-20) also indicate why the
correlation coefficient may be taken as a measure of the degree to which the
scatter points lie close to the regression line. Note, finally, that r is a measure of
the Jinear relation between X and Y; it is an inappropriate and misleading statistic
if the relationship is nonlinear. Suppose, for example, that X indicates a firm’s
rate of output and Y the average variable cost per unit of output. Traditional
theory postulates a U-shaped curve, which would generate points in all four
quadrants of Fig. 2-3, with an r? tending toward zero.
Least squares is just one possible method of estimation. Other estimators may
easily be defined. For instance, we might order the sample data by increasing size
of X and pass a line through the first and last points. Or one might average the
lowest two points and the highest two points and pass a line through these
averages. Applying the first principle to the data of Table 2-1, we have
xX Mu
Lowest point 1 3
Highest point g 17
giving an estimated slope of 14/8 = 1.75, which happens to coincide with the
26 ECONOMETRIC METHODS
least-squares slope. To make the line pass through the lowest point we have
3=a+1.75
= a= 1.25
This also ensures that the line passes through the upper point. The second
principle gives
X vg
f(a, b)
This is defined as the joint sampling distribution of a and b.
As with any bivariate distribution, we can integrate to obtain the marginal
distributions f(a) and f(b), which are the sampling distributions of a and b,
respectively. Concentrating on f(b), we can imagine some distribution such as
that shown in Fig. 2-4. This is a picture of the various sample values of b that
would be obtained by the repeated application of the least-squares method to
successive samples of n observations. The true parameter, B, that is being
estimated is unknown, but it is reasonable to expect some sample estimates to be
above 6 and others to be below.
There are three features of a sampling distribution that are crucial to the
assessment of an estimator. These are the mean, the variance, and the mean-
squared error. The mean value
E(b) = [of(b) db
indicates the average value that would be yielded by the estimator in repeated
applications. The bias of the estimator is defined as
bias(b) = E(b) — B (2-21)
If the bias is zero, the estimator is said to be unbiased. If the bias is nonzero, the
estimator is said to be biased. The variance of the distribution
S(0)
Figure 2-4
28 ECONOMETRIC METHODS
smaller the variance of the sampling distribution, the greater is the precision of the
estimator, that is, the greater is the chance of a sample estimate lying within some
specified interval about the true value. If we are comparing two estimators which
are both unbiased but have different variances, one would naturally prefer the
estimator with the smaller variance. If we consider the class of all unbiased
estimators and can find one with a smaller variance than any other, it is said to be
a best unbiased estimator.
A more difficult choice problem arises in comparing two estimators if both
are biased and also have different variances. If one estimator has a larger bias but
a smaller sampling variance than the other it is intuitively plausible to consider a
tradeoff between the two characteristics. This notion is given formal expression in
the mean-squared error,
where x;
W, =re (2-25)
tfIn evaluating E({b — E(b)}[E(b) — B]) the rules are the same as for summations. Any factor
which is a constant may be moved to the left and put as a multiplier in front of the expectation sign.
The expectation of a constant is that same constant. Thus
(0) == 25 (2-28)
2
9,
var(b) 2-29
var(a) = E{(a - a) }
= X°E{(b — B)) + E(w) — 2X E{(b — B)a)
LG 4 02
=e
exe act
= 6, i + ee (2-31)
Meee
In this derivation E{a?} = 62/n since w@ is the mean of a random sample of n
30 ECONOMETRIC METHODS
drawings from the wu distribution which has zero mean and variance o,;. The
. . . . . . os
=0
The covariance between a and b is
cov(a, b) = E{(a — a)(b — B)}
Satan (2-32)
Formulas (2-29), (2-31), and (2-32) all involve the unknown oa”. To make the
formulas operational, this is replaced by its estimated value s* = Le?/(n — 2).
leak
var(a) e— 2 ee =
—— ——
] 16 12
-os|z+H|-2 0.3
We might use the same data to estimate the sampling variance of the
slope estimated by passing a line through the lowest and highest points. Table
2-1 shows that the smallest X value occurs at the third observation and the
largest at the fifth observation. Thus the alternative slope estimator is
ji we sas
NS iN
]
er ee)
Hence E(b’) = B
so that the alternative estimator is also unbiased. Its variance is
2
0,
ap
Replacing a? by the same estimate as in the least-squares case, the estimated
sampling variance is
var(b’) = + = 0.0156
+ The alternative estimator does not provide any means of estimating 07, and we have only been
able to estimate var (b’) above by using the s” based on the least-squares residuals. However, this does
not affect the calculation of the efficiency of b’ since the proper definition is
var (b)
Efficiency of b’ = are)
AE
02/32
32
AQT
32 ECONOMETRIC METHODS
Lw;c; = Bee
Ex
] :
cao using Eq. (2-34) (2-37)
LW? = els
ye
Thus
Lw,(c; — w,) = 0 (2-38)
and so
The proof of the similar result for the intercept is left as an exercise for the
reader.
This minimum variance property of least squares is the main reason for the
widespread use of the technique. It rests on the assumption that the X’s are fixed
in repeated sampling, that the relationship has been correctly specified in Eq.
(2-8a), and that the disturbances have zero mean, constant variance, and zero
covariances. Notice that assumption (2-8d) on the normality of the u distribution
THE TWO-VARIABLE LINEAR MODEL 33
has not been used so far. If we bring this assumption into play, the sampling
distributions of a and b are fully determined. Since a and }are linear functions of
the u’s, they in turn are normally distributed variables. Once the mean and the
variance are known, a normal distribution is completely specified. Thus
‘y2
o~ Mavei|tae | (2-40)
and
oe
bn 8.28 (2-41)
It still remains to justify the expression given in Eq. (2-16) for the estimator
of the disturbance variance. The ith residual is
A
Cea
=ait BX us; — a — bx;
Sita len) (Dia8 )XG
From the expression for a just before Eq. (2-30),
a-a=iu-—(b—B)X
Thus
e, = (u,— &) — (b- B)x;
Notice that z, being a sample mean, cannot be set at zero, even though E(u;) = 0.
Squaring and summing,
Thus
E(Xe?) = (n — 2)o?
So
E| Le 2 )=02
n—2 u
34 ECONOMETRIC METHODS
and the estimator proposed in Eq. (2-16) is unbiased. The distribution of this
estimator will be established for the general case in Chap. 5.
and
Le? is distributed independently of f(a, b)
Concentrating first of all on inferences about B and recalling that the ¢
distribution is given by the ratio of a standard normal variable to the square root
of a x” variable divided by its degrees of freedom,
biap ~ N(0,1)
——£_
G/N Bee
from Eq. (2-41) so
giving
(2-16), and ty 9); is read off as the 25 percent point of the ¢ distribution with n — 2
degrees of freedom. In general a 100(1 — e) percent confidence interval for B is
given by
B+ t,95/(Ex? (2-45)
To test the null hypothesis that 8 has some Rpecined value Bp, that is,
Hy: B= Bo
against the alternative hypothesis that B has some value other than Bo, that is,
HH: BB
we insert By in Eq. (2-43) and then have the conditional statement. If the null
hypothesis is true,
DSBS
~ t(n
— 2)
s/\Ex2
This gives the sampling distribution of b under the null hypothesis, as shown in
Fig. 2-5. If the null hypothesis were true, 95 percent of all sample values of b
would lie within fp9), standard errors of Bp, that is, inside the symmetrical region
about 8) shown in the figure. If our sample b is found in either tail of the
distribution, then either
In such a case we deliberately choose the second interpretation and thus follow
this procedure. Reject H, at the 5 percent level of significance if
b— By
> lo025
s/\Zx2
ORO a ee
s/\ix*
cba ae
b =
s/\Xx?
The null hypothesis most frequently tested is
H,: B=0
This is referred to as testing the significance of X. If the hypothesis is true, the X
variable plays no role in the determination of Y. Nonetheless when sample values
of X and Y are drawn from such a population and the least-squares formula is
applied, we will usually find nonzero values for b. These can arise from the
fluctuations of random sampling even though the underlying £ is truly zero.
The appropriate significance test follows directly from replacing 8, by zero in the
results above. Thus reject Hp: B = 0 at the 100e percent level of significance if
b
> t, 2
Sex i
The test statistic is now simply the ratio of b to its estimated standard error. Most
computer programs for regression analysis print out the value of b and either its
estimated standard error or else the ratio of b to its estimated standard error, and
this is often referred to as the sample ¢statistic.+ If the sample f¢ statistic is
numerically greater than the preselected critical value of t, we accept the alterna-
tive hypothesis and conclude that X plays a significant role in the determination
of Y. The presentation of sample ¢statistics is thus directly useful for making
significance tests. If, however, one wishes to test a null hypothesis that B has some
value other than zero, one requires the standard error (s.e.) for substitution in the
appropriate test statistic. This can be obtained from the computer printout as
b
carpe ant
7 In the rest of the text we will normally drop the distinction between the true standard error and
the estimated standard error, as it is usually clear from the context which concept is implied.
THE TWO-VARIABLE LINEAR MODEL 37
V(r]
2) (2-46)
piel Ae
Ss een
leek:
at va +
eee | (2-47)
Rey Ne
and the hypothesis
Hy: a= a
would be rejected at the 100e percent level of significance if
SS tay)
Ss aliens.
a Ba
Tests on 0; may be derived from the result stated in Eq. (2-42). Using that
result one may, for example, write
n — 2)s?
PrxBox = ( =e = x30s| = 0.95 (2-48)
u
which merely states that 95 percent of the values of a x? variable will lie between
the values that cut off 25 percent in each tail of the distribution. This is illustrated
in Fig. 2-6. The critical values are read off from the x? distribution with n — 2
P(x?)
a iz,
Xo .005 X0.975
Figure 2-6
38 ECONOMETRIC METHODS
degrees of freedom. The only unknown in Eq. (2-48) is «7, and the contents of the
probability statement may be rearranged to give a 95 percent confidence interval
for 62 as
(n — 2)s? - (n — 2)s?
2 2
X0.975 X0.025
Example 2-4 From the data of Tables 2-1 and 2-2 we have already computed
a=] var(a) = 0.3
b = 1.75 var(b) = 0.0125
Thus
$.€. (a) = v0.3 05471
s.e.(b) = ¥0.0125= 0.1118
Since n = 5, from the ¢ distribution with 3 degrees of freedom,
tooos = 3.182
Thus a 95 percent confidence interval for a is
1 + 3.182(0.5477)
that is,
—0.74 to 2.74
and a 95 percent confidence interval for B is
1.75 + 3.182(0.1118)
that is,
1.39 to 204
The intercept is not significantly different from zero since
a
s.e.(a) 0.5477 = 1.826 < 3.182
while the slope is strongly significant since
b Lea
s.e.(b) = 0.1118 = 15.653 > 3.182
that is,
0.16 to 6.94
The test for the significance of X (Hy: = 0), derived in the previous section,
may also be set out in an analysis of variance framework, and this alternative
approach will be especially helpful when we treat problems of multiple regression
in Chap. 5.
From Eq. (2-41) we have the result
eDiciB-’ ss N(O, 1)
0,/ {Ex
From the definition of the x? variable in App. A-7 we then have
(6 a B) oe x7 (1)
0,/LX*
and since
we?
Faget 2)
independently of b
Source of Degrees of
variation Sum of squares freedom Mean square
(i) (ii) (iit) (iv)
x ESS = Ly? = b*Dx7 l ESS/1
= bY xy
Residual RSS = Le? n-2 RSS/(n — 2)
Total TSS = Ly? n-1
Example 2-5 Table 2-4 shows the analysis of variance for the data of Table.
2-2.
The sample F statistic is
Sample F = = = 245.0
Table 2-4
extensively in later work. To establish the equivalence, recall that the ¢ test for the
significance of B is as follows. Reject Hj: $= 0 at the 5 percent level of
significance if
“b
> tonos(#'— 2)
s/\Lx?
From Eq. (2-50) the F test procedure is as follows. Reject Hy: 6 = 0 at the 5
percent level of significance if
bx
=a
Ye2/(n BS 2) os
0.95 ( )
The second test statistic is seen to be the square of the first. Thus
SampleF = (sampler)
It is also shown in App. A-7 that the F variable with (1, r) degrees of freedom is
the square of a ¢ variable with r degrees of freedom. Thus
wearin Leer
"0-2 te
Again, exploiting the relation between the ¢ and F distributions, Eq. (2-52) gives
fli es2
ry(n — 2)
(253)
a=")
and this statistic may be referred to as the ¢ distribution with n — 2 degrees of
freedom to test the significance of the relationship between Y and X. From
Example 2-2 we have
r
ISLESS 6. 9122.5
TRS TORN es
+ It is, of course, superfluous to insert unity as the divisor of the numerator in both Eqs. (2-51) and
(2-52), but it maintains a correspondence between these expressions in models where there is only one
explanatory variable to later models with several explanatory variables.
42 ECONOMETRIC METHODS
2 OR
(1 — 0.9879)
oe Le 15.65
~ s.e.(b) y0.0125
and the square root of the sample F statistic in Example 2-5 is
t= VF = 245.0 = 15.65
Thus all three tests are simply three versions of a single test. The test involving r
may be regarded as a test of the significance of the correlation coefficient, that
based on b as a test of the significance of the regression slope, and the analysis of
variance formulation tests the significance of the explained sum of squares, but
they are just three different ways of essentially asking the same question.
where uy indicates the value that would be drawn from the disturbance distribu-
tion in the prediction period. The prediction error may then be defined as
6 = ig Y
=U
—(a— a) — (b- B)X, (2-55)
Taking expectations
Ee) 20
since E(u,) = 0 and a and b are unbiased estimators of a and B. Thus the
least-squares predictor, Eq. (2-54), is an unbiased predictor. The variance of the
prediction error is then found by squaring Eq. (2-55) and taking expectations:
var(e,) = E(e)
= var(u,) + var(a) + X$ var(b) + 2X,cov(a, b)
since the other two covariances vanish.} On substitution from Eqs. (2-28), (2-31),
and (2-32) this gives
var(e,) = 62
(2-56)
The variance of the prediction error is thus at its minimum value when X, = X
and increases nonlinearly as Xj departs from X. From Eq. (2-55) ey is seen to be a
linear function of normal variables and so is itself distributed normally. Thus
e
- ~ N(0, 1)
0,u
| 1
bee (Xo ao y
n x
a ein) (2-57)
: ppale. a
n yx?
Everything in Eq. (2-57) is known except Yo, and so, in the usual way, we derive a
+ By assumption wy is independent of u,, u2,..., u,, and thus it has zero covariance with (a — a)
and (b — B), since these are each linear functions of u), uz,..., Uy.
44 ECONOMETRIC METHODS
Sometimes interest centers on predicting the mean value of Yo, that is,
E(Y,) =a + BX,
rather than Yp itself, since there is, of course, no way of predicting the value of a
single drawing from p(u). The prediction error is now
ey = E (Yi Y
= ai Gao ie NO eB) Xe
which gives
1 2 (4) 2)= \2
var(eé,)
( o) = 9; 2+ ——_———
x2
iy CX)= \2
The width of the confidence interval in Eqs. (2-58) and (2-59) is seen to increase
symmetrically the further X, is from the sample mean X, as shown in Fig. 2-7.
a+ bX
ou
>| Xo
Figure 2-7
THE TWO-VARIABLE LINEAR MODEL 45
Example 2-6 Assembling relevant results for the data of Table 2-1,
A
Y=1+4+1.75X
n=5
X=4 :
Ber es
S 2= EAD =e 3 0.5
Lx?
= 40
Suppose we require a 95 percent confidence interval for Y given X = 10.
Applying Eq. (2-58) gives
een se
1 + 1.75(10) + 3.182V0.5 /( A : Wess) |
40
that is,
18 Deere 3:26
or 15.24 to 21.76
The 95 percent interval for E(Y|X = 10) is
Teen (Or 34 ys
18.5 + 3.182y0.5 5 is “Frag ae
that is,
18.5 = 2.36
or 16.14 to 20.86
To test whether a new observation (Xp, Yo) may be thought to come from the
structure generating the sample data, one merely contrasts the observation
with the confidence interval for Y). For example, the point (10, 25) gives a Y
value which lies outside the interval 15.24 to 21.76, and one would conclude
that it was unlikely to have been generated by the same structure as the
sample data.
PROBLEMS
var(a) =,
ries,a
stn
Lx?
i=]
Show that no other linear unbiased estimate of a can be constructed with a smaller variance.
2-2 Show that if z, are independent quantities from the same population, with variance o”, then the
sampling variance of
n
b= ye a;Z;
f=)
46 ECONOMETRIC METHODS
is o?L"_,a?. Observations Y, are related to fixed quantities X, and the quantities z, above by the
t
a(¥%o Ys =)
Deduce the sampling variance of this estimate and compare with it with the sampling variance of the
least-squares estimate.
(Oxford University, 1958)
2-3 From a sample of 200 pairs of observations the following quantities were calculated:
If Y = C + Z (where Zis savings), compute the correlation between Y and Z, the correlation between
C and Z, and the ratio of the standard deviations of Z and Y.
(R.S.S. Certificate, 1948)
2-7 The table below gives the means and the standard deviations of two variables ¥ and Y and the
correlation between them for each of two samples.
Number
Sample in sample x uy Se Sy ies
] 600 5 12 2 3) 0.6
2 400 W 10 3 4 0.7
Calculate the correlation between X and Y for the composite sample consisting of the two samples
THE TWO-VARIABLE LINEAR MODEL 47
taken together. Comment on the fact that this correlation is lower than either of the two original
values.
(R.S.S. Certificate, 1955)
2-8 An investigator is interested in the two following series:
X, deaths of children
under
| year, thousands 60 G2 Ole) Sier 5) OOM NO3E Sut 2a AGerao) AS
Y, consumption of beer,
bulk barrels 23 2S 20 Ooo 0 es Os ao Sueno!
have zero mean, show that the covariance of the least-squares estimates of « and £ is zero. Hence, or
otherwise, prove that an unbiased estimator of 8 can be derived by estimating the equation
, Y, = BX,
which is constrained to pass through the origin. What is the variance of this estimator of 8?
(UL, 1973)
CHAPTER
THREE
EXTENSIONS OF THE
TWO-VARIABLE LINEAR MODEL
The obvious limitations of the two-variable linear model are that it is linear and
embraces only two variables. These restrictions limit the variety of statistical
phenomena for which it provides an adequate description. In this chapter we
describe some of the more important ways of extending the range of the model:
First of all we discuss the case of replicated observations for various XY values,
which enables us to test the adequacy of the linear representation against the
alternative hypothesis that the relationship of Y to X is nonlinear. Then we discuss
various types of nonlinearity and the ways in which they may be handled. In
some cases suitable transformations of the variables return the problem toa linear
framework, in which case the simple techniques of Chap. 2 may be applied. In
others, nonlinear relations have to be fitted directly, but the discussion of
nonlinear estimation is beyond the scope of this book. Finally, an introduction to
three-variable regression is provided, preparing the way for a general treatment of
multiple regression in Chap 5.
Suppose now that we have sample data which could be arranged in the form
shown schematically in Table 3-1.
Here we have p distinct values of X and m observations on Y corresponding
to each X observation, giving n = mp sample observations altogether.
This
48
EXTENSIONS OF THE TWO-VARIABLE LINEAR MODEL 49
x x Mean F¥,
x) Far * ats Y,
Xy YY. sie Tay 0)
X; Yi Yi2 Yim ¥
x, Yo%p2 °° Yom Y,
The first step in the analysis of replicated data is to compute the row means
1 m
where Y,, denotes the jth observation in the ith row or class. The scatter diagram
might look like Fig. 3-1 for the case of m= 4 and p= 6, with the circles
indicating individual observations and the black squares the sample mean values
of Y. Clearly, we can fit a least-squares regression to this scatter of 24 observa-
tions by the methods of Chap. 2. In this case, however, that turns out to be
identical to the regression fitted to the six mean points shown on the scatter.t To
see that this is so, consider the formula for the regression slope
b => by
xs
The deviations in this formula are measured from the overall sample means,
+ This result does not hold when there are unequal numbers of observations in each class. See Eq.
(3-6).
50 ECONOMETRIC METHODS
>
Figure 3-1
y- ODSiya
rae y ey
ai
or, using Eq. (3-1),
ees | Sees
——— Yy. 3-2
p oo
It is convenient to denote the X observations by X;;, with the proviso that
X, =X.= °°: = X,,= X; fori = 1,2,..., p. Thus
aol y > ls
E> iy =
TAD oatae eeore
Then
Pom fees
Ext= i=1 j=1xX)
=m¥i=1ze (X,-¥)= \2
ae — —
-¥(%-
8)E(¥,-¥) 63)
remembering that anything which is constant with respect to a particular summa-
tion sign can be moved leftward in front of that summation sign. But
a pas
EXTENSIONS OF THE TWO-VARIABLE LINEAR MODEL 51
since the first term vanishes and the second is a constant with respect to
summation over j. Thus
P
Pye SX (KF ) (3-4)
i=1
Putting Eqs. (3-3) and (3-4) together, the regression slope is
b= eee = a ae ee Xe (3-6)
Sr xx) Ro OKp XO)”
which can be regarded as a weighted regression applied to the class means, the
weights being equal to the number of observations in each class.
In scatter diagrams such as that in Fig. 3-1, the class means will usually not
lie exactly on the regression line, but will be spread around it in the same way as,
but to a lesser degree than, the sample points are spread around the regression
line. We are thus led to pose two related questions.
Note carefully the distinction between p, and Y,, the former being the true but
unknown mean of the distribution from which the Y, observations are drawn and
the latter being the actual mean of those sample observations. .
The variation hypothesized for the m’s allows a very flexible and general
relationship between Y and X. In this context the null hypothesis of no relation-
52 ECONOMETRIC METHODS
SS
ivy
(ii Talia
i,j
ys i (3-11)
The decomposition of the sum of squares in Eq. (3-11) is similar to that given for
the linear regression model in Chap. 2. The left-hand side again represents the
sum of the squared deviations of Y, measured from the overall sample mean. The
first term on the right-hand side is the sum of the squared deviations of the Y’s,
measured now about the relevant class means, and the second sum of squares is
that due to the variation of the class means about the overall mean. We can thus
describe Eq. (3-11) in the form
TSS = RSS oe ESS
residual or error sum of squares “explained” sum of squares due
{Total Se ot Saree y} “unexplained” by the class means to the class means
pean Y,) ee
Oo
+ We now write the double summation L?_,L7_, simply as ©, j- Equation (3-11) is derived by
noting that the cross-product term vanishes, that is,
or
2
eee) ns
=O, (3-13)
Under the null hypothesis the Y, are independent normal variables with mean be
and variance o*/m. Thus
DEBE eh
<I
oO
(3-14)
and so
pte IT
a =o? (3-15)
We state, without proof, that under the null hypothesis the sums of squares in
Eggs. (3-12) and (3-14) are independently distributed. Thus since the F distribution
is given by the ratio of two independent x7 quantities, each divided by its degrees
of freedom, under H)
F= _mE(¥-Y¥)Ap-1)
_= ESs/(p~1) | Fp sal poet)
Bey, = ¥,)'/p(m i 1) RSS/p(m a 1)
(3-16)
The rationale of this test is easily seen from Eqs. (3-13) and (3-15). Under the null
hypothesis the numerator and the denominator of F are independent estimates of
o*, and thus we may expect F to vary randomly about unity. If, however, the null
hypothesis is not true, the between class sum of squares in the numerator of F will
reflect more than just the random variation of the class means about a common
mean, and F will rise in value. The null hypothesis is then rejected if the
computed value of F exceeds a preselected critical F value from the upper tail of
the distribution.
The sums of squares in this F statistic may also be used to define the
correlation ratio n as follows:
n? = = =
Eee me
Ye
my; =1(Y, eee] eal
ot yom
ty uy
|: (3-17)
Lp Ay Y) aC J Y)
This is analogous to the definition of r? in Chap. 2, with the exception |that the
explained sum of squares is based on the variation of the sample means Y,, rather
than the variation of the regression values Y,. The F statistic of Eq. (3-16) may
then be stated equivalently
fee pew t)
(1 — 9°)/p(m — 1)
which has the same structure as the expression involving r* in Eq. (2-52).
54 ECONOMETRIC METHODS
Source of Degrees of
variation Sum of squares freedom Mean square
P _— —
The test defined in Eq. (3-16) is the standard one-way (or one factor) analysis
of variance. It is based solely on the Y observations and is a test of the
homogeneity of the p class means. The only role for the X variable has been to
classify the Y’s into classes associated with a common X value. Thus the X
variables could have been qualitative variables, such as socioeconomic status or
educational level. The data for the test are normally set out in an ANOVA table,
such as Table 3-2.
The second and third columns of the table are additive, as usual. Thus, once
any two of the sums of squares have been calculated, the third follows by using
TSS = ESS + RSS.
We can now carry this analysis one step further and derive a test for the
linearity of the relationship between Y and X. Figure 3-2 shows just one point
(X;, ¥;;) from the X,Y scatter, and Y = a + bX indicates the linear regression
fitted to the data.
Figure 3-2
EXTENSIONS OF THE TWO-VARIABLE LINEAR MODEL 55
Bile
inj
a= iJ(8 yh (Y= 7) + my es
(3-19)
In words, Eq. (3-19) states
sum of squares
Total sum sum of squares due to variation sum of squares
of squares } = ( about class + ( of class means } +( due to linear
in Y means about linear regression
regression
We notice from Eq. (3-19) that r* is the ratio of the third term on the right-hand
side tod, ,(%; — Y)’*, and 7 is the ratio of the sum of the second and third terms
to the same total sum of squares in Y. Thus
yw>r
The equality would only hold if all class means fell on the linear regression line.
The middle term in Eq. (3-19) is thus proportional to the excess of n” over r* and
is the basis of the linearity test set out in Table 3-3.
The sums of squares may be calculated in various ways. If r* and 7* have
already been calculated, the required quantities follow directly. If not, S, can be
calculated first; S, is then the explained sum of squares due to the regression, to
be calculated by the methods of Chap. 2; S, can be calculated directly; and S,
then follows from the fact that
If the class means deviate significantly from the linear regression, S, will tend to
be large in relation to S,, when appropriate allowance has been made for degrees
+ In Fig. 3-2 we have shown, for simplicity, a case where Y;; exceeds Y,, which in turn exceeds Y,,
which again exceeds Y, so that all the deviations on the right-hand side of Eq. (3-18) are positive. In
general, of course, these deviations vary in sign.
+ All three cross-product-terms vanish, Those involving (Y;; — Y,) vanish since Li — ¥,) is zero
for all i. The term involving Y; (Y, - ¥,)( a= Y) is simply properdonal to the covarianceabetaeen the
regression values Y, and the residuals Y,— Y, about the regression line, and thus also vanishes.
56 ECONOMETRIC METHODS
Sb S,/(p eo2) i
way og Be
and the linear hypothesis is rejected in favor of a nonlinear alternative if the
computed F exceeds a preselected critical value from F( p — 2, p(m — 1)).
Example 3-1 We have eight classes (p = 8), each class being defined by a
particular value of X, and five observations (m = 5) in each class, giving 40
observed points on an X, Y scatter. The class means are shown in the final
column of Table 3-4 and display a negative relationship with X, but the class
means do not lie exactly on astraight line.
We now wish to build up the numerical equivalent of Table 3-3 for these
data. The first step is to fit the linear regression of Y, on X,. The eight pairs of
values yield the following numbers:
47.2
37.0
25.4
19.6
14.8
10.4
8.8
st DUN
eee
St
Nee
DT ON
PWN 2.4
EXTENSIONS OF THE TWO-VARIABLE LINEAR MODEL 57
from which
ss = 9,500
(x, - X¥) - 250)°
aE = 1,687.50
oe- Y) = 3,585-
YG - ¥)(¥, 250(165.6
SS — 1,590.00
a
X(¥,- ¥)= 5,036.56 — os = 1,608.64
SO
1590
Bea cea ete ab Aee
Y = 50.144 — 0.9422.X
Turning now to Table 3-3, we compute first of all the overall sum of squares:
S4 = L(Y; - ras
=D (Et)
= 25,620 — ay (828) = 8.480.4
The sum of squares within classes, S,, is obtained by calculating the sum of
squared deviations within each class about the class mean and aggregating
the result over all classes. For the ith class we have
D(%,- FY EM
J
-(EY a
For the first row in Table 3-4, this gives
Regression S; = 7,490.67 1
Class means about regression S,= 552.53 6 92.09
Within classes S3 = 437.20 32 13.66
Total S, = 8,480.40 39
Thus
Ss, = 5(1,498.133) = 7,490.67
The remaining sum of squares, S,, can now be obtained by subtraction, and
Table 3-5 is prepared.}
} There is a subtle point concerning the interpretation of r? in this example. We have seen that in
the special case, where there is a constant number of observations per class, the regression slope may
be calculated by considering Ft: the 40 observations (X;,, Xi) Y; or the eight observations (X,, Y,).
There are, however, two distinct r?.One relates to the ee of Y, on X; and the other to the
regression of Y;; on X;,. It is intuitively clear that the former r? must exceed the latter since the class
means will lie closer to the regression line than the raw data. In Table 3-3 S, is expressed as r2S,, and
this r? is the one relating to the raw data. By definition, it is
¥ (%, — ¥) = 5(1,687.50)
=8,437.50
i
r> = 0.883292
Then
EEA ar eeAo0eT
LGR Shyse Sees ae vO
The 5 percent critical value from F(1, 40) is 4.08, and the 1 percent value is
7.31, so this is a highly significant sample statistic, and the data would lead us
to reject decisively the hypothesis of a zero coefficient for X. The test statistic
OO
40
Y = 50.144 — 0.9422x
OO
800—
20
5 10 15 20 25 30 35 40 45 50
60 ECONOMETRIC METHODS
and the | percent critical value from F(6, 30) is 3.47. This sample statistic is
also significant, and we conclude that the true relationship is probably
nonlinear. The scatter, regression line, and class means are shown in Fig. 3-3.
The relation
Y=at+BpxXt+u (3-21)
is Jinear in the parameters a and £ and also in the variables X and Y, and, as we
have seen, application of the least-squares principle results in two simultaneous
equations, which are linear in the estimates a and b and thus easy to solve. The
relation
logY=a+BX+u (3-22)
may be written
Z=a+BX+u (3-23)
where
Z = logY (3-24)
The scatter from Eq. (3-22) in the X, Y plane will be nonlinear. However, the
transformation defined in Eq. (3-24) will yield a linear scatter in the X, Z plane,
and so the techniques of Chap. 2 may be applied directly to Eq. (3-23). This is an
example of a transformation of a variable, changing a relationship which is
nonlinear in the original variables into one which is linear in the transformed
variables. As will be seen below, simple transformations of one or both variables —
can deal with a wide variety of nonlinear relations.
Another approach is to introduce additional terms in X on the right-hand
side of the relation. If Y denoted this average variable cost per unit of output and
X the rate of output, the U-shaped cost curve of economic theory might be
depicted as
Y=at+BX+yX*+u (3-25)
In very rare cases, economic theory may indicate the appropriate transformation
of variables. As Zarembka has pointed out, the constant elasticity of substitution
+ The late Sir Julian Huxley once defined God as a “personified symbol for man’s residual
ignorance.” In a similar vein the disturbance term might be regarded as the econometrician’s stochastic
symbol for his residual ignorance, and then, as is sometimes done with God, the inscrutable and
unknowable may be ascribed the properties most convenient for the current problem.
62 ECONOMETRIC METHODS
Y = (a,K® +a,L°)””
gives
ye’? = a, K® + a,L° (3-29)
Thus each observation on output should be raised to the power p/v, and each
observation on capital and labor inputs should be raised to the power p. This is
an example of power transformations of the variables, and although Eq. (3-29) is
linear in the transformed variables, it still poses difficult estimation problems.
Some special cases of power transformations, however, yield simple estimating
procedures, and we now turn to these.
Returning to two-variable relations, let us denote a transformation of the Y
variable by Y“”. This symbolism indicates that the transformation depends only
on a single parameter A,. Likewise, let us indicate the transformation of the X
variable by X2). A very general form of transformation has been proposed by
Box and Cox, namely,+
y"-1
aad hr
Sa
oees A, + 0
a (3-30)
In Y A, =0
and similarly,
xX — ]
cael Xr 0
XO2) = Ay Dh (3-31)
In X A, =0
At first sight these seem needlessly complicated transformations, and one might
well ask, why not use the simple power transformation Y*', X*2. An examination
of Fig. 3-4 indicates the rationale behind the Box-Cox transformation. Fig. 3-4a
shows the simple power transformation Y* for two illustrative values of Y,
namely, 10 and e = 2.84128. The transformed variables cross at the (0, 1) point,
and to the left and right of that point their ordering is reversed. The simple power
transformation is thus unsatisfactory since different values of X would not
preserve the ordering of the data. Fig. 3-4b shows the graphs of Y*/A for the
same two values of Y, and now the ordering is the same for all values of A, but a
discontinuity occurs at A = 0. Finally, Fig. 3-4c shows the graph of the general
transformation
y* = 1
rv
and now the ordering is the same for all values of A, and there is no discontinuity
at A = 0. If we substitute A, = 0 in Eq. (3-30), we obtain Y“” = 0/0, which is
indeterminate. However, the application of L’H6pital’s rule shows thatt
lim YOU = InY
4,70
Suppose now that the transformed variables fit the linear model, that is,
YOU = ay + BXO” + u (3-32)
This model has five basic parameters, namely, aj, B, A,;, Az, and o/. In this
section we will consider only some special cases corresponding to particular
values of A, and A).
Case 3-1: A, = 1 = A, (Linear model). Combining Eqs. (3.30), (3-31), and (3-32)
now gives
Y=a+BX+u
where
a=l+t+a,—8
This is the simple linear model of Chap. 2.
} For a definition of L’Hopital’s rule see, for example, A. C. Chiang, Fundamental Methods for
Mathematical Economics, 2d ed., McGraw-Hill, New York, 1974, p. 420. The application of the rule to’
Y® states that
By e eda) (laa)
jor! ai Loli Cae OR)
= lim
ae (Y*InY )
=InY
This development uses the result that (d/dd,)(Y*') = Y™'In Y. To derive this from first principles,
consider a general function
z=Iny=xlna
Then dz_dzady_\lw
—= =
Ore Gly Obs yn abe
dz
Ge
But
u 7 Ina
Th us aes aS
ye Inag=a*lna
EXTENSIONS OF THE TWO-VARIABLE LINEAR MODEL 65
ORaIBre1
Case 3-2: , = 0 = A, (Log-log model). Combining Eqs. (3-30), (3-31), and (3-32)
now gives
nY=a,+PlnX+u (3-33)
Again, all the techniques of Chap. 2 may be applied, once the original data have
been transformed to logarithmic form. Notice that although Eg. (3-33) is specified
in terms of logarithms to base e, one may take logarithms to base 10 in carrying
out the empirical work. The estimate of a, will be affected by the choice of base,
but that of B will not.
Ignoring the disturbance term in Eq. (3-33), the relationship between Y and X
is
Yor Anxe (3-34)
where
In Ay = ao
From Eq. (3-34)
dY
a
AX BA, X B-1
va x
4
ex
e&
B>0 B<0
O Se @ >X
Figure 3-6
(3-33) is fitted to the data the regression slope is a point estimate of the elasticity.
If 8B= —1, Eq. (3-34) gives
XY = Ay (3-35)
which is a rectangular hyperbola. If Y denoted the quantity purchased and X the
price per unit of some commodity, then Eq. (3-35) would represent a demand
curve with constant elasticity of —1 and a constant total expenditure on the
commodity, whatever its price.
oanB (3-37)
Thus the proportionate rate of change in Y per unit change in X is a constant and
equal to B. The function is only defined for positive values of Y. Ignoring the
disturbance in Eq. (3-36), we may rewrite the function as
Y= extBhx (3-38)
and its general shape is shown in Fig. 3-6. The intercept is given by e%, and the
slope is positive or negative, depending on the sign of B.
+ Equation (3-36) is a widely used specification in human capital models, where Y denotes earnings
and X years of schooling. The specific functional form is derived from theoretical considerations by J.
Mincer, School, Experience and Earnings, Columbia University Press, New York, 1974.
EXTENSIONS OF THE TWO-VARIABLE LINEAR MODEL 67
Average annual
output
(1,000 net tons)
Decade ie Log Y Xt
A special case of Eq. (3-38) occurs when X denotes time and the function
then describes a variable Y which displays a constant proportionate rate of
growth (8 > 0) or decay (B < 0).
Example 3-2 The first step in any empirical application of Eq. (3-36) is to
check visually on whether the constant growth assumption seems warranted.
This may be done either by plotting log Y against time or, equivalently, by
plotting Y against time on commercial semilog paper. The first procedure
applied to the data of Table 3-6 gives the scatter diagram shown in Fig. 3-7,
and it is clear that the relationship is approximately linear.
log y
A
Sle
5.0 Sane
logy = 4.4510 + 0.3760¢
4.5
4.07
68 ECONOMETRIC METHODS
¥ Suppose your friendly local bank manager offers interest of 10 percent per year on time deposits
compounded annually. Then $100 deposited now grows in successive years to $110, $121, $133.1, and
so on. In general, interest at 1007 percent per year on an initial investment of $100 will give a sum of
$100(1 + r)” after n years. Suppose further that the manager agrees to your suggestion that, instead
of adding interest of 10 percent once a year, 5 percent should be added twice a year. After 2 years of
graduate school your $100 would now grow to $100(1.05)* = $121.55, so that your perspicacity
has
paid off to the tune of 55 cents. You now become greedy and suggest that | percent be added 10 times
EXTENSIONS OF THE TWO-VARIABLE LINEAR MODEL 69
can select the origin and the units of measurement for time at our conve-
nience. In the final column of Table 3-6 we have chosen the midpoint of the
1871-1880 decade as the origin and measured time in units of 10 years. The
normal equations for fitting a linear regression of log Y on ¢ are
Llog
Y = na + byUt
Ltlog Y = aXt + byt?
Since the t’s have been chosen so that Lt = 0, these equations simplify to
se LYlog Y
n
ie Lrlog Y
vr?
For the data in the table, these equations give
a 31.1575/7 = 4.4511
b 10.5285/28 = 0.3760
Thus the regression estimate of Eq. (3-42) is
ae
log Y = 4.4511 + 0.3760t
The r? is 0.9945, which confirms the story of the scatter diagram that the
constant growth curve fits this series very well. To find the estimated growth
rate,
log(1 + ¢) = 0.3760
1 4¢'= 2.3768
Thus
$= 1.3768
Since ¢ was measured in units of 10 years, this gives the estimated rate of
growth per decade as 137.7 percent. The annual growth rate is found from
(1 + r)'° = 2.3768
a year, or perhaps } percent 20 times a year. You have, in fact, discovered the principle that
2 3) r\”
G+y<(14+5) <(14+ 5) <-> <(142) <---
Unfortunately it is no magic device for unlimited increases in your wealth. The sequence has a limit,
namely,
where
giving
r = 0.0904
or just over 9 percent per annum. From Eq. (3-43) the corresponding
continuous rate 6 is 0.0866.
¥=a+p(y)+u (3-44)
The slope is dY/dX = —B/X?. Thus if B is positive, the slope is everywhere
negative, and conversely it is positive when B is negative. Since 1/X — 0 as
X — o, a denotes an asymptotic value for Y. The shape of this function is
indicated in Fig. 3-8.
Fig. 3-8a indicates a typical shape of a Phillips curve, and many Phillips
curves have been estimated by regressing the rate of wage change on the
reciprocal of the unemployment rate. The estimate of a indicates the asymptotic
floor for wage change; if it turns out negative, the Phillips curve cuts the
unemployment axis. Fig. 3-8b has often been used to represent expenditure
functions, where Y denotes expenditure on some specified commodity or service
and X denotes total expenditure (or income), the data typically coming from a
cross section of households. This particular application only makes sense when a
is positive and £is negative, for a indicates the asymptotic level of expenditure. It
follows that income or total expenditure has to reach some critical level — B/a
before anything is spent on this commodity. Thus Eq. (3-44) cannot serve as a
general model for all types of consumption since we need to allow for finite
consumption of some commodities at values of X close to zero. This particular
difficulty can be overcome by the choice of yet another pair of A values.
B>0 Sa—0
(a) (b)
a¥
AK of@ (
8) aye
2 Je (3-47)
aAXtee
(B28 eer
New, aXe
Hence there is a point of inflection where X = 8/2. To the left of this point the
slope increases with X; to the right of it the slope diminishes. As X —> 0,
Y > e*. Substituting X = 8/2 in Eq. (3-46) gives the value of Y at the point of
inflection as 0.135e% or, in other words, about 13 percent of the asymptotic value
of Y. The general shape of this function is shown in Fig. 3-9.
If X represents time, Eq. (3-46) then pictures a growth curve which starts at
zero and approaches an asymptotic level. Rewriting Eq. (3-47) in an equivalent
form gives
Adie SBe
VodXee =¥2
that is, the rate of growth in Y per unit change in X is inversely proportional to
the square of X. Thus the rate of growth falls off sharply after low values of X, as
shown in Fig. 3-9. This can be a disadvantage of the curve, which may more than
offset the ease of fitting.
A similar curve, which also has an upper asymptote at some finite level and a
lower asymptote at zero, but has a more symmetrical shape between the two, is
the /ogistic. This curve cannot be derived from the general linear relation (3-32) by
a suitable choice of values for A, and A,. Nonetheless it has been widely used in
fitting growth trends, and so a brief account of it is given here. The logistic
equation is
Yall (3-48)
Leiber =
where a, b, and k are parameters to be determined. We have written Y as a
function of time ¢, as this is by far the most common practice, but in some
applications it is quite feasible to replace t by some independent variable X.
From Eq. (3-48) it is clear that
Yok as t— ©
and Y-0 as t> -o
so that k is the upper asymptote and zero the lower asymptote. The first
derivative of Eq. (3-48) is
Giana
—=— Gee Y(k -- Y) (3-49)
-4
Thus the rate of change of Y with respect to ¢ is proportional to the current level
Y and also to the distance still to travel to reach the saturation level k. The first
derivative is positive for all values of t. The second derivative may be written
e a dY
(3-50)
Setting this to zero gives a point of inflection at
ee ee
a
Thus when Y < k/2, the “large” value of k — Y dominates the “small” value of
Y in Eq. (3-49) and causes dY/dt to increase. As Y increases toward kKy2=the
relative balance of the two forces changes so that dY/dt reaches a maximum
value when Y= k/2 and thereafter declines steadily as Y rises toward the
saturation level k. The typical shape of the logistic curve is shown in Fig. 3-10,
the main contrast with Fig. 3-9 being that the point of inflection occurs at half the
saturation value rather than at a much lower level. The logistic curve is frequently
used as a plausible approximation to the growth of any “population,” whether
bacterial, animal, human, or economic, where growth is thought to be positively
related to the size of the existing population and negatively related to the current
distance from a saturation level.
EXTENSIONS OF THE TWO-VARIABLE LINEAR MODEL 73
idys a
yar 72 (3)¥
If time is measured in constant units, the left-hand side of this equation is
approximately AY/Y, the proportionate rate of growth of Y. Thus one might fit
the linear regression
Tie
z Nyt=a-(f)¥
P a
+e, (3-51)
which yields point estimates of a and k. To obtain an estimate of b, we note that
Eq. (3-48) can be rearranged to give
hay
b= 7 ec. | (3-52)
Thus a value of 6 can be computed for each Y,, given a and k. An estimate b may
then be obtained by averaging some or all of these computed Db values, or
alternatively by substituting Y and ¢ in Eq. (3-52).
There are two main difficulties with this simple procedure for estimating the
logistic function.+ First of all we have point estimates of the parameters, but
inference procedures are difficult, especially for b and k. Second, there is evidence
that the procedure is unsatisfactory compared with a direct estimation by
nonlinear methods.
+ See F. R. Oliver, “Methods of Estimating the Logistic Growth Function,” Applied Statistics,
1964, pp. 57-66. This paper describes an iterative program for computing the (nonlinear) least-squares
estimates. A further paper, F. R. Oliver, “Notes on the Logistic Curve for Human Populations,”
Journal of the Royal Statistical Society, vol. 145, 1982, pp. 359-363, gives formulas for the asymptotic
standard errors of the least-squares estimates.
74 ECONOMETRIC METHODS
and a, 8, y, and 6 are parameters. From Eq. (3-53), the average cost (AC) and the
marginal cost (MC) are obtained, respectively, as
and
d(TC
MC = a Js B + 2yO + 38Q? (3-55)
With certain restrictions on the parameters, these functions will have the conven-
tional shapes shown in Fig. 3-11.
Relation (3-54) shows AC asalinear function of the variables
Q, aQ > aide Oe
EXTENSIONS OF THE TWO-VARIABLE LINEAR MODEL 75
Le AC,MC
O = iO ane O) >Q
Figure 3-11
+ Again, to keep the notation simple, we are not distinguishing between b; as a variable parameter,
as in Eq. (3-60), and 5; as the least-squares estimator, defined in Eq. (3-62).
EXTENSIONS OF THE TWO-VARIABLE LINEAR MODEL 77
Thus the regression plane passes through the point of means. Summing Eqs.
(3-63) over the sample observations and dividing by n gives
Y = b, + bX, + bX,
Comparison with Eq. (3-65) shows directly that
A
ye=ay,
and so from Eq. (3-64),
e@=0 (3-66)
Thus the mean of the regression values of Y is equal to the mean of the actual
sample values or, in other words, the sum of the least-squares residuals is
identically zero.
Result (3-66) may also be obtained directly from the condition
d(RSS) _ 0
Ob am,
for
0(RSS) : =
ab, =e we b, X,,) = —2 0 e,=0
t=] =A
The equality of the other two partial derivatives to zero gives the important result
that the least-squares residuals are uncorrelated with X,, X,, and Y. For
a(RSS ee
“= ies —2)° X,,¢,=0 (3-67)
2 t=1
and
a(RSS
(RSS)
ab, ie —20X,,e, = 0 (3-68)
Further
rz, ea (Zyp)°
eye ae)
But
EVV = Ly (Pre) = Dy
Thus
r2, = ry" = R?
Wy: yy?
R2 = zy?
Ly?
The exposition so far has been slanted toward the calculation of estimates by
hand or on a desk calculator. This may seem uncalled for in an era when there is
EXTENSIONS OF THE TWO-VARIABLE LINEAR MODEL 79
w (%)
ib
O
Cle Oo
oO
ae oO O
sere
oO
4 Oo
& O
3 —
2
00
O dA Uu
4.0 5.0 6.0 |‘ Figure 3-12
Example 3-3 Figure 3-12 shows a scatter diagram of w, versus u, for the first
16 quarters of the data, that is, from 1954: 2 to 1958: 1. It is suggestive of a
nonlinear negative relationship.f Fig. 3-13 shows the scatter of w, plotted
against the reciprocal of unemployment. The slope is now positive and the
scatter approximately linear. The simple regression fitted to these data gives
+ The nonlinearity is, in fact, very slight. A linear regression of w, on u, has an r? = 0.8192, which
is negligibly smaller than the r? of 0.8229 for the linear regression of w, on 1 /u,.
80 ECONOMETRIC METHODS
w (%)
7
olga
0.15
— !
0.20
ae u,
0.24 Figure 3-13
We see that the coefficients change substantially and the fit to the longer
period is much worse than that to the shorter period. This is just the first
example of how regressions can often change markedly when sample data are
extended, and in Chap. 6 we will outline methods of analyzing such changes »
and testing for structural changes in the hypothesized relationship.
Perry introduced the lagged inflation rate as an additional explanatory
variable, assuming it to be a good proxy for the expected rate of inflation.
The resultant multiple regression for the early period 1954: 2 to 1958: 1 is
We see that for the early period the addition of the lagged inflation rate gives
no improvement in the regression: the square of the multiple correlation
coefficient is identical to the third decimal place with the square of the simple
correlation coefficient, and the coefficient of the inflation rate is negligible in
size and perversely signed. For the complete period, the addition of the
EXTENSIONS OF THE TWO-VARIABLE LINEAR MODEL 81
The b’s are point estimates of the hypothetical 8’s, and all the usual inference
questions arise, such as those discussed in Chap. 2. There is little point in working
out these questions explicitly for the three-variable case, for in Chap. 5 we will
develop all the required inference procedures for the general case of k variables.
However, before leaving the three-variable case there are some additional alge-
braic relations to develop, which will also contribute to our understanding of
more complicated relationships later.
With three interrelated variables Y, X,, and X, there are three simple
correlation coefficients denoted by
Tio. 713, and rr,
where the subscript | refers to the Y variable. The techniques of Chap. 2 would
also enable us to compute various regression slopes such as
b, = slope of regression of Y on X,
b,, = slope of regression of Y on X,
by3 slope of regression of X, on X,
and there are, of course, three further regression slopes where the order of the
subscripts on b is reversed.} The first question to explore is the connection
between b, and b,, the slopes of the regression plane, and the simple b’s, the
slopes of the various two-variable scatters. Solving Eq. (3-73) for b, gives
PLA tae ay oe
2
Ex2Ex?2 — (Lx>x5)°
Dividing top and bottom by Lx3Lx; then gives$
eins by3b3. _ M2 = M1373 51 (3-76)
b
7 1 = by by l-rA 52
Similarly,
ie bi3 — By2bo3 _ M13 Miaha3 Si (3-77)
: 1 — by3b3y l-r, 53
+ In the regression X, on X; the residuals are measured in the X, direction, while in the regression
of X; on X, the residuals are measured in the X3 direction.
¢ This follows directly by applying Eqs. (2-14a) and (2-19), that is,
Ux34X a Ux2X3 Ss]
Lyx Dia12 = Tosa
Dia = Ex? ’ b32 ? 23 > 12
ty)
?
“x3 xe
82 ECONOMETRIC METHODS
la eS ame
X(y oe b13X3)(x, - bo3X3)
(3-78)
VE(y im bisxs)s E(x, = bas
EXTENSIONS OF THE TWO-VARIABLE LINEAR MODEL 83
This latter expression may also be derived from first principles by finding the
simple correlation between
= 0 =D
Ely —bisxs)2 ~bass) =EW(%2 ~bass) since x, is uncorrelated with
X(x2 — by3x3) L(x, — by3x3) the residuals (x, — b,3x3)
= 13 \82(T2 — 113%23)
where s, denotes the sample standard deviation of the Y values, s, the sample standard deviation of
the X, values, and so on. In the denominator L(y — b,3x3)° is the residual sum of squares in the
regression of Y on X3. From Eq. (2-20) this may be written as ns?(1 — rj3). Likewise L(x — by3x3)*
is the residual sum of squares in the regression of X, on X3 and may be written as ns3(1 — 73). Thus
the denominator of Eq. (3-78) is ns,s2y1 — ri, yl 15, and Eq. (3-79) follows.
84 ECONOMETRIC METHODS
from Eq. (3-73). Likewise, b, is just the simple regression slope obtained when the
residuals (y — b,,x,) are regressed on the residuals (x; — 53x).
Finally, the multiple correlation coefficient R,5; is a second-order correlation
coefficient, and it may also be expressed in various ways in terms of lower-order
coefficients. For example, using Eq. (3-74) gives
. Dix, Vat Dg XGy
Ribs = eee eae
Ly
Substituting for b, and b, from Eqs. (3-76) and (3-77) gives
he ps ria + 13 — 2 ishs
Ri23 = 1 2
(3-81)
mala)
The buildup of the explained sum of squares may also be looked at in sequential
fashion, and this is illuminating for the analysis of variance treatment in Chap. 5.
Suppose one first regressed Y on X,. Then
ry? = ESS due to the regression of Y on X,
and
ry? = i) = RSS from the regression of Y on X,
This latter quantity is the sum of squares still to be explained. Regressing the
residuals (y — b,,x,) on the X, residuals, namely, (x, — b3,x,), then gives
Aggregating the ESS at each stage gives a total explained sum of squares in Y as
rdy(I = rp)
EXTENSIONS OF THE TWO-VARIABLE LINEAR MODEL 85
where Yy7(1 — rj) is the RSS after Y has been regressed on els) hes
indicates the proportion of the variation left after the simple regression on i
which is explained by adding X; to the set of explanatory variables. A similar
development and interpretation may be made by starting with the simple regres-
sion of Y on X; and then adding X, to the explanatory variables.
Table 3-7 Wage change, unemployment, and price change in the United States}
For the complete sample period the zero-order correlation coefficients are
ry = 0.6508 43 = 0.3567 —r = 0.0726
r2, = 0.4235 =r, = 0.1272
Thus
|. 0.6508 — 0.3567(0.0726) atieaed
a
[1 — (0.3567)"] \/[1 — (0.0726)"|
0.3567 — 0.6508(0.0726)
hy2 = ee = 0.4087
[1 — (0.6508)?] [1 — (0.0726)"|
and rh, = 0.4498 7r2,,. = 0.1670
and Rj>; = 0.5197. In this example the two explanatory variables are practi-
cally uncorrelated, so there is little difference between the zero-order and the
first-order correlation coefficients. Unemployment alone explains over 42
percent of the variation in wage change, price change alone accounts for
about 13 percent, and the two variables jointly account for about 52 percent.
Unemployment accounts for 45 percent of the variation unexplained by price,
and price accounts for about 17 percent of the variation unexplained by
unemployment.
PROBLEMS
3-1 Given five observations u_,, u_,, Uo, u;, and uy at equally spaced points of time ¢ =
— 2, —1,0,1,2, show how to fit a parabola to the observations by least squares and show that the
value given by the parabola at time t = 0 is
Time,
h Firmness
Estimate the parameters in a linear regression of firmness on time. Give standard errors of the
estimates and test the adequacy of a linear regression to describe the results.
(R.S.S. Certificate, 1955)
3-3 Discuss briefly the advantages and disadvantages of the relation
v; = a + Blog vo
as a representation of an Engel curve, where v; is expenditure per person on commodity i and vg is
income per person. Fit such a curve to the following data, and from your results estimate the income
elasticity at an income of £5 per week.
Does it make any difference to your estimate of the income elasticity if the logarithms of vg are
taken to base 10 or to base e? Explain carefully.
(Manchester University, 1956)
3-4 Response rates at various levels of ratable values.
The data relate to a survey recently conducted in England. Estimate the constants in the regression
equation
all variables are expressed as deviations from their sample means. Consider the following alternative
procedures for estimating . ‘
(a) Calculate the estimates B and y ina regression of y on x, and x}.
(b) Regress y on x, and calculate the regression residuals y,*; regress x, on x, and calculate the
regression residuals x*,; regress y* on xf to obtain an estimate b of £.
Show that the two procedures give the same result, that is, B = b.
Show that the regression residuals given by each procedure, that is,
3-7 Outline the properties of the following functions, and sketch their graphs:
(a) y=at Blnx
x
(b) y= oe
ert Bx
Coa eens
Find transformations which linearize functions (b) and (c), that is, for each function find a pair of
transformations f(x) and g(y) such that g() is a linear function of f(x) and the a, 8 parameters
may be estimated.
3-8 Your research assistant reports the following results in several different regression problems. In
which cases could you be certain that an error had been committed? Explain.
(a) R253 = 0.89 and Rj 34 = 0.86
(b) ri = 0.227, r>, = 0.126, and R7,, = 0.701
(c) (Ux?)(Ly?) — (Lxy)? = — 1,732.86
(University of Michigan, 1980)
3-9 Sometimes variables are standardized before the computation of regression and correlation
coefficients. Standardization is achieved by dividing each observation on a variable by its standard
deviation, so that the standard deviation of the transformed variable is unity. If the original relation is,
\
Say,
Y = B, + By)X_ + B3X3 + u
and the corresponding relation between the transformed variables is
Y* = By + By XZ + B3XF + u*
where
Ye = /Sey eX) oe AAS, i= 2,3
what is the relationship between 83, 83, and 8,, B;? Show that the partial correlation coefficients are
unaffected by the transformation.
CHAPTER
FOUR
ELEMENTS OF MATRIX ALGEBRA
It is clear from the last section of Chap. 3 that it would be excessively tedious and
complicated to build up to the general case of k-variable regression in a stepwise
fashion. Fortunately, by the use of matrix algebra we have a compact and
powerful way of treating the problem, and we shall see that the detailed results of
the previous two chapters are merely special cases of a few simple matrix
formulas. The rest of this chapter presents the elements of matrix algebra that are
necessary for following the treatment in the remainder of the book. Most or all of
this chapter may be skipped by those with adequate previous knowledge of
matrices. For those whose knowledge is somewhat rusty it may hopefully serve as
a useful review, but every attempt has been made to make the material accessible
to a student with no prior knowledge of matrices, for the subject is so fundamen-
tal to modern econometrics (and economics) that no serious student can afford to
be without it.
Suppose our theory suggests that a dependent (explained) variable Y is a
linear function of k — 1 independent (explanatory) variables X,, X;,..., X;,.
+ Nonetheless students with no prior knowledge of matrix algebra are likely to get indigestion if
they attempt to work through all the material in this chapter before proceeding with the rest of the
book. The topics are introduced approximately in the order in which they will appear in subsequent
chapters. Thus the students should interact between this chapter and those that follow, learning
enough from Chap. 4 to proceed with Chaps. 5, 6, and so on, and returning to Chap. 4 as necessary.
Summaries of results have also been inserted at various stages in the chapter.
89
90 ECONOMETRIC METHODS
Y, 1 X, Xx eee Ne 11 Ps uy
Y, & Ag Aa es Be x e (4.2)
¥, 1 yo X3,, AP B, Uu,
or
y=XBp+u (4-3)
where the four boldface symbols correspond to the four sets of elements that have
been enclosed in square brackets in Eq. (4-2). These symbols indicate vectors and
matrices. For example,
1 X), X3; Xx
Xo eee Kea
ace
1]X, X3,
ee ae
Xu
is another special case, namely, a matrix with just one row.t We may thus look at
the X matrix in two ways, as an ordered collection of column vectors or as an
ordered collection of row vectors. Each column vector, apart from the first,
denotes the sample observations on a particular explanatory variable. Thus, for
example,
denotes the sample observations on the variable X,. The first column is a
collection of units and, as we will see, is required in order to incorporate the
intercept 8, into the regression. Using this notation for the column vectors, we
could express X as
bat] |
Rex 7) PX ea Xe (4-4)
X= : (4-5)
—* s —_—
We have inserted vertical and horizontal lines in Eqs. (4-4) and (4-5) to emphasize
that the first is a representation of X in terms of column vectors and the second a
representation in terms of row vectors. In practice one usually writes these
expressions more compactly as
S)
S2
X= [xy yer a) ‘
n
it being clear from the context which are row and which are column vectors.
+ We will adhere to the convention of indicating a matrix by an uppercase boldface letter and a
vector by a lowercase boldface letter. ;
+ We have indicated the row vectors by the letter s since they correspond to sample points. A more
common notation is to let x;. indicate the ith row of the X matrix and x. ; the jth column.
92 ECONOMETRIC METHODS
Equations (4-2) and (4-3) are equivalent ways of stating Eq. (4-1). For this to
be so, operations on matrices must follow certain simple rules, which we will now
describe.
The right-hand side of Eq. (4-3) indicates two elementary operations on matrices,
namely, multiplication and addition.
Matrix Multiplication
Matrix multiplication is achieved by repeated applications of vector multiplica-
tion. The multiplication of an n-element row vector into an n-element column
vector is defined as follows:
that is, corresponding elements are multiplied together and the results summed.
As a numerical example,
a, b,
Suppose, however, we define the vectors a and b as column vectors, that is,
a, b,
em) b,
ey bee (4-7)
a, b,
The above multiplication definition cannot apply directly to a and b since they are
both column vectors. We thus define an operation of transposition, which turns
ELEMENTS OF MATRIX ALGEBRA 93
aa a
ou ue ey | ajb, a,b, --- ajb,
s beth, 9-2 b=) a,b, aby 9 2) a,b (410)
Re Sorel | a,,b, a,b, a,b,
Example 4-1
Example 4-2
[ne 1(1) + 6(2) 1(2) + 6(0) 1(3) + 6(4)
|: IL =| 0(1) +.1(2) 0(2) +10) 0(3)
+1(4)
jem 1(1) +1(2) 12) +10) -1(3)
+1(4)
[gard |
BA = 2~-0 4
oe: 7
As the examples and the definition in Eq. (4-10) make clear, the order in
which matrices are multiplied is of crucial importance:
This operation is only possible if the inner products of rows of A and columns of
B exist, that is, if the number of columns inA is equal to the number of rows in
B. In this event the matrices are said to be conformable. The simplest check is to
write down the order of the two matrices to be multiplied, as in Eq. (4-9), and it is
seen that the common index n disappears to give a product matrix of order
m X p. The product BA would only exist if p = m so that the inner products of
the rows of B and the columns of A could be formed. Note carefully that the
definition of matrix multiplication involves the inner products of the rows of the
first matrix and the columns of the second.
A special case of Eq. (4-10) occurs when one of the matrices is simply a
vector. For example,
—- a — a,b
— a, —||! a,b
: Deas eal oes c (4-11)
| :
— a a,,D
and ¢ is an m X 1 (column) vector.
Returning now to Eq. (4-2), the right-hand side incorporates the multiplica-
tion of a matrix by a vector, and applying the above rules gives
xg = |fit
By
Became et
+ By Xo Bat eee
Xe (4-12)
By HeRG Aon bsg et eee
which is just an n-element column vector.
Matrix Addition
The right-hand side of Eq. (4-2) or Eq. (4-3) is now seen to consist of the addition
of the two vectors XB and u. This addition is simply achieved by adding
corresponding elements. Thus the operation is only defined for vectors with the
ELEMENTS OF MATRIX ALGEBRA 95
a, | b,, Garb.
The definition in Eq. (4-13) is readily extended to the addition of matrices. Two
matrices A and B can be added together only if they are of the same order m X n.
The sum matrix is also of the same order, and each element in it is simply the sum
of the corresponding elements in A and B.
Applying Eq. (4-13), the right-hand side of Eq. (4-3) reduces to an n-element
column vector,
A’ =
(nXm)
96 ECONOMETRIC METHODS
that is, the first row of A has become the first column of the transpose, the second
row of A the second column of the transpose, and so forth. The definition might
equally well have been stated in terms of the first column of A becoming the first
row of A’, and so on. Clearly, A’ is of order n X m.
Example 4-3
that is,
Ay = a,; fori + j
This property can only hold for square matrices (m = n), since otherwise A’ and
A are not even of the same order.
Example 4-4
Li 4
A=] -1 0 3) =A’
4 5 2
CC
le
= a,b;
rah
But
ELEMENTS OF MATRIX ALGEBRA 97
A(B + C) = AB + AC (4-18)
To see this, let a, denote the ith row of A and b, and ¢; the jth columns of B and
C. The i, jth element on the right-hand side of Eq. (4-18) is then the scalar
a,b; + a,¢
and by the application of the distributive law for scalar algebra this is clearly
equal to the inner product of a; and the vector b; + ¢,, the jth column of B + C,
which gives the i, jth element on the left-hand side of Eq. (4-18).
98 ECONOMETRIC METHODS
-2|} 2 ales ie —4 af
ara A =a 0. 5
There are some square matrices of particular importance. First is the unit or
identity matrix of order n X n,
with units down the main, or principal, diagonal and zeros everywhere else. As we
shall see, it plays in matrix algebra a role similar to that of unity in scalar algebra.
As one illustration,
IA=AI=A
that is, pre- or postmultiplication by I leaves any matrix unchanged, as may
readily be verified by multiplying out IA and AI. Thus the unit matrix may be
entered or suppressed at will in matrix expressions. For instance,
YoY
= Vie
= (ei
A diagonal matrix is like the identity matrix in that all off-diagonal terms are
zero, but now the diagonal elements are scalar quantities, one of which at least is
nonzero. The diagonal matrix may be written
OO 0
ie OA 5aeeO 0
=a 120Sec Gnomesae aa e
0 Coe)
2 0 0
ie i and 3 —4 )
0 0 5
A special case of Eq. (4-19) occurs when the )’s are all equal. This is termed a
scalar matrix and may be written
rA 0 0
0 A 0|}=AI
Galpie ea ham‘
A0=0
Similarly, we may have null row or column vectors.
Partitioned Matrices
Writing the matrix X in the form
Re Xo, ae Xe!
as in Eq. (4-4), is a special example of a partitioned matrix. The elements on the
right-hand side are not scalars but vectors. In general a partitioned matrix
contains submatrices as elements. The submatrices are obtained by partitions of
the rows and columns of the original matrix. For example,
4 0 2!-1 A A
A=| 6 5 Sis, al = ee al (4-20)
us 3 a 0! 5 21 22:
where
AM iOe a
Ay, i Saal Ds | |
(4-21)
AS = lao 20] Ay =5
The dashed lines indicate the partitioning, yielding the four submatrices defined
in Eqs. (4-21).
+A square nonsymmetric matrix is idempotent if it satisfies A? =A, but we will only meet
symmetric idempotent matrices in this book.
100 ECONOMETRIC METHODS
Our previous rules for the addition and multiplication of matrices apply
directly to partitioned matrices provided the submatrices are of appropriate dimen-
sion. For example, if A and B are both written in partitioned form as
AC
ae and B=
i =
a Ay B,, By
then
A,, + By, Ay t+ By
A+B=
A», ztB,, A» a: B,,
provided A andBare of the same overall order (dimension) and each pair A; ,, B;,
is of the same order. As an example of the multiplication of partitioned matrices,
A A
AB mee ie ie ay
B,, B,,
A; A)
1. The scalar, dot, or inner product of two n-element column vectors a and bis
a’b = bD’a = L?_,a,b,.
2. The typical i, jth element in the product AB, where the matrices are
conformable for multiplication, is D.a,.b
Set Stee S
ic
3. The typical element in A + B is a, 7 Op.
4. A=B means a,, = b,, for all i, /.
5. (AB)’ = B’A’, (ABC) = C’B’A’.
6. (A+ B)+C=A+(B+O).
7. (AB)C = A(BC).
8. A(B + C) = AB + AC.
9. IA= AI=A.
10. The typical element in cA, where cis a scalar, is ca, i
11.A+0=0.
12. AO= 0A = 0.
ELEMENTS OF MATRIX ALGEBRA 101
Returning again to the linear model y = XB + u, this may be written in the form
B,
B,
Yor X75 Kyat X el ial Gee
B,
that is,
and, treating the elements of b as variables, we have to minimize e’e with respect
to b. This requires some elementary results on matrix differentiation.
Matrix Differentiation
If f(b) contains, say, k different b’s, then we may partially differentiate f(b) with
respect to each 5, in turn, obtaining k partial derivatives. Arranging these partial
derivatives in the form of a column vector gives the general definition
a[ f(b)
ab,
a[ f(b)]
al f(b)]
Ta fe (4-26)
aL f(b)]
These derivatives might equally well have been arranged as a row vector. The
important requirement is consistency of treatment and ensuring that vectors and
matrices of derivatives that have to be added and multiplied are of appropriate
order.
Suppose f(b) is a linear function,
f(b) = a/b
=10 Dias Dye a a,
where the a’s are given constants. Application of Eq. (4-26) then
ay
d(a’b) —A(b’a) ea)
Ab Go. oe alee oes (4-27)
on
41; 4p Ai,
A=]. ay ay,
Gi, Ar, AkK
Then
+. Gy,
De
0(b’Ab
ae ) = 2(a,,b, ta a>,b, Steen iet Gee) = 2a,b
where the a’s indicate the rows of A. Collecting these partial derivatives in a
column vector,
oH a,
,
0(b’Ab) =, a,b
: es a;
ob : :
b = 2Ab (4-28)
ae ay
Equations (4-27) and (4-28) give the standard results on the differentiation of
linear and quadratic forms. Notice the parallel with the differentiation of scalar
functions in that the power of the variable is reduced by | so that it disappears on
differentiation of the linear form and appears linearly on differentiation of the
quadratic form.
These two results may now be applied directly to minimize the residual sum
of squares defined in Eq. (4-24).
aWX'y)
ob
_ yy
using Eq. (4-27), since X’y is just a known k-element vector. Also
0(b’X’Xb) _ ,
masstee 2(X’X)b
(ee)
ob
_ _ ayy + 2X'Xb
For a stationary value of the sum of squares all k partial derivatives must be zero,
that is,
d(e’e) _
ae
104 ECONOMETRIC METHODS
and so
(X’X)b = X’y (4-29)
These are the normal equations for the least-squares regression, and include the
equations for the two- and three-variable cases already derived in Chaps. 2 and 3.
eas
5 lex
hes
Thus
oe ee and xy =|
XG Ee DEDE
So Eq. (4-29) gives
nb, + bX = LY
UX + bE XA ee
which are identical with Eqs. (2-13).
and
ai
Xiy = 3| AG
DOXGY
which, on substitution in Eq. (4-29), yield Eq. (3-62).
2d component
This may be pictured as a directed line segment, as shown in Fig. 4-1. The arrow
denoting the segment starts at the origin and ends at the point with coordinates
(2,1). The vector a may also be indicated by the point at which the arrow
--[l
terminates. If we have another vector, say, b,
the geometry of vector addition is conceived as follows. Start with a and then
place the b vector at the terminal point of the a vector. This takes us to the point
P in Fig. 4-1. This point defines the vector ¢ as the sum of vectors a and b, and it
is obviously also reached by starting with the b vector and placing the a vector at
its terminal point. The process is referred to as completing the parallelogram, or
as the parallelogram law for the addition of vectors. Clearly, the coordinates of P
cneve
LFD]-L3
are (3,4), and
so that there is an exact correspondence between the geometric and the algebraic
treatments.
»Ai]-[f
Now consider scalar multiplication of a vector. For example,
gives a vector in exactly the same direction as a, but of twice the length. The
scalar multiplier may also be a negative number. For example,
cae
These two vectors are shown along with a itself in Fig. 4-2. Clearly, all three
terminal points lie on a single line through the origin, that line being uniquely
defined by the vector a.
106 ECONOMETRIC METHODS
2d component
Ist
component
Figure 4-2
efi
--23ff]-$U
may be expressed as
ee
1 3 NE OAS
caffe
may be expressed as
giving A, = 3 andA, = 0.
ELEMENTS OF MATRIX ALGEBRA 107
reff
may be expressed as
eo? +22]
giving A, = 2 and A, = 2.
1. If v, and vy, are any two vectors in the space, then v, + vy, is in the space.
2. If v is in the space and Ais a scalar constant, then Av is in the space.
The set of vectors is said to be closed under addition and scalar multiplication, for
these operations do not produce a vector outside the space.
Let us denote the two-dimensional space by the symbol 7. This vector space
consists of all real two-element vectors. Clearly, any vector in the space can be
expressed as a linear combination of the two vectors a and b. Our specification of
a-[] me a[’
a and b, however, was arbitrary. Consider another pair of vectors
These are usually described as unit vectors, and again any vector ¢ in R” may be
expressed as a linear combination of these vectors, only now the determination of
the A’s is particularly simple. The three previous numerical examples in this case
give
2 (5) 6] +3|°| on 6 8
EE] aw 9 [f
As an illustration of these definitions, suppose
These vectors lie on the same ray through the origin, and the linear combination
3a — b yields the zero vector. However, if we revert to the original a and b
vectors, namely,
Pie male
a= 1| and b i
it is impossible to find a pair of A values, other than two zeros, such that
A\,at+A,b=0
We can easily find a pair of A values to reduce the first element to zero, but the
same combination will not reduce the second element to zero. A basis for R7 is
thus defined to be any linearly independent pair of two-element vectors: It is clear
from the geometry of the two-dimensional case that the representation of a given
vector in terms ofa given basis is unique, that is, there is one and only one pair of
X,, A, values which satisfy
c=A,a+A,b
Given a basis a,b for ®*, we have seen that any vector c in ®? may be
expressed as a unique linear combination of the basis vectors. Thus the vectors a,
b, and ¢ are linearly dependent, for the equation
A\,at+A,b-—c=0
holds for nonzero A’s. We might ask whether any arbitrary vector v in R” may be
expressed in terms of the expanded set of vectors a, b, and c. The answer is, of ©
course, yes, but the coefficients will not be unique. For example, suppose that the
fl EL Ll
a, b, and ¢ vectors are
v = 2a + 2b + 0c
but there are infinitely many others. Rewriting the general linear combination
v=A,a+A,b+A,c
in the form
v—A,¢c=A,a+A,b
any arbitrary value can be assigned to A,, and the left-hand side is then some
ELEMENTS OF MATRIX ALGEBRA 109
|[b||* = b’b
Substituting in Eq. (4-32) gives
a,b, a,b,
cos 8§= ——————— + aaee ee
2d
component
Ist
component
110 ECONOMETRIC METHODS
that is,
a’b
cos 9 = ————— (4-33)
va‘a Vb’b
where it is understood that we take the positive square roots to indicate length.
There are two important special cases of Eq. (4-33). When a and b are
linearly dependent, we may write
b=da
where A is some appropriate scalar. The right-hand side of Eq. (4-33) then reduces
to unity, giving @ = 0°. When a and b are at right angles to each other, 0 = 90°
and cos 6 = 0, giving a’b = 0. Conversely, when a’b = 0, 0 = 90°. Two vectors at
right angles are said to be orthogonal. Thus two vectors are orthogonal if and only
if
a’b = 0
If we take just two of these vectors, say, e, and e,, then all linear combinations of
e, and e, constitute a vector subspace in ®*, namely, the horizontal plane, since
the third component in each spanning vector is zero. More generally, any two
three-element vectors, say,
1 5
a=|2 and li ==|)
3 1
span or generate a plane surface, as indicated in Fig. 4-4, by the plane containing
Oab.
The set of all real n-element vectors constitutes the space ®”. Each vector in
§” may be expressed as a unique linear combination of some appropriate set of n
linearly independent vectors. To see that the linear combination must be unique,
suppose that a vector v can be expressed as two different linear combinations of
the basis vector v,,¥,,..., v,, namely,
V=Ay,
+A +--- +A nn
and
3rd
component
4
2d
component
Ist
component
Figure 4-4
Oe BV Wo Ge (A eye
ieepg oe Ba ee A i ne
Figure 4-5
Xb [xj x. 2 XE ix boxe
by.
which is then a vector that lies in the column space of X. Choosing different b
vectors gives, in turn, different Xb vectors. To each such Xb vector there
corresponds a vector of residuals e, so that the equation
Y=Xbte
as Fig. 4-5 shows, gives y as the sum of two vectors, of which one, Xb, lies in the
space spanned by the columns of X and the other, e, lies outside that column
space.
We wish to choose the b vector so as to make the point given by the tip of the
Xb vector as close as possible to the tip of the y vector or, in other words, to
minimize the length of the e vector. This is achieved by making the e vector
perpendicular to the hyperplane generated by the columns of X. Thus e must be
orthogonal to any linear combination of the columns of X. We have
e=y
— Xb
and Xc is any arbitrary linear combination of the columns of X. Thus the
orthogonality condition gives
X’y — X’Xb = 0
ELEMENTS OF MATRIX ALGEBRA 113
The next problem is how to solve Eq. (4-35) for the desired least-squares
coefficients b. From the original definitions the dimensions of Eq. (4-35) are as
follows: X’X is a square matrix of order k X k, and b and X’y are all k-element
vectors. Thus Eq. (4-35) expresses the X’y vector as a linear combination of the
columns of X’X, and b indicates the coefficients of that linear combination. If the
columns of X’X are linearly independent, they constitute a basis for R*, and any
k-element vector, such as X’y, may then be expressed uniquely in terms of the
basis vectors. In other words, Eq. (4-35) has a unique solution for the b vector.
The solution of Eq. (4-35) for b may be expressed in terms of an inverse
matrix. The meaning of an inverse matrix may be developed as follows. Let A be
a square matrix of order n and let the n columns of A form a linearly independent
set. Does a square matrix B of order n exist such that
AB =I? (4-36)
+ A vector such as e, which is orthogonal to every vector on the hyperplane generated by the
columns of X, is said to be normal to the hyperplane—hence the term normal equations.
114 ECONOMETRIC METHODS
The answer is yes. Letting b, denote the first column in B and equating first
columns on both sides of Eq. (4-36) gives the vector equation
Ab, =e, (4-37)
where e’, =[1 0 0 --- OJ. Since the columns of A are linearly independent,
the vector e, can be expressed as a unique linear combination of those columns.
Thus b, is uniquely determined. By a similar argument each column of B is
uniquely determined, and so there is a matrix B satisfying Eq. (4-36).
We shall see later in this section that if the n columns of A are linearly
independent, then so are the n rows. Then by a similar argument, a square matrix
C of order n can be found such that
CA = I (4-38)
for each row of C is uniquely determined as the coefficients of a linear combina-
tion of the rows of A. Thus Eqs. (4-36) and (4-38) are both true. Postmultiplying
Eq. (4-38) by B gives
CAB = IB=B
But
CAB = CI=C
using Eq. (4-36). Thus
C=B
Thus if the n columns (and rows) of A are linearly independent, a unique square
matrix of order n exists, called the inverse of A, and denoted by A~', such that
AA'=A'A=] (4-39)
If we assume that the k columns of X’X are linearly independent, then the
inverse matrix (X’X)~' exists. Premultiplying both sides of Eq. (4-35) by this
inverse gives
Rank of a Matrix
Consider any arbitrary matrix A of order m X n. The columns of A define n
vectors in ®”. Likewise, the rows in A define m vectors in ®”. Let r denote the
ELEMENTS OF MATRIX ALGEBRA 115
It is obvious that the rank of a matrix cannot exceed the number of columns
or the number of rows, whichever is the smaller. That is,
e(A) < min(m, n) (4-41)
When p(A) = ™, we say that the matrix has full row rank, and when p(A) = n,
that it has full column rank, but, of course, in any specific case, row rank and
column rank are identical, and we speak unambiguously of the rank of the matrix.
Notice that it follows directly from the definition of rank that the rank of the
transpose of A is equal to the rank of A, that is,
Ai | Ay } r rows
a eae ee Se
A>, | Ax } m — r rows
ers es
r n-r
columns columns
Thus Aj, is a square nonsingular matrix of order r. Consider now the set of
homogeneous equations
Ax = 0 (4-43)
where x denotes a column vector of n unknowns. The equations are said to be
homogeneous because of the 0 vector on the right-hand side of Eq. (4-43). If the
equations read Ax = b, for b = 0, they are said to be nonhomogeneous. Clearly,
and
if x, and x, are two distinct solutions to Eq. (4-43), then c,x, + Xx, is also a
solution.
Thus the set of solutions to Eq. (4-43) constitutes a vector space called the
nullspace of A. Our immediate concern is to establish the dimension of this
nullspace (that is, the number of linearly independent vectors which span the
subspace). Let us drop the last m — r rows from A and partition x conformably
with the columns of A. This gives
x, = —AyApx, (4-45)
The x, subvector is arbitrary or “free” in the sense that we can specify the n — r
elements in x, at will, but for any such specification the subvector x, is
determined by Eq. (4-45). Using Eq. (4-45), the general solution vector to Eq.
(4-44) may be written
AGA
I tPF
-
(4-46)
X9
The matrix in Eq. (4-46) has n rows and n — r columns. The n — r columns are
linearly independent. This fact is guaranteed by the presence of the I,,_, sub-
matrix, whose columns are necessarily linearly independent. Thus Eq. (4-46)
expresses all solutions to Eq. (4-44) as linear combinations of n — r linearly
independent n-element vectors. But any solution to Eq. (4-44) is also a solution to
Eq. (4-43), for the rows that have been discarded from Ato arrive at Eq. (4-44)
are linear combinations of the rows of [A,, Aj]. Any discarded row may thus
be expressed in the form
e[A,, Ay]
where ¢’ is some appropriate row vector of r elements. Postmultiplying by x gives
e[A,, Ay]x=0
since x satisfies Eq. (4-44). Thus each solution x holds for the discarded rows, and
Eq. (4-46) defines the solution vector for Eq. (4-43). Thus the nullspace of A has
dimension n — r. This gives the important result that for an m X n matrix A with
rank r
Number of columns = rank + dimension of nullspace (4-47)
n=r+(n-r)
118 ECONOMETRIC METHODS
The nullspace is sometimes referred to as the kernel of A and its dimension as the
nullity. Thus the result may also be stated as
Number of columns = rank + nullity
xy
12a emai
Le 2 eee Reer itO (4-48)
2 4 Raa a 4
0
The rank of the matrix is seen to be 2 since rows | and 2 are clearly linearly
independent, as are rows | and 3, but all three rows are not linearly
independent, since
row1 + row2 — row3 = 0
Discarding the third row gives the set
x]
Loe 3 ae sl eee)
| 2a x3 =a eae)
x4
Columns | and 3 are linearly independent, so we rewrite Eq. (4-49) as
|a3 ie 241 X>
(ee ee5 ene E LS
Solving this pair of equations for x, and x, gives
Xp ie het
x3, = eS ONE
and the solution vector to Eq. (4-49) may be expressed as
cane
Veg ie
So | ||| (4-50)
0 ©
YW
vl
=
The matrix in Eq. (4-50) has two linearly independent columns, and so any
solution to Eq. (4-49) may be expressed as a linear combination of two
linearly independent four-element vectors. The solution vectors thus form a
two-dimensional subspace in R*. There are infinitely many solution vectors
since the vector |
.
solution defined by Eq. (4-50) is also a solution to the initial set of equations
(4-48). This may be seen by noticing that each column vector in Eq. (4-50)
ELEMENTS OF MATRIX ALGEBRA 119
0
and
if
2
0
[2 4 4 5] as =0
2)
1
Since any solution vector x is a linear combination of these two column
vectors, then x satisfies the third equation in Eqs. (4-48). Since it already
satisfies the first two equations, it is a solution to Eq. (4-48). The nullspace of
A thus has dimension 2, which is equal to the number of columns in A minus
the rank of A. Each vector in the nullspace is orthogonal to each row in A.
There is a seeming element of arbitrariness in the partitioning that we
applied to Eq. (4-49) and also in the choice of the row of A to be discarded.
But this is apparent, not real. For example, suppose we partition Eq. (4-49) as
which gives
ee(| ec
Xg= 2x,+ 4x,
with solution vector
1 0
- 0 ee :
elas 6 ie (oy)
2 4
Equation (4-51) again defines the nullspace of the matrix A in Eq. (4-48). It
has dimension 2, and the columns in the matrix of Eq. (4-51) are linearly
independent. This, however, is the same nullspace as defined by Eq. (4-50),
for each column vector in Eq. (4-51) may be expressed as a linear combina-
tion of the column vectors in Eq. (4-50):
1 23 x
0 =i 1 +) 0 SeAieesOs
3 1 0 2 a8 l Nye
2
2 0 1
and
0 -2 4
1 =i 1 0 =>), ,=1,
a 1 0 +X 2 a l A,=4
2
4 0 1
120 ECONOMETRIC METHODS
Thus the nullspaces are the same. Discarding the first or the second equation
from A in the initial stage would also make no difference to the determination
of the nullspace.
We may note here a particular application of Eq. (4-47), which will be very
useful in the treatment of identification in Chap. 11. If A is m X n and has rank
n — 1, then the dimension of the nullspace of A is 1, that is, all solutions to
Ax = 0
lie on a single ray through the origin. Thus if
Xai Se end
is a solution, then so is
Ex! =" [cx aia eee nes)
for any constant c.
Result (4-47) also yields simple proofs of some important theorems on the
ranks of various matrices. We notice that the crucial matrix for the least-squares
vector in Eqs. (4-35) and (4-40) is X’X. The first important theorem states that
To prove the rest of theorem (4-52) we merely note that p(X) = p(X’), and
the above proof immediately gives
p(AQ) = p(Q’A’)
= p(A’) by the above proof
= p(A)
and finally
p(PAQ) = p(A)
follows directly from the previous results.
Both previous theorems involve special cases of the multiplication of one
matrix by another. In Eq. (4-52) a matrix was multiplied by its transpose. In Eq.
(4-53) multiplication was by a nonsingular matrix. Our final theorem on rank
relates to the perfectly general case of the multiplication of one rectangular matrix
by another conformable rectangular matrix. Let A be m X n and let B ben X s.
Then
p(AB) < min[p(A), p(B)] (4-54)
that is, the rank of the product AB is less than or equal to the smaller of the ranks of
the constituent matrices.
122 ECONOMETRIC METHODS
1. A is nonsingular.
2. A has rank n.
ELEMENTS OF MATRIX ALGEBRA 123
To study the jo of the inverse matrix let us begin with the 2 x 2 case.
Denote A and A~' as follows:
S| a 1 a "| re a1 a|
42, 422 Oo OD?
So far we have regarded matrices mostly as collections of vectors and paid little
attention to the individual elements. The standard notation is to use the first
subscript of an element to indicate the row in which that element appears and the
second subscript to indicate the column. The definition of the inverse gives the
general equation
AA~! =] (4-55)
Specializing this equation to the 2 X 2 case and taking just the first column from
each side of the equation gives
an eae |e
Qo, 40 || @21 0
Treating the elements of the inverse as unknowns, the solution of this pair of
equations gives
a2
MNS 11422 a ka 12421
aa
— 41
Sel sk 11422
Waa ala
12421
Similarly, equating the second columns in Eq. (4-55) and solving gives
— 412
[2.2 ata 1242)
11422
a)
22 a
a SS —_—————
411492 — 41242)
ee 1 Eee
aitEtLD Ay | 4-56)
— a (
=
A\14n7 — 4424) | ZI 11
and it may readily be checked that indeed AA! = AA”! =I. Each element in
A~' is a function of the elements in A, and even for the 2 X 2 case certain
important features of A~' are apparent. First, each element in the inverse has a
common divisor, namely, a,,45) — 4,74 ,. This is a function of all the elements in
A. It is a scalar quantity and is defined as the determinant of A. For the 2 Xx 2 case
we thus have .
The two expressions on the left of Eq. (4-57) are alternative ways of indicating the
determinant. The final expression on the right means
Ds a Q\q42B 7.
a,B
sum of all possible products of the elements of A, taken two at a time,
with the first subscript in natural order 1,2 and a, B indicating all
possible permutations of 1,2 for the second subscript, each product
term being affixed with a positive (negative) sign as the number of
inversions of the natural order in the second subscript is even (odd).
There are only two possible permutations of 1,2, namely, 1,2 itself and 2, 1.
There is one inversion of the natural order in 2, 1 since 2 comes before 1. Thus the
terms in the expansion are simply
411422 — 4129)
The numerators of the elements in A~' could have been produced by the
following two rules:
1. For each element in A, strike out the row and column containing that element
and write down the remaining element prefixed with a positive or negative
sign in the pattern
a,B,y
There will be 3! = 6 terms in the expansion, since that is the number of possible
permutations of 1, 2, 3. Half will have a positive sign and half a negative sign. The
explicit expression is
JA] = @)14y2433 + 41747343) + G1347)43) — 441473432 — 412471433 — A)349)45,
(4-59)
As a check on the signs we may notice, for instance, that in the third term in Eg.
ELEMENTS OF MATRIX ALGEBRA 125
|
Gy O23
Mann 433
rather than a scalar element. We, in fact, replace a,, with the determinant of this
submatrix, appropriately signed, and similarly for the other elements. The general
rules for determining the elements of A~! in the 3 x 3 case may now be stated.
Let M;,; be the determinant of the 2 x 2 submatrix obtained when row i and
column j are deleted from A. M,, is termed a minor. Further define
Pe Nery
C,, = ( 1) M;;
C,; denotes a cofactor and is simply a signed minor. Thus the sign of M,; does not
change if i + / is an even number and does change if that sum is odd.
The rules then become as follows:
A5n G93).
ir ~ |432 agi 447433 ~ 473439
a a
Ces = i ast = — (4,43) — 4)243;)
Example 4-9
1 Sa
A Ne lite? eee
Died ad
Replacing each element by its minor gives the matrix
Fi 4 E | E a
ans m1 5 2 4
E ‘| | ‘ : 3 -|-1 = z
ic 5d OMS hia apeeg Ns he
erst lever plees
Jami Lot fe
Signing the minors gives the matrix of cofactors as
On=3 0
ec 2
— eal
Transposing gives the adjugate matrix
6 Ie 5)
=O eteS 3
0 ot onal
Expressing the determinant of A in terms of the elements in the first row gives
JA] = 4 Cy, + ayCyp + 4)3C\3 = 1(6) + 3(—3) + 4(0) = —3
Thus the inverse matrix is
ELEMENTS OF MATRIX ALGEBRA 127
For the nth-order case the rules for obtaining A~! are essentially those
already stated for the third-order case. The determinant is defined as
Properties of Determinants
The following properties are stated for the determinants of nth-order matrices,
but they will often be illustrated for the 2 x 2 case. To economize on space,
proofs will not always be given.
: |e
b= ad ZS be Jal =|,
Ge eG
The numerical values of these two terms are identical; the crucial question is the
sign. To determine the sign of the second term, the first subscripts must be put in
natural order and the number of inversions in the second subscript determined.
This gives
bi pb.4b3y 2a bs
and, compared with the corresponding term in |A|, one inversion has been
introduced or removed, so that this term (and each and every term) changes sign
in |B| as compared with |A|. If we interchange rows i and j, which are separated
by, say, r rows, reordering the b elements in any term to put the first subscripts in
natural order will involve 2r + 1 changes, where each change introduces a new
inversion in the second subscript or removes an existing inversion. This is
illustrated below, where only the first subscripts on the b’s are shown.
ee
15:4 1O;42 aaa 5-15;
r elements
Since 2r + 1 is an odd number, the sign of the term changes, and so |B| = —|A].
Property 1 then ensures that interchanging any two columns will also change the
sign of the determinant.
3. If a matrix has two or more identical rows (or columns), its determinant is zero.
a b
=ab-—ab=a
aan
From property 2, interchanging identical rows would change the sign of the
determinant. But the new matrix is identical with the old, and so its determinant
is unchanged. This gives
|A| = -|A|
so that
|A| =0
4. Expansions in terms of alien cofactors vanish. By this is meant an expression
such as
jn~in
where the elements of row j are multiplied by the cofactors of the elements of
row i.
This is exactly the expression we would obtain for the determinant of a
matrix whose rows i and / are identical. By property 3, that determinant is zero.
|B| = (a, zt AG}, )Cy iv (a; ete Nan) Go Sine? ae (4, as Aaj, )C;n
6. If the rows (or columns) of A are linearly dependent, |A| = 0, and if they are
linearly independent, |A| + 0. If the rows of
uae eb
a ie a
are linearly dependent, there exist nonzero scalars \,, \, such that
A\,a+A,c=0
\,5+A,d=0
Thus
C= A,
ie and d=-— AAe
Example 4-10
reserskveaye
A=]1 2 ] 1
2 4 -6 -—10
The rank must be at least 2, since although
ee 2aRe
f 3|=0
there are plenty of nonvanishing second-order determinants that can be
formed from the elements of A. For example,
; el
1 3
=-12 Z
4
]
_y|7 -4
and so on. Notice that these are the determinants of second-order sub-
matrices obtained by deleting any one row and any two columns from A.
There are four possible third-order determinants to evaluate. Deleting the
fourth column,
2 3 0 0 2 i 2
] ie el 2 ] = 2) |= 0
2 a>) —6 Z A os 6
The first step in this evaluation has been to subtract row 2 from row 1, which
by property 5 does not alter the value of the determinant. This gives an
expansion, using Eq. (4-63), in terms of the first row, which now contains just
a single term. To evaluate
1 3 4
1 ] 1
7a Comet ()
we might subtract row 1 from row 2 and we also subtract twice row 1 from
row 3 to get
1 3 4
0 outiedks = lias Rl=°
0 -12 -18
The two other third-order determinants may similarly be seen to be ZeTO, SO
that p(A) = 2. Alternatively, we might have spotted that
not hold between the three columns. The linear dependence between the three
columns is expressed by
2-column | — 1 - column2 = 0
or )
2- column| — 1- column2 + 0- column3 = 0
The important point is that a set of vectors is linearly dependent even if some
(but not all) of the coefficients in the linear combination are zero.
Zee 0 0
ee ay, Ay O 0
Cae C323 0
Gn) a2 an3 nn
or upper triangular, as in
a, 0 0
|A| = 4);|432 433 0
an2 an3 ann
Expanding the new determinant by its first row and repeating the process n times
gives
|A| = 19970 °° * Any
10 0
|AJ =|0 1 0;=1
eceree '
These properties follow directly from the definition of the determinant in Eq.
(4-62), where it is seen that each term in the expansion is the product of n
elements, one and only one from each row and column of the matrix.
9. The determinant of the product of two square matrices is the product of the
determinants.
|AB| = |A| + |B|
This rule is only of interest when A and B are both nonsingular. If either is
singular, AB is singular and both sides of the equation are zero. If A is
nonsingular, repeated applications of property 5, that is, additions of multiples of
rows and columns, can produce a diagonal matrix D, such that |D| = [A].
Example 4-11
Die OT 2
4 with |D| = —2 =
|A|
If these steps are performed on the matrix AB, the result is a matrix DB with
|AB| = |DB| by property 5. This statement in general requires that only row
operations have been performed on A to obtain the diagonal matrix D. This
is always possible. The first step in the example is equivalent to premultiply-
ing A by
ELEMENTS OF MATRIX ALGEBRA 133
E, = L. 1
; 6 1|
The sequence of operations is then described by premultiplication by a single
matrix
Saleen |
F-%F,-| 72
The simplest proof is to multiply AB by the suggested inverse and see that the
unit matrix results, since we already know that the inverse matrix is unique:
ABB 'A~! = AJIA~! =
and similarly,
B-'A~'AB =I
This technique is sometimes useful in deriving inverse matrices, namely, guess at a
plausible inverse and check by multiplication to see whether it works. The above
result extends readily to products of three or more matrices. Thus
(ABC) | = C7'B-'A7!
The warning must again be inserted that this result only holds when the
constituent matrices are nonsingular. Students occasionally produce “‘ miraculous”
proofs by applying this theorem to rectangular matrices.
2(At et =A
that is, taking the inverse of the inverse reproduces the original matrix.
(A“!)(AT')' =]
Premultiplying by A gives the required result.
3. (A) 1 =A)
that is, the inverse of the transpose equals the transpose of the inverse.
134 ECONOMETRIC METHODS
We have
AAS I]
Transposing,
(A-')'A’ =I]
Postmultiplying by (A’)~!,
(A“TYA(A)' = (A)
Thus
(Anty = (a)
]
4. |A|A7'| | = ——
iA]
that is, the determinant of A“' is the reciprocal of the determinant of A.
AA t=]
gives
JAI? [AT = 1
5. The inverse of an upper (lower) triangular matrix is also an upper (lower)
triangular matrix.
a, 90 0
A= la, ay 0
43, 432 33
By inspection it is seen that three cofactors are zero, namely,
&: 0 0 yi lit-Owee 10 a, O
Cy re Az, 33)” C3 = as, OP Cy. rary an 0
Thus
Gowen 0
A=! = |A| Ci Cy 0
C3 G3 Cay
ne A 11
| A |
Ax, Ax
ELEMENTS OF MATRIX ALGEBRA 135
—B yA Ai By
where
These formulas are frequently used. The first form, Eq. (4-65), is the simpler
if we are interested in an expression that involves just the first row of the inverse.
Conversely, Eq. (4-66) is the simpler for expressions involving the second row.
The derivation of the formulas is straightforward but tedious. Let
Aahes =|
B,, By
where the B,,; submatrices have the same dimensions as the corresponding A,;
submatrices. Postmultiplying A by A”! gives the matrix equations
A,B, + A,B, =1
A, By. + A,B, =0
A,B), + A.B), = 0
A,B, + Ax
B =I
where the unit matrix has been partitioned conformably with A. The third
equation in this set gives
B,, = —Ax A,B, (4-67)
Substituting this in the first and solving for B,, gives
F. =1
B,, = (Ai, a A, Ay Ad) (4-68)
A similar treatment of the second and fourth equations yields
Bi) = —Aj'A,) By (4-69)
=1
and B,, = (Ay 7 A») Ai Ai) (4-70)
These four expressions are seen to constitute, respectively, the first and second
columns in the two alternative formulations of A7'.
To derive the remaining columns in Eqs. (4-65) and (4-66) we multiply out
A~'A = I to obtain
B,,A\, + By,A2, = 1
B,,A,. + B,,A. = 0
B,,A,, + B.A, = 0
B,,A,, + B,A =I
136 ECONOMETRIC METHODS
a,, 0 0
A=|090 ay 0
jot pee hee
1
— 0 0
a);
‘A 1
Aol= 0 = 0
A)
0 0
a
0 I im |Ai,| (4-78)
An Ala tan
A A
(4-80)
This follows from the same argument used to establish Eq. (4-78). We can now
find the determinant of a block-triangular matrix.
a A1
| Aa
Az, Ax)
where A,, and A,, are square and nonsingular. Define
B, =
[Av ae
AGt and B, =
I
be
|
|! I | : ee I
Then
BAR|!A,, — Ap 12442249]
Az JA 0
Cramer’s Rule
This inordinately long section on the solution of equations may be rounded off by
the derivation of Cramer’s rule for the solution of a set of n nonhomogeneous
equations in n unknowns. The set of equations may be written
Ax =b (4-84)
where, by assumption, A is a square known matrix of order n and nonsingular, x
is a vector of n unknowns, and b is a known n-element vector. There is an
unfortunate clash of notation between conventions in algebra and conventions in
Statistics. The normal equations for the least-squares vector are
(X’X)b = X’y
Here (X’X) is a known matrix and X’y a known vector, each depending on the
empirical data in a given problem, and b denotes a vector of unknown coefficients.
It is too late in the day to resolve this conflict; the student must maintain
sufficient intellectual agility to interpret the symbols according to the context.
Returning to Eq. (4-84), the solution vector is written
x=A'b
Substitution for A~! from Eq. (4-64) gives
Cy Cy Cu |} 5
pya | Crp Cy Cro || 2.
TA AM et eee eee
Cin Gy Cin || On
ELEMENTS OF MATRIX ALGEBRA 139
Thus
1
xy = ja 2G WD5C3j oh ve C1)
Similar results hold for each element in x. The ith element is thus the ratio of two
determinants, the denominator being the determinant of A and the numerator the
determinant of the matrix obtained from A by replacing the ith column of A by b
and leaving the other n — 1 columns unchanged.
giving
X3 = ]
El ek 11 8 =a)
° 23 14 -—10
Thus
2 4a xy 15
OT 2: Ieee oleae
0 0 0.5 }| *3 0.5
This gives an upper triangular system, which is solved for the x’s by back
substitution. The third equation gives directly
x3, =1
The second equation
= 9X5 4S 22K = DNS or Sky a5
then gives
x, =3
ELEMENTS OF MATRIX ALGEBRA 141
gives
x, =2
In the elimination method the inverse A~' is never calculated at all. The
calculations are fast and simple compared with the first two methods, but it
does not shed light on the theoretical properties of the inverse.
The previous section was concerned with the solution of the set of equations
Ax = Db (4-84)
This section is concerned with solutions of
Ax = Ax (4-85)
where A is a known square matrix of order n, x is an unknown n-element column
vector, and A is an unknown scalar. This problem will arise in a number of places
later in the book. It is known as the eigenvalue problem. In contrast with Eq.
(4-84) there are now two unknowns, a vector and a scalar. Solutions will come in
pairs; to each A there will correspond an x vector. The A’s are known as
eigenvalues, latent roots, or characteristic roots and the x’s as eigenvectors, latent
vectors, or characteristic vectors.
For n = 2, Eq. (4-85), written out in full, becomes
(a,, —A)x, + ax, = 0
yx, + (ay — A)x, = 0
which may be put back in matrix form as
(A —XI)x =0 (4-86)
Equation (4-86) is equivalent to Eq. (4-85) for any n. If the matrix A — AI is
nonsingular, the only solution to Eq. (4-86) is the trivial x = 0. Thus for a
nontrivial solution to exist, the matrix must be singular or, in other words, have a
zero determinant. This condition gives
|A —AI| =0 (4-87)
which is known as the characteristic equation for the matrix A. This gives a
polynomial equation in the unknown A. Each root or eigenvalue A; may be
substituted back into Eq. (4-86) and the corresponding eigenvector x, obtained.
For the 2 X 2 case it is easily seen that the characteristic equation is
= (ay, + ay )A + (441422 — 442431) = 0 (4-88)
with roots
142 ECONOMETRIC METHODS
Example 4-13
4 x
eel
Thus
(a-an= [45% a
and the characteristic equation is
’-— 5A =0
with roots
A, =5 and A, =0
For A, = 5, substitution in Eq. (4-86) gives
9) =A ce =0=> x, = 2X5
Thus one element in the eigenvector is arbitrary, and so ifx satisfies Eq.
(4-86) for some A, then so does cx, where c is an arbitrary constant.
It is
conventional to normalize the vector by setting its length at unity, that
is,
making
Xp xe =]
which, with x, = 2x,, gives
2
xX, =
v5 corresponding toA, = 5
ee
5
ELEMENTS OF MATRIX ALGEBRA 143
Mo
v5
Xo =
2.
v5
is the eigenvector corresponding to A, = 0.
It is seen that the eigenvectors are orthogonal, xx, = 0. If we assemble
the eigenvectors in a matrix X,
ena Agee
v5 v5
ca ae as
KS — =
v5 v5
and then form X’X, we obtain the result that
and Ay = ux + Ay
Premultiplying the first equation by y’ and the second by x’ gives
y’Ax = dx’y — py’y
x’Ay = px’x + Ax’y
When A is symmetric, y‘Ax = x’Ay (a scalar equals its transpose). Subtracting the
first equation from the second then yields
0 = w(x’x + y’y)
Since the eigenvectors must be nontrivial, x’x > 0 or y’y > 0 (or both), so
pw =0
that is, there cannot be a complex eigenvalue. Real eigenvalues in turn generate
real eigenvectors, that is, y = 0.
3.* If an eigenvalue has multiplicity k (that is, is repeated k times), there will be
k orthogonal vectors corresponding to this root.+
ee)
A=
OZ 0
OF Uae
The characteristic equation is
(1-A)(2-A) =0
with roots
A, =1 with multiplicity 2
0 Ob ts
The multiple root gives
Ord. OLX; 1 0
OF One, =0=>x, =~x,|9| +x, 0
Obed Oi 01h ee 0 ]
The root with multiplicity 2 thus yields two orthogonal eigenvectors e, and e3.
4.* The nth-order symmetric matrix A has eigenvalues X,, X4,..., Aq, possibly not
all distinct.} Properties 2 and 3 then guarantee a set of n orthogonal eigenvec-
tors X,,X,-..,; X,; Such that
x)x, = 0 pe Tet ea een (4-93)
that is, although X was constructed as a matrix with orthogonal columns, its row
vectors are also orthogonal. Thus an orthogonal matrix is defined by
X’X’ = XX = I (4-99)
Example 4-14
ie 0
A=| 202/72
Oyo Gl
The characteristic equation is then
[aera 2 0
2° 2d 2 |=0
0 v2 .1-A
thatis,
(1 =) )(1 oA) ae ero
with roots
A, =1 A,=-1 A, =4
st
Osriez Oe ex 0 V3
i (A ~T)x=}2) 1 -y2 ||] eo 0 xa] ae
2 ollx] lo 2v3
=
i D> 0 3 0 oe
Ag= ols. CA Dx= 2 > 3. “V2 35 |= 0 =>X,= Vio
Ce ey eiies 0
v5
wee
= 2 O ll x, 0
A; =4: (A — 4I)x = 2 ey | 0|}>x,= V5
0 y2 -3]1% 0 2
v5
ae ees
v1l0.— V5
Thus X= balay
2 iy cians
vl0.— v5
cin eaion
Vee uany15)
ELEMENTS OF MATRIX ALGEBRA 147
The reader can check numerically that the rows of x all have unit length and
are pairwise orthogonal (as, of course, are the columns).
Premultiplying by x’,
x,Ax, = A,x'x, =A,6,, using Eq. (4-95) (4-101)
Equation (4-101) displays the i, jth element in X’AX, and collecting for all i, 7
gives Eq. (4-100). An alternative proof illustrates a useful exercise in matrix
manipulation.
= XA
Premultiplying by X’ then gives Eq. (4-100). We should not conclude from this
result that only symmetric matrices can be diagonalized. If for any matrix A there
are n linearly independent eigenvectors and we arrange them as the columns of a
matrix X, then
X-'AX=A (4-102)
The contrast with Eq. (4-100) is that the columns of X are not necessarily of unit
length, nor are they necessarily orthogonal.
6. The sum of the eigenvalues is equal to the sum of the diagonal elements (trace)
of A.
This property is true for any matrix, but the proof is particularly simple for
symmetric matrices. Denote the trace of a (square) matrix A by
tr(A) = a,, + a,+-:-+4,,
For two matrices, A of order m X n and B of order n X m,
tr(AB) = tr(BA) (4-103)
148 ECONOMETRIC METHODS
Thus
Thus
This result is again true for any matrix, but the proof is very simple for
symmetric matrices. We note first that when X is an orthogonal matrix,
Sr ece (4-106)
for
XX = I= |X’| + |X| =1
but |X| = |X’|
Thus |X| = +1
Returning again to
X’AX =A
[X"| + |A] + |X] = |A|
Thus [A] =A,A,-=A, (4-107)
ELEMENTS OF MATRIX ALGEBRA 149
Ax = Ax
Premultiplying by A,
A’x = \Ax = 0’x
which establishes the result. We may note, in passing, a very useful application of
this result in analyzing the stability of dynamic systems. Suppose y, denotes a
vector of the values taken by a number of economic variables in time period f,
and suppose y, can be expressed in terms of the previous values by the system of
equations
Y= AY (4-109)
Even if the original specification of the system involves lags of more than one
period, an appropriate definition of new variables can produce a derived system
of the type of Eq. (4-109).+ Successive substitution in Eq. (4-109) gives
y,= A’Yo
where y, denotes initial values of the variables. Provided A has a linearly
independent set of eigenvectors,
xX'AX=A
or A=XAX'
Thus AP = XAXT'XAX b= XNV’-X"!
So At = XA'X™!
and the elements of y, are seen to be linear combinations of the tth powers of the
eigenvalues of A. Thus if the system is to be stable, we need
|A,| < 1, b= licen
10. The eigenvalues of A~' are the reciprocals of the eigenvalues of A, but the
eigenvectors of both matrices are the same.
Ax
= Ax
+ See G. Chow, Analysis and Control of Dynamic Economic Systems, Wiley, New York, 1975, pp.
21-35.
150 ECONOMETRIC METHODS
Premultiply by A7!,
or
By property 9
Cx x
But when Ais idempotent,
Atx = Ax =x
Thus
A(A — 1)x = 0
and since any eigenvector x is not the null vector,
A=0 or A= 1
We have already introduced quadratic forms briefly in Sec. 4-2 and have
seen that
there is no loss of generality in considering only symmetric matrices.
For a 2 x 2
symmetric matrix A and a two-element column vector x, the quadratic form
is
MAN Giri dy i a
For a third-order matrix
WAX = 4),x7 + 2a,x,x, + 2413X\X3
La Xe ay wax
Ste A33X3 2
ELEMENTS OF MATRIX ALGEBRA 151
ON Geeks ae 2dn xe
fey nee
Definitions
If x’Ax > 0 for all x * 0, the quadratic form is said to be positive definite and A is
said to be a positive definite matrix.
If x’Ax > 0 for all x + 0, the form and matrix are positive semidefinite.
Reversing the above inequality signs defines negative definite and negative
semidefinite matrices, respectively. If a form is positive for some x vectors and
negative for others, it is said to be indefinite.
It is important to have tests for positive definite matrices.
To prove the necessary condition assume x’Ax > 0. For any eigenvalue A,
Ax; = AX;
Premultiplying by x’, gives
x’ Ax, = A.x)x;
baa ee
=A;
Since x’Ax > 0 holds for any x + 0, it holds for each eigenvector, and so A, > 0
for all i. To prove sufficiency we assume all A; > 0 and show that x’Ax > 0. Since
a symmetric matrix has a full set of n orthogonal eigenvectors x,,X5,..., X,, any
nonnull vector x may be expressed as a linear combination of the eigenvectors
KUMI Cok ee a COX
Thus Ax = c, Ax, + c,Ax, + -:: +c, Ax,
= CyAGX, td C>A5X5 ct nae. stg Ch
x’Ax (5X4 + CyX > ate ol te ex) (eiAGx, ta €A5X>5 + Pthe sta CoN Xe
factored into
A= ANZA?
Ar
where A? =
ro
ie
Substitution in Eq. (4-112) gives
A =XA'2A12x' = (XA'/2)(XA!/2)/
3. If Ais n X m with rankm < n, then A’A is positive definite and AA’ is positive
semidefinite.
4. If A is n X m with rankk < min(m,n), then A’A and AA’ are each positive
semidefinite.
5. If A and B are positive definite matrices and A — B is also positive definite,
then B-' — A~' is positive definite.+
toady ee ee ae elt
Bah
and the dx; indicate arbitrary changes in the x;. For small dx, the first-order
differential gives the approximate value of the resultant change in y. Denoting the
vector of partial derivatives by f and the vector of differentials by dx,
fi dx
ie dx
f= ay. = i dx = ;
ox ; ‘
a dx,
dy = f’ dx (4-114)
If y has a stationary value at a point
collate
then dy = 0 for all points in the neighborhood of x*. For such points dx + 0, and
so from Eq. (4-114) the necessary condition for a stationary value is
f=0
that is, all partial derivatives are zero at the stationary point.
A stationary point may be a maximum, where the value of the function is less
at all points in the neighborhood of x*; a minimum, where the value of the
function is greater at all points in the neighborhood of x*; or a saddle point,
where the value of the function increases in some directions from x* and
diminishes in others. One may distinguish between these possibilities by means of
the second-order differential dy. The second-order differential may be found by
totally differentiating the first-order differential. It is an approximation to the
change in dy as we move away from the point x*. Clearly, for a maximum value
dy will decrease from zero to some negative value, so d*y will be negative, and
conversely for a minimum value d7y will be positive. For a saddle point d*y will
be positive for some dx and negative for other dx.} Totally differentiating Eq.
} It is possible, but extremely rare, to have d*y = 0 for some dx. Such complexities are ignored
here.
ELEMENTS OF MATRIX ALGEBRA 155
(4-113) gives
0
d*y “ays Liat hak, re ef dx. Pax,
0 |
0
++. + a Thax yay cee fax, |an.,
eens dx?
where
gy
hij = hi = (8x, fori * j, alli,7
and dx? indicates the square of the differential dx,. The second-order differential
is thus seen to be a quadratic form in dx. The matrix of the quadratic form is the
symmetric Hessian matrix of second-order partial derivatives, which we will
denote by
Sere ee eae ue
Tigihel ss ES
and we may write
d*y = dx’F dx (4-115)
Thus d’y is positive or negative as F is positive definite or negative definite. To
summarize, the conditions for a maximum or a minimum at a point x* are as
follows:
First-order Second-order
condition condition
: 0 a7y . : ie
Maximum f=—=0 F = —~ is negative definite
dx dx?
2
Minimum f= ove 0 F= oy is positive definite
Ox ax2
Constrained Extrema
In finding stationary values of y = f(x,, X,,..-, X,) the x’s were assumed to be
independent variables. Thus we could specify n arbitrary differentials
dx,, dx,..., dx, In some problems, however, the x’s may be subject to one or
more constraints, and we have to find a maximum or minimum value of y subject
to the constraints. We will assume for the moment that the function hasa single
maximum or minimum value and state the problem formally as follows:
8, (x)
(Ca
8m (X)
and a column vector of m Lagrange multipliers,
A,
A,
Nal ee
r
Using these we define a new objective function as
Op Of, TeOe
ox dx ax (W’a(x))
9 = =0 (4-1 17)
@
aN g(x)
Oe
3x (N’B(x))
Since
we have
d(N’g(x))
Ox,
_ , 98,
Pee pases
dg, ocean
Og ROE
ve oe
98)
Ox;
; 98,
of
ax = Ox;
. a
Pela an
Bn
0g;
<f_@n=0
i (4-118)
g(x)
=0
The second equation in Eqs. (4-118) ensures that the stationary value satisfies
the constraints. To distinguish between maxima and minima, we must still
examine whether the quadratic form in Eq. (4-115) is negative definite or positive
definite, but now only for dx vectors which do not violate the constraints. Totally
differentiating the jth constraint
o (Xk aes = 0
gives
0g; 98; 98; dx,
Os aes Peg Re emat Ax
n
158 ECONOMETRIC METHODS
There is a similar condition for each constraint. Thus the dx vectors which do not
violate the constraints are given by
G'dx =0 (4-119)
In many cases the F matrix consists only of constants, and so its definiteness can
be established independently of any x values.
PROBLEMS
4-1 Expand (A + B)(A — B) and (A — B)(A + B). Are these expansions the same? If not, why not?
How many terms are in each?
4-2 Given
3 4 l 2
a=(! ; al p= (1 = | c-|-i]
= : l 2 229 4
Calculate (AB)’, B’A’, (AC)’, and C’A’.
4-3 Find all matrices B obeying the equation
Oe il LO O71
f 2 |e4 ki 0 |
4-4 Find all matrices B which commute with
oe Ome
NG [2 |
to give AB = BA.
4-5 Write down a few matrices of order 3 X 3 with numerical elements. Find first their squares and
then their cubes, checking the latter by using the two processes A(A”) and A?(A).
4-6 Prove that diagonal matrices of the same order are commutative in multiplication with each other.
4-7 Let
QO 4
J=/0 1 0
OO)
Write out in full some products JA, where A is a rectangular matrix. Describe in words the effect on A.
Do the same with products of type AJ. Find J?.
4-8 If
OY tk @
YeIlO @ fl
@) @ @
find V* and V*. Examine some products of the type VA, VA, and V’A.
4-9 Given
aoa ae eee)
IWS ||2 8 and selene
T @& 1 Or @ A
Calculate |A|, |E|, and |B|, where B = EA. Verify that |B| = |E||A|.
4-10 Show that
1 l l
a gd 2Cla Ceo Oh)(Gad) CD)
a* b2 Ce
ELEMENTS OF MATRIX ALGEBRA 159
4-11 If (x), y,) and (x3, y)) are points on the x, y plane, show that the equation
XPV al
x y ll=0
X, yn 1
© |
Al
al-
Sl-
ale
t+
o
is orthogonal, that is, that Q’ = Q-!.
4-14 If the u; are normal variables with
E(u;) =0 eer
E(ujuj)=0 ~ij7
show that E(u’Au) = o7tr(A).
4-15 Given
et)
peels
Peet —
NO
We
Compute
A = (I, — X(X’X)'X’)
Show that A is idempotent and determine its rank. Find the characteristic roots and the
associated characteristic vectors of A, and hence obtain the orthogonal matrix which diagonalizes A.
4-16 A and B are nonsingular matrices of the same order. Prove that AB and BA possess identical
characteristic roots. Show also that no such matrices can be found to satisfy the equation
AB — BA=I
5 -6 -6
A= —] 4 2
m4
4-18 Examine the following quadratic forms for positive definiteness:
WIN
Wie
Wily wiry
wl
BIN Bl—
wiry
WIN
_ Of (X)
Bs ax
1S a matrix of the same order as X such that
CHAPTER
FIVE
THE k-VARIABLE LINEAR MODEL
Let x denote a vector of random variables X,, X,,..., X,,. Each variable has an
expected value
p= EM) ie pin
Collecting these expected values in a vector p, gives
E(X;) by
E( xX, 2
eH : is : cen
BCI lhe
161
162 ECONOMETRIC METHODS
The application of the operator E to the vector x means that £ is applied to each
element of x. The variance of X;, by definition, is
(X,
— 4)
E(x ~ w)(x- w= E ee [2% = my)O% a) OG —
(4, - 1)
+ Alternative expressions for the variance-covariance matrix are cov(x) and V(x).
THE k-VARIABLE LINEAR MODEL 163
involving random variables. Since Yis a scalar random variable, E (Y7) > 0. Thus
cae > 0
and 2 is positive semidefinite. But
E(Y?)=0=Y=0
which, from Eq. (5-3), means that the X deviations (X, — p,), (Xo = pS) ne ee
— p,,) are linearly dependent. Thus
2 is positive definite, provided no linear dependence exists among the X’s.
The n random variables will have some multivariate probability density function
(pdf) written
DIS) = Pl BOG Xe)
which is simply some formula or rule giving the likelihood of various combina-
tions of X values. The most important multivariate pdf is the multivariate normal.
The univariate normal distribution is specified once its mean p and its variance o”
are given. The multivariate normal is similarly specified in terms of its mean
vector p and its variance matrix 2. The formula is
p(x) = eatin
1
Qmie
xp] — 5(x— wy E-"(x - w) (5-4)
A compact shorthand statement of Eq. (5-4) is
x ~ N(p, 2)
to be read, “the variables in x are distributed according to the multivariate
normal law with mean vector p and variance matrix 2.” When n = 1, = = o? and
Eq. (5-4) becomes
1
P(X) = v270 exp|— 55 x= 4)
which is the familiar univariate normal density. When n = 2, if we use p to denote
the correlation between X, and X,, the variance matrix becomes
0;2 po,0,
Sa with |2| = 0707(1 — p’)
Dory Bo:
Notice that |=| > 0 unless p* = 1, so that the variance matrix is positive definite
unless there is perfect linear correlation between the two variables, in agreement
with the general result above. Substitution in Eq. (5-4) gives
(X,, %) 1 o| 1 (+ =H)
P , a eer ay. Ntk mame ah)
te 270,051 — p* 2a Ge) Oily
An especially important case of Eq. (5-4) occurs when all the X’s have the
same variance o” and are all pairwise uncorrelated.+ Then
= =o’l
with
[2] = [21111229]
Making these substitutions in Eq. (5-4) gives
+ The assumption of a common variance is only made for simplicity. All that is required for the
result is that the 2 matrix be diagonal.
THE k-VARIABLE LINEAR MODEL 165
that is,
x’x ~ x?(n)
for x*(n) is the sum of the squares of n independent standardized normal
variables.
Suppose now that
x ~ N(0, 671) (537)
The variables are still independent and have zero means, but each X has to be
divided by o to yield a variable with unit variance. Thus
XX
St 7 aegis
nae 2)
age)
that is,
eh ~ x7(n) (5-8)
or x’(o71)_
‘x ~ x2(n) (5-9)
Equation (5-9) shows explicitly that the matrix of the quadratic form is the
inverse of the variance matrix.
Suppose now that
x ~ (0,2) (5-10)
where & is a positive definite matrix. The equivalent expression to Eq. (5-9) would
now be
x= 'x ~ x*(n) (5-11)
This result does in fact hold, but the proof is no longer direct since the X
variables are no longer statistically independent. The trick is to transform X’s
into Y’s, which will be independent standardized normal variables. Since 2 is
positive definite, by Eq. (4-111) there exists a nonsingular matrix P such that
= = PP’
166 ECONOMETRIC METHODS
which gives
2) = (Po) Pa and PS (Be) (5-12)
Define an n-element y vector as
y=P 'x
The Y variables are multivariate normal since they are linear combinations of the
XS,
E(y) = P“'E(x) =P '0=0
and var(y) = E{P” 'xx’(P7')’}
=iP => (Bea):
on
from Eq. (5-12). Thus the Y’s are standardized normal variables and
VV AT)
But
yy = x(P))
Ps y= x7 ix
from Eq. (5-12). So
xD 'x ~ x?(n)
which is the result anticipated in Eq. (5-11).
Assume again
x ~ N(0,1)
and now consider the quadratic form x’Ax where A is idempotent with rank
r <n. If we denote the matrix of eigenvectors of A by Q, then
]
]Pst
r terms
Q’AQ=A= 1 (5-13)
0
me n — r terms
0
where A will have r units and n — r zeros on the main diagonal. Define
y = Ox
Thus
x = Qy
since Q is orthogonal. Then
E(y) =0
THE k-VARIABLE LINEAR MODEL 167
it need not be square or symmetric. If the variables in Ax and Lx are to have zero
covariances, we require
E{Axx’L’} = o*AL’ = 0
or equivalently
LA =0 (5-15)
The first basic assumption of the model is that the vector of sample observations
on Y may be expressed as a linear combination of the sample observations on the
explanatory X variables plus a disturbance vector, that is,
Y B, u,
ie | | | B, u,
y= X=/]X, xX, X; 5 = u=
| | |
ne B, Uu,
The central problem is to obtain an estimate of the unknown 6 vector. To make
any progress with this we need to make some further assumptions about how the
observations on Y have been generated.
+ An outline of the various reasons for the introduction of the disturbance term has already been
given in Sec. | of Chap. 2.
THE K-VARIABLE LINEAR MODEL 169
specific set of numbers for family income, size, and composition. Let s, denote a
row vector consisting of these numbers. Then
E(Y,)= s,B
is the average, or expected, level of travel expenditure for this type of family.
However, if we observe the actual travel expenditure of a family with these
characteristics, it may be greater than the expected level, and the expenditure of
another family with the same characteristics may well be less than the expected
value. Or if we observe the travel expenditures of the same family in different
periods of time, these may be expected to fluctuate around the mean value.
However, if the theorist has done a good job in specifying all the significant
explanatory variables to be included in X, it is reasonable to assume that both
positive and negative discrepancies from the expected value will occur and that,
on balance, they will average out at zero, that is,
E(u,) =0
Similar considerations apply to each row of X, and so we have
E(u,) 0
E(u,) 0
E(u)= | =|.
E(u,) 0
3. E(uu’) = o7I
4. p(X) =k
This assumption states that the explanatory variables do not form a linearly
dependent set. For example, if we had just two explanatory variables, X, and X,,
and this assumption was not fulfilled, there would then exist an exact relationship
X3 =c¢, + ¢,X, (5-18)
which, combined with the hypothesized
Y=8B, + BX,
+ BX, + u (5-19)
gives
Y = (B, + B3c,) + (B, + Bycy) X_ + u (5-20)
The constants c, and c, can be determined exactly, and we can estimate the
intercept and slope of Eq. (5-20), but there is no way to obtain estimates of the
three 8 parameters.
5. X is a nonstochastic matrix.
L
This assumption at first sight seems incongruous. It means that if we take
another sample of n observations, the X matrix of explanatory variables remains
unchanged, the only source of variation then being in the u vector and hence in
the y vector. However, the social sciences are notoriously difficult for being
observational and nonexperimental so that in general the X variables are not
subject to experimental control by the social scientist. There are three main points
to be made about this assumption. First of all, in spite of the remarks above, there
are cases where the X data can be controlled. In a cross-section survey, the sample
design may call for the inclusion of certain numbers of families with specific
characteristics, and sampling is continued until these specifications are met.
Second, even if it is not in fact feasible to control the X data precisely, it is still
useful to be able to make statistical inferences which are conditional on the X
values actually present in the sample. In this light it is very much an assumption
of convenience in that it simplifies dramatically the derivation of several basic
statistical results. Third, once these simple results have been derived, it is possible
to weaken the assumption to allow the X variables to be stochastic, but distributed
THE k-VARIABLE LINEAR MODEL 171
independently of the disturbance term, and then see what modifications of the
earlier results are required.
The most frequently used estimating technique for the model outlined in Sec. 5-2
is least squares. The hypothesized model is
y= Ap eu (5-22)
Let b, denote any arbitrary k-element vector. This in turn serves to define a
vector of errors, or residuals,
e, = y — Xb, (5-23)
The least-squares principle for choosing by, is to minimize the sum of the squared
residuals e,e,. From Eq. (5-23)
The necessary condition for a stationary point requires that we set Eq. (5-24)
equal to the 0 vector. Denoting the resultant OLS solution for b, simply by b gives
(X’X)b = X’y (5-25)
These are referred to as the OLS normal equations. Assumption 4 ensures that X’X
is nonsingular. Thus an equivalent expression for b is
b = (X’X) 'X’y (5-26)
The vector of OLS residuals is likewise denoted by e, where
e=y— Xb (5-27)
Using this expression to substitute for y in Eq. (5-25) gives
(X’X)b = (X’X)b + X’e
xje 0
xe 0
Thus Xe=/] |=|].|/=0 (5-28)
172 ECONOMETRIC METHODS
This is a fundamental OLS result. The first element in this equation gives
e=0
that is, the residuals from the OLS regression always have zero mean, provided
that the equation contains a constant term. The remaining elements in Eq. (5-28)
state that the residual has zero sample correlation with each X variable.
To establish that the stationary point does indeed correspond to a minimum
of the sum of squares, differentiate Eq. (5-24) once again with respect to b to
obtain
d*(e,ex)
ab,
—_———. (X’X)
= 2(X’X (5-29)
5-29
From Sec. 4-7 this gives a minimum provided X’X is positive definite. To establish
this, let d be any nonnull k-element vector, and consequently define an n-element
vector ¢ as
c = Xd (5-30)
The assumption that X has full column rank ensures that ¢ is nonnull; otherwise
Eq. (5-30) would express a linear dependence between the columns of X. Thus
c’c = d’X’Xd > 0
and X’X is positive definite.
Returning to Eq. (5-26),
h = (XX) -X’y
and substituting
y=Xfp+u
gives
Thus
var(b) = E{(X’X) 'X’uu’X(X’X) '}
(X’X) 'X’o7IX(X’X) ' from assumptions 3 and 4
o?(X’X) !
(5-33)
since I may be suppressed at will and the scalar o* moved in front or behind
matrices. The elements on the main diagonal of Eq. (5-33) give the sampling
variances of the corresponding elements of b, and the off-diagonal terms give the
sampling covariances. The most important result in least-squares theory is that no
other linear unbiased estimator can have smaller sampling variances than those of
the OLS estimator in Eq. (5-33). OLS estimators are thus said to be best linear
unbiased estimators (b.1.u.e.), that is, to have minimum variance within the class of
linear unbiased estimators. This result is known as the Gauss-Markov theorem.
The following proof is somewhat roundabout, but it has the advantage of
establishing a further important result at the same time. Let ¢ denote an arbitrary
k-element column vector of known constants and define a scalar quantity pw as
w= cB (5-34)
If we choosec’=[0 1 O --- OJ, then yu = £,. Thus we can use Eq. (5-34) to
pick out any single element in B. Or if we choose
coma Kop eas el, Mya
then
w= E(Y, 41)
which is the expected value of the dependent variable Y in period n + 1,
conditional on the X values in that period.
We wish to consider the class of linear unbiased estimators of w. Thus define
a scalar m which will serve as a linear estimator of uw, such that
m=ay=aXB+ a'u (5-35)
where a is some n-element column vector. The definition ensures linearity. To
ensure unbiasedness we have
E(m) a’XB + a’E(u)
a’XB
= cB
only if
aX=c’ (5-36)
From Eggs. (5-35) and (5-36),
var(m) = E{a‘uu’a}
=o0°a’/a
which derivation uses the fact that since a’u is a scalar, its square can be written
174 ECONOMETRIC METHODS
as the product of its transpose and itself. The problem is thus to choose a to
minimize a’a subject to the k side conditions a’X = ce’. Define
= a’a — 2 (X’a — c) (5-37)
Here X is a column vector of k Lagrange multipliers, and the side conditions
(5-36) have been transposed to make the multiplications in Eq. (5-37) conform-
able. Differentiating
and —dg
Oy =
2(X’a — c) =0
, — => (5-39)
5-39
a = X\ = X(X’X) 'c
and so the desired minimum variance linear unbiased estimator of c’B is
m=aly
= ¢(X’X) “X’¥y
= c’b (5-40)
that is, the unknown B is replaced by the OLS b. It follows directly that +
1 , Lovack 1
A= tell Ped 1
OU Neer coer Raewean
1 real 1
fy =|[¥,; Y, *= “¥s}then
—iy=Y
me Ame
and Ay =y-iY=
Y,-Y
Thus premultiplying any column vector of observations by A produces a vector
showing those observations in deviation form. Two special cases are
Ai = 0 (5-42)
or, more generally, premultiplying any vector of identical elements by A gives the
zero vector. Second,
Ae =e (5-43)
for the residuals have zero mean, and are thus already in deviation form. It is
easily verified that the A matrix is symmetric idempotent.
The OLS estimator b and residual vector e are connected by
y=Xbt+e (5-44)
If we partition the X matrix as
X=[x, X,]
where x,(= i) is the usual column of units and X, the n xX (k — 1) matrix of
observations on the variables X,, X,,..., X,, we can rewrite Eq. (5-44) as
y =x,b,+ X,b, +e (5-45)
where b’ = [b, 4] indicates a conformable partitioning of the b vector into the
intercept b, and the subvector b, of slope coefficients. Premultiplying Eq. (5-45)
by A gives
Ay = AX,b, + e
using Eqs. (5-42) and (5-43). Premultiplying this by X’, yields
X’, Ay = X’,AX,b, (5-46)
176 ECONOMETRIC METHODS
for X5e = 0 from Eq. (5-28). Finally, using the symmetric idempotency of A
means that Eq. (5-46) is equivalent to
by
Yi=3| lax X,]| 2
by
or
b, = Y—b,X,
— b,X, — ++: — b,
X, (5-48)
The sum of squared deviations in the dependent variable, denoted by TSS, is
TSS = y’Ay
This may be decomposed into an explained sum of squares (ESS) and a residual
sum of squares (RSS) in the manner of Chaps. 2 and 3. Return to
Ay = AX,b, +e
Transposing and multiplying,
y’Ay =b,X,AX,b, + e’e (5-49)
(TSS) (ESS) (RSS)
since the cross-product term vanishes in view of X’e = 0. The multiple correlation
coefficient R, »;..., for the k-variable case may then be defined in a number of
alternative ways. The basic definition is
ESS e’e
ick = ass TVS Yay oy
In view of Eq. (5-49) this is equivalent to
Ro biee
X5,AX,b, biX5A
ey (5-51)
y’Ay y’Ay
where the second expression follows from Eq. (5-46). Alternatively, we may start
with the complete OLS regression
y= Xb+e
THE k-VARIABLE LINEAR MODEL 177
Example 5-1 To help fix some of these concepts, here is a brief numerical
example. The numbers have been kept artificially simple so as not to obscure
the nature of the operations with cumbersome arithmetic. The sample data
are
3 alc ee)
] Pr 4
Vo =a1e8 and Kee lS 26
3 [S24
5 bea 6
where we have already inserted a column of units in the first column of X.
From these data we readily compute
Sacro 20
OX peo) en and X’y =| 76
PS NE eAVAD 109
The normal equations of Eq. (5-25) are then
Selon 25 iD, 20
15.558) tbs) |e= ar
25. Sl “12976; 109
Rather than invert (X’X) we will solve these equations by the elimination
method. In the first step subtract three times the first row from the second
and five times the first row from the third. This gives the revised system
Selene onl leas 20
Oy s10e) koyl Rb arn G
O26 oc tlies 9
Next subtract six-tenths of row 2 from row 3 to get
5s 15) 2954 4) be 20
0 10. 46.8 |\tbal=s lado
Cig 500 4a ee ~0.6
The third equation gives 0.46, = —0.6, that is,
gives
b, = 2.5
fs alle]=[
The observant reader will notice that these are the second and third equations
obtained in the first step of the elimination method above.+
Thus the solutions for b, and b, will coincide with those already ob-
tained. Likewise, b, will be the same as before, for the final equation in the
back substitution above is readily seen to be
Thus the elimination process applied to (X’X)b = X’y is, in fact, equivalent to
transforming the data into deviation form and proceeding in two-step fash-
ion.
To calculate R* we note from the Ay vector that
TSS = y’Ay = 28
so that the regression has accounted for almost 95 percent of the variance of
Y. As a check we may calculate the explained sum of squares from b’X’y by
subtracting the correction for the mean,
20
b’X’y = [4 2.5 -13] a = 106.5
109
nY2 = 5(4)° = 80
Thus ESS = b’'X’y — n¥? = 26.5
in agreement with the previous calculation.
Estimation of o7
Finally in this section we derive an estimator of o*, the variance of the dis-
turbance term. As the values of u are not directly observable, it seems plausible to
base an estimate of o” on the residual sum of squares e’e. The only question is
what should the divisor be, and this can be settled by requiring the estimator to
be unbiased. We have
e=y — Xb
=y — X(X’X) 'Xy
= [1- x(x’x)'x’ly
= My (5-54)
where
Taking expectations
E(e’e) = E(u’Mu)
= E({tr(u/Mu)} since u’Mu is a scalar
= E{tr(Muu’)} —_ from Eq. (4-15)
=o’ trM by assumption 3
From Eq. (5-55)
tr(M) = tr(1) — tr[X(X’x)~'x’]
tr(1) — tr[(X’X)~'x’x]
=k
Thus if we define
= (5-57)
it follows that
E(s?) =o?
and we have found the desired unbiased estimator. The square root s is often
referred to as the standard error of the estimate, and may be regarded as the
standard deviation of the Y values about the regression plane.
So far we have not used the assumption that the u’s are multivariate normal, but
this now becomes necessary. We now make the twin assumptions
u ~ N(0, 071)
and X is nonstochastic with rank k
BSP (0,1)
Oye
From Eqs. (5-59) and (5-60),
(n—k)s?
= ) ie x?(n a k)
independently of b;. Thus we can proceed directly to form a ¢ variable, that is,
bp al)
oa, sy(n = k)
or
A= b; i B,
~t(n—k) fori=1,2,...,k (5-61)
s\a;;
Result (5-61) may be used to test an hypothesis about B, or set up a confidence
interval for B; in the usual way. However, we will not pursue the details further at
the moment as it is more efficient to develop a general set of inference procedures,
of which tests on asingle coefficient are just one particular application.
RB =r (5-62)
where R is a known matrix of order q X k with q < k, andr is a known q-element
vector. We also assume R to have full row rank, that is, that there are no linear
dependencies between the hypotheses. It is extremely important to understand the
THE k-VARIABLE LINEAR MODEL 183
of order (k — 1) X k and
0
0
r=].
0
of order (k — 1) X 1. This is equivalent to the joint hypothesis
B, 0
B; 0
By 0
that is, that the set of explanatory variables X,,X4,..., X, has no influence in
the determination of Y. This is a very important hypothesis. The test of this
hypothesis is often referred to as a test of the overall relation. Notice that the
hypothesis does not include B, = 0, since that involves the additional implica-
tion that the mean level of Y is zero. Our usual concern is whether the
hypothetical explanatory variables help to explain the variation of Y around
its mean value, but the actual level of the mean is of no particular importance.
184 ECONOMETRIC METHODS
It is thus clear that a procedure for testing the general hypothesis RB = r will
be extremely useful and powerful, since various specifications for R and r will
cover a range of questions.
To develop such a test procedure, we first of all replace the unknown B
vector in Eq. (5-62) by the OLS vector b, obtaining the vector Rb. The more this
vector departs from r, the greater is the doubt cast on the hypothesis. The
problem is to determine the sampling distribution of Rb and devise a practical
test procedure. First of all, we see directly that
E(Rb) = RB (5-63)
and
<Feeo ~ x(n k)
independently of b, and hence independently of Rb. Thus we can form an F ratio,
and the unknown o? will cancel out. The basic result is thus, if RB = ris true,
ith element
one wishes to test the hypothesis that 8, assumed some specified value,
B; = Bio
linear combination of the rows of R. Since R has full row rank, v + 0. Thus
7/R(X’X)'R’z = v(X'X) ‘v
But (X’X) is positive definite by assumption, and so (X’X)~' is positive definite, since its eigenvalues
are the reciprocals of the eigenvalues of (X’X). Thus
v'(X’X) 'v>0
and so R(X’X) 'R’ is positive definite.
186 ECONOMETRIC METHODS
F = (b, ce Bi)
———
sa,u
This is, of course, the same result as that already derived by a different route in
Eq. (5-61), since t?(n — k) = F(1,n — k).
F
__b4(X4AX,)by/(k — 1) (5-70)
e’e/(n — k)
From the decomposition of the total sum of squares in Eq. (5-49) above this is
seen to be
. ESS/(ke71}
RSS/(n = k) (5-71)
THE k-VARIABLE LINEAR MODEL 187
b,
y=[X, X,] +e=X,b,+ X,b+e (5-73)
b,
We will now show that this numerator has a very fundamental and important
interpretation in terms of sums of squares. Suppose y is regressed just on the
subset of variables in X,. Let e, denote the resultant vector of residuals. From Eq.
(5-54) we have
e,= M,y
where M,, is exactly the matrix just defined in Eq. (5-75).
Thus
1. Regress y on the variables X, which are not in the subset, and measure the
residual sum of squares e/e,.
2. Carry out the complete regression and measure the residual sum of squares
e’e. The difference efe, — e’e measures the reduction in the residual sum of
squares due to adding X, to the regression.
3. The mean square (e’e, — e’e)/s, associated with the subset, is then contrasted
with the overall mean square e’e/(n — k). If the resultant F value exceeds a
preselected critical value, the hypothesis that the variables in X, have zero
effect on Yis rejected.
The previous test for the joint significance of all the explanatory variables
may also be seen to be of the same form as Eq. (5-76). That test was based on
ESS/(k — 1)
RSH SK)
THE k-VARIABLE LINEAR MODEL 189
a WAy = ee)/(k = 1)
e’e/(n — k)
and y’Ay, which is the sum of the squared deviations of the Y values, can be
interpreted as a residual sum of squares when Yis regressed only on a vector of
units i, for replacing X, in Eq. (5-75) by i gives
Meee ait
n
This is the A matrix of Eq. (5-41), which transforms a variable into deviation
form. Thus ee, becomes y’Ay in this case.
The test of a single coefficient is merely a special case of the test of a subset.
Thus the ¢ or F test for the significance of a single coefficient may also be
interpreted in a sums of squares context. The test of
Hy: B; =a)
amounts to
3. Compute the reduction in the residual sum of squares from step | to step 2
and contrast with e’e/(n — k).
Confidence Intervals
Confidence intervals for a single B coefficient may be readily determined from the
result on the ¢ distribution in Eq. (5-61). Joint confidence regions for two or more
parameters may also be determined. From Eq. (5-65) we have
[R(b — B)]'[o?R(X’X)
'R’] '[R(b —B)]~ x2(q)
and, as usual,
,
ee
oO
= mean)
independently of b. Thus
p_Be
R(b — SOROS
B)]’|R(X’X) = 'R’| -1[R(b —
e’e/(n — k)
Appropriate specifications of R in Eq. (5-78) will yield confidence regions for
various groups of parameters. For example, setting R = J, and equating the
expression in Eq. (5-78) to some critical value F, gives a condition on the
unknown B vector from which a joint confidence region may be determined.
ESS
/(ee) e205 31)
a RSS/(w#—)) = 17.67
15/6 = 3)
THE k-VARIABLE LINEAR MODEL 191
From the tables of the F distribution, Fy ;(2, 2) = 19.00, so that the sample F
falls short of the 5 percent critical value. Even though the sample R? is
numerically high, the sample size is so small that it fails to reach significance.
. Testing the significance of X,
where a,, is the ith term on the main diagonal of (X’X)~'. We do not need,
however, to invert the 3 X 3 matrix X’X. In the development of Eq. (5-70) we
showed that the right lower k — 1 submatrix in (X’X)~! is given by
(X‘,AX,)~', which is simply the inverse of the matrix of sums of squares and
xan['9§
products of the variables in deviation form. For this example we have
. 10 6
Thus
, -1_ 1 —1.5
Ce ns he be
giving a, = 2.5. Further, s* = e’e/(n — k) = 1.5/2 = 0.75. Finally, sub-
stituting — 1.5 for b, and 0 for B, gives the test statistic
=
ee Val
¥0.75 ¥2.5
which is insignificant.
Alternatively, we may show that the same numerical value for the test statistic
comes from the stepwise reduction in the residual sum of squares. It is again
simpler to work with the data in deviation form. Regressing Y on X, gives an
estimated regression coefficient of
ae
vs 10
Hy: B, + B;
=0
From the general formulation
RB=r
this gives
R=[0 1 1] and’ r=0
with g = 1. The appropriate test statistic is given by the general result in Eq:
(5-68), namely,
b, + b,)°
aye) =2.66
0.75(0.5)
which falls well short of any usual critical value for F(1, 2). Thus the data are
not inconsistent with the hypothesis that 8, + 8, = 0.
and
no-8)=|7]- [2]=|3-8
[R(x’x) 'R’] | = fe ‘|
Substitution in Eq. (5-78) gives
he
asm tsa’? SIL iS fl 1*5
—1.5 — B,
This defines the 95 percent confidence ellipse for 6, and B,, which is sketched
in Fig. 5-1. The ellipse is centered at the estimated point b, = 2.5, b, = —1.5.
There is a strong negative covariance between the two estimates and the
origin lies just inside the ellipse, in agreement with the result of test 1 above.
b,
¥,=[1 10 10]| 6,| =Rb
b;
THE k-VARIABLE LINEAR MODEL 195
es ~ N(0,1)
o\1 + R(X’X) 'R’
Replacing the unknown o by
s=ee/(n—k)
then gives
ae ~t(n—k)
sy1 + R(X’X) 'R’
and so a 95 percent confidence interval for Y; is
We also have
5” =1 0:75
and
toons (2) = 4.303
3.66 to 24.34
or
This is a prediction interval for Y,, the value of Y in the forecast period.
Sometimes an investigator prefers to set up an interval for E (¥;), that is, the
mean or expected value of .Y in the forecast period, the reason being that Y,
contains the disturbance u,, which is essentially unpredictable. We have
Y,= RB + u,
Thus E(Y,) = RB
and the forecast error would now be defined as
¥, + to.o258/R(X'X)'R’ (5-80)
The numerical implementation of Eq. (5-80) gives
14 + 4.30370.75 V6.7
or 4.36 to 23.64
X=[i AX,]
THE k-VARIABLE LINEAR MODEL 197
n 0 i'y
0 X‘,AX, X’, Ay
0 i'y
I
X|-
So (X,AX,)"|| X,Ay
Yj
(X,AX,) 'X,Ay
where we used the result that i/AX, = 0 (the sums of sample deviations being
identically zero). The covariance matrix is
1
a 0
var(b) = o7| ”
0 (X,AX,)
The point forecast may be written
Y,= Y +x,b,
where x;=[x2, **: X,,] is a row vector of the X deviations in the
forecast period and b, is a (k — 1)-element column vector of the OLS
regression slopes. Thus
E(¥,) = £E(¥)+x,B,
and
var ¥,) = var(Y ) a x -E{(b, mtB, )(b, = B,)’}x’,
= o?|+ + x,(X,AX,)
x4
since the matrix var(b) above shows that Y and b, are distributed indepen-
dently. For the problem in hand,
Y=4 X,=3 X,=5
and so
x,=[7 5]
; -1_ 10 —-—1.5
ob %) wee |
and s? = 0.75. Thus the estimated var(Y;) is
0.75(0.2) + 0.75[7 oifea? a Te = 0.75(6.7)
198 ECONOMETRIC METHODS
R= [0 Xe tare exes|
the values that the forecaster thinks will be obtained in the forecast period. The
true value Y; is given by
Tee Paes
and the point prediction will now be
Y, = Xb
Thus the forecast error is
cia Yaa,
For simplicity we will drop the f subscript, since there is no ambiguity, and write
the forecast error as
e=u~X(b—B)—(R~
x)
= (UX (DiBaaRee es) (5-81)
If we assume that the forecaster makes unbiased forecasts of the X¥ values, that is,
E(&) =x
and, in addition, that there is zero covariance in the population between forecasts
of x and estimate of B from the sample data, then
E{%'(b — B)} = 0
and so
E(e)=0
Hence the variance of the forecast error is found by squaring Eq. (5-81) and
THE k-VARIABLE LINEAR MODEL 199
Alternatively, the forecaster may have subjective assessments that a forecast value
is very likely to be within, say, 5 percent of the true value, which in turn implies a
figure for the variance.
The remaining practical difficulty about the use of Eq. (5-83) is that we can
no longer determine exact confidence intervals using the ¢ and normal distribu-
tions. The reason is that even if normality is assumed for & as well as u, the
forecast error in Eq. (5-81) is not normally distributed since it involves %’(b — B),
which is the sum of products of normal variables. One may follow the suggestion
of Feldstein to use the Chebyshev inequality to determine an outer-bound
forecast interval.+ The practical procedure is as follows. Letting s; denote the
square root of the estimated value of Eq. (5-83) we can state:
The probability that the observed value of Y in the forecast period will fall
outside the interval Y; + cS, does not exceed 1/ Ca
The researcher can set the value of c to make 1 /c” equal to 0.05 or whatever
is desired. The Chebyshev inequality strictly involves the true o,, but it is a very
conservative statement and unlikely to be seriously affected by the replacement of
o; by sy. If the distribution of y were sufficiently well behaved to be unimodal
and symmetric, the probability of Y;, lying outside the interval i + cs would not
exceed 4/9c?.
PROBLEMS
LY = 20 XX, = 30 LX, = 40
LY? = 88.2 UXP = "92 DEX a168
LYX, = 59 LYX, = 88 LX, X, = 119
Estimate the regression of Y on X, and X>, and test the hypothesis that the coefficient of X, is zero.
5-3 Let
7M. S. Feldstein, “The Error of Forecast in Econometric Models when the Forecast-Period
Exogenous Variables are Stochastic,” Econometrica, 39, 1971, pp. 55-60.
THE k-VARIABLE LINEAR MODEL 201
1. e,; one;
2. yonX
Prove that:
(a) The slope b = e},e;/e;e; from regression 1 and the multiple regression coefficient b; from
regression 2 are identical.
(b) The residuals from the two regressions are identical.
(c) The simple correlation between e,; and e; is the same as the partial correlation between y and
x, 1n regression 2.
5-4 The following regression equation is estimated as a production function for Q:
logQ = 1.37 + 0.632 log K+ 0.452 log L
(0.257) (0.219)
R?=0.98 — cov( bg, b;) = 0.055
and the standard errors aregiven in parentheses.
Test the following nullhypotheses:
(a) The capital K and labor L elasticities of output are identical.
(6) There are constant returns to scale.
(University of Washington, 1980)
Note: The problem does not give the number of sample observations. Does this omission affect
your conclusions?
5-5 Consider a multiple regression model for which all classical assumptions hold, but in which there
1s no constant term. Suppose you wish to test the null hypothesis that there is no relationship between y
and X, that is,
Ho: B= = By =0
against the alternative that at least one of the 8’s is nonzero. Present the appropriate test statistic and
state its distribution (including the appropriate number(s] of degrees of freedom).
(University of Michigan, 1978)
5-6 One aspect of the rational expectations hypothesis involves the claim that expectations are
unbiased, that is, that the average prediction is equal to the observed realization of the variable under
investigation. This claim can be tested by reference to announced predictions and to actual values of
the rate of interest on three-month U.S. Treasury Bills published in The Goldsmith-Nagan Bond and
Money Market Letter. The results of least-squares estimation (based on 30 quarterly observations) of
the regression of the actual on the predicted interest rates were as follows:
r= 0.24 + 0.94 r*+e,, RSS = 28.56
(0.86) (0.14)
where r, is the observed interest rate, and 7;* is the average expectation of 1, held at the end of the
preceding quarter. Figures in parentheses are estimated standard errors. The sample data on r* give
Carry out the test, assuming that all basic assumptions of the classical regression model are satisfied.
(University of Michigan, 1981)
5-7 Consider the following regression model in deviation form:
Vp = ByX1,+ BoxX2, + u,
202 ECONOMETRIC METHODS
C,= B, + BY,
+ B3C,_, + u,
(University of Michigan, 1981)
5-9 Prove that R? is the square of the simple correlation between y and y, where y = ».(0... Gz
5-10 Prove that if a regression is fitted without a constant term, the residuals will not necessarily sum
to zero, and R?, if calculated as 1 — e’e/(y’'y — nY?), may be negative.
S-11 A researcher wishes to estimate the regression of y on X without an intercept term, that is, X
does not contain a column of Is. Unfortunately, the regression program at hand automatically
computes an intercept term. Douglas M. Hawkins suggests that the program can be “tricked” into
estimating the correct intercept free regression by entering each data point twice—once in its correct
form (y;,x;) and once with the opposite sign (—y,, aXe)
Prove that:
(a) The “trick” regression and the correct regression (with intercept suppressed) yield the
same
coefficients for X.
(b) The residual sum of squares from the “trick” regression is exactly double the value
from the
correct regression.
Compute the ratio of the standard errors of the two regressions.
(American Statistician, 34, Nov. 1980, p. 233)
5-12 (a) Prove that R* increases with the addition of an extra explanatory variable
only if the F
(= 17) statistic for that variable exceeds unity.
(b) Prove that the partial
r=
A F =
t*
FF redheads
where1 is the value of the statistic for testing the significance of the coefficie
nt of the X; to which the
partial r is related, and df is the number of degrees of freedom in the regression.
5-13 Let the regression equation be partitioned as
y= X,B, + XB, +e
Let b, and b, be the usual least-squares estimators. Suppose that E(e)
= X,y, that is, the mean vector
of the disturbances is a linear combination of some of the regressors.
Prove that b, is biased but b, is
unoviased.
(University of Michigan, 1981)
THE k-VARIABLE LINEAR MODEL 203
5-14 Suppose that the m X 1 vector x; denotes m observations on the ith individual (i = 1,..., p)
and x; is the corresponding vector of deviations from the ith sample mean. Let the x, and x; vectors
be “stacked” to give mp X | vectors
x= (x, x, --: x5]
and v= [k R ¥]
Find a matrix D such that Dx = x.
CHAPTER
SIX
FURTHER TOPICS IN THE
k-VARIABLE LINEAR MODEL
In Chap. 5 we have described the procedure for testing the hypothesis that the
elements of the population vector B obey the set of q (< k) linear restrictions .
embodied in the relations
Hy: RB=r
If Ho is not rejected, one may wish to reestimate the model, incorporating
the
restrictions in the estimation process. One important reason for such
reestimation
is that it will improve the efficiency of the estimates. This produces
an estimator
b, which then satisfies
Rb, =r (6-1)
For example, if the hypothesis of constant returns to scale is not
rejected for a
production function, the reestimation process would yield a produc
tion function
with estimated elasticities which sum to unity.
We must first of all show how to derive an estimator b, which
satisfies Eq.
(6-1). Second, we will use this estimator to cast new light on
some of the test
procedures of Chap. 5, and third, we will look at some important
applications of
the new estimator.
The assumed model, as before, is
y=Xp+u
FURTHER TOPICS IN THE k-VARIABLE LINEAR MODEL 205
= (XX) 'Xy
this equation may be solved for \ as
+ To keep the notation as simple as possible, we have not distinguished between the vectors b, and
X which appear in the objective function, Eq. (6-2), and the specific vectors that emerge as the
solutions to Eqs. (6-3) and (6-4).
+ Provided the restrictions RB =r are true, the variance-covariance matrix of the restricted
least-squares estimator may be shown to be
var(bs) = 02{(X’X) | — (xx) 'R'[R(X’x)'R’]'R(X’X)“'}
See Problem 6-6. We should also note that in some problems it may be simpler to obtain b, by
imposing the restrictions directly on the problem rather than by substituting in Eq. (6-5). For example,
suppose the data are already in deviation form and we wish to estimate
y = Box.
+ B3x3 + u
206 ECONOMETRIC METHODS
— (exes— e'e)/q
Het e’e/(n — k) ic)
where ee, denotes the restricted residual sum of squares derived from the vector
b,, which satisfies the q restrictions Rb, =r, and e’e denotes the unrestricted
residual sum of squares from the usual OLS regression. We have already derived
this result for one particular application in Eq. (5-76), but the derivation leading
up to Eq. (6-8) is perfectly general and applies to all cases.
To summarize, the test of the hypothesis that the elements of B obey aset of g
(< k) linear restrictions embodied in
Hy: RB=r
may be carried out by computing the unrestricted OLS vector b and the residual
vector e and then calculating the F statistic, Eq. (5-68),
b, — B
re (bet b)’X’X(b, — b /a
yeR(bsTih) (6-9)
e'e/(n — k)
or equivalently
F= (exe, — e’e)/q
e’e(n — k)
One of the most useful applications of these formulas is in tests of structural
change.
Y,, Lane 8 0 B u,
=---|=|-------<-- Ve ea (E11)
Te OP 10s 1 PEA Aa atae2 Un +1
Yr? OO Ae By Uy, +2
Oi ie)coy, eee
ny +ny Unitny
where the wartime observations have been listed first and the peacetime
208 ECONOMETRIC METHODS
ay
b= |i]a, = (xx) xy
b J
b,
i (GX) 0 aca
0 (XEX3) ssa
(XX) 1,
if: (6-13)
(X’,X,) X5Y>
These estimates are seen to be identical with those obtained by applying OLS
separately to Eqs. (6-10a) and (6-106). One merely sets the data up in the
form of Eq. (6-12) and a single regression will produce all four regression
parameters. Using Eq. (6-13) one can then obtain the vector e of ny +n,
residuals, and e’e gives the unrestricted residual sum of squares.
Now set up the null hypothesis of no structural change. This may be |
formulated as
a) a
qs Hi 7% a ae
or, putting it in the RB = r framework,
a)
li, 60. 3h EO 0
Ho: K [nO | a, -(¢]
B,
so that
R = [I - I] and ‘r=0 (6-15)
+ When using computers the student must take care to understand the propertie
s of the program
being used. If the program automatically estimates an intercept, feeding
in the block-diagonal X
matrix would produce a linear dependence between the column of units supplied
by the computer and
the first and third columns of X. Thus one must either feed in X as
it stands and suppress the
automatic intercept, or else allow the automatic intercept and modify the
X matrix in a way to be
discussed later in this section.
FURTHER TOPICS IN THE k-VARIABLE LINEAR MODEL 209
Peacetime data
1 ] 2
3 1 4
3 ] 6
5 1 8
6 Wess lO)
NG wee 1
if bala
9 Lat
9 eels
pl Be
erates eae 60
XY, ela 292 lee
yiy, = 61 Y2¥2 = 448
210 ECONOMETRIC METHODS
ae — 0.062500
ye Ei m (X/X,) XY, ae 0.437500
b, (Xexey 1X25; 0.400000
0.509091
Thus the estimated regressions are
Y = —0.0625 + 0.4375X wartime
and Y = 0.4000 + 0.5091X peacetime
These point estimates give the wartime function a smaller intercept and lower
slope than the peacetime function. The residual sum of squares from the
wartime regression is
ere; = yiy, — bi Xiy,
= 61 — [-0.0625 0.4375]| oF = 61 — 60.3125
= 0.6875
Similarly for the peacetime regression
giving
(sents 15
(X4Xx)
».¢ xX =
145 te ane eee
f =>
sel
The restricted coefficient vector is
Y = —0.0698 + 0.5245.
FURTHER TOPICS IN THE k-VARIABLE LINEAR MODEL 211
— 0.062500
= ut 0.437500 as — 0.462500
Reg Lege! ~ 0.400000 eee
0.509091
and
r = 0. Thus
SS
3.3969/2
3,1602/11 5.91
A as before
f
212 ECONOMETRIC METHODS
Example 6-2: Tests of change in the regression slope Example 6-1 showed
how to test the hypothesis
i a, a,
PE tienes
The restricted and unrestricted models are pictured in Fig. 6-1.
Sometimes the investigator is more interested in testing for the homo-
geneity of the regression slope, the values of the intercept term being of no
particular importance. The null hypothesis is now specified as
Hy: By = 8 (6-17)
The @ parameter is free to take on different values in the two subperiods. For
instance, in simple Keynesian theory the size of the national income multi-
plier depends only on the marginal propensity to consume £ and not at all on
the intercept a. Thus the H, in Eq. (6-17) is equivalent to asking whether the
income multiplier is the same in each subperiod. The restricted and unre-
stricted models may then be set up as follows:
Restricted Unrestricted
a)
. a 3
vie 0 x, A aw eto a 0 9071) 8, coe
Y2 0 i, x, B Y2 Oe OF exe
B,
(6-18)
where i, denotes a column vector of n, units, i, a column vector of nN, units,
x, a column vector of the n, observations on wartime income, and X, a
column vector of the n, observations on peacetime income. OLS may then be |
applied directly to each model in Egs. (6-18) and H, tested by comparing the
aS (a 2,8)
(a1,8))
(a,8)
Ue Xi ov ~ X
(a) (d)
Figure 6-1 (a) Restricted model; (b) unrestricted model.
FURTHER TOPICS IN THE K-VARIABLE LINEAR MODEL 213
residual sums of squares from the restricted and unrestricted models in the
usual way. The two models are shown in Fig. 6-2.
The unrestricted model in Eqs. (6-18) is exactly the same as that in Eq.
(6-12), so we already have
e’e = 3.1602.
For the restricted model,
sea la ee gs
eG Ls hex
so
>
=e
x O = OG
(a) (bd)
Figure 6-2 (a) Restricted model; (6) unrestricted model.
214 ECONOMETRIC METHODS
p- X’x
and iat — KKK)
exe, = VF Weep eyes
KF
Computing the deviations for the two subperiods and evaluating these
expressions gives
203
b= Ale 0.4951
and
2
ee, = 104 -— (203)"
410
3.4902 as before
Example 6-3: Testing for structural change in the intercept The null
hy-
pothesis is now
Hy: a, =a, (6-19)
We must be very careful in the specification of the restricted and unrestr
icted
models. By analogy with Example 6-2 it might seem reasonable to specify
the
FURTHER TOPICS IN THE k-VARIABLE LINEAR MODEL 215
restricted model as
y my beky ix 1 “elles
| : hi 0 Fe 4 uae Gey)
with the unrestricted model as before. Model (6-20) imposes a common
intercept but specifies different slopes. If the functions have different slopes,
they must intersect at some X value. There may be cases where it is relevant
and important to test that the intersection occurs at X = 0, as is implied by
specifying Eq. (6-20) as the restricted model. However, this is not usually the
case, and the most common practice is to test Hy, subject to the assumption of
a common regression slope. Thus the restricted and unrestricted models
become
i, x ise oe ey et
y2 In Xo yp 0 i, x, B
Notice that the unrestricted model in this example is the restricted model of
Example 6-2 [see Eqs. (6-18)], and the restricted model here is the same as the
restricted model in Example 6-1. Thus from our previous calculations the
relevant sums of squares are
ee, = 6.5565
and e’e = 3.4902
Thus the test statistic for Hj: a, = a, conditional on a common 8, is
6.5565 — 3.4902
F=~3.4902/12 = 10.54
and Fj 99(1, 12) = 9.33 so that the difference in the intercepts is significant at
the 1 percent level. The models are shown in Fig. 6-3.
4 Yi
x O Fe
(a) ()
Y Ys
' f
Y,
Y, d
Ya an
Y,
“1
O > xX Oo Xe
(a) | (b)
Figure 6-4
é a
I spe a hs Xj ns re differential intercepts,
Yy> OI xX, B common slope
O47 |e
Yoo Ay OF OR, differential intercepts,
Ill = : +
Y> 0) Oi, xy differential slopes
B,
Fitting each model by OLS produces a residual sum of squares. There are
three basic tests on the differences between the various residual sums of squares.
These are the following:
e Test of differential intercepts—model I contrasted with model II
e Test of differential slope coefficients—model II contrasted with model III
e Test of differential regressions—model I contrasted with model III (slopes and
intercepts)
The tests outlined above have implicitly assumed that the disturbance vari-
ance o” is the same in each period. Schmidt and Sickles have investigated the
effect of departures from this assumption on the significance level of the test.f For
equal-sized samples there are modest increases in the true significance level over
the nominal level, even for very large departures from the assumption of equal
variances. For instance, with n, = n, = 25 the true significance level only rises to
0.059, compared with a nominal value of 0.05, when one variance is 100 times the
other. If the X variable is a linear trend, the true significance level rises to 0.063
for a tenfold increase in the variance and to 0.084 for a one-hundredfold increase.
When the sample sizes are unequal, the true significance level shows a greater
departure from the nominal level, and it may now be less or greater than the
nominal level. Full details are given in the reference.
+P. Schmidt and R. Sickles, “Some Further Evidence on the Use of the Chow Test under
Heteroscedasticity,” Econometrica, 45, 1977, pp. 1293-1298.
218 ECONOMETRIC METHODS
follows:
X, 7 [i, XT]
a
Il Yu ie ee pee Ota a differential intercepts,
f ¥ | O15 01 XS TR . differential slopes
By
where we have partitioned the k-element B vector as
a
B, 3
STB a
By,
Application of OLS to each model will yield a residual sum of squares (RSS) .
with an associated number of degrees of freedom as indicated by
Model I RSS, n—k
Model IT RSS, Ne ike ab
Model III RSS, R= 2k
where n = n, + n, indicates the total number of observations in the com-
bined samples. The test statistics for various hypotheses are then as follows:
oe (exes — e'e)/q
e’e/(n — k)
where g indicates the number of coefficients in the subset. Formally the
restricted model is set up as
Bi,
Yi Xi 0 ue B
= +u 6-26
Ki |0 XX, Xx» B, ( )
D
Example 6-5: Tests of structural change (n, < k) A special problem arises if
one of the subperiods has fewer observations than the number of parameters
to be estimated in the model. Let us assume that we have n, (> k)
observations in one subperiod and n, (< k) observations in the other. There
is no difficulty about the restricted model in which one set of k parameters is
estimated for the n (= n, + n,) sample observations, namely,
3 be X,
b,
+ ey
Y2 X,
[3
be fitted and will have a residual vector
7 This is only a heuristic proof. For an exact derivation of Eq. (6-27), see F.
M. Fisher, “Tests on
Equality between Sets of [Link] Two Linear Regressions: An Expository Note,”
Econometrica,
28, 1970, pp. 361-366. An alternative proof is given in Sec. 10-1.
FURTHER TOPICS IN THE k-VARIABLE LINEAR MODEL 221
y
i, Xf
y2 i Qa
I eee X3 fa +u
o * we
Y, Pp P
Y) oa
: i, 0 OF xe
it ‘i = 7) 0 Xi] - | t+u
0 0 1, XS a>
Yp B*
Oy
Eo)
; i, 0 Oe Xt 0 Dal eee
Tt erty, os eal NOs Se a Onl laa
: 0 i ee 0 AS Bx
.
P
By
differential intercepts, differential slope vectors
where
LW2. ar
222 ECONOMETRIC METHODS
Table 6-1
Class
l 2 3 4
Observation Y X x XG VG X va XE
l 22 29 30 15 12 16 23 5
Pp. 22 20 32 9 8 31 25 25
3 20 14 26 l 13 26 28 16
4 24 21 26 6 25 35) 26 10
5 12 6 37 19 7 12 23 24 Ys X
Ny 0 0 dix a
(134)°
eye, = 88 — 55, — 26.925
i (117)
e794 = 26.897
eels
/
ang
0 ——
8
giving
RSS, = 195.8
The various tests may be set up in an analysis of variance framework as
shown in Table 6-3.
The test for a common regression slope is
7. BSS
= RSS)/3 _ 184 |
7 RSS,/12 163 eke
224 ECONOMETRIC METHODS
og eal
li-6 0 9- g ZOE
ee
here
p 0
Z- 0 € I Z- 81
ee EO
8— L Z II Z1- Z8E
a
€ LLI
On) ee
I- S— 0 ZI9- 902
a
SseID
ee
re
¢ I- 6- y- 6 p07
a
Z LU
ee a
Se a) 0 a p- $- L r6
oy
eS
= Il Z p- € Z1- 67
hy)
I rel
ay
re
Z Z 0 P g— 88
Ap ere
a
Ce
LE
hawx
a uoNeAIISqO
ee ie
7-9
=
FIQUL By)
1)
i I é € p ¢ ia li ii
FURTHER TOPICS IN THE k-VARIABLE LINEAR MODEL 225
Table 6-3
Mean
Model Residual sum of squares Degrees of freedom square
and Fo 99(6, 12) = 4.82, so that this too is a highly significant result, but it
would appear that the significance is due to variation in the intercepts and
not in the slopes.
Dummy variables have already made their appearance in the previous section, but
we have not explicitly labeled them as such. For example, the unrestricted model
in Eqs. (6-21) specifies a consumption function which has different intercepts, but
a common slope, in wartime and peacetime periods. The specification is repeated
here
: a)
= if , au %}+u (6-28)
yz 0 i, X, B
where the subscript 1 refers to wartime and the subscript 2 to peacetime. This
model may be written as
Y,= 0,D,,+ a)D,,+ BX,+u, t= 1,2,...,0 (6-29)
D,, and D,, are dummy variables whose sample values are given in the first two
226 ECONOMETRIC METHODS
Equation Equation
(6-29) (6-30)
Wartime intercept a, vA
Peacetime intercept Q> Yi + Y2
The choice between the two estimation procedures is of no great importance, but
it is very important to be clear about precisely what is being tested in either
model. For instance, testing the significance of D, in Eq. (6-30) is, in effect, testing
the hypothesis
Hy: a,—a,=0
which is testing whether the peacetime and the wartime intercepts are significantly
different, whereas testing the significance of D, in Eq. (6-29) is asking whether the
peacetime intercept is significantly different from zero.
FURTHER TOPICS IN THE k-VARIABLE LINEAR MODEL 227
The dummy variables may also be allowed to interact with the X variable.
Consider
Y =a,D, + aD, + B,(D,
+ B,(D,X)
X)+4 (6-32)
where the subscript ¢ has been omitted for simplicity. Equation (6-32) implies two
separate relations, namely,
Y=a,+B,X+u wartime function
Y=a,+fB,X+u peacetime function
Thus performing a single regression of Y on D,, D,, D,X, and D,X with the
general intercept suppressed is equivalent to fitting separate regressions to the two
subperiods. An alternative formulation of Eq. (6-32) is
Thus we see that, in the two-variable model, the tests for homogeneity of
intercepts and homogeneity of slopes are equivalent to tests of the significance of
single coefficients in an appropriately specified regression equation using dummy
variables.
Dummy variables may also be usefully applied in more complex models. For
the data of the last numerical illustration we may specify
different from zero. This, of course, confirms the homogeneity of regression slopes
established earlier by the F test. Imposing the assumption of a common regression
slope gives the revised regression
Y =13.4882 + 12.8969D, — 9.1726 D, + 5.742D, + 0.3621X
(4.79) (4.68) (—3.42) (2.20) (3.04)
with R? = 0.7341 and 15 degrees of freedom. All three dummies are significantly
different from zero at the 5 percent level, thus establishing that the intercepts in
the second, third, and fourth classes are different from the intercept in the first
class, again in agreement with the earlier F test on intercepts. One advantage of
this type of dummy variable setup is that in cases where the tests examine the
Joint significance of a subset of variables the dummy variables can indicate which
variables may have made the most important contribution to the overall signifi-
cance of the group.
The dummy variables specified above play an important role in describing
temporal effects (where the classes refer to different time periods), spatial effects
(where the classes refer to different regions or countries), industrial effects (where
the classes refer to industries), and so forth. Suppose we have qualitative variables
such as
e Education (none, grammar, some high school, high school diploma, some
college, college degree, advanced degree, foreign education)
¢ Marital status (unmarried, married 1 year, 2 years, 3 years, 4 years, 5—9 years,
10—20 years, over 20 years)
¢ Sex (male, female)
e Race (white, black, other)
Only the last two are truly qualitative variables. Education might be treated as a .
cardinal variable, measured by years of formal education, and likewise, duration
of marriage is a cardinal variable. In both cases, however, we may use groupings
of a cardinal variable to define a qualitative variable. If a qualitative variable is
thought to influence some dependent variable, we may use the categories of that
variable to classify the sample observations into various classes, and the preceding
method of analysis applies. There are, however, some slight complications if we
wish to use two or more qualitative variables in a single equation.
and
ee 1 if observation relates tosexj, j= 1,2
0 otherwise
Suppose we then wish to examine the relationship between hours spent in reading
nonfiction Y and these two qualitative variables. It is instructive to examine first
of all what happens if we have only one set of dummy variables in the model. A
linear model for the influence of E on Y would be written
a, %
a,)= yy
a3 Ne
so that the OLS regression coefficients are simply the mean values of Y in each of
the educational classes. If we used the alternative formulation
a Y,
eee
BV i
i = =
a3 Y, 1
+ See Problem 6-7.
230 ECONOMETRIC METHODS
Educational level
E,
7 We might choose any one of the six cells to be represented by p. The differenti
al effects would
then be measured from that cell, but the numerical estimates of the conditional means will be
invariant to the starting position. See Problem 6-8.
FURTHER TOPICS IN THE k-VARIABLE LINEAR MODEL 231
12, 14
10 20
OLS, and from these we obtain unique estimates of the expected values in Table
6-5.
R
a, |+u (6-39)
N
Tw)oO
Pm
pe
ep
et
ee SO
[Link]
or [OS
Ore
OE
IS OS)
eS =)
SS
SIS
Interaction
The main drawback of Eq. (6-38) and the estimates to which it gives rise is the
built-in assumption that the differential effect of each factor is constant across the
levels of the other factor. Thus Table 6-7 shows that hours for S, are 7.35 lower
than for S,, irrespective of the level of education. Conversely, E, shows 3.31 more
232 ECONOMETRIC METHODS
hours than E,, and E, shows 12.54 more than £,, irrespective of whether we are
in the S, row or the S, row. This implies the absence of any interaction between
the two factors. If, however, it is to be expected that the differential sex effect
varies with the level of education, then an interaction effect exists, and we need to
see how to incorporate it into the model and estimate it.
Returning to Eq. (6-38), we would now expand the relation to read
There are only two possible interaction variables in this case, and they are found
by multiplying each E level by each S level. The conditional expected values are
now shown in Table 6-8.
The first row is the same as in Table 6-5, but the second row incorporates the
interaction effects. Thus the sex differential is
B, for E,
By + ¥p for E,
By +7; for E,
B+ Q, b+ a
p+ a, + Bo + yp w+az+ B+ ¥5
FURTHER TOPICS IN THE k-VARIABLE LINEAR MODEL 233
1333 13.00
0.50 10.00
Referring back to the data matrix in Eq. (6-39), the data matrix for this problem
would now be
E, E; So ES) E3S>
1
1
|
1
tea
ee 1 1
1 1
1 |
is ise
1 ea |
Thus
10Ne O24 GL 117
aoe EGhe (uetdee 0 36
Oe nee2 ae Oe 40
ee eed UAT Bee a Hl See Unt oet
Tete colts Ie O 10
es Os eel Otel 20
The OLS equation is now
Y = 13.33 — 0.33E, + 6.67E, — 12.835, + 9.83( E,S,) + 12.83(£;S,)
and substitution in Table 6-8 gives the estimated number of mean hours shown in
Table 6-9.
Compared with the previous regression, where no interaction effect was
incorporated, we now have a large negative sex effect (— 12.83) at E,, which is
reduced to —3.00 at E, and eliminated completely at E,. This last result is an
automatic consequence of our data, where in the interests of simplicity we had
only one observation in each of the £; cells and also in the E,, S, cell. The
regression values, with interaction, then coincide with these observations. This has
also distorted the estimate of the E, differential effect to give a small negative
number (— 0.33) for S,, but the calculations do illustrate the principles involved.+
+ This section has only dealt with dummy variables on the right-hand side of the equation. For a
discussion of the application of dummy variables to the left-hand-side variable see Sec. 10-5,
Qualitative Dependent Variables.
234 ECONOMETRIC METHODS
VO OVO
OV Sn Ona)
OFS.O eae 0
D=s(0O
08 02
1 0250-550
Ome? OVO
OF 052.0: el
This is the sample matrix for four dummy variables defined by
yo = My (6-42)
where
MD = 0 (6-44)
The series y* cannot serve directly as a deseasonalized series for two reasons.
First of all, it sums to zero, and it would seem plausible to require a deseasonal-
ized series to have the same sum as the original, unadjusted series. Second, as
FURTHER TOPICS IN THE kK-VARIABLE LINEAR MODEL 235
Yy,
Y,
b=|_
Y,
Y,
where Y, (i = 1,..., 4) is the mean of all ith-quarter Y values. Thus y* merely
consists of deviations of the Y values from the quarterly means. But if the series
contains trend and/or cyclical components, the elements of b will be an amalgam
of trend, cyclical, and seasonal effects. Thus subtracting b year by year from the
actual Y values will not yield satisfactory estimates of a deseasonalized series. The
remedy is to introduce into the regression a polynomial in time of sufficiently high
order to represent the trend and cyclical components, so that the coefficients of D
will be a more satisfactory estimate of the seasonal component. Thus one
computes the regression
y =Pa+Db+e (6-45)
where
1 12 er 12
2 22 2?
Paine B32 3?
4 4? 4P
4n (4n)° (4n)?
The deseasonalized series would now be defined as
y* =y — Db (6-46)
Jorgenson has argued that if the P and D matrices are properly specified, then a
and b will be best linear unbiased estimates of the systematic and seasonal
components, since Eq. (6-45) is then a straightforward example of ordinary least
squares.} The estimates of a and b are given by
Seasonal component
Method by by b; bg
y= (x p][,']+e
The OLS coefficients are then given by
c, = XX XD|~'[X’y
|
b, | |DX DD] [Dy eee)6-52
Applying Eq. (4-68), the first element in this inverse matrix is
+™M. C. Lovell, “Seasonal Adjustment of Economic Time Series,” Journal of the American
Statistical Association, 58, 1963, pp. 93-1010. The basic result goes back to R. Frisch and F. V.
Waugh, “Partial Time Regressions as Compared with Individual Trends,” Econometrica, 1, 1933, pp.
387-401.
238 ECONOMETRIC METHODS
Cy © (6-55)
This result is, of course, symmetrical with respect to X and D, and the D matrix
need not consist of dummy variables; it is merely any subset of explanatory
variables. However, Lovell is concerned with seasonal adjustment, and D is then
appropriately an n X 4 matrix of quarterly seasonal dummies.
Two further basic results from Lovell are that the regressions
y= X°c, cs e3
Thus
c, = (X’MX)_'X’My
These results raise some further questions. We have already seen that if D is
merely a matrix of seasonal dummies, then y* and X°, defined in Eqs. (6-54), are
not properly deseasonalized series. On the other hand, if properly deseasonalized
series are obtained by using the transformation matrix T defined in Eq. (6-50),
this matrix, though idempotent and orthogonal to D, does not have the symmetry
property used in the above proofs. Furthermore, many official series are not
deseasonalized by least-squares methods at all, but by moving average or other
methods. Thus the Lovell results cannot be expected to hold exactly when y* and
X“ indicate properly deseasonalized series. Nonetheless some experimental calcu-
lations with various equations from the Oxford econometric model of the United
Kingdom indicate agreement to several decimal places between estimated coeffi-
cients, whether the regression has been run with raw data and dummy variables or
with deseasonalized variables produced by moving average methods or by least-
squares regressions on D or on[P_ D].} The years covered by the model showed
fairly steady growth and negligible cyclical oscillations. One would not expect
such close agreement if the cyclical effects were very strong, and in practical work
one should not allow this theorem to be a substitute for careful thought about the
proper specification of the relationship.
6-5 MULTICOLLINEARITY
Le altel ae
ee Ombvcetyi Wika eae oe ls 30?
# as ca ae | Pes
+ A. Georgopoulou and J. Johnston, “Seasonal Adjustment of Economic Time Series,” University
of Manchester, discussion paper.
240 ECONOMETRIC METHODS
In case 1 the two explanatory variables are orthogonal and the coefficients of the
X’s in the multiple regression equation would be the same as those given by the
simple regressions of Y on each X in turn. Orthogonal variables may be set up in
experimental designs, but they are the exception, not the rule, in economic data.
Cases 2 and 3 display increasing correlation between the two explanatory
variables, as evidenced by the increasing numerical value for the off-diagonal
(covariation) term. This is also reflected in the dramatic fall in the value of the
determinant. This is described as a situation of collinearity (or multicollinearity)
between the explanatory variables. Three important effects are illustrated in the
sequence of matrices:
by 0.952 A aly
0.9b,+b,=2.9 ? 3
—_ = a
Now suppose the X; variable is somewhat more highly correlated with X, and
we have normal equations for case 3 as
b, + 0.99b, = 2.8 ee
O99 b= ae aa
The only numerical change between the two sets of equations is a 10 percent
(or less) increase in two coefficients, yet the solution values change dramati-
For simplicity these three important points have been illustrated for the case
of two explanatory variables. It is important to establish that similar results hold
for the k-variable case and to discuss how multicollinearity may be detected and
what may be done about it. However, before doing that, we will discuss the
limiting case of exact, or complete, multicollinearity.
Kok Ei “|
eae
with |X’X| = 0 and p(X’X) = 1. This is simply a breakdown of the assumption
that X has full column rank, and so we cannot obtain the unique OLS vector
defined by
b= |= (X’X) 'X’y
The normal equations
(X’X)b = X’y (6-58)
however, will admit an infinity of solutions for
X’y = Ex2y| 1]
Qa
+ This is only a hypothetical example, but the literature of applied econometrics is full of examples
of small changes in the data base producing substantial changes in estimated coefficients. For one
example, see J. Johnston, “An Econometric Model of the United Kingdom,” Review of Economic
Studies, 29, 1961, pp. 29-39.
242 ECONOMETRIC METHODS
so that the rows of X’y exhibit the same linear dependence as the rows of X’X.
The set of equations in Eq. (6-58) is thus consistent, and there is an infinity of
solution vectors. Taking the first equation in Eq. (6-58), we have
Ex7(b, + ab;) = Expy
and the second equation is
abx3(b, + ab,) = abx,y
Both equations reduce to
b, + ab, =
x7y
2
(6-59)
2
Thus no matter which arbitrary solution to Eqs. (6-58) we take, the linear
combination b, + ab, will always have the same numerical value. We then define
B, + aB; as an estimable function, where we notice that the a in the estimable
function is the parameter defining the linear dependence between x, and x3.
The same result may be derived by writing the model in deviation form as
y = Bx. + Bx, + (u— @)
and substituting Eq. (6-57) to get
y = Bx,+ (u-@) (6-60)
where
B = B, + af, (6-61)
The B parameter may be estimated by applying OLS to Eq. (6-60) to give
poy (6-62)
Ds
which is the same expression as that already obtained in Eq. (6-59). The expected |
value of y for a given x, (and x;) is
E(y)= ByX_ + B3x3
= (B, + aB;)x,
= Bx,
Thus E(y) can be estimated uniquely since 8B can be estimated uniquely by Eq.
(6-62).
Table 6-11
Estimate of
b; by b> 5 2b; E(y|x2 = 20)
0 0.5 0.5 10
l —1.5 0.5 10
—] DS 0.5 10
110) 20.5 0.5 10
with solution
7 >= 20D;
B 10
Taking some arbitrary values for b, gives Table 6-11.
The linear combination b, + 2b, is invariant to the solution chosen for
the normal equations, and it is readily seen to be equal to
z DX y Ce
b 0.5
re, al?
Likewise, the regression value for any given x, is invariant to the normal
solution vector.
To summarize, even though £, and £, cannot be estimated, a certain
linear combination of B, and f, can be estimated, and E(y) can also be
estimated for any given x, value.
The nature of the problem may also be illustrated geometrically. In Fig. 6-5a
the standard OLS case is shown. The x,,x, vectors are not perfectly collinear,
and they span a two-dimensional subspace in &”. Dropping a perpendicular from
(a) (d)
Figure 6-5
244 ECONOMETRIC METHODS
9=XB (6-70)
FURTHER TOPICS IN THE k-VARIABLE LINEAR MODEL 245
and the estimates in Eqs. (6-69) and (6-70) will have the usual OLS properties.
The operational procedure would be to identify the largest submatrix in X with
full column rank, denote it by X,, and substitute in Eqs. (6-69) and (6-70). If
there is more than one such submatrix, } will be invariant to which is chosen.
As in the case of two explanatory variables, an alternative procedure is to
derive any solution by to the normal equations _
(X’X)by = X’y
and compute
§ = Xby (6-71)
The numerical values for § in Eqs. (6-70) and (6-71) will be identical. One needs
to determine the linear dependencies in the X data (as in the W and Z matrices)
in order to determine the precise linear combinations of £ coefficients that are
being estimated in B,, but such combinations are not usually of any economic
significance. Furthermore, the use of Eq. (6-70) or Eq. (6-71) for forecasting
outside the sample observations rests on the same linear dependencies holding
among the X’s in the forecast period.
Near Multicollinearity
The prevalent case in so much econometric work, especially with time series data,
is one of high but not exact multicollinearity. This raises three questions:
Effects
Provided the X matrix has full column rank, the OLS estimates exist and will still
be the best linear unbiased estimates. This property, however, is now cold comfort
since the sampling variances of the estimates increase alarmingly with rising
collinearity. To prove this in the general case, partition the X matrix as
», [x, X;]
Then
* Xi XA,
ee Nine XK
Applying Eq. (4-68), the leading term in (XX) is
where
M, = 1 - X,(X;X,)°X;
Thus the sampling variance of the OLS estimate of ; is
ue (6-72)
2
0
var(b,) =
LOR O02a0 OO 0
1. Ove) 0 OP 0) 0 0 0
ORO! 0 0 1
det = 50
1 -
RSS,
—
TSS,
Then Eq. (6-72) may be rewritten as
o2
var(b,) =
RSS, ~ TSS,(1 — R?)
Letting b,, denote the estimate cf B, in the orthogonal case,
var(b,,)= TSS,
Thus, if TSS, is held constant, the magnification of the sampling variance with
increasing collinearity is given by
var(b)) nil
(6-73)
var(®,,) 1 — R?
The orthogonal case is not meant to be a feasible target, but is used as a
248 ECONOMETRIC METHODS
b
RCC i bs eh ahs
var(b;,)
A common result is to find regressions possibly with a very high overall R*, but
100 |-
80
60
40F
with some (or many) individual coefficients apparently insignificant. The high R*
arises when the y vector is close to the hyperplane generated by the x, vectors and
the apparently insignificant coefficients arise because the x ; vectors are nearly
linearly dependent. It is also possible to find a high R? and highly significant ¢
values on individual coefficients, even though multicollinearity is serious. This can
arise if individual coefficients happen to be numerically well in excess of the true
value, so that the effect still shows up in spite of the inflated standard error
and/or because the true value itself is so large that even an estimate on the
downside still shows up as significant. The multicollinearity would likely show up
in varying parameter estimates as some sample observations are dropped or
added. For any regression, however, comparison of the R?’s shows which
coefficients are likely to be most seriously affected by collinearity.
Detection
Computer programs often print out |X’X|. As our numerical examples illustrate,
the determinant declines in value with increasing collinearity, tending to zero as
collinearity becomes exact. While a useful warning signal, we have no calibration
scale for assessing what is serious and what is very serious, and again it gives no
guide to the relative effects on individual coefficients. Similar remarks apply to the
computation of the eigenvalues of X’X. Since
[XX] =A\Ap--- Ay
a small determinant means that some (or many) of the eigenvalues will be small.
But again knowledge of the eigenvalues is of little direct help in assessing effects
on individual coefficients.
The most useful single diagnostic guide is the R?’s, as shown above. In a
sense TSS, determines the minimum sampling variance that might be achieved for
b,; in that in the orthogonal case
ae
Any collinearity in the sample data will raise all sampling variances, but the
relative magnifications for different coefficients will be indicated by a comparison
of the R?’s.
Belsley, Kuh, and Welsch suggest the combined use of two diagnostic tools to
detect which coefficients are most likely to be affected by the collinearity.t The
first statistic is the condition number of the X matrix, defined by
r max
K(X) =
r min
+ The precise relationship between var(b;) and the A’s is derived below.
+ D. A. Belsley, E. Kuh, and R. E. Welsch, Regression Diagnostics, Identifying Influential Data and
Sources of Collinearity, Wiley, New York, 1980, chap. 3.
250 ECONOMETRIC METHODS
where A,,,, and A,,, denote the maximum and minimum eigenvalues of X’X,
respectively. If the X matrix has been standardized so that each column has unit
length, then «(X) is unity when the columns of X are orthogonal and rises above
unity with collinearity between the columns. Various applications with experi-
mental and actual data sets suggest that condition numbers in the range of 20 to
30 are probably indicative of serious collinearity problems, and a fortiori for
numbers in excess of that range. A condition index may be computed for each
eigenvalue, starting at unity for A, = A,,;, and rising to «(X) for A; = A,,.x- Thus
a given data matrix may yield one or more condition indexes in excess of a
“danger” level. The second and related diagnostic tool is the regression coefficient
variance decomposition. If X isn X k and V is the k X k matrix that diagonalized
X’X, then
(X’X)V = VA
where A is the diagonal matrix of the eigenvalues of X’X. Thus
var(b) = o?(X’X)| = o2VA~'V’
and
One2 ©:2 v;2
var(b,)
( ) = 07) =!
r, +2
v5 4... 4 8
rx i= aeek
where 0,;, U;2,---, U;, are the elements of the ith row of V. From this one may
compute the proportions of var(b;) associated with each A. The two-step procedure
recommended by Belsley, Kuh, and Welsch is
1. Compute the A,’s and identify any A; which gives a condition index in excess
of the “danger” level (say, 20 to 30).
2. For each of those selected A,’s inspect the proportions of the sampling
variance of each 5, associated with that eigenvalue. Coefficients with propor- .
tions in excess of, say, 0.50 are likely to have been adversely affected by the
collinearity in the X matrix. Reference should be made to Belsley, Kuh, and
Welsch for detailed examples of the technique.
Remedies
This is a sequential approach using one set of cross-section data and another set of
time-series data, rather than a joint (or simultaneous) set. The latter requires that
the observations relate to a common set of decision units.
The general framework for the incorporation of prior estimates of some
parameters may be set out as follows. Partition X, B, and b as
+ Notice that this operation involves taking expectations over two different sets of data. E(u) refers
to expectations over the current sample data and E(b,) to expectations over the data underlying the
prior estimate b,.
252 ECONOMETRIC METHODS
7 See A. E. Hoerl and R. W. Kennard, “Ridge Regression: Biased Estimation for Non-Orthogonal
Problems,” Technometrics, 1970, pp. 55-68, for an exposition of the theory; and A. E. Hoerl and
R. W. Kennard, “Ridge Regression: Applications to Nonorthogonal Problems,” Technometrics, 1970,
pp. 69-82, for two illustrations.
+ See P. Schmidt, Econometrics, Marcel Dekker, New York, 1976, pp. 48-55, for this result and a
very useful discussion of the theory of ridge regression.
§ For a definition of MSE see Chap. 2, pages 27-28.
{| For further discussion see the series of papers in Journal of the American Statistical Association,
75, 1980, pp. 74-103.
FURTHER TOPICS IN THE k-VARIABLE LINEAR MODEL 253
Another approach to improving the MSE involves the suggestion that one or
more explanatory variables be dropped in order to improve the MSE of the
remaining coefficients. To illustrate the approach consider just a three-variable
model,
y = Bx. + Bix, +-u (6-80)
where the variables have been expressed in deviation form.} Let us denote the
coefficients resulting from the application of OLS to Eq. (6-80) as
b,.; = OLS estimate of B,
b\32 = OLS estimate of B,
From the properties of the OLS model we know these estimators are unbiased
and their sampling variances are$
o2
var(b
( 123) SS rx3(1 fn r3) (6-81)
ee
var(b,39) = (6-82)
eal _ 7)
Clearly, as 73 gets close to unity, both sampling variances increase dramatically.
Now consider the simple regression of y on x, and denote the slope coefficient by
yx
by = 5
LXS
Substituting for y from Eq. (6-80) gives
Yx,u
bi, = By + bs)
B; + a (6-83)
Dx
where
b =
x53
eee
a x5
+ Strictly speaking, when the relation is written in deviation form, the disturbance term is u — i,
but this slight complication has no effect on any of the derivations in which we are interested and so
may be ignored.
+ These are derived from the general formula
which gives
a7 DxF i o
var(b\23) =
Lx3Dx} — (Lx_x3)° tsEx3(1 = 3)
where r>, is the simple correlation coefficient between x, and x3. A similar derivation yields the result
for var( 5,32).
254 ECONOMETRIC METHODS
Thus b,, is a biased estimator of 8,, unless x, and x, are orthogonal so that
b,, = 0. However, comparison of Eqs. (6-81) and (6-85) shows that b,, has a
smaller sampling variance than b,,;. The possibility then exists of a tradeoff
between bias and variance. The crucial question is under what conditions b,, may
have a smaller MSE than 5), 3.
As shown in Eq. (2-23),
MSE = sampling variance + square of bias
Thus
2
MSE(b,,) = sear
0
b2, B2
2
and
o2
MSE(6,,
3) =
Exel ‘a ra)
A little algebra then showst
MSE(b,,)
MSE(b,,) |* ee)
——~—*~ = 14+ 73(7? - 1) 6-86
where
2 = a a
Bs = —3_
Bs (6-87)
0° /Ex3(1—rZ) — var(b,39)
This 7° statistic is the ratio of the square of the true (but unknown) B, to the true
(not the estimated) variance of b,,,. From Eq. (6-86), if 7? < 1,
7 Notice that, in this case, var(b,7) is the variance about a biased expectation. From first principles
OF (=) os
xe xs
1Sp eS 6-88
57x (il ~ 1) ( )
OLS is preferable to any of the COV estimators unless the researcher has a
strong prior belief that tT < 1.4
1+ 7?
and substituting this value of A in Eq. (6-90). Feldstein’s simulation experiments
show the WTD estimator to be generally superior to the various COV estimators
in his study, but to be inferior to OLS when |7| > 1.5. Thus exhaustive study of
the three-variable case suggests that, even in the presence of high correlation
between x, and x3, the best procedure is probably the straightforward OLS
regression of y on x, and x3. Only if the investigator has really strong prior beliefs
that B, is less than /var(b,,,), should x, be dropped from the regression. Even
this nonstartling advice to drop a variable when you are fairly sure its coefficient
is “small” is only helpful if the investigator is mainly interested in the other
coefficient, B,.
Even though these results on the three-variable case are not very helpful,
considerable work has been done on extensions of the approach to the k-variable
case. As we have already seen in Chap. 5, setting a coefficient or group of
coefficients at zero is a special case of imposing a set of linear restrictions on the
coefficients. Thus the question arises whether the imposition of a set of restric-
tions will result in estimators which are better in some MSE sense than the
unrestricted OLS estimators, even though the restrictions may not, in fact, be
true.
The first problem is the generalization of the MSE criterion to a number of
estimators. Consider the usual linear model
y=Xfp+u
with the set of g (< k) restrictions embodied in
RB =r
As seen in Eq. (6-5), the estimator embodying these restrictions is
b, = b + (X’X) 'R’[R(X’X)
'R’] ‘(r — Rb)
where
b = (X’X) ‘'X’y
is the unrestricted OLS estimator. We may define the MSE matrix for b, as
and similarly for MSE(c’b). Thus Eq. (6-93) requires that the MSE of any linear combination of the
elements of b, be no greater than the MSE of the same linear combination of the elements of b.
FURTHER TOPICS IN THE K-VARIABLE LINEAR MODEL 257
A 20° aa:
(6-95)
As in the previous simple case, this condition involves the true but unknown B
vector and the unknown o?”. If these are replaced by their OLS estimators and the
resultant value of the statistic in Eq. (6-95), denoted by A, it is easy to see that
F=
2%,
+ C. Toro-Vizcarrondo and T. D. Wallace, “A Test of the Mean Square Error Criterion for
Restrictions in Linear Regression,” Journal of the American Statistical Association, 1968, pp. 558-572;
T. D. Wallace and C. E. Toro-Vizcarrondo, “Tables for the Mean Square Error Test for Exact Linear
Restrictions in Regression,” Journal of the American Statistical Association, 1969, pp. 1649-1663;
T. D. Wallace, “Weaker Criteria and Tests for Linear Restrictions in Regression,” Econometrica, 40,
1972, pp. 689-698; J. Goodnight and T. D. Wallace, “Operational Techniques and Tables for Making
Weak MSE Tests for Restrictions in Regressions,” Econometrica, 40, 1972, pp. 699-709.
£ Notice that dropping x, from the model
y = Box.
+ B3x3+u
is equivalent to imposing the restriction
Ont
B, =
|B;
and with these specifications of r and R condition (6-95) becomes
ae
2var(bi32) 2
or tT? < 1 as derived in Eq. (6-87).
258 ECONOMETRIC METHODS
where
F ae (e,ex 8 e’e)/q
e’e/(n — k)
is the sample statistic, defined in Eq. (6-8), for testing the null hypothesis
H,: RB=r
When H, is true, A = 0, and F has the central F distribution with g, n — k degrees
of freedom. The test of H, is made, as we have seen, by comparing the sample F’
with a preselected critical value from the central F distribution. The basic result
of Toro-Vizcarrondo and Wallace (1968) is that when Hy is not true, the F
statistic, defined above, follows the noncentral F distribution with degrees of
freedom g, n — k and noncentrality parameter A, defined in Eq. (6-95). Thus the
test for the improvement in MSE is to compare the sample F statistic with a
critical value from the noncentral F distribution with X = 0.5. Critical points of this
distribution are tabulated in Wallace and Toro-Vizcarrondo (1969). The practical
procedure is as follows:
1. Compute the usual F statistic, based on the difference in the residual sums of
squares from the restricted and unrestricted regressions.
2. If F > F(q,n — k)oos, say, in the table by Wallace and Toro-Vizcarrondo,
reject the hypothesis that the restricted estimators are better in MSE. If the
sample F is less than the critical value, use the restricted estimators.
The above procedure is for the strong MSE criterion, embodied in Eq. (6-93).
Wallace (1972) has shown that the weaker MSE criterion (6-94) will be satisfied if
AS5Z
Noncentrality parameter
A=0 A = 0.5 A=q/2
Strictly speaking the term specification error covers any mistake in the set of
assumptions underpinning a model and the associated inference procedures, but it
has come to be used particularly for errors in specifying the data matrix X.f
There are two problems involved in specifying X. The first is knowing which
variables (such as income, relative prices, etc.) to include, and the second is in
what mathematical form each variable is to be included. So far we have blithely
assumed such knowledge to be readily available. In practice it is not. Economic
theory can normally indicate the set of explanatory variables corresponding to
any assumed model (utility maximization, cost minimization, etc.), but theory
cannot usually indicate the precise form of the relationship. In less favorable
situations where there is no clearly articulated theory there may be no clear guide
to relevant explanatory variables. On top of all this one may not be able to obtain
measurements on appropriate variables and, hence, have to use proxy variables in
their place.
To establish the effects of misspecification of X, let us suppose that the true
model is
y=XBp+u (6-96)
+ See H. Theil, “Specification Errors and the Estimation of Economic Relationships,” Review of
the International Statistical Institute, 25, 1957, pp. 41-51.
260 ECONOMETRIC METHODS
with
E(u)=0 and = E(w’) =o7I
The model specified by the investigator is
y=X,B+u (6-97)
where, of course, some variables may be common to both X and X,. The
investigator thus computes the estimated coefficient vector
b, = (XiX,) Xsy
Substituting for y from Eq. (6-96) gives
Case 6-1: Exclusion of relevant variables Suppose that the X, and X matrices are
Rei (X) Xre ex Ay
x= [x, Xp. 0 XX yp x;,| =[X, X,]
The investigator has correctly included the first r explanatory variables but
mistakenly omitted the remaining k — r variables. It follows directly that
M, =1— XX XX,
is a symmetric idempotent matrix of rank and trace equal to n — r. The residual
FURTHER TOPICS IN THE k-VARIABLE LINEAR MODEL 261
sum of squares is
RSS = y’My
Writing Eq. (6-96) in partitioned form as
y = X,B, + X,B, +u
and substituting in RSS gives
= 07(n—r) + BiX5M,X,B,
and so
RSS 1 ros
E| ma | = o2 apw B2X2Mi X28 (6-99)
The matrix of the quadratic form in Eq. (6-99) is the matrix containing the sums
of squares and the cross products of the residual vectors obtained when each
excluded variable in X, is regressed on the set of included variables X,. Apart
from a constant divisor it is a variance-covariance matrix and thus positive
semidefinite, so Eq. (6-99) establishes that the residual variance estimated from
the specified regression of y on X, will, on average, overestimate the true
disturbance variance. As in Eq. (6-98), the bias involves the true but unknown
coefficients of the excluded variables. The bias in the regression coefficients would
disappear if the included and excluded variables were orthogonal, X/, X, = 0, but
the estimated disturbance variance would have expectation
E| RSS ]
]meh=o°+
esta. ye
oerrarBy XX 2B, >o 2
iat,
so that faulty inferences would still be made.
Case 6-2: Inclusion of irrelevant variables The X, and X matrices could now be
specified as
X, = [X, X,]
Xi [X,]
where X, isn X k and X, (the matrix of irrelevant variables) is n < s. When each
true variable in X, is regressed on [X, X,], the least-squares fit will force the
coefficient of that same variable on the right-hand side to unity and all other
coefficients to zero. Thus
ot I
(XEXe) XX = |‘|
262 ECONOMETRIC METHODS
It would seem from the discussion of these two cases that it is more serious to
omit relevant variables than to include irrelevant variables since in the former
case the coefficients will be biased, the disturbance variance overestimated, and
conventional inference procedures rendered invalid, while in the latter case the
coefficients will be unbiased, the disturbance variance properly estimated, and the
inference procedures will be valid. This constitutes a fairly strong case for
including rather than excluding variables from a regression equation. There is,
however, a qualification to this view. Adding extra variables, be they relevant or
irrelevant, will lower the precision of estimation of the relevant coefficients. This
point has already been illustrated for a simple model in the previous section on
multicollinearity. Suppose the true model is
y= Bx, +u (6-100)
and the assumed model is
y = Bx. + Byxz4+ u (6-101)
FURTHER TOPICS IN THE K-VARIABLE LINEAR MODEL 263
var(.b,) = Eee Ak
yx3(1 = 13)
whereas the correct sampling variance, under Eq. (6-100), is o*/Lx3. More
generally if X, indicates the set of true explanatory variables, X, the set of
irrelevant variables, and if y were regressed just on X,, the variance matrix for the
estimated coefficients would be
o*(X,X,) | (6-102)
When y is regressed on [X, X,], the variance matrix for the coefficients of the
variables in X, is
Case 6-3: The general case The general case relates to the mistaken use of the X,
matrix instead of the X matrix, as specified in Eqs. (6-96) and (6-97). The residual
sum of squares from the regression of y on X, is
e’e = y’Myy
where
M, = 1— X,(X4X4)
X%
Substituting for y from Eq. (6-96) gives
e’e = (XB + u)’M, (XB + u)
= uM,u + B’X’M,XB + 2B’X’M,u
and
the variance-covariance matrix computed from the residual vectors obtained when
each variable in X is regressed on X,, is positive semidefinite. Thus the expected
value of the residual variance computed from the regression of y on X, will
exceed o7 and would only fall to «? when X, = X. This provides a rationalization
for the common practice of searching among regressions to find the minimum
residual sum of squares (or maximum R?’), though, of course, in any specific
application sampling fluctuations might yield a lower residual sum of squares for
X, than for X.
The result obtained in Eg. (6-98) that specification error leads to biased
estimates of the population parameters must be interpreted with care. Suppose,
for example, that y indicates observations on the rate of inflation, X the set of
explanatory variables in a “fiscalist’” theory of inflation, and X, the set of
explanatory variables in a “monetarist” theory. A fiscalist will estimate Eq. (6-96)
and a monetarist Eq. (6-97). Monetarists will have little interest in the “news”
that their monetary coefficients are biased estimates of the coefficients of fiscalist
variables, nor would fiscalists be interested in the reverse information. Even if one
model really is the “true” model, the substantial correlation existing among
economic data may well help the “wrong” theory to put up a reasonably good
statistical showing. We are touching on the very difficult problem of the choice
between alternative models, which we will discuss in some more detail in Chap.
12?
PROBLEMS
X=[X, X,]
where X, is n X k, and X, is n Xk. Show that the upper left-hand block in (X’X)~' may be ©
expressed as
7 —1
(X,M,X,)
where
M, =I XK, (X5X5)Xs
Give a least-squares interpretation of M,X, and hence of X{M,X,.
6-2 The following estimated equation was obtained by OLS regression using quarterly data for 1958
to 1976 inclusive:
Ex}; = 12 Lx;X3; = 8
Ex3; = 12 LyiX2; = 10
Yy7=10 Ly, x3, = 8
(a) Estimate B,, 83, their standard errors, and R?.
(6) Test the hypothesis that B, + B; = 1.
(c) Suppose now that you wish to impose the a priori restriction that 8, + B; = 1. What is the
least-squares estimate of 8, and its standard error? What is the value of R? in this case? Compare
these results with those obtained in (a) and comment.
(UL, 1979)
6-5 A set of cross-section data on family income y and expenditure c is partitioned into subsets of
observations, relating to families headed by:
1. Manual workers
2. Salaried workers
3. Self-employed
A regression of log c on log y is computed for each subsample and for the full sample, yielding:
A
B : s? 10
Here B is the slope coefficient (standard errors in parentheses), s* is the residual variance, and T is the
sample size.
Test the hypotheses that:
(a) The elasticity of c with respect to y is the same for all occupational classes.
(b) Its value is unity.
Interpret your results and give some possible explanations for the observed differences.
(UL, 1979)
266 ECONOMETRIC METHODS
var(by) = 0°{(X’K)'
—(XX) 'R'[R(X’X)'R’] 'R(X’X)'}
ox = = | a
aye | Y,
as|}=|%- Y,
a} Yerdi
where Y, denotes the mean value of Y in the ith educational class.
6-8 Rework the estimation problem based on the data in Table 6-6, using any other cell as the starting
position and confirm that one obtains the same numerical estimates of the expected number of hours
as those given in Table 6-7.
6-9 Prove the result on MSEs stated in Eq. (6-86).
6-10 Derive the result given in Eq. (6-91).
6-11 The set of restrictions RB = r, with appropriate partitions of R and B, may be reformulated as
R\B, + R2B, =r
where R, is g X q and nonsingular and R, is q X (k — q). Show that the restricted estimator b,,
defined in Eq. (6-5), may be obtained in two stages, namely:
(a) Regress the vector (y — X,Rj 'r) on the matrix (X, — X,R; 'R,) to obtain an estimate b, of
Bo.
(b) Substitute this estimate in
B, = RK;'(r — RB.)
to obtain an estimate of B}.
CHAPTER
SEVEN
MAXIMUM LIKELIHOOD ESTIMATORS AND
ASYMPTOTIC DISTRIBUTIONS
Chaps. 5 and 6 have set out the main features of the k-variable linear model. It is
very important to emphasize that the results obtained so far depend upon the
particular set of assumptions made in specifying the model. It will be helpful to
review those assumptions and results very briefly as this sets the stage for the
remainder of the book, which is concerned with the many problems that arise in
econometrics when various assumptions underpinning the simple model that we
have considered so far have to be revised and extended.
The k-variable linear model with n sample observations was specified as
y=X 8B
+ U4
yen ey ay d
(n X 1)(n X k)(k X 1) (nx1)
with two crucial sets of assumptions, namely, assumptions about the X matrix
and assumptions about the disturbance vector u, that is,
or
3. u~ N(0, 071)
267
268 ECONOMETRIC METHODS
The combination of assumptions 1 and 2 yields the result that the OLS
estimator b = (X’X)~!X’y with var(b) = 07(X’X)' is a best linear unbiased
estimator of B. The development of inference procedures required an assumption
about the form of the distribution of the disturbance term, and the combination
of assumptions 1 and 3 resulted in a comprehensive set of exact, finite sample
inference procedures—tests of coefficients, confidence intervals, analysis of vari-
ance procedures, tests of structural change, and so forth.
The above assumptions are very restrictive, and parts of Chap. 6 examined
some issues relating to the X matrix. Sections 6-3 and 6-4 indicated various
applications resulting from the incorporation of dummy variables among the X’’s.
Section 6-5 examined the problems that arise when the X variables are highly
correlated, and Sec. 6-6 discussed the problems involved in specifying the X
matrix, that is, in knowing which variables in what functional form should
comprise the columns of X. None of these issues violates the basic assumption
that X was nonstochastic. It is, however, very important to relax this assumption.
Also important is the relaxation of assumptions about the disturbance term. We
will see in Chap. 8 that many real-world situations would preclude var(u) from
having the extremely simple form set out in assumption 2, and it is important to
develop appropriate estimators for these more complicated situations. We also
need to ask what are the effects of removing the normality assumption for the
disturbance term. Finally we note that when X is nonstochastic, there is no
question of any statistical dependence between the X’s and the u’s, but when the
nonstochastic assumption is removed, this now becomes a possibility to be
investigated, and, in fact, this particular problem has generated some of the major
developments in econometric theory.
In tackling this broader range of complex problems, the least-squares princi-
ple alone cannot always yield an appropriate estimator. We need, therefore, to
introduce the powerful maximum likelihood principle. Furthermore, in many of
the new problems it proves excessively laborious and often impossible to derive
exact finite sample results, but it is possible to derive results which hold in the
limit, or asymptotically, as the sample size becomes infinitely large. Thus we need
a simple introduction to asymptotic theory, and this is attempted in Sec. 7-2,
followed by an introduction to maximum likelihood estimators in Sec. 7-3.
Convergence in Probability
A basic result in elementary statistics states that, if the x’s have been drawn at
random from some distribution with mean p and variance o?,
o2
Ee =o Toands var) ae
Thus x,, is an unbiased estimator of » for any sample size, and the variance tends
to zero as n increases indefinitely. It is then intuitively clear that the distribution
of X,, whatever its precise form, becomes more and more concentrated in the
neighborhood of pu as n increases. Formally, if one defines a neighborhood around
pf. as p + «, the expression
Pri — e < X, < po + 2) = Pri|x, — p| < 2}
indicates the probability that x, lies in that interval. The interval may be made
arbitrarily small by suitable choice of e. Since var(x,,) declines monotonically with
increasing n, there exists a number n* and a 6 (|8| < 1) such that for all n > n*,
Pri{|x, — p| <e}> 1-8 (7-1)
The random variable x, is then said to converge in probability to the constant wp.
As n increases, the probability of x, lying in a specified interval becomes larger,
that is, 6 becomes smaller. Thus an equivalent statement is
lim Pr{|x, — p| < e}= 1 (7-2)
plim x, =p (Ges)
where plim is an abbreviation of probability limit. The sample mean is then said
to be a consistent estimator of the population mean p. By a similar argument the
reader may easily show that, in the two-variable regression, b, is a consistent
estimator of B, since it was shown in Chap. 2 that
These two examples are very simple in that the estimators are unbiased for all
sample sizes. Suppose we have another estimator, m,,, of . such that
Cc
270 ECONOMETRIC METHODS
lim E(m,) =
n— oo
plim(x~!) = (plim
x)7!
ol) fate
whether or not x and y are independently distributed. Probability limits may also
be extended to vectors and matrices. It simply means taking the probability limit
of each element of the vector or matrix, provided of course that such probability
limits exist. Operation with these probability limits is again extremely simple. For
example,
plim(AB) = plimA - plimB
plim(A~!) = (plim A)!
As an illustration, recall the OLS coefficient vector from the basic model in Chap.
5, namely,
b = (X’X) 'X’y
= B + (X’X) Xu
=e (2xx) (2x)
ae
The matrix
consists of the mean squares and mean cross products of the explanatory
variables. If the X matrix is constant in repeated samples, thent
im (Lxex) = (ex]
n>o \Nn n
If the explanatory variables are stochastic, it can be shown that the sample
moments will converge in probability to the population moments. Thus we write
whit)
plim(—5X.,u,
The element inside the first parentheses is 7. Since E(#) = 0 and var(z) = o*/n,
it follows that plim(#) = 0. For the ith element
E(—EX,u,] = 0
n
which holds both for the case where X is fixed and also for stochastic X on the
assumption of zero covariance with u. Also
In view of Eq. (7-5), the probability limit of £X7/n is a constant. Thus the
probability limit of the variance is zero. Repeating the argument for the other
terms,
et ielee t \es
plim(—X'u} = 0
and so
: NEES Niwan We e e
plimb = B + plim(—xx| plim(—-X’u
=B+2°'-0
=8
which proves the consistency of the OLS estimator.
Convergence in Distribution
Return again to the sample mean x, If the population from which the x’s are
drawn at random may be characterized by
x ~ N(p, 07)
then x,,, being a linear combination of normal variables, has a normal pdf. Thus
2 07
x,~N [H,al
and f(xX,,) is normal for every n. The limiting distribution is found by examining
what happens to f(X,,) as n goes to infinity. Since var(X,,) goes to zero, the whole
mass is concentrated on the point p in the limit and the distribution is said to be —
degenerate. A simple transformation of x,, however, can lead to a limiting
distribution which does not collapse on a single point. Consider
Zn Vk)
Clearly, E(z,,) = 0 and var(z,) = 0”. Thus
f(z,)
is N(0,67) — foranyn
that is, the limiting distribution and all finite sample distributions are identical
since the parameters of the distribution do not involve n.
The real application of these ideas comes in situations where finite sample
pdf's either cannot be derived at all or are very difficult to derive and manipulate,
but a tractable limiting distribution can be obtained. The limiting distribution may
then be taken as an approximation for the unknown or intractable finite sample
distribution. As an illustration, suppose the random variable X has mean p. and
variance o7, as before, but the distribution of X is no longer normal. A
MAXIMUM LIKELIHOOD ESTIMATORS AND ASYMPTOTIC DISTRIBUTIONS 273
Thus irrespective of the form of f(x), the limiting distribution of z, 1s still normal,
though the quality of the approximation to any finite pdf will be influenced by the
extent to which f(x) departs from normality. Alternative ways of expressing this
result are
ae AN (1 =| (7-7)
with
o2
asy var(X,,) = an
ae ox an Dia ae Sa sua
P (2102)"/* P hes
and so
Equation (7-9) involves both the observations on y and the unknown parameters
B and o*. Writing p(y) in the form L(y; B, 0”) soe neees that it is the
probability cap for the y’s, given the parameters B and o? . Alternatively,
writing it as L(B, 0°; y) stresses that for given y it can be fecarticd as a function
of the parameters. It is termed the likelihood function and is conventionally
denoted by the symbol L.
The ML principle is to choose as estimators of B and o? the values which
maximize the likelihood function, given the sample data y. Letting 0’ = [B’ 0°]
denote the vector of unknown parameters and 6 the ML estimator, 6 is obtained
as the solution of the equation
OL
00
=0 (7-10)
In practice the derivation of the ML estimators is often simplified bymaximizing
the log of the likelihood function, that is, by finding 6 as the solution to
a(InL ys
(7-11)
305% i
Since
A(inL) 1 OL
00 Taeieo
the same vector 6 is obtained as the solution to Eqs. (7-10) and (7-11) for any
BO;
Taking the natural logarithm of the likelihood in Eq. (7-9) gives
apes
(In L)
Se 2X’y
/
+ 2X’xB) = 32l (XY, — XXB)
ywA\)
=
—
7-12
AO)= 5 + al ~ XB)
d(In L)
0
(y-xB) =1 an x8) = 1)
The ML B is seen, in this case, to be idence with the OLS b. The estimate of 07,
however, differs from the unbiased s* of Eq. (5-57) by the factor (n — k)/n,
which illustrates the fact that ML estimators are nos necessarily unbiased. In une
application B is an unbiased estimator of B, but 6? is a biased estimator of o?
276 ECONOMETRIC METHODS
Properties of ML Estimators
ML estimators have a number of desirable properties, some of which hold for
finite samples and some of which only hold asymptotically. Of the finite (small
sample) results one of the most important is the following:
L(6ly) = T(x)
Let 6 denote an unbiased estimator . Then the Cramer-Rao theorem states
Thus the MVB, for any 6,, is given by the ith element on the principal diagonal of
R-'(8).+
As an illustration of this result, let us return to the k-variable linear model.
The first-order derivatives of the likelihood function were given in Eq. (7-12).
Differentiating these again we obtaint
OF (in Tey, Mal,
ORORT g2
a(InL)_ n__ (y— XB)‘(y
—XB)
(02) 204 o°
iD
-1{ 2002)
dB op’
_ Ly
0
2 2
ee |eME) eZ ea
d(02)” 2o0 o° 204
since
since
[8 OAC XX) 0
R-!
\ 4= 204 (7-17)
oO 0 ia
n
We see immediately that the ML (OLS) estimator of B attains the Cramer-Rao
MVB, since var(B) = var(b) = 07(X’X) !,which is identical with the top left-hand
+ Derivations of the Cramer-Rao MVB may be found in P. G. Hoel, op. cit., pp. 362—365;. L. D.
Taylor, op. cit., pp. 209-213, and M. G. Kendall and A. Stuart, op. cit., Chap. 17.
+ Note that Eq. (7-12) contains B and a? because the first-order derivatives had been equated to
zero. We now ignore the equalities, replace B by B, o* by o”, and differentiate again.
278 ECONOMETRIC METHODS
submatrix in R~!. The same result does not hold for either the OLS or the ML
estimator of 0”. The OLS estimator is
3l (ee)
,
FpNe 2 kaka)
Thus
2 0° 2
SS 7a ae (aah)
Recalling that the variance of a x* variable is equal to twice its number of degrees
of freedom,
O 20°
var(s?) = Rien, Es
which, for any finite n, is somewhat greater than the variance term given in R~!.
There is, in fact, no unbiased estimator of o” which can attain the MVB. The
derivation of the variance of the ML estimator, 6? = e’e/n, is left as an exercise
for the reader, but in any case it is a biased estimator of or
A second important feature of ML estimators is their invariance property,
which holds for any sample size and may be stated as follows:
plim6 = 6 (7-18)
§ ~ AN(0,R-') (7-19)
where R has already been defined in Eq. (7-16) as
d7InL
a =2 30 00”
+ Reference may be made, for example, to Kendall and Stuart, op. cit., for a comprehensive
statement of the underpinning assumptions and the derivation of this and other results.
MAXIMUM LIKELIHOOD ESTIMATORS AND ASYMPTOTIC DISTRIBUTIONS 279
Thus the ML estimators, besides being consistent and asymptotically normal, are
efficient in that the asymptotic variance matrix reaches the Cramer-Rao lower
bound.
In this section we will relax two of the assumptions underpinning the k-variable
linear model, namely, the normality of the disturbance term and the nonstochas-
tic nature of the X matrix.
Nonnormal Disturbances
Let us retain assumptions | and 2 of Sec. 7-1, that is,
but dispense with the assumption of a normal distribution for the u’s. Under
assumptions | and 2 the OLS b is still a best linear unbiased estimator of B with
variance matrix o*(X’X)~'. Moreover, as already shown, b is a consistent estima-
tor of B. Thus even when the disturbances are nonnormal, OLS is still a very
acceptable technique for deriving point estimates. The difficulty is that the various
exact inference procedures outlined in Chaps. 5 and 6 are no longer strictly valid
since their derivation depended on the assumption of normality. However, one
may conjecture that the procedures are reasonably robust for moderate depar-
tures from normality.t More importantly the tests can be given a large sample
justification. This requires the use of two theoretical results.
First, if X is nonstochastic of full column rank k, E(u) = 0, var(u) = o7I, the
elements of X are uniformly bounded, and lim,_,,(1/n)X’X = &, a finite,
symmetric, positive definite matrix, thent
1
- (X’u) ~ AN(0, 72) (7-20)
n
+“... it has been shown that these tests are not very sensitive to departure from normality. If the
errors are not normally distributed but have a variance, it is generally true that only trivial errors are
made in the powers or the levels of significance if we retain the formulae which are strictly applicable
in the case where the errors are normal.” E. Malinvaud, Statistical Methods of Econometrics, 2nd ed.,
North-Holland, Amsterdam, 1970, p. 99. See also Malinvaud’s discussion on pp. 296-302. Additional
references on this topic are P. Schmidt, Econometrics, Marcel Dekker, 1976, pp. 55—64; A. C. Harvey,
The Econometric Analysis of Time Series, Wiley, New York, 1981, pp. 112-117; and G. G. Judge,
W. E. Griffiths, R. C. Hill, and T. C. Lee, The Theory and Practice of Econometrics, Wiley, New York,
1980, Chap. 7.
+ For a proof see P. Schmidt, op. cit., pp. 56-60.
280 ECONOMETRIC METHODS
Second, let (X,,, Y,,} denote a sequence of pairs of random variables, where X,,
has a probability limit and Y, a limiting distribution, that is,
plim X, =c
and
ee
then
»—p=(4xx) (4xu} n
maak
n
Thus
vn (b — B) = [Exx) [xu]
vn
This is seen to be in the form of Eq. (7-21), and a direct application of Eq.
(7-22) gives the result that Vn (b — B) has a limiting normal distribution with zero
+ See C. R. Rao, Linear Statistical Inference and Its Applications, Wiley, New York, 1965, pp. 101
ff., for this and other important limit theorems.
MAXIMUM LIKELIHOOD ESTIMATORS AND ASYMPTOTIC DISTRIBUTIONS 281
Coy Sy =o S|
Thus we may write
Stochastic X Matrix
The explanatory variables in an econometric relation are not usually subject to
control by the economic researcher, the secretary of the U.S. Treasury, or anyone
else. Rather they are mostly the outcome of the functioning of some
economic/social system. Let us, therefore, characterize the X;, (i = l,..., k;
t = 1,..., n) as possessing some multivariate density function g(X). We will make
two crucial assumptions about this density function, namely:
independent of all past values, the current values, and all future values of all
explanatory variables. This strong assumption would be violated if a lagged
value of Y, such as Y,_,, appeared among the explanatory variables, for u,_,
influences Y,_,, so Y,_, is not independent of u,_,. Furthermore u,_,
influences Y,_,, which in turn influences Y,_,. Thus Y,_, is dependent on
U,_1,U,_7,---, but is independent of u,, u,,,, and all later u’s.f
3. u~ NO, 071
Thus the log likelihood becomes
8 o*E(X’X)' 0
Rela = 0 20+ (7-25)
ae
so that the asymptotic variance matrix for B(= b) is o7E(X’X)~', which is the
MVB.
Turning to the small sample properties it is easy to show first of all that b is
an unconditionally unbiased estimator. From Chap. 5,
Thus
E(b) = B + E{(X’X) 'X’u)
= B+ E{(X’X)” 'x’\ - E(u) since X and u are independent
=B since E(u) = 0
For the variance-covariance matrix
var(b) = E{(b — B)(b — B)’}
= E{(X’X) 'X’uu’X(X’X) |}
= E,{ Eyx(X’X)
u|x "X’uu’X(X"X) "|
where E,,,, indicates the expected value in the conditional distribution of u given
X, and £, indicates the expected value in the marginal distribution of X.+ Thus
var(b) = E,{0?(X’X)")
= 0°E(X’'X)' (7-26)
This shows that the finite sample variance attains the Cramer-Rao lower bound
and differs only from the corresponding formula in the stochastic case in that
(X’X)~! is replaced by E(X’X)~!.
Formula (7-26) may be established in an alternative and instructive fashion.
It was shown in Eq. (5-33) that the variance matrix for b, given some X matrix,
was o°(X’X) |. Emphasizing the conditional nature of this variance we can write
var(b|X) = 07(X’X)|
or letting S = X’X,
var(b|S) = o*S~!
Now suppose that the random X’s can, in principle, generate a finite number of S
matrices S,, S,,..., S,, with probabilities p,, p,,..., p,,- Then the unconditional
variance matrix for b, determined from the marginal distribution, is
o°E(X’X) |
To summarize the position so far, when the X’s are stochastic but indepen-
dent of the u’s, the ML (= OLS) estimators for the fixed X case are still ML for
the stochastic case, and thus all the conventional test procedures are still justified
asymptotically. For finite samples the OLS (ML) b(B) is unconditionally unbi-
ased, and var(b) attains the Cramer-Rao MVB. Moreover the conventional
estimator s*(X’X)~' is an unbiased estimator of that MVB.
The only remaining question concerns the finite sample validity of the
conventional inference procedures. The basic result is that all confidence interval
statements and significance levels derived from the usual formulas are still correct,
but the probabilities of type IJ errors and the widths of confidence intervals will
be different. Confidence intervals and hypothesis tests are derived by calculating
probabilities from the sampling distribution of some appropriate test statistic
under H,. The test statistic is in general some function of the sample observations
and may be denoted by f(y, X). If, for example, the test statistic is found to follow
the ¢ distribution, then we may make the probability statement
Pr{ —t, 2 < tly,
X) < ta} = 1-4 (7-27)
This statement may be used to derive a confidence interval or, equivalently, to
determine acritical region for a test of Hp at the a level of significance.
The statement in Eq. (7-27), however, has been derived under the assumption
of a fixed X matrix, and so it is a conditional probability statement. We need to—
find what unconditional probability statement can be made about f(-) when the
stochastic nature of X is allowed for. Let A denote the event
A: = bestia anete 7,
and let us suppose that the pdf for X gives a finite number of X matrices,
Xsse e
with associated probabilities
Thus a confidence interval computed in the usual way will have the same
confidence coefficient for random X as for fixed X, and a hypothesis test will have
the same significance level in each case. The assumption of a discrete distribution
for X is only a simplification, and it is clear that the argument carries through for
continuous distributions.+ Notice that it has not been necessary to derive the
sampling distributions of the estimators to establish the above result. These
distributions will be more complicated than those already obtained in Chap. 5. As
an illustration consider the ML estimator ® of the slope coefficient B in a
two-variable model. Letting x denote the vector of sample observations on the
explanatory variable, we know that the conditional distribution of B, given x, iS
f(Blx) = {a2
See
When xis stochastic, the marginal distribution of is
PROBLEMS
7-1 Derive the mean and the variance of the ML estimator 6” = e’e/n of the disturbance variance for
the regression model y = XB + u with u~ N(O, 071). (Hint: If w ~ x?(r), then E(w) =r and
var(w) = 2r.)
7-2 Prove that the OLS estimator s* = e’e/(n — k) and the ML estimator of o7 in the regression
model of Problem 7-1 are both consistent.
7-3 Suppose that we have n independent observations y,, y2,.--, ¥,, Say, incomes, drawn by simple
random sampling from a Pareto distribution which has the following pdf:
10,000°
P(yla)=————_—y = 10,000; « > 0
y
where x, is nonstochastic and ¢, €,...,€, are independently and identically distributed. The
distribution of e, is
f(e,) = Ne OT? eos eee
+ A complete derivation of this and other results is given in F. A. Graybill, Theory and Application
of the Linear Model, Duxburg Press, Mass., 1976, chap. 10.
286 ECONOMETRIC METHODS
Suppose A is not known. Set up the likelihood function for y,, y3,..-, ¥, and describe a way to obtain
ML estimates of a, B, and X.
(University of Michigan, 1977)
7-5 Consider the uniform density
f(X) = lfa 0<X<a
What is the ML estimator for a? Does it attain the Cramer-Rao lower bound? Compare the
asymptotic efficiency of the ML estimator for a with the alternative estimator derived from using the
sample mean, and prove the consistency of that estimator for a.
(University of Chicago, 1975)
CHAPTER
EIGHT
GENERALIZED LEAST SQUARES
1. To indicate some of the more important cases in which assumption (8-1) may
not be fulfilled
2. To determine the properties of OLS estimators if they are (perhaps inad-
vertently) applied, even though the underpinning assumption about the
disturbances is not valid
Ww . To develop tests of whether assumption (8-1) has broken down
Similarly if Y denotes profits and X is some measure of firm size, the same
property is to be expected. The specification of the disturbance variance matrix
would then be
a; 0 0
E(u’) =) 0) tops PO (8-2)
Qu 0 0.
which is the standard case of heteroscedasticity. Formulation (8-2) still assumes
that the disturbances are pairwise uncorrelated.
Suppose, to take a different example, that an investigator is studying the
relationship between wage change and the level of unemployment and that he
measures wage movements in terms of four-quarter overlapping changes. That is,
the annual rate of wage change in quarter ¢ is specified as
WwW, — W-4
W,—4
where w, is the level of the wage index in quarter t. The observed change in the
index, w, — w,_4, is the result of some groups securing a wage change in quarter
t — 3, some others in quarter ¢t — 2, and so forth. If one assumes one fourth of the
labor force to secure a wage change in each quarter, the dependent variable is an
average of these separate changes, and the disturbance term in the macrowage
equation is similarly an average of the separate quarterly disturbances and so
might be specified as
Dea (eee They ete eat) (8-3)
where the e’s indicate the disturbances in the wage change equations for the
separate groups. Let us assume
E(e)=0 and E(ee’)=o/1 (8-4)
It then follows from Eq. (8-3) that |
E(u?) = 402
which does not depend on f, so that the {u,} series is homoscedastic. Further
E(u,u,_\) = 0,
E(u,u,_>) = %60,
E(u,u,_3) = 766,
and
E (uu) =0. = forse
=4
Thus the variance matrix following from Eqs. (8-3) and (8-4) is
1
E(uu’) = raed
(8-5)
GENERALIZED LEAST SQUARES 289
This again is a departure from Eq. (8-1), but in contrast with Eq. (8-2) there is just
one unknown in Eq. (8-5), namely, 02. This is an example of autocorrelated
disturbances. The autocorrelation arose from temporal aggregation over individ-
ual disturbances, which were themselves uncorrelated over time.
To continue the wage change model, it was customary in many early studies
of the Phillips curve for researchers to employ four-quarter overlapping changes.
These, however, have some unfortunate side effects. In consequence it is now
more customary to specify a model of the form
(1+ pL+p'L?+---)e,
+ Note that we retain our convention of using u to indicate the disturbance term in Eq. (8-6). It
does not indicate the unemployment rate.
¢ The lag operator may be treated as a scalar for purposes of algebraic manipulation. For any
nonzero constant a we have (1 — a) '=1+a+ a* +--+. Replacing a by pL gives the result
stated above.
290 ECONOMETRIC METHODS
that is,
u, =e, + pe, + pe,» + °°" (8-8)
Squaring both sides of Eq. (8-8) and taking expectations,
02
var(u,) = E(u?) = ; age lp| <1 (8-9)
The right-hand side does not involve ¢, thus the {u,} series has a constant
variance, 0” = o7/(1 — p’).
Using the definition of u, in Eq. (8-8) and the properties of e, assumed in Eq.
(8-4), it is simple to establish that
E(u,u,_1) = po?
E(u,u,_>) a p°o*
and, in general,
If p were known, this expression, like Eq. (8-5), would involve only one unknown.
There are many other ways, as we shall see later, in which nonspherical
disturbances may arise, but these three examples illustrate some important
patterns. The general nonspherical disturbance matrix may be specified as
E(uv’) = 0° (8-12a) |
or E(uu’) = V (8-125)
The choice of specification depends on whether or not we wish to single out an
unknown scalar, which multiplies all the elements in the matrix as in Eqs. (8-5)
and (8-11). In either case, since we are dealing with a disturbance matrix, 2 and V
are assumed to be positive definite matrices.
b =8 + (X’X) 'X'u
Thus E(b) =8B
so that OLS is still unbiased. The variance matrix is given by
var(b) = E{(b — B)(b — B)’}
E{(X’X) 'X’uu’X(X’X) '}
O2( XX) XOX). (8-13)
Thus the conventional formula o7(X’X)~' no longer measures the sampling
variances of the OLS estimators, and any application of it is potentially mislead-
ing. More importantly, even if one could use Eq. (8-13) to estimate the sampling
variances, the substitution of these numbers in the conventional ¢ formulas and
confidence interval formulas is strictly invalid since the assumptions used in
deriving those inference procedures no longer apply. For the same reason the
optimal minimum variance property of OLS no longer holds. We will illustrate
these points for various specific departures from spherical disturbances in later
sections and also discuss various specific tests for departures from spherical
disturbances. Now it is more important to turn to the development of a more
appropriate estimator.
Comparison of Eqs. (8-16) and (8-17) shows that the appropriate T is given by
haw
and it easily follows that
Om [= P’- Ip- ]
a (P~ ea 1
= TT
Applying OLS to Eq. (8-14) then gives
b, = (X’'T’TX) 'X’TTy
= (x‘2~-'x) 'xQ-'y (8-18)
with the variance-covariance matrix given by
var(by) = 07(X’Q>'X)| (8-19)
The estimator b, is defined to be the generalized least-squares (GLS) or Aitken
estimator. Since Eqs. (8-15) and (8-16) imply that Eq. (8-14) satisfies the assump-
tions required for the application of OLS, it follows that by, is a best linear
unbiased estimator of B in the model y = XB + u with E(uw’) = 07Q.
Alternatively Eq. (8-15) may be written
E(Tuu’T’) = TVT’
and setting TVT’ = I gives T’T = V_' so that the GLS estimator may also be
written as
b, = (XV
'X) Ux’v7y (8-20)
with
L = p(y|X) = ae
DR
(27) o"|Q|'7
xp — <5(y— XB)@""(y — xB)
+ Note that the transformation matrix T, defined in Eq. (8-16), differs by a scale factor from that
defined by TVT’ = I.
GENERALIZED LEAST SQUARES 293
b, = (X-'X) 'x/Q-'y
as in Eq. (8-18).
An unbiased estimator of o* may be derived from the application of OLS to
Eq. (8-14). It is
ee (Ty — TXb,)’(Ty — Txb,)
n—k
_ (y
Xb,)’T’T(y
— —Xb,)
n—k
_Lae
yQ°'ty — bx’ Q"'y
ae (8-22)
On the assumption of normality for the disturbance term all the inference
procedures of Chaps. 5 and 6 carry through for this model. Thus the test of
Hy: RB=r
is based on
having the F(q, — k) distribution under the null hypothesis, where b, is the
GLS estimator defined in Eq. (8-18) and s” the variance estimator defined in Eq.
(8-22).
The above formulas are only operational if the elements of Q are known. In
some exceptional cases this may be so, but in most practical cases it is not. We
must therefore proceed to the development of operational procedures for such
cases, but there is, in fact, no single procedure which is generally applicable. One
must look for the procedure which is best suited to the features of each specific
problem in turn, and that is done in the remaining sections of this chapter.
8-4 HETEROSCEDASTICITY
where the u, are homoscedastic with zero covariances. However, suppose we only
have access to data which have been averaged within m groups, where n, indicates
the number of observations in the ith group. The form of the model appropriate
to the data is now
and clearly
var(u;) = — i=1,...,m
Thus
a 0 0
ta
a7 0 = 07 eae Sn as 0 (8-23)
nS
Os0 =n
where Q is known and the GLS estimator can easily be computed.
Example 8-1 We have taken the same X, Y data as in Example 2-1, only now
it is assumed that they relate to group means. The n, column indicates the
number of observations in each group. The overall means are easily computed
from
Seen)
X= oa mes0 he 4.04
ea 00
Y= a ras) ae 8.00
which are almost identical with the simple means of 4 and 8 in Table 2-1. We
assume that Eq. (8-23) is the appropriate assumption about var(w), that is,
var(u) = 0?Q = 0” ie
GENERALIZED LEAST SQUARES 295
Thus
ny 12
Nn» 6
Qu! = = 1]
10
£5 11
It may then be seen that
ee acy 1 Ny me
EK cae dl ds > :
ee X, - cries
Ns 1 X,
un, &n,X,
un,X, &n,X?
i
and X’2>'y = es
un, X.Y,
Formula (8-18) for the GLS estimator now simplifies to
ae es rn;X,Y,
which is a form of weighted least squares. Applying the data from Table 8-1
gives
506,4 + 2026,, = 400
202b, + 1254b,4= 2388
with solution b,, = 0.88 and b,, = 1.76. To obtain the sampling variance of
these estimates, substitute for Q~' from Eq. (8-23) in Eq. (8-22) to obtain for
Table 8-1
2 4 2 24 48 48 96 192
3 7 6 18 42 54 126 294
1 3 11 11 33 1] 33 99
5 9 10 50 90 250 450 810
9 Iba 11 99 187 891 1683 3179
this example
bn iY,
(n os k)s? aa ny? as ar by «|
un, X.Y,
ns,
Thus the data of Table 8-1 could have been recorded as
the experiment. For dosage X,,n, plots are chosen, and Le iawn)
denotes the resultant set of n, yields: A model for the linear effect of fertilizer on
yield would then be specified as
ys Oe BX u,, Lesa of ole oy ert i (8-24)
Denoting the vector of disturbances in the ith application by u;, we make the
conventional assumptions
E(uj)=0 and E(uwi)=0671, i=1,...,m (8-25)
Thus Eg. (8-25) allows the disturbance variance to be different in different
applications, but assumes homoscedasticity and zero covariances within applica-
tions. However, an additional assumption is now required to cover the relations
between disturbances in different applications. We assume these to be uncorre-
lated, that is,
E(u.) = 0 Dare
ee loser: (8-26)
The complete model may now be written
Yi x, uy
¥
=| X
ial+
Ihe mM)
(8-27)
where
ee
X,=
nex,. ieee mn
ie
A more compact form of Eq. (8-27) is
y=XBp+u
where y’ = [y; y; °°: ¥,,], and so on. Assumptions (8-25) and (8-26) produce
a block-diagonal form for var(u), namely,
a I,,, 0 oss 0
var(u) = 0 vg ves 0 (8-28)
‘ neta: : a
Notice that each X, submatrix has only unit rank, since the same dosage is
applied to all plots within the group, but the X matrix has full column rank.
Model (8-27) is a special case of a more general model, which may be written
yi XxX, u, |
Se eS Efe (8-29)
Yin xXm um
298 ECONOMETRIC METHODS
Test for the equality of variances. In the case of replicated data, model (8-27), and
u, ~ NCO, GLa) a standard test for the equality of variances is available. The
hypothesis of homoscedasticity is
Ay: o¢=o0;=--- =O"
The test is conducted as follows.
where v; = n; — 1, and
n;
Y. = ie ii;
U Nn;
where
va 2 (apa)
i=1
and the quantity
Q’ = vins* — = vIn s?
i=l
Under the null hypothesis Q’ will be approximately distributed as x*(m — 1).
However, the approximation will be improved by dividing Q’ by the scaling
+ The zero covariances incorporated in Eq. (8-28) may be an oversimplification for this model. See
Sec. 8-6, Sets of Equations, and Sec. 10-3, Pooling of Time-Series and Cross-Section Data.
+See M. G. Kendall and A. Stuart, The Advanced Theory of Statistics, vol. 2, Griffin, London,
1961, pp. 234-236.
GENERALIZED LEAST SQUARES 299
constant
] elem!
C=
ane y; |
to give
eee
2="C
. If Q > xho5(m — 1), say, then the null hypothesis would be rejected at the 5
percent level of significance.
Example 8-2 The data of Table 6-1 do not fit this test exactly since the X
variable is not constant within each class. However, we will ignore this
discrepancy and use the Y data of Tables 6-1 and 6-2 to illustrate this test.
We have pv; = 4 for each i and
p= y= 16 m =4
=|
Thus
where
b, = (X/X,) . 'X/y,
1. Fit the OLS regression of y on X and obtain the vector e of OLS residuals.
2. From e compute
n 2
2 ee ete:
n
and the series
ee
Saris faile sen
6
3. Specify the variables in the vector z,. Notice that the functional form h(-) in
Eq. (8-30) does not have to be specified, merely the variables in the linear
combination z,a. Then fit the regression of g, on z’, and compute the
explained sum of squares (ESS) from the regression.
4. The quantity Q@ = ESS/2 is, under the null hypothesis, asymptotically distrib-
uted as x*(p — 1). Thus if Q > x{5(p — 1) one would reject the hypothesis
of homoscedasticity at the 5 percent level.
The Goldfeld-Quandt test. An especially simple and finite sample test, which is
applicable if it is thought that one of the X variables is the basic explanation of
7 T. S. Breusch and A. R. Pagan, “A Simple Test for Heteroscedasticity and Random Coefficient
Variation,’ Econometrica, vol. 47, 1979, pp. 1287-1294.
GENERALIZED LEAST SQUARES 301
The power of the test will depend, among other things, on the number of
central observations excluded, and will clearly be small if c is too large (so that
RSS, and RSS, have very few degrees of freedom) or too small (so any possible
contrast between RSS, and RSS, is reduced). A rough guide is to set c at
approximately n/3.£
The Glesjer test. None of the previous tests yields any specific estimate of the
form of heteroscedasticity which could then be inserted in var(u) to help derive
the GLS estimator. A test which helps in this direction is that due to Glesjer.§ It
is suggested, however, only for the case where a single variable Z is presumed to
determine the heteroscedasticity. The Z variable may, of course, be one of the
explanatory X variables in the structural relation. The test proceeds as follows.
+S. M. Goldfeld and R. E. Quandt, “Some Tests for Homoscedasticity,” Journal of the American
Statistical Association, vol. 60, 1965, pp. 539-547; or S. M. Goldfeld and R. E. Quandt, Nonlinear
Methods in Econometrics, North-Holland, Amsterdam, 1972, Chap. 3, for a more general discussion.
+See A. C. Harvey and G. D. A. Phillips, “A Comparison of the Power of Some Tests for
Heteroscedasticity in the General Linear Model,” Journal of Econometrics, vol. 2, 1973, p. 312.
§ H. Glesjer, “A New Test for Heteroscedasticity,” Journal of the American Statistical Association,
vol. 64, 1969, pp. 316-323.
q Alternatively one might use e? as the dependent variable.
302 ECONOMETRIC METHODS
As it stands, this relation is nonlinear in 6), 6,, and A. Glesjer suggests trying
regressions for a few specific values of h, such as 1, —1, +. The estimated slope
coefficient 5, is then used to test the hypothesis that 4, is zero, although the
conditions required for the validity of the usual significance test will not, in fact,
be satisfied by this regression.t Acceptance of Hy: 5, = 0 implies homoscedas-
ticity and its rejection, heteroscedasticity.
1. Example 8-1 has illustrated a simple case of GLS estimation for grouped
data, where the Q matrix was known.
2. Another simple case occurs where one of the explanatory variables de-
termines the heteroscedasticity. This may have been determined by a Glesjer
type regression or postulated on a priori grounds. Suppose the heteroscedas-
ticity is modeled by
0, = 0° X;, Ca en al (8-32)
Xj ) oar 0
VAT(
41) =O 1s Ogee N pas meonn
Osman O Xe
1
xee 0 0
T = 0 val
ies 0.
(8-33)
and the inference procedures of Chap. 5 could then be validly applied to the
transformed variables in Eq. (8-33). Notice, however, that B; is estimated by
the intercept in the transformed relation, and the original intercept By is
estimated by the coefficient of 1/ X;. Equivalently, the diagonal matrix
Ql = diag ek ss
Lies eke
may be inserted along with the original y, X data in Eqs. (8-18) and (8-19).
3. In cases 1 and 2 we have assumed the elements in the Q matrix to be known
exactly. In many realistic cases these elements have to be estimated and the
estimates then substituted in the GLS formulas. This is sometimes referred to
as a two-stage Aitken estimator (2SAE), or as a feasible GLS (Aitken)
estimator. For example, in the case of replicated data the within-group,
sample Y variances could be estimated and substituted in Q. Or if a
Glesjer-type assumption postulated
0, = 6) + 6X,
and these parameters were estimated from
2 = 8) + 6,X, + error
the estimated disturbance variance matrix would be
var(u) = diag(S, + §,X,,6, + 8,X,,..., 6) + 8,X,) n
where X, takes on the values 1, 2, 3, 4,5. Let b denote the OLS estimator of 6 and
b, the GLS estimator and let us assume further that the nature of the hetero-
scedasticity is
0,Dig = 0°X; eee:
+ For a summary and detailed references see G. G. Judge, W. E. Griffiths, R. C. Hill, and T. G
Lee, The Theory and Practice of Econometrics, Wiley, New York, 1980, Chap. 4.
+ For the condition under which there is asymptotic equivalence see P. Schmidt, Boman.
Marcel Dekker, New York, 1976, Chap. 2, or H. Theil, Principles of Econometrics, Wiley, New York,
1971, Chap. 8. Unfortunately these conditions need to be checked out for each specific application.
304 ECONOMETRIC METHODS
mene)
var(b)
=aie (8-34b)
Thus in this illustration, the efficiency of the OLS estimator ranges from 56 to 83
percent of the GLS estimator, depending on the postulated range for the
heteroscedasticity. Finally, we may note that in the heteroscedastic case, and in
other cases where GLS estimation is appropriate, there is no unique measure of
goodness of fit. A measure may be based on weighted sums of squares, using On!
(or V_') as a weighting matrix, or on sums of squares of the transformed vector
Ty, although the latter is inappropriate if the transformed relation does not have
an intercept term. For details of these and other measures the reader should
consult the article by Buse.
8-5 AUTOCORRELATION
Definitions
The autocorrelation, which is the focus of this section, is that of the {u,} series.
There may or may not be autocorrelation in the explanatory variables, but for the
moment we are only concerned with possible autocorrelation in the disturbance.
term. When present, it results in some or all of the off-diagonal terms in the var(u)
matrix being nonzero. This in turn destroys the optimal properties of OLS and
gives rise to another application of GLS.
We assume, as usual, zero mean for the series, that is,
E(u,) =0 for all ¢
The autocovariance at lag s is defined by
y,= E(u,u,,) s=0,+1, +2,... (8-35)
At zero lag we have simply the constant variance of the series
Yo = Eu? = of
The autocorrelation coefficient at lag s is defined by
+ A. Buse, “Goodness of Fit in Generalized Least Squares Estimation,” The American Statistician,
vol. 27, 1973, pp. 106-108.
GENERALIZED LEAST SQUARES 305
We note that the y’s and p’s are symmetrical in 5 and have been assumed to be
independent of the ¢ subscript, that is, these coefficients are constant over time
and depend only on the length of lag s. The variance matrix for the disturbance
term may then be written as
Yo a Y2 haem ele
var(u) =|] ¥ Yo 7% Wes
Yn=1 Yn-2 nes Yo
l P| P, Pas
a a P) | P| se pt Pn—2 (8-37)
Pn—1 Pn —2 Pn—3 |
u, = f(u, U3)
Uy = f(u,, U5, U4, Us)
and so on, and we would have some nonzero terms in the off-diagonal positions in
var(u).
Estimation of var(u) as in Eq. (8-37) from any finite sample is impossible
since the number of unknowns exceeds the number of observations, nor is any
relief afforded by increasing the number of observations, as it brings a concom-
itant increase in the number of unknowns. The practical procedure is to secure a
reduction in the number of unknown parameters by postulating some structure
Qs
A
1.0
UES) =
U, a pu,_| ty e; || = ]
where the coefficient of the lagged term is denoted by ¢ as we now wish to use p
to denote an autocorrelation coefficient. This is a first-order AR(1) process, and
the result already established in Eq. (8-10) gives us the autocorrelation function
(ACF) of the process as
p,=¢ s=0,1,2,... (8-38)
Thus the autocorrelations decay exponentially and will oscillate in sign if ¢ is
negative. The graph of the autocorrelation function is called the correlogram, and
a typical correlogram for the AR(1) process, with positive @, is shown in Fig. 8-2.
The AR(2) process is defined as
+ A stationary process has a constant and finite mean and variance and a set of covariances which
are independent of time and are functions only of the lag length.
+G. E. P. Box and G. M. Jenkins, Time Series Analysis Forecasting and Control, revised edition,
Holden-Day, San Francisco, 1976, p. 58.
GENERALIZED LEAST SQUARES 307
Se
ca 8-4]
P| 1 = 5 ( )
=o, > +
$1 >
Pr = ie (8-42 )
2 ie (1 cs >) ‘ o,.
i (8-43)
u
-@3]
-¢)°1
(1 + @)[( B
rewritten as
u,=(1+ OL )e,
giving (1+ OL) ‘u, =e,
or (1 — 6L + 67L? —---)u,
=e,
or = Ou), — Oreo Ou ee tre.
which is an infinite AR process with the restriction that the coefficients are given
by the successive powers of 6 with alternating signs. Similarly the AR(1) process
may be written
(1 — oL)u, = &,
or u,=(1-$L)'e,=(1+@L
+ @L? +---)e,
or Ue Er Dey + ge, 5+ se
Va = a :
HE tess
and
es ] if units i, j are contiguous
? 0 otherwise
In this formulation p is a scalar indicating the overall strength of the autocorre-
lation, and the weights w,, are essentially dummy variables which allow any
disturbance to be affected by contiguous disturbances. The matrix formulation of
7 See, for example, R. L. Martin, “On Spatial Dependence, Bias and the Use of First Spatial
Differences in Regression Analysis,” Area, vol. 6, 1974, pp. 185-194.
GENERALIZED LEAST SQUARES 309
Eq. (8-49) is
= eWut+e (8-50)
and for Fig. 8-1 the W matrix would be
10g SF 0:0 0
pb 20 go
Pe ee Oba Ome
DSTA ie, Osada
Oe Omi ac)
Os Oe ea tw) m0
From Eq. (8-50) the variance matrix of the disturbance term is
var(u) = o2(I — pW) ‘(I — pW)’! (8-51)
In Eq. (8-51) W is generally a known matrix, but p and o? are unknown scalars.
However, we will not consider the resultant estimation problems here.+
|
|
|
|
|
| |
| L >~X
x; Xx) Figure 8-3
Q='\) p
Si
I Dian
: cer sie ah ae oe: : (8-54)
GENERALIZED LEAST SQUARES 311
l =i 0 Me 0 0 0
—p 1+ 9° a) see 0 0 0
ees) te i et ne DE ete
0 0 0 —p l+p”? —-p
0 0 0 0 = ]
Substituting Eq. (8-54) in Eq. (8-53) for the model of Eq. (8-52) gives
var(OLS b)
== f2 eal n—-1
Dek oae =o
Sees Wake
ep tee =e
t=1%7 ai ae yy
(8-56)
If p were known and GLS was applied to Eq. (8-52), then substitution of Eq.
(8-55) in the general formula
var(b,) = 62(X’Q-'X)"'
gives
and negative p can moderate the dramatic declines in efficiency shown in the
right-hand side of the table. These calculations are, of course, only illustrative, but
they indicate the possibility of a serious loss in efficiency if OLS is applied in the
context of autocorrelated disturbances.
A second problem with the application of OLS is that the conventional
formula on the computer for var(b) will, in this example, estimate o7/L7x?,
whereas Eq. (8-56) shows that this is no longer the true variance. As the sample
size gets very large, the ratio of the conventional formula to the true variance is
given by
Leak
eae Ok
Thus the proportionate bias that the conventional program will impart to the
estimation of the true sampling variance of the OLS 6 is, in the limit,
—20N
Asymptotic proportionate bias = rear (8-59)
Table 8-3 shows values of this statistic for selected values of p and A. Again it is
instructive to consider the table in two halves. A positively autocorrelated
disturbance in conjunction with a positively autocorrelated {x} series implies
underestimation of the sampling variance by the conventional OLS formula. If
p = A = 0.9, the estimated variance will only be about one-tenth of the correct
number, which would cause a serious overestimation of t statistics and significance
levels in conventional inference procedures. On the other hand, different signs for
p and A will cause an overestimation of the sampling variance. Casual empiricism
p
rN — 0.9 S05) a2) 0.2 0.5 0.9
0 0 0 0 0 0 0
0.2 43.9 22D 8.3 ale) SUS SNS
0.5 163.6 40.0 DoD, Sal Se — 40.0 SOD)
0.9 852.6 163.6 43.9 30:5 = 6271 SOO)
GENERALIZED LEAST SQUARES 313
n n 2 n 2:
es 1x; ak, ee }
E(e’e)
,
= o;2, [xvee
Fete
aN
If p gue \ have the same sign, then s* will have a downward bias as an estimator
of [Link], for instance, p = 0.9 = A and n = 101,
E(s*) = 0.9150;
Thus when p and J have the same sign, this bias accentuates the bias analyzed in
Table 8-3. It is clear that autocorrelated disturbances are a potentially serious
problem, and it is very important to be able to test for their existence.
Thus even if the null hypothesis is true, so that E(uu’) = 071, the OLS residuals
will display some autocorrelation, for the off-diagonal terms in M do not vanish.
More importantly M is a function of the sample values of the explanatory
variables, so that it is impossible to derive an exact finite sample test on the e’s
which will be valid for any X matrix that might ever turn up.
Durbin-Watson test. These problems were treated in a pair of classic and path-
breaking articles.t The Durbin-Watson test statistic is computed from the vector
of OLS residuals e = y — Xf. It is denoted in the literature variously as d or DW
and is defined as
eee =. ext
a=
8-61
Biaaes (
Figure 8-4 indicates why d might be expected to measure the extent of first-order
autocorrelation. The mean residual is zero, so the residuals will be scattered
around the horizontal axis. If the e’s are positively autocorrelated, successive
values will tend to be close to each other, runs above and below the horizontal
axis will occur, and the first differences will tend to be numerically smaller than
the residuals themselves. Alternatively if the e’s have a first-order negative
correlation, there is a tendency for successive observations to be on opposite sides
of the horizontal axis, so that first differences tend to be numerically larger than
the residuals. Thus d will tend to be “small” for positively autocorrelated e’s and
“large” for negatively autocorrelated e’s. If the e’s are random, we have an
in-between situation with no tendency for runs above and below the axis or for
alternate swings across it, and d will take on an intermediate value.
7 J. Durbin and G. S. Watson, “Testing for Serial Correlation in Least Squares Regression,”
Biometrika, vol. 37, 1950, pp. 409-428; vol. 38, 1951, pp. 159-178.
GENERALIZED LEAST SQUARES 315
(a)
If the sample value of d exceeds 2, we wish to test the null hypothesis against
the alternative hypothesis of negative first-order autocorrelation. The appropriate
procedure is to compute 4 — d and compare this statistic with the tabulated
values of d, and d, as if one were testing for positive autocorrelation. The
original DW tables covered sample sizes from 15 to 100, with 5 as the maximum
number of regressors. Savin and White have published extended tables for
6 <n < 200 and up to 10 regressors.} The 5 percent and 1 percent Savin-White
tables are reproduced in App. B-S.
There are two important qualifications to the use of the Durbin-Watson test.
First it is necessary to have included a constant term in the regression. Second, it
is strictly valid only for a nonstochastic X. Thus it is not applicable when a lagged
dependent variable appears among the regressors, and indeed it can be shown
that the combination of a lagged Y variable and a positively autocorrelated
disturbance term will bias the Durbin-Watson statistic upward and thus give
misleading indications.t Even when the conditions for the validity of the Durbin-
Watson test are satisfied, the inconclusive range is an awkward problem, espe-
cially as it becomes fairly large at low degrees of freedom. A conservative
practical procedure is to use d,, as if it were a conventional critical value and
simply reject the null hypothesis if d < d,. The consequences of accepting Hy
when autocorrelation is present are almost certainly more serious than the
consequences of incorrectly assuming it to be absent, which is one reason for the
procedure.§ Second, it has been shown that when the regressors are slowly
changing series, as many economic series are, the true critical value will be close
to the Durbin- Watson upper bound.
When the regression does not contain an intercept term, d is bounded by
dy Sas,
where d, is the upper bound of the conventional Durbin-Watson tables.
Farebrother has provided extensive tabulations of both lower and upper 1 percent
and 5 percent significance points for d,,.|
+N. E. Savin and K. J. White, “The Durbin-Watson Test for Serial Correlation with Extreme
Sample Sizes or Many Regressors,” Econometrica, vol. 45, 1977, pp. 1989-1996.
¢ M. Nerlove and K. F. Wallis, “Use of the Durbin-Watson Statistic in Inappropriate Situations,”
Econometrica, vol. 34, 1966, pp. 235-238.
§ A comprehensive Monte Carlo study relevant to this question is J. K. Peck, “The Estimation of a
Dynamic Equation Following a Preliminary Test for Autocorrelation,” Cowles Foundation Discussion
Paper, no. 404, September 9, 1975. After studying the properties of regression estimators following
different significance levels for d, the author recommends using a significance level much more
likely
(than the conventional levels) to reject Hy when it is true. This is in the same spirit as
using d,, as the
critical value.
| H. Theil and A. L. Nagar, “Testing the Independence of Regression Disturbances,”
Journal of
the American Statistical Association, vol. 56, 1961, pp. 793-806; and E. J. Hannan
and R. D. Terrell,
“Testing for Serial Correlation after Least Squares Regression,” Econometrica,
vol. 36, 1968, pp.
133-150.
|| R. W. Farebrother, “The Durbin-Watson Test for Serial Correlation when There
Is No Intercept
in the Regression,” Econometrica, vol. 48, 1980, pp. 1553-1563.
GENERALIZED LEAST SQUARES 317
where the e’s are the usual OLS residuals. Wallis derives upper and lower bounds
for d, under the assumption of a nonstochastic X matrix. The 5 percent
significance points are tabulated in App. B-6. The first table is for use with
regressions with an intercept, but without quarterly dummy variables. The second
table is for use with regressions incorporating quarterly dummies. As shown in
Chap. 6, one may employ a constant term and three quarterly dummies or use
four quarterly dummies without a constant term.
Further significance points at 0.5, 1.0, and 2.5 percent levels are provided by
Giles and King.§ The same authors also point out that if one is testing Hy) against
the alternative hypothesis H,: 9, < 0, the test statistic 4 — d, may be correctly
referred to the critical values 4 — d, ,, and 4 — d, ,, where d, ,, and d, , are the
5 percent values tabulated by Wallis, only in the case where seasonal dummies
have been included among the regressors. For the case where an intercept but no
seasonal dummies have been employed, these critical values are inappropriate and
the authors provide a revised set.
+™M. L. King, “The Durbin-Watson Test for Serial Correlation: Bounds for Regressions with
Trend and/or Seasonal Dummy Variables,” Econometrica, vol. 49, 1981, pp. 1571-1581.
+K. F. Wallis, “Testing for Fourth Order Autocorrelation in Quarterly Regression Equations,”
Econometrica, vol. 40, 1972, pp. 617-636.
§ D. E. A. Giles and M. L. King, “Fourth-Order Autocorrelation: Further Significance Points for
the Wallis Test,” Journal of Econometrics, vol. 8, 1978, pp. 255-259.
q M. L. King and D. E. A. Giles, “A Note on Wallis’ Bounds Test and Negative Autocorrelation,”
Econometrica, vol. 45, 1977, pp. 1023-1026.
318 ECONOMETRIC METHODS
Durbin tests for a regression containing lagged values of the dependent variable.
As has been pointed out, the Durbin-Watson test procedure was derived under
the assumption of a nonstochastic X matrix, which is violated by the presence of
lagged values of the dependent variable appearing among the explanatory vari-
ables. Durbin has derived a large sample (asymptotic) test for the more general
case.} Consider the relation
1. Fit the OLS regression denoted by Eq. (8-66) and note var(b,).
2. From the residuals compute ror, alternatively, if the Durbin-Watson statistic
has been computed, we may use the approximation r =~ 1 — d/2.
3. Substitute in the formula for A, and if h > 1.645, reject the null hypothesis at
the 5 percent level of significance in favor of the hypothesis of a positive
first-order autocorrelation.
4. A similar one-sided test for negative autocorrelation can be carried out for
negative h.
The test breaks down if it should happen that n - var(b,) > 1. Durbin showed
that an asymptotically equivalent procedure is the following.
1. Estimate the OLS regression of Eq. (8-66) and obtain the residual e’s.
2. Estimate the OLS regression of
et one r= A yee ee, st
+ L. G. Godfrey, “Testing Against General Autoregressive and Moving Average Error Models
when the Regressors Include Lagged Dependent Variables,’ Econometrica, vol. 46, 1978, pp.
1293-1302; and T. S. Breusch, “Testing for Autocorrelation in Dynamic Linear Models,” Australian
Economic Papers, vol. 17, 1978, pp. 334-355.
320 ECONOMETRIC METHODS
a 1&;
bee De ie
ESS=e'[E, xX] sc
E,
GD
Using e’X = 0, this simplifies to
Thus l= Eas
62
2. Regress €, on €,_,,..., € t—p? and x, (that is, the ¢th row of X).
Since there are only n values of e available, this regression might be carried
out using only the last n — p observations. The Breusch-Godfrey procedure sets
€o, @_},--- at zero. Asymptotically it does not matter which route is taken, and it
is a moot point whether it matters in finite samples.
1— OF) 50 Or 0)
T= —p ene 0 an0
0 —p 1 Ot, 50
ND Ace eae ae
Ty 0X, i x x a2
Y; — pY, eee a(1 — p) eA :
: =i fal X; — px. B tiles (8-73)
Me oi Pa
iL :
En
so that only n — 1 transformed observations are used in the OLS estimation. The
variables in Eq. (8-73) are sometimes referred to as quasi first differences, and
the intercept term being estimated is now a(1 — p). The variance matrix of the
disturbance term in Eq. (8-73) is an (n — 1) X (n — 1) matrix,
var(e) = (1 — p?)o7I
Application of T to Eq. (8-72) gives the transformed model
67 21 Vere
x2 —p
| 1 ji-pe jl-p-X, a é
ere
= en Nee la] + ‘i (8-74)
Sonat ere
ep jee : ie En
G5. C,laeece)
for an AR(1) scheme. Thus var /1 —p- u,| =o.
If p were known, GLS estimation could be achieved by applying OLS to Eq.
(8-74), or the process could be approximated by using Eq. (8-73). The difference
between the two procedures can be important when the sample size is small. The
extensions to include additional explanatory variables and higher-order AR
processes are simple. Additional X’s are treated in exactly the same way as the
single explanatory variable in the example. If the disturbance term followed an
AR(2) scheme,
ty ae ee t=3,...,0
Special transformations would also be required for the first two observations.+
The assumption, however, of a known value for p is unrealistic. It is a
parameter to be estimated along with a, 8, and 07. Lagging Eq. (8-72) one period
and subtracting from Eq. (8-72) gives
Y,=a(l — p) + BX, — BpX,_, + pY,_, + «, f= Dome Ween (S275))
The disturbance in Eq. (8-75) satisfies the assumptions required for OLS. How-
ever,
Le? = f(a, B, p) (8-76)
ela Pl,
ay OCIS p) + BCX, — pX,_1+)
&;
and
lat 4X) ipYep Oa Ag) te,
Starting with any value for p, the quasi first differences in the equation of step 1
could be computed, and OLS applied to it would then yield estimates of a and P.
These estimates in turn can be used to compute the Y, — a — BX, series. Regress-
ing this series on itself lagged one period in the equation of step 2 yields a revised
estimate of p, which can then be fed back into the equation of step 1, and the
process continues.
This is known as the Cochrane-Orcutt iterative process, and versions of it are
incorporated in almost all social science computer packages.t There is a variety of
+ For details see F. B. Lempers and T. Kloek, “On a Simple Transformation for Second-Order
Autocorrelated Disturbances in Regression Analysis,” Statistica Neerlandica, vol. 27, 1973, pp. 69-75.
+D. Cochrane and G. H. Orcutt, “Application of Least Squares Regressions to Relationships
Containing Autocorrelated Error Terms,” Journal of the American Statistical Association, vol. 44, 1949,
pp. 32-61.
324 ECONOMETRIC METHODS
starting positions and rules for termination. If the initial value of p is set at zero,
step | is then simply the OLS regression of Y, on X,, which yields the OLS
residuals e, = Y, — a — bX,. In step 2 e, is regressed on e,_,, without an intercept
term, to obtain an estimate r of the first-order autocorrelation coefficient. Alterna-
tively r may be computed from the Durbin-Watson statistic, which is a routine
output in an OLS package, as
r= 1—td
The estimated r is then used to compute the series {Y, — rY,_ ,) and {iy 7 Agen
which are used in a repeat of step 1. The process may be stopped any time the
Durbin-Watson statistic in step 1 indicates random residuals. This frequently
occurs after one complete iteration. Alternatively one can stop the process after
successive estimates of the parameters differ by less than some prescribed amount.
It is clear that step 1 in the Cochrane-Orcutt process is the use of model
(8-73) and the associated transformation matrix T,. A modification of the process
is to use model (8-74), where: the first term gets explicit treatment. This is often
referred to as the Prais-Winsten method.} The Prais-Winsten modification may be
expected to improve the efficiency of the estimation, especially in small sample
sizes. Yet another modification is to use a method suggested by Durbin for
obtaining the initial estimate of p.t This is to fit Eq. (8-75) by OLS without
worrying about the nonlinear restriction and take r as the coefficient of Sere
Monte Carlo study by Griliches and Rao suggests that a two-step estimator
consisting of the Durbin estimate of p followed by the Prais-Winsten treatment of
the transformed variables performs somewhat better than any of the other
variants over a fairly wide range of parameter values.§
The two-step Durbin estimator extends easily to more than one explanatory
variable and to higher order autoregressive schemes. For example, suppose the
model is
Yea By + Xa He Bex ee
with u, = >\U,_. + ou,_, + &,
Combining the two equations gives
¥e= $,¥,-, + &Y0 FBX, +: - + + BEX $B, Xo, 7
OUP Xie $2 By Xy4_9 Se $B, X,, 1-2 +1 — 6; - )B, + e;
Let $, and ¢, denote the coefficients of Y,_, and Y,_, when this regression
is '
fitted by OLS. The transformed variables are then computed as
(Ye OY 1oon eal (Xe ey ee
Lie 2s a es bo
and OLS is applied to these to obtain estimates of the B’s.
7 |detT| = 1 — p*
+C. M. Beach and J. G. MacKinnon, “A Maximum Likelihood Procedure for Regression with
Autocorrelated Errors,” Econometrica, vol. 46, 1978, pp. 51-58.
+ See App. A-9, Change of Variables in Density Functions.
$26 ECONOMETRIC METHODS
and so
A more recent study by Park and Mitchell confirms the main findings of
Harvey and McAvinchey and adds some additional findings.§
2. They also investigate how well the various estimators perform in hypothesis
testing by looking at the number of type I errors in 1000 trials at the 0.05
significance level. The results are only reported for positively autocorrelated
disturbances, but the message is very clear. All estimators seriously under-
estimate standard errors, making estimated coefficients appear to be much
more significant than they actually are. This is, of course, to be expected for
OLS, but it is also fairly substantial for two-stage Prais-Winsten (2SPW),
iterative Prais-Winsten (ITERPW), and Beach-MacKinnon ML (BM). For a
sample size of 20, p = 0.8, and GNP as the trending explanatory variable, the
number of type I errors reported are OLS (449), 2SPW (251), ITERPW (246),
and BM (258). These numbers should be contrasted with an expected range
of 37 to 63. Thus it would be advisable to apply more stringent significance
levels than usual in testing coefficients in models with autocorrelated dis-
turbances.
Y= at BX, + u, pata
ee
AT = n—1
; In(277) —
n—-1
Ino? — ales be
7 20, t
and o, are
dln ES le
Le,
0a a,
diInL _ ]
Goel ee oX,_1)&,
Op €
dln L a ee
Lu; = 18;
dp a,
dln L Lk n—-1 1
do2€ 20, DGe
where all summations are over ¢ = 2,..., n. Setting these derivatives to zero and
solving for the parameters gives conditional ML estimators (conditional, that is,
on X, which is taken as fixed). We may note in passing that the first three
equations give \
L(Y, 6Y,,) = (n= la + BEX, 16X,5,)
(x "¥ 6Y,_,)(X, ae 6X,_,) a aX X, A 6X,_,) si pLex aq pXRa)s
p=
By Oe Be Oss BX iy)
Ye Sills Bx
which are the equations of the iterative Cochrane-Orcutt process, the first two
being the least-squares equations on the quasi first differences and the third the
first-order autoregressive coefficient of the estimated residuals.
Turning to the second-order partial derivatives
PnL_ _(n—1)(1-p)
da? 0
07 In L l 2
ad pre Pee)
071 DYE ESM ae
dp” Ce
d*In L Sele!
Se AT yy?
d(02) 20, 0,
07 In L l=)
da Op Etoy ne Exp p Xe)
d7In L l
da dp care G2 ee ai (1 a p)Xu,_,)
Oana
Nieie ] ele De
0a do, 02
GENERALIZED LEAST SQUARES 329
d7In L 1
OB dp ga eet — Ue, — eX.)
eine: 1
dp aoe ae gael %,5 oxX,_;)
07 ln L 1
os
ede
Taking the negative of the expectations gives the information matrix
= oxy)
(n= 1)@= eyo © = p)ELY 0 0
ear ote erie x Li)" 0 0
0 (n= 1)o? 0
ee o2 0
0 0 fia
0
202€
The crucial feature of this information matrix is its block-diagonal nature.
Asymptotically the estimates of the regression parameters a and £ are distributed
independently of the estimate of the autocorrelation parameter p and of the
estimate of 0”. Referring back to Eq. (8-73), the data matrix for this regression is
given by the (m — 1) X 2 matrix X,, where
at l—p 1h". oes 1—p
mee PAS py GSS pay re OG a Xe
with unknown parameters a and B. The 2 X 2 submatrix in R is easily seen to be
X4,X,. Since the asymptotic variance matrix is given by R7' and since this has the
same diagonal form as R, we have
A
Qa
asyvar |=o (X-Xe).
B
which is consistently estimated by the usual least-squares procedures, justifying
the remark at the end of the previous section. We also see from R~! that
1—p°
asy var(6) =
no
remembering that o7 = o7/(1 — p’).
If the model
Y,=a+ BX, +4, uU,= pu,.4 + &,
has been estimated from n sample observations, the best prediction of Y in period
n + 1 is no longer
x, staal =a+t bX, n+1
where a and bare estimates obtained by any of the above methods, since this
prediction sets the disturbance term at zero, and the AR(1) process implies
330 ECONOMETRIC METHODS
E(u,,,,) = eu, Both elements in pu, are unknown, but might be estimated by
ri, =r(Y, —a— bX,)
The suggested predictor is then
Y¥,,,=a+bX,,,+ 1,
n
This, in fact, would be a best linear unbiased predictor if » were known and rset
equal to p since it is the predictor that would emerge from the relation
Ye pYeie SP) HRC ee eee
which may be rewritten as
Y= BX pe re = BAe ae
giving the predictor.}
yen+1 = CDA eee Ue
The expenditure equation (8-79) is nonlinear in the a’s and b’s. However, if the
logarithm of the ratio Z,/Z 18 taken,
M M
ingZ in Z; = A; 4 bn| | - bl Fe + Uj, (8-80)
and = tence,
This is clearly an estimable equation. Given r commodities, there are r(r — 1)/2
such equations, but most are redundant. As an illustration, for commodities i and
k we have
In Z, — In Z, = A,, + btn| | = bain|5] Tate (8-81)
P: P,
Subtracting Eq. (8-80) from Eq. (8-81) gives
In Z,~IZ, = Ay +n 5M -bain|5M J+
a Pi,
_ a,b;
where A, = or
and Os, Ey Ey
Thus of the three possible equations for commodities i, 7, and k only two are
independent. Given any pair of equations, the third follows by subtraction. For r
commodities there are just r — 1 independent equations, and for estimation
purposes one may select any set of r — 1 independent equations. Thus one might
write the system
M, \ M,
In-Z,,— In Z,, = Ay, + 0, In| —— | — dyin Pitti.
; Pi, 12
M, M,
; Pi, P3,
The sample observations on the first equation in Eqs. (8-82) may then be
written as
y, = X,B, + u,
where
| | | ae €11 — a
©]
; M M 12 22
X,=|i in(
5 | In | B,=] 4, u, =
| f
332 ECONOMETRIC METHODS
Likewise, defining Y,, = In Z,, — In Z;,, the sample observations on the second
equation in Eqs. (8-82) may be written as
y. = X,B, + u,
where
oil muses
! z |M ae Eyn 889
X,=|i in|
> | In|5) Bp,=| 5, u,=
P Ps
b;
| | | Ein — ©3n
y; X, B, u;
y xX B u
a oar ec \aolee (8-84)
OF aS
y=XB+u (8-85)
Because of the block-diagonal form of X the application of OLS to Eq. (8-85),
treated as a simple regression, would be exactly equivalent to the application of
OLS to each of the m equations in Eqs. (8-83) separately. However, the applica-
tion of OLS to Eq. (8-85) would not be optimal for two reasons. First of all the u
vector is not homoscedastic. From the structure of the u’s
U6) ie. (me eee
Thus var(u;) = var(e,) + var(e,,,) — 2cov(e;, €,,,)
Even if the original e’s are contemporaneously uncorrelated,
var(u;) = var(e,) + var(e,,)
,
and var(u,) = var(e,) + var(e;,)
Thus the u’s would only be homoscedastic if the e’s were homoscedastic,
but
there is no a priori reason for the disturbance variances in the various expendi
ture
equations to be equal.
A second reason for the nonoptimality of OLS is that the off-diagonal
terms
in var(u) will not be zero.
E(u,u;) x E(e, rh €41)(& Be a)
Suter ye ee iy
Z= |), 97m
the variance matrix for the u vector in Eq. (8-85) may be written}
var(u) = V=Ze@I (8-86)
Thus a set of demand equations should almost certainly be considered as a group
and estimated by GLS because of the nature of the variance matrix of the
disturbance term. In addition, theoretical considerations will impose restrictions
across equations. For example, in the addilog demand system above the second
parameter, 5, in each B, vector is constrained to be equal across all m equations.
This constraint has not been imposed in the specification (8-84). Implementation
of that system would allow a different estimate of the coefficient b, to be made for
each commodity. One may wish to test for constancy of b, across commodities
and to reestimate the system with constancy imposed.+
A second illustration of sets of equations with cross-equation restrictions and
connections between the various disturbances is found in sets of “share” equa-
tions approximated by transcendental logarithmic functions, which has recently
become the dominant methodology in the estimation of various substitution
elasticities, especially in the field of energy economics.§ Consider a production
function
Q=f(X, X%,.-., X,)
where Q denotes the rate of output and X; (i = 1,..., 7) the rate of input of the
ith productive factor. If one assumes the firm to face a given set of factor prices
+ See Eqs. (4-76) and (4-77) for the definition of a Kronecker product and its inverse.
+ For an illustration of the estimation and testing of three different demand systems see R. W.
Parks, “Systems of Demand Equations: An Empirical Comparison of Alternative Functional Forms,”
Econometrica, vol. 37, 1969, pp. 629-650.
§ See, for instance, E. A. Hudson and D. W. Jorgenson, “U.S. Energy Policy and Economic
Growth,” Bell Journal of Economics, vol. 5, 1974, pp. 461-514; E. R. Berndt and D. O. Wood,
“Technology, Prices and the Derived Demand for Energy,” Review of Economics and Statistics, vol.
57, 1975, pp. 259-268; and J. M. Griffin and P. R. Gregory, “An Intercountry Translog Model of
Energy Substitution Responses,” American Economic Review, vol. 66, 1976, pp. 845-857.
334 ECONOMETRIC METHODS
P,,..., P,, one formulation of the firm’s decision problem is to choose the input
mix to minimize the cost of producing a given output Q. This gives rise to a set of
factor demand functions
X= fF(Pie sees
0), > sie
Denoting the optimal inputs by X*, the optimal (minimal) cost level is
c* = PAX =f (Pie On)
oC =r
OP.l
} This result is an application of Shephard’s lemma. (R. W. Shephard, Theory of Cost and
Production Functions, Princeton University Press, Princeton, NJ, 1970, p. 170.) The lemma may be
illustrated for a two-factor production function Q = f(X,, X2). Suppose the firm is required to
produce some stated output Q at minimum cost, given factor prices P, and P,. If we define
Beh
dg
=0 (1)
dp
pr ~f% X%) = 2'= 0
The solution of these equations gives the cost-minimizing factor demands
X* and X*, expressed as.
functions of P,, P;, and Q. The minimum achievable cost is then given by
Ge PX eh Xe (2)
Differentiating Eq. (2) partially with respect to P, gives
SOE ae OI ee
GPa TOP,
Thus Shephard’s lemma requires that
P, 0X; fe OXF ni
OP, OP,
and, similarly, that
OX* OXF ot
Prop, + 2 OP,
Further
ac* P, aa P, X7*
OP Cre rcs
ES dlnC* = PX Lae
Cit? Gt oe
where S; denotes the cost share of the ith factor, that is, the proportion of total
cost absorbed by the ith factor. Since C* depends on the factor prices and output,
the cost shares will be functions of the same variables, that is,
Q=f(K,L,
E, M)
where the inputs distinguished are capital K, labor L, energy E, and materials M.
Assuming constant returns to scale plus exogenous factor prices P,, P,, P;, and
Py, and imposing symmetry on the second-order partial derivatives, gives the
translog cost function
nC =a, + n@e a,inP, +a, in/P, + a,in Peay,in Py,
]
dX, = Aoi dP,
where A is the determinant of the 3 x 3 matrix of coefficients on the left-hand side of Eq. (3). Thus
ox} OX a
P OP, + P, aP, = _ (Ahh aaP,f?)
Differentiating InC with respect to the logs of the prices gives the cost share
equations
Sx = ay + By,in 2, + p,,in P, + Bp, Pe tbe eae
S, = a, + Bein Pe Bein P, + Bplt Pee eee
S; =o; + By,ln PP, +6, ,In P, + Bp, Prt Bay bay
Sy Oy. + Bey
ln Pe By 5,0, 5B py, eee ans ee
Since the shares must sum to unity,
Ay ta, ta; + ay = 1
and the B’s sum to zero in each column (and row). Imposing the rowwise B
constraints on the first three share equations gives the system
P P P
Se ae Bexln(5 . Brctn|5 |a Breln|5 |
P P P
S, =a, + Brctn|5 |+ Balm 5"|a Bretn|5 | (8-87)
Py Py Py
P P P.
S,; = a, + Beetn|5 + Bretn|5 |+ Beeln|5 |
Because of the symmetry in the 8’s there are just nine independent parameters in
this system. Estimation of these, in conjunction with the summation conditions on
the a’s and £’s, will yield estimates of all the coefficients of the cost function
except a.
For the translog cost function the Allen partial elasticities of substitution are
given by
B.,+ SS.
6, = i+]
J S,S;
ee ee
and se fa eles
S2
115 = 8/5;
Since the four shares sum identically to unity, one must expect nonzero contem-
poraneous covariances between disturbances in different equations, and there is
also no a priori reason to expect the same disturbance variance in different share
equations. However, this system differs in one major aspect from the addilog
demand functions in Eqs. (8-82). In Eqs. (8-87) the same set of explanatory
variables appears in each share equation, but that is not true in Eqs. (8-82), and
we will return to the significance of this point below. At the next level
of
disaggregation a production function could be specified for the energy sector with
various specific fuels as inputs and the parameters estimated from a set of energy
cost share equations. There have been many applications of this cost share
approach in recent years. However, a word of caution is required. As the
GENERALIZED LEAST SQUARES 337
derivation made clear, a basic assumption underlying the derivation of the share
equations is that in each observation period in the sample there has been afull
and complete adjustment of the input mix to the factor prices ruling in that
period so that the minimum cost level C* is achieved. This is an implausible
assumption for many production processes, and actual cost shares probably
represent various Jagged adjustments to changing factor prices. The assumption of
instantaneous adjustment is likely to produce seriously biased estimates of the
various elasticities.
by = (X’V~'X) 'x’v-ly
From Eqs. (8-86)
where o'/ denotes the i, jth element in 2~'. Substituting for V~! in the formula
for b, gives
eg Xiy,
Nyx Ry’yX 3 Cee X ee el ein
BS alenee us: Rete WEE) oh een (8-88)
eX X ork X oF en
oxy:
1. Apply OLS separately to each equation in Eqs. (8-83), obtaining the vectors
of sample residuals e,,e,,..., €,, where
e7e;
i
+A. Zellner, “An Efficient Method of Estimating Seemingly Unrelated Regressions and Tests for
Aggregation Bias,” Journal of the American Statistical Association, vol. 57, 1962, pp. 348-368.
338 ECONOMETRIC METHODS
Ty = (TX)B + Tu
where
TT=V'!
Making the appropriate substitution of Ty for y and TX for X in the OLS test
7 In the two illustrative examples the X; matrices had an equal number of columns, but there is no
need to impose such a condition generally. The exposition also assumes an equal sample size in each
regression, but this is merely a simplification and need not be imposed generally.
+ See Problem 8-2.
§ J. Kmenta and R. F. Gilbert, “Small Sample Properties of Alternative Estimators of Seemingly
Unrelated Regressions,” Journal of the American Statistical Association, vol. 63, 1968, pp. 1180-1200.
GENERALIZED LEAST SQUARES 339
(r — Rb,)’[R(X’V-'x)
'R’] ‘(x — Rb,)/q
eV 'e/(n — k)
(8-90)
where b, is the GLS estimator, g is the number of restrictions embodied in the
null hypothesis, and e = y — Xb,. Under the null hypothesis this statistic follows
the F(q, n — k) distribution.
The SURE model specified in Eqs. (8-84) to (8-86) gives a special case of this
Statistic. There are m separate equations with n observations on each, giving
N = mn observations in all. There are k; variables in X,, and the estimation of the
unrestricted model, Eq. (8-84), will thus yield estimates of K = Lk, parameters.
Finally the V matrix has the special form shown in Eq. (8-86). Thus the test
Statistic becomes
bi? — 52 = 0
bi) — 62 =0
Thus R=
S _ o So aS So | = So
a
and
If the null hypothesis is not rejected and one wishes to reestimate the system
with the constraint imposed, one may take the formula for the restricted OLS
estimator, given in Eq. (6-5), and replace X by TX to obtain
Aj)
Aj;
y) i 0-10 x) x, 0 0 Ais u,
Yo) = (0 7h Omex, 0 aX, 0 b, | +]u,] (8-93)
y3 0 Ou x, 0 0 Xl Oe u;
by
b,
1. Compute the s,; from the OLS residuals, as described above, and hence
obtain 2.
2. Compute the elements of =~! and substitute in Eq. (8-88) to compute b,.
Wo. Using by compute a new set of residuals e, = y — Xb,.
4. Partition e, into the subvectors corresponding to each equation and use these
subvectors to compute new s,> thus starting the process over again.
PROBLEMS
8-1 Derive the results on the efficiency of the OLS estimator under the two forms of heteroscedastic-
ity, given in Eqs. (8-34a) and (8-34d).
8-2 Prove that the SURE estimator in Eqs. (8-88) reduces to the application of OLS to each equation
separately if
(a) 9,;; = 0 for all i + j
or
(b) Xj = X,=-:-- =X,,
8-3 Specify the R matrix and r vector for testing the symmetry conditions in the set of equations
(8-87).
8-4 Consider the four cost share equations prior to Eqs. (8-87) and explain how to test the full set of
summation restrictions (on a’s and B’s) and symmetry conditions. Which, if any, of these restrictions
might be satisfied exactly by the estimated coefficients?
8-5 Consider a heteroscedastic model (for which all other classical assumptions hold)
Y,,=a+ BX, + uj; b= Nose (m
> 1)
Geereile),
; Namal
where
ny
ad ey;
n.!
Determine E(s7).
(University of Michigan, 1981)
ie) 2 eee
yi 10 -1 1-1
(a) Find the best linear unbiased estimates of the parameters a and B.
(6) Test the null hypothesis
Hy: a=B8
against the alternative H;: a = B.
(University of Michigan, 1980)
CHAPTER
NINE
LAGGED VARIABLES
We will use the term “lagged variables” to cover the inclusion on the right-hand
side of the regression equation of lagged-values of the explanatory variables, the
X’s, and/or lagged values of the dependent variable Y.
X,
J lege i
Assuming the lag pattern to persist through time, any Y, is seen to be built up as
the sum of effects from current and previous values of X. Thus the lagged effect
assumed above would generate the relation
Y, =p Ts 8yX,at: 5, X,_, + 6, X,_> i 6; X,_3 atUu,
where we have also allowed for an intercept and a disturbance term. In practice
one does not usually have any strong a priori information about the maximum
length of lag, and one formulates the general relation
Y= e+ D(L)X, +4, (9-1)
where D(L) is a polynomial of some degree s in the lag operator, that is,
D(L)=6,+6,L+---+6L' (9-2)
If X has remained constant at some level X for s periods, then, apart from
disturbances, Y will have reached an equilibrium value
Yop
+ D)X
where D(1) indicates the value of the polynomial when L is replaced by unity, and
is simply the sum of the 6, coefficients, namely,
D(1) = y6;
i=0
If X changes in period ¢ by an amount AX, and is then held constant at the new
level, Y will gradually adjust from Y to a new equilibrium. The changes are
Period t Bote Lay eh
Change in Y by AX, 5,AX, 6,AX,
The coefficient 6) (= AY,/AX,) thus represents the impact multiplier for X.
Partial sums of the 6’s indicate intermediate multipliers. The 6’s may also be
standardized by dividing by their sum D(1). Partial sums of the standardized 8’s
then indicate the proportion of the total effect achieved by a certain period. For
example, knowledge of the 6’s enables one to estimate how many periods must
elapse before, say, 90 percent of the total effect is achieved. An important concept
is that of the median lag, which is the number of periods required for 50 percent
of the total effect to be achieved. When all the 6’s are positive, another useful
Statistic is the mean lag defined as
ya 0t0y Oy ees a eas
so) 89 $0, FO, +> + 8
From Eq. (9-2) it is seen that differentiating D(L) with respect to L gives
DL )= 64 20, b--2-= $56,124
Thus Mean lag = DAY)
D(1)
As an illustration suppose an estimated version of Eq. (9-1) yields
D(L) = 0.10 + 0.25L + 0.35L? + 0.15L3 + 0.0514
LAGGED VARIABLES 345
Period 0 l 2 3 4
+See C. E. P. Box and G. M. Jenkins, Time Series Analysis: Forecasting and Control, tevised
edition, Holden-Day, San Francisco, 1976, pp. 53-54.
+Note that(1 — a@,L)~'=1+a,L+a7Ll?4+---.
346 ECONOMETRIC METHODS
D(1) = fy + “ufoA
_ Bot Bi
|e a,
B(1)
Clearly, extending the power of the B(L) polynomial would extend the number
of “free” 5 coefficients before the exponential decline sets in. The mean lag may
also be derived from the A(L), B(L) polynomials. Since D(L) = B(L)/A(L),
DULY BCE) ACE)
DCB) CBRE) Sea)
and so
Koyck scheme of declining exponential weights.f The simple Koyck scheme has
the coefficients on the X’s declining exponentially from the start, that is,
6, = a,6,_, i i cee
This corresponds to the specification
A(L)=1-a@L and B(L)=82,
and the relationship may be formulated as
¥, = p+ OX, + a6) X,_, + a76)X,_. +--+ + u, (9-10)
or, equivalently, as
(1 — a, L)(¥, — #) = BX, + ©,
which may be written
Vo Gl Oy) chee
19 ataby Aaa (9-11)
Equivalence between Eqs. (9-10) and (9-11) requires
By = 85
and v, = (1 — a,L)u, =u, - a,u,_, (9-12)
Thus if the original disturbances {u,} in Eq. (9-10) are serially independent, the
transformed disturbances {v,} in Eq. (9-11) are serially dependent, which has
implications for the estimation procedures to be considered in Sec. 9-2. Since
A(1)=1—a,, A(1) = —a,, B(1) = By, and B’(1) = 0, the mean lag for the
simple Koyck process is a,/(1 — a@,). As has already been indicated, raising
the degree of the B(L) polynomial, while retaining A(L) = 1 — a,L, increases
the number of “free” coefficients before the Koyck exponential decline comes
into play.
So far we have considered the distributed lag effect of just a single explana-
tory variable. Suppose there are two explanatory variables, each with a Koyck lag.
There may be no a priori reason to expect an identical decay parameter in each
lag. Thus the relation might be formulated as
Ye dt BX, Pa BX eo pkey ct oe yz,
+ 05yZ,_, + a3yZ,_, +++ +4, (9-13)
or Ln peereet ee
which gives
YAS pet (a, ta, )Y, we, ¥_5 + BX, 56%, + ¥Z,— o,yZ2,.5
(9-14)
where u* = w(1 — a )(1 — @,)
and OR Uy ioe (a, aH Oy) U,_| + A ,U,_4
so that, compared with the single-variable Koyck scheme in Eq. (9-11), we have
two lagged values of Y and lagged values of each explanatory variable. For
estimation purposes the essential point to notice about Koyck schemes is that
they may be formulated either with only lagged values of explanatory variables on
the right-hand side, as in Eqs. (9-10) and (9-13), or with lagged Ys appearing on
the right-hand side, as in Eqs. (9-11) and (9-14). The former have nonlinear
restrictions on the parameters combined with presumably “well-behaved” dis-
turbance terms, while the latter have a dramatic reduction in the number of
right-hand side variables, but “complicated” disturbance terms and sometimes
restrictions on the coefficients [as in Eq. (9-14) but not in Eq. (9-11)].
Adaptive Expectations
Lagged dependent variables may also appear among the regressors in various
expectational models. A firm may base its production rate Y, not on the current
sales rate X,, but on the expected, permanent, or trend sales rate X*. Thus one
may specify
Y,=a+t BX* + u, (9-15)
where a disturbance u, has been included to allow accidental over- or under-
achievement of the production target. Equation (9-15) is not usually statistically
operational since there is a dearth of published information on expected or
forecast sales rates and similar variables. It is therefore customary to add an
auxiliary hypothesis about the formation of expectations, and one of the most
widely used (if not, indeed, abused) schemes is that of adaptive expectations,
which is that expectations get updated each period on the basis of the latest
information about the actual value of the variable. The formal specification is
Xe Xe = (1 A) CX, NE ee 0 er el (9-16)
In this formulation X* indicates the expectation formed at the end of period ¢,
when the information about the current level X, has become available. If expec-
tations were formed at the beginning of the period, X, in Eq. (9-16) should be
replaced by X,_,. If A = 0 in Eq. (9-16), the expected value adjusts period by
period to the current observation and all previous history is irrelevant. If A = 1,
an expectation, once formed, continues unchanged, irrespective of current or
earlier observations. The intermediate and more realistic case of A being a positive
fraction means that expectations get adjusted each period by some proportion of
the discrepancy between the latest observation and the expectation for that
period. Low values of A imply substantial adjustments in expectations, and large
values imply slowly changing expectations.
Equation (9-16) may be reformulated as
(TL WA "(les Ai)X
aN
or P,bs a To AL (9-17)
Partial Adjustment
Another process which can generate lagged dependent variables among the
regressors is that of partial adjustment. Consider the adjustment of gasoline
consumption to a substantial price rise such as that engineered by OPEC in
1973/1974. Initially the scope for economies in consumption, even in the face of
very substantial price rises, was limited by such factors as
In the short run, economies could be made in shopping and vacation trips,
car pooling on work trips, and so forth. In the longer run, one expects adjust-
ments in the more fundamental factors, such as the fuel efficiency of the vehicle
fleet. Such adjustment has its own costs and, in any case, must take time to be
achieved. Thus one may postulate Y*, the optimal consumption rate appropriate
to a gasoline price of X,, with income and other factors being held constant, as
*-a + BX,
t (9-19)
For reasons such as those suggested one would not expect actual consumption Y,
to adjust completely to X, in period ¢. Instead, a partial adjustment process is
frequently specified as
ii)-(Fe)"~
The partial adjustment process would then have to be specified conformably as
Y y* =I
a ts
The partial adjustment process specified in Eq. (9-20) has been widely used in
applied work because of the simplicity of the resultant estimating equation, such
as Eq. (9-25). Nonetheless it implies a pattern of adjustment that may sometimes
be implausible. Suppose X had been constant at X sufficiently long for Y to have
settled at the desired level, Y= a + BX. In period t we assume X to become
X +AX and then to remain at the new level indefinitely. The new desired Y is
given by Y = a + B(X +AX), and the adjustment to that level implied by Eq.
(9-20) for a A value of, say, 0.5 and a negative B is shown in Fig. 9-1.
In the first period one-half of the total desired adjustment is achieved; in the
second period one-half of the remaining adjustment is accomplished, and so
LAGGED VARIABLES 351
Figure 9-1
forth. Thus the maximum adjustment is achieved in the first period, and each
successive adjustment is a fraction A of the previous adjustment. This might be a
plausible reaction pattern for, say, the consumption of broiler chickens in
response to a significant price change, but it is less plausible for the consumption
of gasoline since that consumption is mediated through durable equipment.
A further difficulty with the simple partial adjustment process arises when Y*
is a function of more than one explanatory variable. Suppose, for example, the
optimal level of energy consumption depends on both the relative price of energy
and the level of output in the economy. Applying the partial adjustment process
to actual energy demand imposes the same adjustment parameter on each
explanatory variable. Even if the form of the adjustment process is similar for
each variable, the speed of the process may well be different. Thus at given prices,
one might expect energy consumption to move more or less in step with output,
but to react much more slowly to price changes.
Y= aap x (9-26)
This is not an operational equation since there are no direct observations on the
variables. However, the adaptive expectations hypothesis may be used to explain
X* and partial adjustment to explain the adjustment of Y to Y*. Thus combining
Eqs. (9-17) and (9-21) with Eq. (9-26) and allowing the A parameter to be different
+ See M. Friedman, A Theory of the Consumption Function, Princeton University Press, Princeton,
NJ, 1957.
352 ECONOMETRIC METHODS
Let us begin with the estimation of the distributed lag function (9-1), that is,
Yi pot 0g ch OX ict On a aes (9-28)
where, for simplicity, we restrict consideration to the lagged values of a single
explanatory variable. We usually cannot expect theory to indicate the maximum
length of lag, but one would ordinarily expect significance tests on the 5’s to give
some indication both of the maximum lag length and of any delay in the initial
transmission of an effect from X to Y. The validity of such significance tests
depends on the properties of the disturbance process {u,} and the associated
estimation methods. If E(u) = 0 and var(u) = o7I, then, in principle, OLS would
be an appropriate estimation technique. In practice, however, its application is
likely to be plagued by collinearity between the regressors, leading to great
imprecision in the estimates of the 6’s.
Almon Lags
A general strategy for dealing with this collinearity and the associated imprecision
is to reduce the number of parameters to be estimated by the assumption of some
pattern for the 6’s. The Koyck scheme of Sec. 9-1 is perhaps an extreme example
of such a pattern. The Almon lag scheme provides a more flexible method for
reduced parameterization.t Under the Almon scheme one rules out the direct
7S. Almon, “The Distributed Lag between Capital Appropriations and Expenditures,”
Econometrica, vol. 30, 1962, pp. 407-423.
LAGGED VARIABLES 353
6) 6)
+— — Le vy
—=6
Si
(a) (2)
Figure 9-2
approach of attempting to estimate all (s + 1) 6’s and assumes instead that the
6’s can be approximated by some function 6, = f(i), as in Fig. 9-2b. The basis of
the approximation is Weierstrass’s theorem, which states that a function continu-
ous in a closed interval may be approximated over the whole interval by a
polynomial of suitable degree, which differs from the function by less than any
given positive quantity at every point of the interval.}
As an illustration suppose we postulate a third-degree polynomial, that is,
CD) =a,+ ait Osh + Ont
Then approximately
8) = f(0) = a
6, =f(l) =a, t+a,t+a,+
a,
6, = f(2) 3 ao aie 2a, ote 4a, ate 8a, (9-29)
+R. Courant, Differential and Integral Calculus, vol. 1, 2d edition, Blackie & Son, Glasgow, United
Kingdom, 1937, p. 423.
354 ECONOMETRIC METHODS
yield estimates of the 6’s from Eq. (9-29). The sampling variances and covariances
of the §’s can be computed from those of the &’s and significance tests carried out
on the 5’s. Defining W, as the matrix of coefficients in Eq. (9-29),
Orne 0
anor | ] 1
Woes | eer.
Byes 27
hy4 vs)
Bs3
where the subscript 3 indicates the use of a third-degree approximating poly-
nomial. Equation (9-29) then becomes
5 = W,a (9-31)
and, given 4,
5 = W,4a (9-32)
The matrix form of the original equation (9-28) is
y=ip+ X8+u
Using Eq. (9-31),
y =in + XW,a + u
An OLS regression of y on [i XW,], where XW, is the matrix of observations on
the “new” regressors shown explicitly in Eq. (9-30), gives the estimated coefficients
| plant ixXw, |-'| ivy
&| | WsX’'i WiX’XW, WiX’y
with
as4 .
var(&) = 02 |W,xX’XW, — TW5X'HXW, | (9-33)
MS, = 0
But Aé; = 6, — 8,_,
A’6, a (4, Oe (On 8-2)
= 0, 20; jt 0,25
A°5, = 6, — 36,_, + 36,_, — 8,_,
A*6, = 6, — 46,_, + 66,_, — 46,_,+6,,
Thus the assumption of a third-degree polynomial places a set of linear restric-
tions on the 6’s. The full set of restrictions is
6, — 46, + 66, — 46, + 6) =0
6, — 46, + 66, — 46, + 6, =0
oP Bese eraehvelel ce. (ef a°h¢. ise) 0/0, ie)is) orled euiem eo)Seuuey (s)fe Ta 6 (9-35)
64S
Ss RY
65.) 4b + 8 =O
(Sy eees WO
=o :
te Ott Ol Py) sar O41 3
=a — a (i
—1) — a)(i — 1)’ — «(i — 1)’
= (a, — a, + a;) + (2a, — 3a3)i + 30,17
The second difference of 5; is found by repeating the first difference operation. Thus
: 2
A*8, = (2a) — 3a3)i + 3a3i* — (2a) — 3a3)(i — 1) — 3a,(i - 1)
= (2a, — 643) + 603i
The degree of the polynomial in i decreases by | with each differencing. The third and fourth
differences are then
M36, = 6a;
and Mos 0
356 ECONOMETRIC METHODS
No restriction involves the intercept term w. Thus the restrictions (9-35) may be
expressed as
R35 |=0 (9-36)
where R, is the.(s — 3) X (s + 2) matrix
0 | = 4 Gea 4 1 0 Sas 0
0 0 bo 4 6 a4 1 vee 0
0
A second-degree approximating polynomial would imply the set of s — 2 linear
restrictions given by
[
R.[5|-°
where R, is the (s — 2) X (s + 2) matrix
Orman Sear ] 0 ee 0
0 Os Ss ] ves 0
R,= : : (9-37)
0
Notice that the nonzero elements in the rows of the R matrices are given by the
appropriate set of binomial coefficients with alternating signs.¢ If r denotes
the degree of the approximating polynomial, the nonzero elements in R, are the
coefficients of L in the polynomial (1 — L)’*', but in reverse order. However,
since the restriction sets linear combinations of the 6’s equal to zero, we can
multiply the rows of R, by —1 and get the coefficients in natural order.
For a given maximum lag s the sequential procedure for finding a suitable
degree of the approximating polynomial would be as follows.
1. Start with a polynomial of fairly high degree, say, the fourth or fifth.
2. Set out the corresponding R matrix and test the null hypothesis
LL
Hy: R|5|=0
l 3 3 ] second degree
] A ate ibs 4 1 third degree
where an internal element in any row is the sum of the pair of elements immediately above.
LAGGED VARIABLES 357
Under the null hypothesis the resultant test statistic has the F(s — r,n — 5 —
2) distribution, where r is the degree of the approximating polynomial.
3. If the null hypothesis is rejected, the initial polynomial has not been of
sufficiently high degree.
4. If the null hypothesis is accepted, proceed to the next lower degree and test
the new set of linear restrictions, proceeding in this way until the null
hypothesis is rejected.
If the null hypothesis is accepted, say, for R, but rejected for R,, the
appropriate procedure is to find a third-degree approximating polynomial. This
may be done by computing the four “new” regressors specified in Eq. (9-30),
estimating the a’s by OLS and then using Eq. (9-32) to estimate the 6’s.
Alternatively one may use the formula for the restricted estimator given in Eq.
(6-5), and inferences may be made by using the variance matrix given in the
footnote to Eq. (6-5).
The above procedure is conditional on some assumed value for the maximum
lag s. It may be repeated for various values of s and a judgment made by looking
at the overall fit and the significance of the higher-order 5’s.
An implication of the Almon procedure, which does not seem to have
attracted much attention, is that it is likely to yield biased and, indeed, incon-
sistent estimates. Write the original model, Eq. (9-28), for simplicity as
y=Xd+u (9-38)
If the 6’s do not lie exactly on the approximating polynomial, then a formula
such as Eq. (9-31) has to be amended to
8=Wa+v (9-39)
where v is an r X 1 vector of errors involved in the use of an rth-degree
approximating polynomial. Notice that v is independent of time and is a vector of
unknown constants, which does not vanish with increasing sample size. Substitut-
ing Eq. (9-39) in Eq. (9-38) gives
y = XWoa + (Xv + u) (9-40)
In Eq. (9-40) there is obviously some correlation between the explanatory
variables XW and the expanded disturbance term Xv + u, which would lead one
to expect inconsistency in the estimation of a and hence of 6. Looking directly at
the estimator of 6,
§ Wa
w(W a W’X’y
w]= w{(5xx}8
1 1
i Xu} n n
Assuming
ee a5
358 ECONOMETRIC METHODS
and
plim|
hi (x
— X’u =0
we have
plim§ = W[W’2,,W] 'W’S,,8 (9-41)
Substitution of Eq. (9-39) in Eq. (9-41) gives
plim§ = 8 + W[W’S,,W] 'W’S_v (9-42)
so that the Almon estimator is inconsistent unless the unknown 6’s lie exactly on
the chosen polynomial, in which case v = 0. The finite sample bias of the Almon
estimator can be serious if one fits a polynomial of too low degree. This bias,
combined with the smaller sampling variation (as compared with unrestricted
OLS), can sometimes give sampling distributions for the Almon estimators which
fail to contain the true 6 parameter altogether or else have it located near an
extremity of the distribution.
Computer packages with Almon lag estimators usually offer the facility of
including end-point restrictions such as 6_, = 0 and/or 6,, , = 0. Since 6_, is the
notional coefficient of X,,, and that variable has no effect on Y,, it might seem
sensible to incorporate that end-point constraint. As Dhrymes and Schmidt and
Waud have pointed out, that is a fallacious argument.} Setting 6_, = 0 implies a
restriction on the a’s and hence on the 8’s, which in turn is a restriction on how
X,, X,_},---, X;_, affect Y,. For a second-order polynomial the implied restriction
is
My — a, tay=0
Such arestriction could, of course, be tested by estimating the a’s and using the
variance matrix in Eq. (9-33). The purpose of the Almon polynomial is to give a.
good approximation to the unknown 6’s over the interval 0 to s. Its behavior if
extrapolated outside that interval is irrelevant. The second end-point restriction,
8, ,, = 0, may not produce much distortion in the approximation if the coefficients
are decaying with increasing lags but, again, it implies a restriction on the a’s and
6’s, and there seems little valid reason for imposing it.
rewritten as
Y=
+ 8(X, 4 0X. + +: + at 1X) + a'8(X
+aX,
, to) +a,
or Y,=pe+ OxX* + a'y + u, (9-44)
where
AP SX aXe aX,
0
Le Aga he u
y=/]1 X¥ a? |/s]+u (9-45)
i. y% oe a" Y
This matrix is symmetric and so only the upper triangular portion has been
shown. The unknown parameters in Eq. (9-46) would be replaced by their
estimated values, and the inverse would give the estimated variance matrix for the
parameters.
(1 — B,L)Y, = B, + BX, + u,
giving
Y= a+ B(X, + BX. + BFX,at <1) +O, (9-48)
where a=
B,
[ees
and o= (1 = BiE\gea,
If X were held constant at some level X and Ydenotes the corresponding level of
Y, then
E(Y )= By B, x
ete ees
provided |,| < 1. If |6;| = 1, E(Y) would explode. In practice a {Y,} series may
have explosive tendencies, which are held in check by various “floors” and/or
“ceilings.” A model of such a process would be highly nonlinear, and the
statistical treatment of such models is still in its infancy. We therefore impose the
constraint
|B3| < 1 (9-49)
LAGGED VARIABLES 361
It is also clear from Eq. (9-48) that expressions such as (LY,7/n) and GAY)
will involve linear combinations of quantities such as
] ]
eke yekmir
] ]
and year pele Ce ere
The additional assumption is then made that the X, are bounded and that the
above quantities have finite limits as n tends to infinity.
The model of Eq. (9-47) may be written in matrix form as
y=ZBp+u (9-50)
where
te eX Y
Leis Y,
+If it is not, the effective sample size is n — 1, and the statistical inference procedures are
conditional on Y, with n— 1 observations, rather than conditional on Y) with n observations.
Asymptotically, of course, it makes no difference.
+ For a complete derivation see E. Malinvaud, Statistical Methods of Econometrics, 2d edition,
North Holland, Amsterdam, 1970, pp. 540 ff.
362 ECONOMETRIC METHODS
B = (ZZ) 'Zy
=B+ (ZZ) ‘Zu
Thus
(b-B)-(52z)
-
vn(B-B) 1
= [|,272} ——el Zu (9-53)
9.53
Using Eqs. (9-51), (9-52), and (7-24) gives
vn (B — B) ~ AN(0, 073;,')
or B ~ AN(B, 02722] ZZ
(9-54)
Thus even without the assumption of normality for the u’s the OLS estimators
will be consistent and asymptotically normally distributed. The unknown variance
matrix in Eq. (9-54) can be consistently estimated by the usual formula s?(Z’Z) ~!.
If, in addition, the u’s are normally distributed, the estimators are also ML and
efficient. These results extend simply to the general case of various lagged Y
values and several X’s. Thus there is substantial justification for the continued
use of OLS in relationships containing lagged dependent variables, provided the
disturbance term is serially independent. The estimators will, however, be subject
to finite sample bias, and one should also recall the problems of testing for
autocorrelated disturbances in this case.
+H. B. Mann and A. Wald, “On the Statistical Treatment of Linear Stochastic Difference
Equations,” Econometrica, vol. 11, 1943, pp. 173-220, especially pp. 185-190.
+ See Sec. 8-5.
§ Note that this is different from the error structure in Eq. (9-11) associated with the Koyck lag;
the latter [an MA(1) error] is considered below.
LAGGED VARIABLES 363
This new assumption has an important effect. From Eq. (9-55) it is seen that u ra
influences u,, but from Eq. (9-47), in period ¢ — 1, u,_, influences Y,_,. This sets
up a dependence between u, and Y,_, in Eq. (9-47), that is,
E(Y,_\u,) = 0
From Eqs. (9-47) and (9-55) it follows that}
po, 2
plim|“ZY, 1}
“T= Be (9-56)
The consequence is that the application of OLS to Eq. (9-47) will yield incon-
sistent estimates of all parameters. This is so because
plim(B) = 6 + =; - plim(—-Z/u]
and
Ae (es
plim yee 0
aL a ecrc i 0
plim(= u| = plim(-EX,u,} = poz
plim(—Y,_.«,] s
] pea
Instrumental Variables
Consider
y=Zpt+u
Premultiply by Z’ to give
Z'y = Z'ZB + Zu (9-57)
The OLS estimator b of Chap. 5 may be obtained from this equation simply
by setting Z’u = 0, giving
Z'y = Z'Zb (9-58)
On the assumption that
In the present model the assumption that plim((1/n)Z’u) is the zero vector
cannot be sustained. Suppose, however, that one can find an n X k matrix W
containing variables which are thought to be contemporaneously uncorrelated
with the disturbance term. That is, we assume
W’y = (W’Z)byy
=B+2,)-0
=B
so that the IV estimator would be consistent.
The variables in W are referred to as instruments. Some of them may simply
be variables from the original Z matrix. In the present model there is no need to
replace X, since it is already assumed to be independent of the disturbance term.
In addition to being uncorrelated with the disturbance term, the instruments
should not be totally uncorrelated with the explanatory variables since W’Z
would then be a null matrix and the estimating technique would break down. If,
in fact, W’Z is “nearly” null, the IV technique will give very poor results.
In Eq. (9-47) we need just one instrument, and it is customary to select X,_ 1
as the instrument for Y,_,. The appropriate matrices are then
1 XX Les
We | Pi lle Xo tvs
Lag ae Dey oer
LAGGED VARIABLES 365
by =B + (WZ) 'Wu
Thus
bp )= [Gwz) [zw
n
+ Should the values Xj and Yo not be available, the first row is dropped from W and X, the
summations run from t = 2 to t = n, and n is replaced by n — | in the formula for byy.
+ Contrast the assertion by P. J. Darymes, Econometrics—Statistical Foundations and Applications,
Harper and Row, New York, 1970, p. 297: “All IV estimators, no matter what the choice of
instruments, are unbiased and consistent.”’ This statement comes after a passage in which the only
explicit assumptions relate to probability limits. The IV estimators are consistent. A possible
explanation of the incorrect assertion about unbiasedness is given in App. A-8, Expectations in
Bivariate Distributions, where the matter is discussed in detail. The same type of error can also affect
the derivation of results about finite sample variance matrices, as in formula (6-4-12) of Dhrymes.
366 ECONOMETRIC METHODS
plim(—W'] = 2,,,
Vn (By — B) ~ AN(0, 6,
2,22
yyBuz)
or |
Maximum-Likelihood Estimator
Combining Eq. (9-47) with the AR(1) disturbance process in Eq. (9-55) gives
Given a starting value for p, the transformed variables in Eq. (9-66a) could be
computed and OLS applied to yield estimates of the B’s. These estimates in turn
could be used to compute the transformed variables in Eq. (9-666) and OLS
applied to produce a revised estimate of p with the iterations continuing till
convergence.
Setting up the log likelihood for Eq. (9-65) and differentiating gives the
information matrix}
2
mda Clem OE XS palGlicyee Ri) 0 0
a DX? E(XX*Y* ,) 0
1
2
Bp. ry*2, oa 0
R B; => 30
p °c no, 0
o2 = 2)
. p
ae
20,
(9-67)
where
may be imposed around the p value chosen in the first grid search and a second
grid search applied to obtain a finer estimate of the minimizing p value. This
value and the corresponding 8’s obtained from Eq. (9-66a) constitute the point
estimates, and the asymptotic standard errors can be obtained from Eq. (9-67).
MA(1) Disturbance
Instead of the AR(1) disturbance process assumed in Eq. (9-55), let us now
consider an MA(1) process. As has been shown, this is likely to occur in a simple
Koyck scheme or in an adaptive expectations model. In each of these cases there
is the further significant feature that the parameter of the MA(1) process is also
the coefficient of the lagged dependent variable. The model to be considered is
thus
Y,=a+AY,_,
+ BX,+(u,—Au,) [AL <1 (9-68)
where it is assumed that
u ~ N(0, 621)
Utilizing the existence of the common parameter, this relation may be rewritten
as
LZ, SOA NZ. sete piAy (9-69)
where
Z,=¥,- 4,
Successive substitution for the Z variable in Eq. (9-69) gives
Z,=a(1+A+V4---4+2X7')
+ B(X,+AX,_) + PX. + + + NIX) + ZX
or Y,=a(1+A+--- +7!) + BX* + ZN + u, (9-70)
where now
XP = X, A NX,_4 + VXLg Hee HAY,
which may be computed recursively, for any given A, as
XP =X tN Aeon with X* = X,
Relation (9-70) has a well-behaved disturbance term suitable for ML (or equiva-
lently OLS) estimation, with Z, treated as a nuisance parameter. The data matrix
for OLS estimation would be
1 AtEN
1+A XEN
Oe 1+A+2”
The appropriate procedure is then a grid search over the interval 0 < A < 1. For
LAGGED VARIABLES 369
each value of A, X(A) is computed, OLS applied to Eq. (9-70), and the set of
parameters is chosen which minimizes the residual sum of squares.
The asymptotic standard errors may be obtained from the information matrix
in the usual way. The log likelihood for Eq. (9-70) may be written
1
nL = -— 7in(27)— sino; ~ 392 oti (9-71)
where
u, = Y, — aW, — BX* — ZN
and W.=1+A4---4)01
The unknown parameters in Eq. (9-71) are a, 8, A, Zy, and 0°.It may be oun
that the expected values of the cross ee order partial derivatives involvin07
g
are all zero. Thus inferences about 07 may be made independently of the other
parameters. The ML estimator is
a2_ dt;ae
0,
n
with asymptotic variance 20,1/n. The information matrix for the remaining four
parameters ist
a YW?
t [Link]* DWV, IWR
: B S ae axe xe, exe (0-72)
r a, yy LV, x
a DM!
where W, and X* have already been defined and
Ou,
EX
=a{1+2A+---+(¢-1)N-?] + B[X,_, + 20X,_,
4+-
+ (¢— 1)N~?X,] + tZ,rN-!
For a penultimate problem we return to Eq. (9-27), which represents a
combination of adaptive expectations and partial adjustment. The equation is
MS a A ag) (Nak) YBN Ages
P(A, )(L AZ)X, + (0, = ABU)
Defining Z, = Y, — u,, this may be rewritten as
Zim gia Ng Yo et By kee NoLyay (9-73)
where & = a(l —A,)(1 -—A,)
Bo =P = A) —A3)
Baas
bal ey voeo.
The disturbance series {v,} follows an MA(2) process. Ignoring this complication
for the moment and assuming the v’s to be independently and identically
distributed normal variables, the application of unrestricted OLS to Eq. (9-14)
would not yield the ML estimators since the seven coefficients are functions of
only five parameters. However, the relation may be rewritten as
Y* = p* + BX* + yZ* + v; (9-75)
where
Ve Ye (a op) Xe aaae
XPS X05 X55
LZ, = ZZ pl
The transformed variables in Eq. (9-75) depend on the a,, a, parameters. Given
any pair of a,, a, values and assuming the v’s to be independently distributed,
OLS could then be applied to Eq. (9-75) to yield estimates of u*, 8B, y, and the
residual sum of squares. The indicated estimation procedure would be a two-
dimensional grid search over a,, a, pairs, each parameter being constrained to the
(0, 1) interval.
Alternatively, if one makes the explicit assumption that the v, follow an
MA(2) process and if u ~ N(0, 021), the variance matrix for the v’s is given by
Opt) Ope Uae. vee 0
8, 8 § 8 0 te 0
s O54 f(s bona Oro vee 0
E (vy)
i=07 een ae (9-76)
LAGGED VARIABLES 371
where
8) = 1+ (a, +.a,)° + aa?
Ota (a, iz a,)(1 Ee Ay)
6) = aa,
The appropriate estimation procedure for Eq. (9-75) is then a combination of
GLS and a two-dimensional grid search over a,,a,. For each a,, a, pair the
variance matrix in Eq. (9-76) is computed and then GLS applied to Eq. (9-75).
One chooses the set of parameters that minimizes the weighted sum of squares
e’Q~ 'e, where e is the vector of residuals computed from Eq. (9-75) by using the
GLS estimates and Q is the matrix in Eq. (9-76).
Various models have been considered in this section, and it may be helpful to
summarize them briefly in Table 9-1.
have come to be more extensively employed. As seen in Sec. 9-1, this relation may
be formulated equivalently as Eq. (9-5),
Y=pt
B(L) Xa
A(L)
t
where D(L), B(L), and A(L) are all polynomials in the lag operator, but the
orders of B(L) and A(L) are expected to be small relative to the order of D(L).
The relation (9-5) is known as a transfer function in the time-series literature.}
There are four main characteristics which distinguish time-series estimation
methods from the various estimation procedures described in Sec. 9-2.
1. Before estimating the transfer function, the “input” series (X,} and the
“output” series {Y,} are subjected to sufficient differencing to render both
resultant series stationary.
2. The orders of the A(L), B(L) polynomials are determined empirically from
the data by an identification process and without imposing any a priori
theoretical specifications, such as a set of declining exponential coefficients.
3. The disturbance term in the transfer function is estimated as a general
ARMA process, as described in Sec. 8-5, rather than as a low-order AR or
MA process as in some of the models in Sec. 9-2.
4. The transfer function approach has been most extensively developed for the
single-input case (that is, one explanatory variable with various lagged values),
and there is no firm agreement yet on the appropriate extension to cope with
two or more inputs, each with aset of lags.
Stationarity
The simplest example of a stationary process is the white noise series {€,}, where
the e’s are independently and identically distributed as N(0, 67). It follows from
+ The basic reference is G. E. P. Box and G. M. Jenkins, Time Series Analysis: Forecasting and
Control, revised edition, Holden-Day, San Francisco, 1976, especially Chaps. 10 and 11.
LAGGED VARIABLES 373
Clearly, var(X,) still explodes and X, is not a stationary series. However, AX, =
(1 — L)X, is a stationary series since it is equal to €,. Thus first differencing the
random walk series produces a stationary series, but no finite number of differences
of Eq. (9-77) can produce astationary series if |p| > 1.
Extending the model to a second-order scheme gives
X, = 9 X,_1 + Oy.X-2 + & (9-78)
or p(L)X, =, (9-79)
where p(L)=1-6,L—-¢L’
By analogy with the first-order case we seek conditions on the roots of »(L)
which might distinguish between the stationary case, the explosive case, and the
intermediate case, where differencing might produce a stationary series. The
polynomial may be factorized as
p(L) = (1-¢,L)Q - © L)
and so the roots of the polynomial are c,;' and c; '. From Eq. (9-79)
Xone,
=. See
a
(l—¢,L)d=oL£)*
The term 1/(1 — c,L)(1 — cL) may be expanded in partial fractions as
Ree
Tt ao
Se aeeen
Vea ber) aa py ka)
where d = c,/(c, — ¢), as may be verified by multiplying out. Thus
d head
A Tay eae oe eens
= d(e, + ¢,@,_, + ¢78,25 + --+)
+ (1 — d)(e, + cye,_, + che, +-:-)
and the variance of X, will only be finite and constant if |c,| and |c,| are both
less than unity, that is, if the roots of p(L) lie outside the unit circle. The condition
on the roots may be stated equivalently in terms of the $,, ¢, parameters of Eq.
(9-78) as}
Ip.|
<1
g, + o, < 1
g, — 9, < 1
+ G. E. P. Box and G. M. Jenkins, Time Series Analysis: Forecasting and Control, revised edition,
Holden-Day, San Francisco, 1976, p. 58.
LAGGED VARIABLES 375
PUL) (Leela L)
Eq. (9-79) becomes
(VP GLE BE) XS ep) AX, = &, (9-80)
Even if |c,| < 1, the X, series is nonstationary since the other root lies on the unit
circle. However, it is clear from Eq. (9-80) that the A X, series is stationary as long
as |c,| < 1. If a third-degree polynomial factorizes as
2
p(L)'= (l= eh) L)
then second differencing the X, series will yield a stationary series as long as
leq] <1.
So far we have just considered AR processes of the form y(L)X, = €,, where
€, 1s white noise, and have seen that the condition for stationarity can be
expressed in terms of the roots of »(L). The same conditions hold when the
disturbance of the right-hand side follows an MA scheme, for if we write
p(L)X, = O(L)e, (9-81)
where 6(L) is a finite MA operator,
6(L)=1-6,L—-6,L?—---- we Ogee
then 6(L)e, is a stationary series. It has zero mean, a constant variance, and an
autocorrelation function which is nonzero for the first g lags and zero thereafter.
Thus the stationarity of the X, series still depends on the roots of g(L). The
general form of Eq. (9-81) is
=(1-0L-6,L?-----6L*)e, (9-82)
This is an autoregressive, integrated, moving average, ARIMA(p, d, q) scheme,
where p is the order of the AR polynomial, d is the degree of differencing required
to yield a stationary series (or equivalently, the number of unit roots in p(L)),
and gq is the order of the MA polynomial. The term integrated refers to the reverse
of the differencing operation since the differenced series have to be summed (or
integrated) to retrieve the original series. It did not arise in Sec. 8-5 where ARMA
processes were introduced to model a disturbance series which was already
stationary.
The general ARIMA model of Eq. (9-82) has been found to be a very flexible
tool for the univariate modeling and forecasting of a wide variety of homogeneous
nonstationary series—series that are not explosive, but which may display drift or
apparent short-run trends as well as various irregular oscillations. The univariate
modeling procedure consists of first determining the amount of differencing
required to produce approximate stationarity. Typically it appears that, if
differencing is required, first or at most second differences suffice. Defining
x, = @ aa exe
376 ECONOMETRIC METHODS
which is a transfer function of order (r, s, b), where b > 0 represents any delay in
the transmission of an effect from X to Y. The X and Yseries are appropriately
differenced to achieve (near) stationarity and are also expressed as deviations
from the sample means, if necessary. The problem now is the determination of
the values of r, s, and 6 and the estimation of the consequent a and 6 parameters.
The Box-Jenkins starting point is the calculation of the covariances (current and
lagged) between x and y and the autocovariances of the x series. The solution of a
set of simultaneous equations yields estimates of the 6 coefficients.{ From the
resultant 6 coefficients rough guesses are made of the values of r, s, and b on the
basis of a comparison between the pattern of the 6 coefficients and the theoretical
patterns for various values of r, s, and b. From the 6’s initial estimates of the a’s
and £’s can be derived and an iterative estimation process carried out, with
interaction between the estimation of the transfer function weights and the fitting
of an ARIMA scheme to the disturbance term. If the original disturbance was a
white noise series, any differencing will have produced an MA process in the
transformed disturbances, and if the original disturbance was complicated, the
transformed disturbance will normally be more complicated.
Box and Jenkins also suggest that the efficiency of the above process could be
improved if an ARIMA model was first fitted to the x, series. Denote such a
+ It is assumed that the same degree of differencing has been applied to each series. However, Box
and Jenkins state, “the procedures outlined can equally well be used when different degrees of
differencing are employed for input and output” (op. cit., ftn., p. 378). Consider
Y,=a+t BX,+ u,
First differencing both Y and X gives
AY, = BAX, + Au,
so that the original f coefficient is retained while the intercept disappears. If different degrees of
differencing are applied to each variable, one would no longer be estimating the original f coefficient.
+ These are the coefficients of the various lagged values of X, defined earlier in Eq. (9-1).
LAGGED VARIABLES 377
model by
$(L)x,
=6(L)n,
where 7, is approximately a white noise series. Now multiply through the model
y, = D(L)x, + u,
by 0. '(L)o,(L). The result is
y* = D(L)n, + 2, (9-83)
where
y=Xd+u (9-84)
As seen in Chap. 8, a nonspherical variance matrix for the disturbance term leads
to GLS estimation procedures. The GLS procedure is equivalent to premultiply-
ing Eq. (9-84) by a transformation matrix T and applying OLS to the transformed
data Ty and TX. The matrix T is chosen according to the assumed properties of
the disturbance term so as to make Tu a white noise series. The time-series
approach concentrates first of all on the properties of y and X in Eq. (9-84) and
not on the nature of u. A common differencing procedure is applied to Y, and X,,
followed by a common filter derived from the ARIMA model fitted to X,. The
D(L) polynomial containing the “long” series of 5 coefficients is finally repre-
sented by the ratio of two low-order polynomials which are estimated along with
an ARIMA model for the disturbance term.
It is impossible to give here a detailed operational description of the time-series
procedures.+ However, it is clear that a considerable amount of “judgment’’ is
required at various stages in choosing between different ARIMA and different
transfer function models. Time-series analysts also stress that long runs of
observations, preferably in excess of 100, are desirable, which requires the
+ Reference should be made to G. E. P. Box and G. M. Jenkins, Time Series Analysis: Forecasting
and Control, revised edition, Holden-Day, San Francisco, 1976, or to G. W. J. Granger and P.
Newbold, Forecasting Economic Time Series, Academic Press, New York, 1977. A lucid introduction
to a wide range of time series topics is provided by C. Chatfield, The Analysis of Time Series: Theory
and Practice, Chapman and Hall, London, 1975.
378 ECONOMETRIC METHODS
assumption that the underlying economic structure has been stable for that length
of time. The estimation of even the single-input case is fairly complicated. The
model is perhaps most appropriate to a “black-box” situation where interest
centers on a single input variable which can be controlled in any desired manner
but the researcher has no clearly articulated theory of the relation between the
input and the output.
The approach described above cannot be simply extended to multiple-input
models, since the covariances between the output and any input are contaminated
by the effects of the other inputs, unless the inputs are orthogonal. Spectral
methods are a possibility, but are not yet well developed for this case. A
somewhat different approach for dealing with two or more inputs has recently
been suggested by Liu and Hanssens.f Their approach is a modification of the
corner method for ARMA identification proposed by Beguin, Gourieroux, and
Monfort.£ Much work is proceeding in this field, and it is too soon to assess the
likely practical significance of the methods currently under development.
A final time-series approach that may be noted for the two-variable case is
the prewhitening of both series. Letting y, and x, denote appropriately differenced
series as usual, a separate ARMA model is fitted to each series, denoted by
$,(L)y, = 9,(L)a,,
(9-85)
and ,(L)x, mT 6.(L)i,,
where a, and #,, denote estimated residuals which are approximately white noise
series. Thus y, is prewhitened by the filter 6, \(L)$,(L) to yield a#,,, and x, is
prewhitened by its filter to yield @,,. It is argued that this approach is useful in
cases where there is doubt about the direction of causation. Does x cause y so that
one expects nonzero correlations between y and earlier values of x, or is it the
other way around, or is there joint causation and feedback? The suggested
procedure is to compute the cross correlations at various lags, positive and
negative, between a, and @,,. Inspection of these correlations should lead to a
decision about causation. For example, if causation is thought to run from x to y,
a transfer function model is estimated for d@,, on a,,,xt say,
(L)=y+yL+yl?+--:
to emphasize that they are not the original structural coefficients D( L) connecting
y and x. Finally substituting for u,,and a,, from Eq. (9-85) gives
y= D(L)x,
However, the above is in terms of the true coefficients and has also ignored the
noise terms. In practice the bivariate prewhitening approach requires the estima-
tion of more parameters and greater manipulations of those estimated parameters
than does the univariate prewhitening method. It would be interesting to see
comparative case studies of the results yielded by the two approaches, but there
do not yet appear to be any. An extensive application of the bivariate pre-
whitening approach to various time series of money and interest rates yielded “a
surprising, probably disconcerting, lack of relationship among several variables.” +
Pierce’s main conclusion was, “Extensions of time series modeling procedures of
Box and Jenkins reveal that numerous economic variables which are generally
regarded as being strongly interrelated may with equal validity, based on recent
empirical evidence, be regarded as independent or only weakly related.” A further
study by Haugh and Box illustrated the same approach to the study of the
connection between the GNP X and the unemployment rate Y in the United
Kingdom.} Each series was first differenced to yield x, = X,— X,_, and y, =
3eT -(7)
0 I Z € v ¢ 9 iB 8 6 Ol Il ral €I vl SI
vd 6£€0- v70- 970- €00 €00- 600 910- 810 +270 3800 600 100 670 810 OK) 70-
4nXn
“TL ‘q y8nezyy
pure‘°O “q ‘gq ‘xog UONROYNUEPT,,
Jo omMeUXG Uorssa1Zay painquysiq)
(eT sfopoy] SUNIBUUOD
OM], PWT], ,“SITI2G DUNO,
fo ay]
uDIIUaU JDIIISIIDIS ‘UONDIIOSSP
JOA ‘ZL ALLO“LT
LAGGED VARIABLES 381
The first concern was the direction of causation. Table 9-2 shows the various
lagged cross correlations. A positive lag is here defined as the y series lagging
behind the x series. The asymptotic standard error for r is 0.13. Table 9-3 presents
three summary statistics computed by the author from the data in Table 9-2. The
data in these two tables hardly seem to give any clear indication of the direction
of causation. However the authors of the paper state, “It is concluded that any
feedback effect is of secondary importance, as evidenced by the small cross
correlations at negative lags... . This direction of causation from x to y agrees
with that considered by Bray.”+ Another time series analyst might well interpret
these cross correlations differently and fit a different transfer function to the
residuals.
There is as yet no clear consensus on the relative roles of time-series
techniques and the more orthodox econometric methods. Some mistakenly view
them as competitive rather than complementary. Each is still an “art,” as distinct
from a “science,” in that time-series practitioners have to make various subjective
judgments in the course of their analyses just as econometricians conventionally
“choose” between different regressions and specifications. Investigators with
strong prior beliefs can usually see “patterns” in the data that may be invisible to
more sceptical colleagues.
PROBLEMS
with
e ~ N(0, 0/1)
The 6 parameter is estimated by § = SYdeo 1/2 ).= jasnowe
ge nt OCbetoa): a
(a) plim d = 8 + +55,
286 where $ eee
MS (1 — 3°) . =a ‘
OY,
7 1+ 26¢ ond ef : ed
9-3 Show that a second-degree approximation for the Almon lag implies the restrictions
6 = — Bi = py Boy —1Ba ey
&— Uy Pur)
where
9-7 Derive the information matrix (9-72) and show also that in this model the estimator of 0, is
asymptotically independent of the remaining estimators.
9-8 Consider X, = 2 X,_, + e, where (e,) is a white noise series. Draw some sets of e’s from a table of
random normal deviates and compute the corresponding sample realizations of the process for
t= 1,..., 10, starting each realization off by setting Xy = 0. Satisfy yourself that X, can become
“very large” in both positive and negative directions.
9-9 If u,=(1-0,L—0,L? —--- — 6,L7)e, and {e,) is white noise, derive the autocorrelation
function of the {uw,) series.
9-10 In the rational lag equation
3L
iia ee
L092 20274
determine:
(a) The total multiplier
(b) The mean lag
(c) The coefficients of Xp forj = 0,1, 2,3.
(UL, 1980)
LAGGED VARIABLES 383
TEN
A SMORGASBORD OF FURTHER TOPICS
—
/
xX, ——
where y,_, denotes the subvector consisting of the first r — 1 elements of y. Using
b,_,;
r one may “forecast” y, at sample point r, corresponding to the vector x, of
explanatory variables at that point. The forecast error is
ae xpTae
and, as shown in Sec. 5-4, the variance of this forecast error is
, , al
o7(1 a (XC Xs) x,)
1. Choose a base of k observations. For the moment let this be the first k
observations in the sample, whether it be composed of time-series or cross-
+ Recursive residuals are a member of the general class of LUS residuals (linear unbiased with a
scalar variance matrix). Another important set is the BLUS residuals due to Theil. See H. Theil,
Principles of Econometrics, Wiley, New York, 1971, Chap. 5.
+ As is customary, the first element in each x vector will be unity to accommodate the intercept
term.
386 ECONOMETRIC METHODS
Vari — Xs 1D,
Weil >
Vl 2 Kan (XX) Uae
w ~ N(0,07I,_;) (10-3)
Since we have already seen that each w, is normal with zero mean and variance
o*, the proof of Eq. (10-3) just requires the establishment of zero covariances. The
numerator in Eq. (10-2) may be written
E(u,u,) = 0
E(u,u,)
E(u,u,)
E(u,_,u,) = fe =0
E(u,_,u,)
a
Wy
E(u,_w,_,) = E [std ithe eat paras 27)
a
ea ipee 0: ey)
A SMORGASBORD OF FURTHER TOPICS 387
Multiplying out the right-hand side of Eq. (10-4), remembering that the expecta-
tion of a scalar can also be written as the expectation of the transpose of the
scalar, and using the above results, easily establishes
E(ww,)=0 forallr,s;r#s
and so Eq. (10-3) is proved.
The computation of the recursive residuals might be achieved by using the
conventional OLS formula repeatedly to compute each b vector in the sequence
b,,b,,),--., b,. However, the calculations are simplified by using the following
recursion formulas:
it follows that
XX, = X_|X,_,) + XX,
Eq. (10-5) may then be checked by multiplying the left-hand side by X’.X,, the
right-hand side by X,,_,X,_, + x,x’,, and seeing that both reduce to the identity
matrix.t Relation (10-6) may be simply derived since
(X/X,Jb,= X,y
Be eee,
= XPRpX Dey ax
(XX) ky, Xeb, 7)
Finally, relations (10-5) and (10-6) may be used to derive the following: §
RSS,
= RSS,_, rN + w? r=k+ Vn (10-7)
where
RSS, - (y, aiX,b, )’(y, re X,b, )
+ See R. L. Brown, J. Durbin, and J. M. Evans, “Techniques for Testing the Constancy of
Regression Relationships over Time,” Journal of the Royal Statistical Society, ser. B, vol. 37, 1975, pp.
149-192, for a statement of these formulas and some notes on their history. A useful survey of
recursion formulas for various models is to be found in W. C. Riddell, “Recursive Estimation
Algorithms for Economic Research,” Annals of Economic and Social Measurement, vol. 4, 1975, pp.
397-406.
+ See Problems 10-1 and 10-2.
§ See Problem 10-3.
388 ECONOMETRIC METHODS
for structural change in the case where the second sample contains fewer than k
observations.} Based only on a heuristic proof, it was asserted in Eq. (6-27) that
under the hypothesis of no structural change
- (ee, — ee) /n»
—F ny, 1 =i)
eie,/(n, ak)
where e,e, denotes the residual sum of squares from a regression fitted to all
n, + n, observations and e'e, is the residual sum of squares from a regression
fitted to the first n, observations. From Eq. (10-7) it follows that for a regression
with n observations,
n
RSS ioe
r=k+1
ea (ier)
Since under the null hypothesis the w, are independently and identically distrib-
uted normal variables, the F statistic is seen to be the ratio of two independent x?
variables, each divided by the appropriate number of degrees of freedom, and so’
it has the F(n,, n, — k) distribution.
A second useful application of recursive residuals lies in testing for hetero-
scedasticity.t If the alternative hypothesis to homoscedasticity is that 0,° varies
with Xjns the procedure would be as follows.
1. Order the data according to the values of X; and choose a base of at least k
points from among the central observations.
2. From that base compute a vector w, of recursive residuals corresponding to
the first m observations, and another vector w, of recursive residuals corre-
sponding to the last m observations.§ Since the smallest feasible base is of size
k, the maximum value of m is (n — k)/2.
+ See A. C. Harvey, “An Alternative Proof and Generalization of a Test for Structural Change,”
The American Statistician, vol. 30, 1976, pp. 122-123.
$¢A. C. Harvey and G. D. A. Phillips, “A Comparison of the Power of Some Tests for
Heteroscedasticity in the General Linear Model,” Journal of Econometrics, vol. 2, 1974, pp. 307-316.
§ Notice that there is no problem in computing recursive residuals backward or forward in a
sample from any suitably chosen base, or indeed in adding “new” observations in any order.
A SMORGASBORD OF FURTHER TOPICS 389
3. Under the null hypothesis it follows directly from the properties of recursive
residuals that the test statistic
F= WW, ~ F(m,m)
of (10-8)
ww)
Some sampling experiments by Harvey and Phillips indicate that the power of the
test in Eq. (10-8) compares favorably with that of the Goldfeld-Quandt test
described in Sec. 8-4. They recommend setting m at approximately n/3. An
advantage of the recursive residuals test over that of Goldfeld and Quandt is the
greater flexibility of the former. If, for example, one now wished to test whether
o* varies with some other variable X;, one could simply regroup the existing
recursive residuals according to low and high values of X, and compute Eq. (10-8)
afresh, whereas the Goldfeld-Quandt test would require the computation of two
new regressions.
A third application of recursive residuals is in testing for autocorrelation.} In
a time-series application one may take the first k observations as the base. From
the resultant n — k recursive residuals the conventional von Neumann ratio ist
(10-9)
Ss
z ei (ate)
These points are tabulated in App. B-7. The von Neumann ratio is arithmetically
closely related to the Durbin-Watson statistic, which could, of course, be com-
puted from the recursive residuals. The crucial point, however, is that the
multivariate normal distribution for w specified in Eq. (10-3) satisfies the assump-
tions underlying the derivation of the von Neumann (Press and Brooks) signifi-
7G. D. A. Phillips and A. C. Harvey, “A Simple Test for Serial Correlation in Regression
Analysis,” Journal of the American Statistical Association, vol. 69, 1974, pp. 935-939.
+J. von Neumann, “Distribution of the Ratio of the Mean Square Successive Difference to the
Variance,” Annals of Mathematical Statistics, vol. 12, 1941, pp. 367-395.
§ B. I. Hart, “Significance Levels for the Ratio of the Mean Square Successive Difference to the
Variance,” Annals of Mathematical Statistics, vol. 13, 1942, pp. 445-447.
4S. J. Press and R. B. Brooks, “Testing for Serial Correlation in Regression,” Report no. 6911,
Center for Mathematical Studies in Business and Economics, University of Chicago, Chicago, 1969.
390 ECONOMETRIC METHODS
cance points so that an exact test is available, thus avoiding the inconclusive zone
associated with the Durbin-Watson statistic calculated from the OLS residuals.
Some sampling experiments by Phillips and Harvey suggest that the power of this
test may be increased by forming the initial base from a mixture of the first and
last observations.
Fourth, recursive residuals provide a test of some possible forms of
misspecification.; Since, under the null hypothesis, the recursive residuals are
independently and identically distributed normal variables with zero expectation,
the mean of the residuals divided by its estimated standard error will follow a ¢
distribution. Formally
Ww
Ge ee ae 1) (i0-10)
where
ae Linkt
homed
and
n wir 2
Le w )
=
tke |
As an illustration of the use of this test in specification analysis suppose the
postulated model is a linear relation between Y and X. If the true relation is
convex (concave) and the data are ordered by the size of X, the recursive residuals
would be expected to be mainly positive (negative) and the computed f statistic
will tend to be large in absolute value. In a multivariate situation this specification
test could still be carried out for any single explanatory variable, if it were thought
that the other explanatory variables were correctly specified, but this type of a
priori knowledge is seldom available. Several specification errors might have a
self-canceling effect on the recursive residuals, so this test is not likely to be very
effective in multivariate situations.
Finally Brown, Durbin, and Evans describet an important application of
recursive residuals in testing for structural change over time. The null hypothesis
of no structural change for the model y = XB + wis specified as
Hy) By Bye BB
re =97 =o?
where B, denotes the vector of coefficients ruling in period ¢ and o/ the dis-
turbance variance in that period. It is clear that the null hypothesis would be
violated if the B vectors remained constant but o? varies. This would be the classic
where
RSS
67 = ——
Ok
W, is seen to be a cumulative sum, and it should be plotted against r. As long as
the B vectors are constant, E(W,) = 0, but if the B’s change W,, will tend to
diverge from the zero mean value line. For a forward recursion the significance of
the departure of W, from the zero line may be assessed by reference to a pair of
straight lines which pass through the points
{k, tavn—k} and ({n, +3a¥vn—k}
where a is a parameter depending on the significance level a chosen for the test.
The correspondence for some conventional significance levels is
a= 0.01 a = 1.143
a = 0.05 a = 0.948
a = 0.10 a = 0.850
The lines are shown in Fig. 10-1.
The equation of the upper line in Fig. 10-1 may be determined from
Wi Wisk | avn Kk
t—k n—k
or
Zatzkh)
= avn—-k + ————
Vik
and the equation of the lower line is given by its negative.
The second test statistic is based on cumulative sums of the squared residuals,
namely,
oe 2
ieee
I eT ae (10-12)
ae
The mean value line giving the expected value of the test statistic under the null
hypothesis 1s
hak
Ee Dacarks
which goes from zero at r = k to unity at r = n. The significance of the departure
of s, from its expected value may be assessed by reference to a pair of lines drawn
parallel to the E(s,) line at a distance cy above and below. Values of cy for
various sample sizes and levels of significance are tabulated in App. B-8. Refer-
ence should be made to the Brown, Durbin, and Evans article for practical
illustrations of the technique and for interpretations of various plots. The basic
idea is that instability of the parameters would be indicated if the plot of W, or s,
crossed the significance lines described above. There is some evidence that the
cusum test is less powerful than the cusum of squares test. Some Monte Carlo
experiments by Garbade also suggest that the latter may not be very powerful in
comparison with tests based on variable parameter models.t However, the
explanatory variable in his experiments was random over time, and it would be
interesting to see if the same result was obtained with an autoregressive explana-
tory variable.
+ K. Garbade, “Two Methods for Examining the Stability of Regression Coefficients,” Journal of
the American Statistical Association, vol. 72, 1977, pp. 54-63.
t See D. J. Poirier and S. G. Garber, “The Determinants of Aerospace Profit Rates, 1951-1971,”
Southern Economic Journal, vol. 41, 1974, pp. 228-238; or D. J. Poirier, The Econometrics of Structural
Change, North-Holland, Amsterdam, 1974, Chap. 2.
A SMORGASBORD OF FURTHER TOPICS 393
B, = 6,
B, = 6, + 46, a, =a, — d,a (10-15)
B, = 6, + 6, + 6, a, = a, — 5b
y y
{ t
: a b mae Oi a b ie
(a) (b)
Figure 10-2
394 ECONOMETRIC METHODS
Fitting Eq. (10-14) directly by OLS will yield estimated functions which meet at
the knots, and the estimated a and 8 parameters of those functions can be
determined from Eqs. (10-15). Tests on a’s and B’s imply equivalent tests on the
6’s. Thus testing the significance of 6,(= 8,) is asking whether there is a positive
(or negative) trend in the first period. Testing the significance of 5, is asking
whether the trend slope in the second period differs significantly from that in the
first, and similarly, testing the significance of 6, amounts to asking whether the
trend slope in the third period differs from that in the second. Setting up the null
hypothesis
5, 0
at bs
| ; |
0|
is equivalent to postulating that the B’s and the a’s are the same in all three
periods, that is, that the data may be adequately described bya single linear
trend. This test may be carried out most simply by fitting
y, =at bw, + u,
as the restricted model, the full spline function (10-14) as the unrestricted model,
and calculating the test statistic defined in Eq. (6-8).
An alternative estimation procedure is restricted least squares. Returning to
Eqs. (10-13), the restrictions implied by the join points are
a, + B\a=a,+
B,a
a, + B,b =a, + B3b
which may be set up in the conventional framework as
R B r
ay
B,
1 oe —a 0 Ones “|
0 0 beet o b, -|° (10-16)
a3
B;
Thus the model
ae |
rz,|
an 3 Q,
: eG pel) A feesape By
is! bel ean Y a,
ys aga 2), Polit ar
Y |= ee 10-
¥3 ee ete eon sl | iia ete a,
| |
| oe b+ 1 B,
| |
|
A SMORGASBORD OF FURTHER TOPICS 395
where the empty cells in the data matrix are all zero, is fitted subject to the
restrictions in Eq. (10-16). The appropriate formula is given in Eq. (6-5). The
estimates of the a and 8 parameters will be identical to those derived from the
estimated coefficients of the spline function in Eq. (10-14).+
This simplified example used time as an explanatory variable. The procedure
works equally well for any explanatory variable x with known join points, or
knots, at x,, x,, and so on. A possible disadvantage of the linear spline is that
while the function itself is continuous at the knots, there is a discontinuity or
jump in the first derivative. This may be overcome by the introduction of
quadratic or cubic splines. To illustrate a cubic spline function, suppose we have a
two-variable relation with known knots at x, and x,. Within each subset y is
expressed as a third-degree polynomial in x, namely,
i=] Ks xe
i= 2 My a es
i=3 Xho
1. The panel consists of, say, 1000 households whose savings behavior Y,, is
monitored along with various explanatory variables X,,,, such as income,
family size, and composition over a number of time periods.
2. The panel consists of a set of firms, and the object of study is the size and
timing of their investment expenditures Y,, as a function of the group of.
explanatory variables thought to influence investment.
3. The panel might consist of the 50 states of the United States, and the focus of
investigation are the determinants of the unemployment rate Y,, across states
and over time.
4. The panel consists of the OECD countries, and Y,, indicates the per capita
consumption of gasoline in country 7 in year t. The relevant question is
whether the usual economic variables such as income and relative prices can
adequately explain the variation in Y,,.
The most common way of organizing the data in Eq. (10-19) is by decision
units. Thus let
Yi ei
ae F Xx Xyi X31 Xx i :
y; = {= poe mee oe u,= :
No 2im 3im
im
ki
Ues
denote the data and the disturbances relevant to the ith unit. The data may be
“stacked” to form
Y; X, u,
Tele ols ce | (10-20
where y isn X 1, Xisn X (k — 1), and uisn X 1. The model in Eq. (10-19) may
be expressed as
y= [ix] 8 +u (10-21)
where i is an n X 1 vector of units, a is a scalar, andB=(f, 8B; --- B,J’.
A variety of models has been proposed for time-series and cross-section data,
and most have been fitted to some data set or another. These models may all be
derived from Eq. (10-21) by varying the assumptions made about the systematic
part of the equation and/or the assumptions made about the disturbance vector.
A possible taxonomy of models is indicated in Table 10-1. The meaning of
various terms in the table may not be clear at first sight but will become so as the
models are explained.
Model I(a) is perfectly straightforward. The systematic part of Eq. (10-21)
postulates a common intercept and a common set of slope coefficients for all units
at all time periods. The disturbance assumption is
u,, ~ iid(0,02) for alli,t
where iid means independently and identically distributed. Thus there is no serial
correlation in the disturbances for any individual unit, there is no dependence
between the disturbances for different units, either contemporaneous or lagged,
and the disturbance has a constant variance at all points. The appropriate
estimation method is OLS applied to the stacked data of Eqs. (10-20). If, in
addition, the u;, are assumed to be normally distributed, all the finite sample
inference procedures of Chaps. 5 and 6 are valid.
Model I(b) allows a richer specification for the disturbance term. There are,
in fact, several versions of model I(b) depending upon the precise assumptions
Assumptions about
Vector of slope
Intercept coefficients Disturbance term
Model a B Ui,
I(a) Common for all 7,t Common for all i, ¢ E(uv’) = oI,
I(b) Common for all i, ¢ Common for all i, ¢ E(uu’) = V
I(a) Varying over i Common for all i, ¢ Fixed effects model
II(b) Varying over i Common for all 7,t Random effects model
III(a) Varying over i, t Common for all i, ¢ Fixed effects model
III(b) Varying over i, f Common for all i, ¢ Random effects model
IV Varying over i Varying over i E(uu’) = 671 or E(uu’) = V
398 ECONOMETRIC METHODS
The application of GLS to Eq. (10-21) using Eq. (10-22) would now yield the
[Link].e. of B,
y: ai:Ec
v> i, o 0 X,
ae Lin cons’oop Oe Boe heme aee (10-23)
.
Y,
0 0 In
x,P || 8
or
y=Za+xXBP+u (10-24)
+ See Problem 10-5 and also J. Kmenta, Elements of Econometrics, Macmillan, New York, 1971,
pp. 512-514, for a discussion of this case.
A SMORGASBORD OF FURTHER TOPICS 399
where the definition of Z is obvious from the comparison of Eqs. (10-23) and
(10-24). Define the matrix B as
B = 2(Z'Z) ‘2’
It is easily seen that B is ann X n matrix given by
J, 0 0
B=—|0 J, 0
T7Ua |WR eetcecahtyy
Ge EF Heer
0 oO J,
where
1 Y,
Sl es
Y,
where
iameay
i m ~ ij
1
(Z'Z)
| =diag{m—! m=! --- m")
and premultiplying an n X 1 vector by Z’ serves to sum elements within each
group. Thus Eq. (10-27) implies
NaNO
Oye ts Vip Yh J, ee (10-28)
This model, which is designated as model II(a), is usually known as the fixed
effects model. The fixed effects are the intercepts a,, one for each group. It is
usually assumed that the u vector in Eq. (10-24) is homoscedastic and nonauto-
correlated so that OLS provides b.l.u.e.’s, though GLS estimators could be
constructed on the lines of model I(b). The b vector of Eq. (10-26) is also
sometimes referred to as the “within” estimator, since it is based on the
within-group deviations (Y,, — Y,) and (X;;, — X;;). Equations (10-21) and (10-24)
have already appeared in Sec. 6-2 on Tests of Structural Change. The exposition
in that section was solely in terms of a time-series application where the “groups”
referred to p different subperiods, not necessarily all of the same length. Equation
(10-21) is a restricted version of Eq. (10-24), and tests of the restrictions may be
made in the context of OLS estimates as in Sec. 6-2, or in the context of GLS
estimates as in Sec. 8-6.
Model II(b) is the random effects, or error component, model. Instead of
assuming a set of given (unknown) constants a,,..., a, for the p groups, a single
intercept a is postulated, and the differential intercepts are merged with the
disturbance term. The model is now formulated as in Eq. (10-21), namely,
y = [ix] 8 +u
but the assumptions about u are
Uy = a; + bj,
where the a; are drawn at random from M(0, 62) and the e,, are drawn at random
from N(0, 07). The a, are now increments (positive or negative) to the common
intercept a. To derive the variance matrix of u we note that for the ith group we
may write
Us Gl as.
A SMORGASBORD OF FURTHER TOPICS 401
p p
=o;|p 1 e|=0,7A
p ]
where
62
0, =90,+06,2 and (ise
0,
Since E(u,u’,) = 0,
A 0 0
V=E(u’)=o67/0 A 0
FO tay BHA
= ole @A
The matrix A may also be expressed as
A7! 0 ae 0
The difficulty, however, is that V~' involves the unknown oa? and o2. Before
dealing with this problem it may be shown that the GLS estimates can also be
402 ECONOMETRIC METHODS
veer oe
Leper imp
wheret
j
Coates \aio (10-34)
0; + mo?
+z can thus represent the sample observations on the dependent or an independent variable for
any given unit.
+ This result is stated without proof in J. A. Hausman, “Specification Tests in Econometrics,”
Econometrica, vol. 46, 1978, p. 1262.
A SMORGASBORD OF FURTHER TOPICS 403
The estimation of Eq. (10-21) by GLS using the V~! defined in Eq. (10-31) then
gives
vir Vier
uj, = a; + E,,
Averaging over ¢ for unit i gives
S| I Ri ates I
Averaging then over i gives
Sl RI +€
and the usual decomposition of sums of squares gives
N
Y(u,, * iz)” a Lui i it,)” a X(a = Sl
2
Beak
Recall that if x,, x.,..., x,, are drawn at random from x ~ N(0, o?),+
p{PRa |a
m— | eo
+ The assumption of normality is not required for this result, only that x ~ iid(0, 02).
404 ECONOMETRIC METHODS
Density Expected
Source Sum of squares function Mean square mean square
2
Within groups D(H - a;)” p(m— 1) a= Q€
AS tne tene
re
= =\2 2 2
and the expected between-group mean square is mo? + 62. The u,, in Table 10-2
are, of course, unobserved, but we can estimate the relevant disturbance variances
by substituting estimated u’s in these formulas.
The estimation procedures may be summarized as follows.
1. Fit the basic model, Eq. (10-21), by OLS and obtain the n X 1 vector & of
OLS residuals. Compute also the mean residual a, for each unit, and note
that
a= 0.
2. Compute
3. Compute the quasi deviations y,, = Y,,— cY,, and so on, and apply OLS.
A SMORGASBORD OF FURTHER TOPICS 405
The direct application of the GLS formula requires estimates of p = 0,/o, and of
o,,. These are also obtained from the OLS residuals, The steps are as follows.
1. As in OLS procedure.
2. Compute
s=2=— 1 _ YF(a,- a)
Peey sy
t= 4,
eeDepea
] Oe
t ares)hehe
A ne aa? Be?
Ss, =s2+s
2
b=3
3. Using 6 and s2, compute V and the GLS estimator defined in Eq. (10-35).
The estimation of the variance component from the OLS residuals is not to be
recommended when lagged values of Y appear in the X matrix. Since p = 02/
(0, + 9,’) is constrained to be in the (0, 1) interval, a grid search over this interval
for the ML estimator is a feasible procedure.
The final question with respect to model II is the choice between fitting either
the fixed effects or the random effects model. The choice basically has to be made
by the researcher based on the institutional realities relevant to the problem being
studied. Returning to the examples given at the beginning of this section, suppose
certain monetary /fiscal policies are set in place in an attempt to reduce unem-
ployment rates across the country, and after some time an analysis is made of the
experience of the various states. As a result of historical developments, the states
have variable mixtures of industrial, commercial, private, and public structures.
One would thus expect differential effects across states, which would be modeled
appropriately by the fixed effects assumption. On the other hand, if we look at the
per capita consumption of gasoline in the OECD countries, we will certainly
observe very different levels of the dependent variable in different countries.
However, it is also true that for tax and other reasons the real price of gasoline
has historically been very different in different countries. For sound economic
reasons this may be expected to have Jong-run effects on the size of automobiles
and on per capita gasoline consumption. Inserting dummy variables to allow
different intercepts across countries removes this variation from the data, and the
“effects” of the explanatory variables are estimated solely from the within
+ See P. Balestra and M. Nerlove, “Pooling Cross-Section and Time Series Data in the Estimation
of a Dynamic Model: The Demand for Natural Gas,” Econometrica, vol. 34, 1966, pp. 585-612; and
G. S. Maddala, “The Use of Variance Components Models in Pooling Cross-Section and Time-Series
Data,” Econometrica, vol. 39, 1971, pp. 341-358.
406 ECONOMETRIC METHODS
estimator, Eq. (10-26), which is based on the within-country variation and is not
influenced by the between-country variation. Thus a fixed effects model would be
liable to underestimate the price elasticity. The random effects model would be
equally inappropriate since it would attribute significant variations in consump-
tion to unidentified stochastic factors rather than to price. In this case a more
sensible estimator of long-run price and income effects would be obtained by
computing the between estimator based on country (group) means. Averaging Eq.
(10-19) over groups gives
1. Asm — oo, c > 1, and the GLS (random effects) estimator of B tends to the
fixed effects estimator of B.
2. As a, becomes very large relative to 07, c > 1, and again the random effects
and fixed effects estimators of B will tend to coincide.
3. As o2 — 0, c — 0, and the random effects estimator would tend to the OLS
estimator (X’X) 'X’y.
Returning to the taxonomy in Table 10-1, model III allows the intercept to
vary over units and time periods, while retaining the assumption of a common B
vector for all 7, t. This again may be estimated bya fixed effects or random effects
approach. The former extends Eq. (10-24) to include dummy variables for the
time periods, taking care to use only m — 1 such dummies in order to avoid a
singular data matrix. The random effects model postulates the disturbance to be
Ui Oped fi
where the y’s are assigned at random to the time periods from some postulated
distribution. Just as a, is assumed common to the ith unit for all time periods, so
+ For a lucid and practical discussion of these issues see J. M. Griffin, Energy Conservation in the
OECD: 1980 to 2000. Ballinger, Mass., 1979, Chap. 2.
A SMORGASBORD OF FURTHER TOPICS 407
is y, assumed common to all units in the ¢th time period. The extensions from
model II are relatively straightforward, and we will not go into them here.+
Model IV allows both the intercept and the B vector, or some components of
it, to vary across units. This model has already been studied in Sec. 6-2 under the
simplest possible assumptions about the disturbance term and in Sec. 8-6 in the
context of the SURE model. The random effects version of model IV might also
be extended to allow for time-specific as well as unit-specific error components. In
testing for the stability of the B vector (whether across unit or over time) it is then
especially important to use the procedures of Sec. 8-6 with an appropriately
specified variance matrix for the disturbance.
This topic has already appeared in several places. Sec. 6-2 on structural change
investigated variations in some or all of the parameters of a relation, but it was
known apriori at which point possible structural breaks might have occurred
(peacetime, wartime, and so on). Section 10-2 on spline functions showed how
different functions might be fitted so as to meet at the known join points. Section
10-3 on time-series and cross-section data considered many possible variations in
parameters, but again, as in Sec. 6-2, there were obvious points at which such
changes might be expected. Only in Sec. 10-1 on recursive residuals was there
some discussion of the case where the B vector might change at unknown points.
We must now consider cases where there is no a priori information on the
observational points at which structural changes might have taken place, and in
this brief section we will consider just two possible approaches. The approach of
switching regressions is based on the assumption that there is a known (small)
number of different regimes, but the switching points are unknown. The other
approach is based on the assumption of continuous parameter variation.
Switching Regimes
The simplest case of switching regimes is based on the assumption of just two
different regimes. The switch may depend on time or on a “threshold” value for
some variable, or it may be triggered stochastically. For instance, wage and price
decisions may be different in periods of low inflation and in periods of high
inflation. The pioneering treatment of switching regimes is due to Quandt.§ To
+ Reference may be made to the articles by Maddala and by Balestra and Nerlove already cited,
and also to T. D. Wallace and A. Hussain, “The Use of Error Component Models in Combining
Cross-Section with Time-Series Data,” Econometrica, vol. 37, 1969, pp. 55-72; and Y. Mundlak, “On
the Pooling of Time-Series and Cross-Section Data,” Econometrica, vol. 46, 1978, pp. 69-86.
£See Problem 10-7 and B. H. Baltagi, “An Experimental Study of Alternative Testing and
Estimation Procedures in a Two-Way Error Component Model,” Journal of Econometrics, vol. 17,
1981, pp. 21-49.
§ See S. M. Goldfeld and R. Quandt, Studies in Nonlinear Estimation, Ballinger, Mass., 1976,
Chap. 1, and references therein.
408 ECONOMETRIC METHODS
illustrate the approach suppose we have ¢ = 1,..., 2 sample observations and the
hypothesis is that
Regime 1: y, = a, + B,x, + u,, holds fort < 1*
Regime 2: y, =a, + B,x, + up, holds fort > ¢*
where ¢* is unknown. Assuming the u’s to be normally and independently
distributed with zero means and variances 0) and 93, the log likelihood is
t* ron t*
26? 26? 2
and so
* ere
Inf= 2
in de — ene
ae 2
ee (10-38)
An estimate of the switch point r* could then be made by evaluating Eq. (10-38)
for all possible values of t* and choosing the one that maximizes the likelihood.
With n sample observations and two variables the possible range for ¢* is from
t* = 3 to * = n — 3, implying the calculation of n — 5 pairs of regressions..
Riddell, however, has recently pointed out that the computational burden is
considerably reduced by making use of recursive residuals.+ Consider the set of
forward recursive residuals w,, w,,... . From Eq. (10-7) we have
RSS, = RSS,_,
+ w?
ie
Thus RSS2= 5)
t=3
and
620) = —
2 RSS.
In a similar fashion 6;(t*) can be constructed from the set of backward recursive
residuals. Thus just two passes of a recursive residuals program will generate all
the data required to find the maximum of Eq. (10-38).
pe)
L(&)
where L(&) is the unrestricted maximum of the likelihood function over
the
entire parameter space. In this example it is the antilogarithm of the maximum of
Eq. (10-38) since it is assumed to be known that there is at most one switch point
and the restriction of a single regression (no switch point) has not been imposed.
L(@) is the maximum of the likelihood function over the subspace w C Q2 to
which one is restricted by the hypothesis. In this problem it is the maximum value
of the likelihood for a single regression. Under the hypothesis of no switch
+ For a brief account of likelihood ratio tests see P. G. Hoel, Introduction to Mathematical
Statistics, 4th edition, Wiley, New York, 1971, pp. 211-217.
+R. L. Brown, J. Durbin, and J. M. Evans, “Techniques for Testing the Constancy of Regression
Relationships over Time,” Journal of the Royal Statistical Society, ser. B, vol. 37, 1975, pp. 149-192,
especially p. 161.
§ S. M. Goldfeld and R. Quandt, Studies in Nonlinear Estimation, Ballinger, Mass., 1976.
A treatment of disequilibrium models is beyond the scope of this book. Some important
references are R. C. Fair and D. M. Jaffee, ““Methods of Estimation for Markets in Disequilibrium,”
Econometrica, vol. 40, 1972, pp. 497-514; R. C. Fair and H. H. Kelejian, “Methods of Estimation for
Markets in Disequilibrium: A Further Study,” Econometrica, vol. 42, 1974, pp. 177-190; T. Amemiya,
“A Note on a Fair and Jaffee Model,” Econometrica, vol. 42, 1974, pp. 759-762; G. S. Maddala and
F. D. Nelson, “Maximum Likelihood Methods for Models of Markets in Disequilibrium,”
Econometrica, vol. 42, 1974, pp. 1013-1030; S. M. Goldfeld and R. E. Quandt, “Estimation in a
Disequilibrium Model and the Value of Information,” Journal of Econometrics, vol. 3, 1975, pp.
325-348.
410 ECONOMETRIC METHODS
os (B, ff 01;) + (B, 45 0 ;)Xp, Grea saatiate (B, a5 On;) Xe; Vo Allee nN
(10-39)
The ’s in Eq. (10-39) are unknown constants common to all sample points.
The
v,; are stochastic variables which determine the coefficient vector for the
jth
sample point. The n sample points might, for example, be a cross
section of
households where important explanatory variables may be unobserved,
and their
influence affects slope coefficients as well as the disturbance term.
The reaction of
mortgage debt to, say, the measured rate of interest may well
depend on the
unobserved age of the head of household. There is no need to
insert the usual
equation disturbance term in Eq. (10-39) since it will merge with
v,,. Equation
(10-39) may be rewritten as
Y=
XiBi Wendel lau (10-40)
where uz = XV,
x LL Seat
y= [ei P.)|
Assumptions about the vy, are required to make the model
operational. A simple
set of assumptions is
E(v,) =0 J=1,...57
areas0 0
E(vv') = 0 a, ae Osli—A alee (10-41)
ed) a,
E(vv/) = 0 f=. nti sy
The stochastic elements in the coefficients
are thus assumed to have zero means
A SMORGASBORD OF FURTHER TOPICS 411
and to be uncorrelated between sample points and also between different coeffi-
cients for any given sample point. The last assumption is possibly the least
plausible. If, for example, age has an effect on the reaction of mortgage debt to
the rate of interest, it may have a related effect on the response to income. These,
however, are the assumptions of the original Hildreth-Houck random coefficient
model.
From Eqs. (10-41) the disturbances in Eq. (10-40) have the following proper-
ties:
E(u;)=0 fecal
E(u;) = E(x'vvix, jae leees ai
= x)AX;
E(u,u;) = 0 LJ Sahat ea
3; E(u?) a x Xia,
i=]
= Xa (10-42)
Sane
where x; = [1 XD)2 Sa: Xa2
where X denotes the matrix obtained from X by squaring each element. The form
of this relation suggests that if estimates of the left-hand vector could be obtained,
a regression on X could yield an estimate of a. Looking at the residuals obtained
from the OLS fit to Eq. (10-40), e = y — Xb, we know from Chap. 5 that
Ee?
E(é) =| 2* |= Mo?
Ee?
Thus E(é) = MXa (10-44)
Equation (10-44) leads to the following procedure for constructing a feasible GLS
estimator.
— Fit OLS to Eq. (10-40) and square each residual to obtain the vector é.
2. Regress é on MX, which can be constructed from the original data matrix X,
to obtain an estimated vector &.
~ Substitute & in Eq. (10-42) to obtain estimates o of the variances of the w’s.
4. Using the s? obtain the GLS estimate of B in Eq. (10-40).
¥, = SV (Bae,
ctu ale eee (10-45)
There are p separate units with m sample observations on each. The
X, are all of
order m X k and rank k. The B vector of k coefficients is common
to all units. The
¥, vectors model the stochastic variation of the coefficient vector across
units. For
all i, 7 = 1,..., p it is assumed that
Yi XxX, X 0 0 vy Bi
Y2 X, V> U5
Sele Pree ceo ui 1 Ot ae leila (10-47)
_ ae
1. Compute the OLS vectors for each unit separately, that is,
b, = (X,X,)'Xiy,
and the vectors of OLS residuals e; = y, — X;b,.
2. An unbiased estimator of o,,; is given by
_ ee;
Ce iran ke
3. An unbiased estimator of A is given by
A 5 Thee “Fi
A =—*2 -— ¥5,,(XX,
Dis 1 P 2d . ( )
where
P ee
SiGe L bb - > ebb!
i=] i=l i=1
4. Substitution in Eq. (10-48) gives V which may then be used to derive the
feasible GLS estimator of B in Eq. (10-47) as
b, = (X’V-'X)
UX’ ly
where y and X denote the stacked vector and matrix in Eq. (10-47). The
estimated variance matrix is
est var(by) = OVX)
and the conventional tests on b, would be valid asymptotically.
denote the k xX 1 vector of coefficients for the ith unit, we set up the null
hypothesis
Hot BiB = Peak
This hypothesis may be tested by computing the test statistic defined in Eq. (8-91)
for the SURE model, where the = in that formula is the variance matrix of the u’s
defined in Eq. (10-46), line 1. The R matrix would be set up by reformulating the
null hypothesis as
B, = B,
B, = B;
B, = 8,
However, as shown in Sec. 6-1, the same test statistic can be derived from the
residual sums of squares from the restricted and unrestricted versions of the
model. Under the null hypothesis the restricted model is
Yi x, u;
: Mealae
¥; X u
Yp X, B, u,
Cr
E(uu’) = +2 @I m
Dp
/ l / l /
Cree Be fa im ae
where
h l
b = (=|>—xX’xX.
cae $1
. x/x,] y —X’y.
a xy, i
(10-49)
and all summations are over i = 1,..., p. The residu
al sum of squares from the
A SMORGASBORD OF FURTHER TOPICS 415
unrestricted model is
1 1
ee = Lyi cae Lyi Xib, (10-50)
i
= 5 1 (b, - b)Xy,
= 5+(b,- byX;X,b,
Finally it may be shown thatt
P
1 (10-52)
ee, —ee= > a) — b)’X’X,(b, — b)
j=] 0
where b and b, are defined in Eqs. (10-49) and (10-51). If, in addition to the
assumptions already made, the w’s are normally distributed, then under the null
hypothesis,
pylecesunre
ere AEG 6) Kp) as F[k(p— fe1), p(m—k)]
i
This development has, however, used the unknown o,,. Replacing them by the
estimated values s,,, the same test statistic can be computed, but it will now just
have asymptotic validity. This model has been extended to include lagged
variables and more complicated assumptions about the vy, vectors.{
Oy a Ope Dre:
n al Sey) <a
iO 0 em Sines eara
Onna ome Cneee prac BLOM, cel Soe Peg
ie oe oe ; ‘ Sia
1 Loa hese
l
A SMORGASBORD OF FURTHER TOPICS 417
If the w’s and v’s are normally distributed, the log likelihood is
n n 1 1 pis:
InL= —- 5nd = 5 Ino® = 5 In|2| ar
aa a XB) Q iy = XB)
B = (xQ-'x) 'x'Q-'y
and 6° = (y — XB)'@"'(y — XB)
However, y is unknown, but it is confined to the interval (0, 1), which suggests a
grid search. Substituting B and 6? for B and o? in the log likelihood gives the
concentrated function
In L = constant — sin se 5In|o (10-58)
Maximizing Eq. (10-58) over y yields ¥, which then gives (2. The feasible GLS
estimators are then
b, = (X'0-'x) 'x'O-'y
and s? = (y — Xb,)’07'(y — Xb,)
The asymptotic distribution of by is normal with mean B and variance matrix
o2(X’Q~'X)~!. The asymptotic variance matrix for (y, 0”) is more complicated
and is given in the first of the Cooley-Prescott papers.
The idea of adaptive coefficients can obviously be extended to slopes as well
as intercepts. This is done in the third of the Cooley-Prescott papers. To illustrate
the treatment consider the three-variable model
Bi,
Y, ee [1 X5, X3,] Bo,
Bs,
The assumptions now are
Bi, = Bi + u; (10-59)
Bei eno f= 1.2°3
PAP
it ee try
where the superscript p denotes the permanent part of a coefficient. The Cooley-
Prescott assumptions about u,, and v,, are
~ N(0,(1 — y)o73,
ee Cat iene P=Aeekn (10-60)
v, ~ N(0, yo?S,)
where u, and v, are 3 x 1 vectors. In addition the u, and the vy, are serially
independent, and u, and v, are independent for all s, ¢. The new feature is the
appearance of the 3 X 3 variance matrices >,, and &,,. For the estimation method
to work, these matrices have to be known up to scale factors. Thus they can be
418 ECONOMETRIC METHODS
normalized by setting, say, the element in the top left-hand position to unity.
Writing 2, as
LP OSe rn
r,=|9 o% 9%
0 033 033
implies that u,, is independent of u,, and u;, for all ¢, that the random
components in 8, and £, have variances proportional to 0} and o%4, respectively,
and that these same random components have a covariance proportional to 055. If
one has no reason to expect a nonzero covariance, 044 is set at zero and an
becomes diagonal with only two elements to specify. If one assumes that the
intercept is the only coefficient subject to transitory changes and that the
permanent changes are independent, the matrices become
=,={10 0 0 x, =|9-9%- 0
O00 OP ene
Finally if one assumes the slope coefficients to be constant, the matrices reduce to
jis ie LO)
2,=2,=10 0 0
Om Ose
and Pie Bian
Bf, ae Bite + Ui,
which is simply another way of writing the adaptive intercept case already
studied. The intercept is then B? (= a,), the equation disturbance is
u,,, and
Re
t petty at Dia rs
For the general case of variability in all coefficients Cooley and Prescot
t
Suggest that unless there is special a priori knowledge, one assumes
the matrices
2,, and 2, to be equal. In one practical application the diagonal element
s in this
common matrix were set equal to the estimated sampling varianc
es of the
parameters computed under the assumption of parameter constan
cy. The authors,
however, report that losses in efficiency are surprisingly small,
even for sizable
errors, in specifying 2, and ,.
The general model may now be sketched briefly:
Y= xB feee
eyall
where x, is the k X 1 vector of explanatory variables at time
t, including unity in
the first position to take care of the intercept. The variable-p
arameter assump-
tions are
B=BP+u, BP=BP,+y, ¢t=1,...,n
It then follows that
n+]
Br =p a Ds V,
s=t+]
A SMORGASBORD OF FURTHER TOPICS 419
and so
TXB
aS ,
ie (10-61)
Rael
where wW=x'u,—x, ), Y,
S=tt|
and emphasis is placed on estimating the permanent coefficients for the first
postsample period. The variance matrix for the disturbance term in Eq. (10-61) is
E(ww’) = o?[(1 — y)R + yQ] = 0’ (10-62)
where Ris a diagonal matrix with
ry = Xj2,X;
and Q is defined by
qi; = min(n —i + 1,n—j + 1)x,2,x,
Given =, and &,,, Q depends only on y. Thus a grid search over the (0, 1) interval
will yield a ¥ which in turn gives () and the estimators
b?,, = (XQ7'xX) 'xO-'y
2 e (y a Xb’, ,)'Q7'(y a Xb/?, ,)
and
n
The grid search for 7 is in terms of the concentrated likelihood function Eq.
(10-58), with Q now defined in Eq. (10-62).
We saw in Sec. 6-3 on dummy variables that there was no essential difficulty in
the incorporation of qualitative variables in the X matrix. It is, however, quite a
different matter when the dependent variable is qualitative or categorical in
nature. We may distinguish three main cases.
which is subject to some limit, whether upper or lower, or both. This is also
referred to as the case of censored, or truncated, variables.
Space forbids a treatment of all three cases. We will concentrate on the basic
ideas underlying the binary case, which are also the foundation for any treatment
of the more complicated cases.
t% = [Hoy dw (10-63)
Management would clearly like to know as much as possible about the distribu-
tion f(w). What value of wy, for example, would be required in Eq. (10-63) to
yield a probability in excess of, say, 0.5? If the distribution f(w) remained
constant over a sequence of contracts and various wage increases were subjected
to ballots, estimation of the parameters of f(w) would be a possibility. Alterna-
tively, at a given period in time, one might imagine a government mediator
sampling various groups of workers with a variety of hypothetical wage increases .
in an attempt to chart the f(w) distribution.
The main use of this type of analysis has not been in economics but in
bioassay.f Applications in economics are, however, increasing with the ever
expanding supply of micropanel data. In bioassay a specific dosage z, of, say, a
poison is administered to each member of a population (insect, animal, human).
The responses of the individual members are presumed independent of each other.
For a great variety of reasons the tolerance to the poison varies from individual to
individual and may be described by some distribution f(z). If the tolerance is less
than the dosage, the individual succumbs to the poison. Thus the proportion
of
the population dying at dosage z, is
% = [Ve dz
where F(-) is the cumulative standard normal distribution and yy = (Xo — )/9.
This is shown in Fig. 10-4, which is simply a repeat of Fig. 10-3 with the
horizontal axis translated to y. Inverting Eq. (10-64) gives
(10-65)
PX
Figure 10-3
422 ECONOMETRIC METHODS
™ = Fo) —--———s :
Figure 10-4
Given a value of y, one can read off the corresponding 7. Conversely, given 77, one
can read off the corresponding value of y. The y variable is defined as the normal
equivalent deviate (n.e.d.) or by the somewhat unattractive term “normit.” A
probit is defined as
Probit = y + 5
From Eq. (10-65) there is an exact linear relationship between the n.e.d. and
dosage or, equivalently, between probit and dosage. The n.e.d. will be negative
whenever 7 < 0.5, whereas the probit will almost never be negative.f Fisher and
Yates give a table transforming percentages to probits.+
In a typical experiment dosages X 1, Xz,..., X, are administered to n,,n5,...,
n, Subjects, respectively. The resultant proportions P\, P2,--+, Pg are measured.
The estimation procedure then follows directly from Eq. (10-65).
1. Convert the sample proportions P\> Pz». Pz into n.e.d.’s and plot against
dosage x.
2. If the scatter in step 1 is approximately linear, then fit the regression
NED = a+ bx (10-66)
where§
=p
a = estimate of ——
oO
:
I estimate 1
of —
oO
A simple OLS regression would be unbiased but inefficient since it ignores the
properties of the error structure. A GLS estimator may be obtained by taking
account of the likely nature of the errors. Write the sample proportions as
Dp, a8; Vices £
Thus}
je binomial ;5
n(1— 7
t
=I
F-'(a,) ne te
Seen as Naess,
ap; Pi=7;
Thus
ax Le 1
Ep Vas als ea, (10-67)
GG o&
dF!
where ie
ap; Dim
Returning to
ae | 2
Pi == F(y,)j= beac
See dy
+ The binomial distribution applies since each individual in the ith group is subjected to dosage x;
and hence to a probability 7, of death or whatever. Moreover, individual responses are assumed to be
independent of one another.
+ p = F(y) is a monotonic function, and so is its inverse. We have
dF dF!
dp =—eo dy and ly EE Ip
dy = ——d
Thus
aE ay os A
dp lp dF /dy
424 ECONOMETRIC METHODS
ane at (l—q.
7
eee (10-68)
The regression equation (10-67) thus has a heteroscedastic disturbance given by
Eq. (10-68). Feasible GLS estimators would be achieved by computing a weighted
regression of the empirical n.e.d.’s on dosage x using n,;Z?/p,(1 — p,) as weights.
The next extension to consider is where the stimulus or dosage is not a single
variable but some linear combination of variables. Thus the ith level of the
stimulus might be denoted byt
= eeD
where x; is a column vector of k variables and B is a k X 1 vector of coefficients
presumed constant over all individuals. For example, in the question of whether
or not to purchase a new car in a given year the x vector would include such
variables as income, the relative prices of cars and gasoline, the age of the present
car, and so forth. We still assume that each individual has a threshold level for car
purchase, and we postulate a distribution f(s) over the population, where s
indicates the threshold or minimum stimulus required to trigger a new car
purchase. Thus the probability of a car purchase at stimulus level 5; 1S
7, = ff) ds
If the f(s) distribution were normal with mean p and variance o7, then
7,= F|
=}
Sit:
where F(-) again indicates the cumulative standard normal distribution. The
observed sample proportions p, are transformed into n.e.d.’s, and the appropriate
regression is
ae FED | = F~'(;) + u;
or y, = i
Ss. =
+agyeas
U, LL= — og
x’B
7, = x'B
or p= x Btu;
If this is estimated by OLS, or by GLS taking account of the heteroscedasticity in
u, it may give a reasonable fit to “middle-range” data, but it is doomed to run
into difficulty for extreme values of x’,B since there is nothing in either procedure
to prevent estimated probabilities turning out to be negative or in excess of unity.
The more common and more sensible procedure is to model the probabilities
a, by some distribution function other than the cumulative normal. Perhaps the
most frequently used is the Jogistic.} This may be formulated as
Xj
a, = partgaUsoabona ia (10-70)
1+e%8 1 +e 7x8
Clearly, 7 is constrained to the (0, 1) interval. It increases monotonically with the
stimulus x’B, it equals 0.5 when x’B = 0, and it has a shape similar to that of the
cumulative normal.t It is, however, simpler to work with than the cumulative
normal.
It follows directly from Eq. (10-70) that
T.
| =X;xB (10-71)
10-71
that is, the logarithm of the odds ratio or /ogit is an exact linear function of the
x’s. As before, the observed sample proportions p; = 7, + ¢; follow the binomial
distribution
7m a
[Doe binomial, nN:
1
We seek a relationship between the observed logits and the true logits. Letting
fp) = In ap
|
+ The classic reference is D. McFadden, “Conditional Logit Analysis of Qualitative Choice
Chap. 4.
Behavior,” in P. Zarembka, Ed., Frontiers in Econometrics, Academic Press, New York, 1974,
London, 1970, p. 28, Table 2.1.
+ See D. R. Cox, The Analysis of Binary Data, Methuen,
426 ECONOMETRIC METHODS
f( 2) = f(7;,) PIER oa
1 |pj=,
and |of l
=——_
Op; Pi=7, m(1 — 7)
Thus
Vf ]=x’ B+ u, (10-72)
Dep; ‘
h
where a
fae aan
so that
l
E(u;)=0
j= and — var(u,))=——_—_
Sa (10-73 )
1. Compute the observed logits In[ p,/(1 — p;)| from the sample proportions.
2. Carry out a GLS regression of Eq. (10-72) using the disturbance variances
obtained from Eq. (10-73) by replacing the unknown m1, by p,.
So far in both the probit and the logit approaches we have assumed that there
were several observations at each level of the stimulus so that sample proport
ions
could be computed. In some cases this may be infeasible and we Just have a
single
observation, y = 1 or y = 0, at each xB. The scatter would then look like
Figs
10-5.
f ee
1 2
Z
ye
S
Ov Ko?
a oS
7 oS
YES
eo
x
4
4
Yi
4
Ze
Vz
7
ee
—©-_-©-@ Ss @)— Agee > x8
ya
Figure 10-5
A SMORGASBORD OF FURTHER TOPICS 427
n= Pr(y, = 1) =
Ss.
e I
ee ‘ (10-74)
and 1 — a, = Pr(y, = 0) eae
Suppose that r responses and n — r nonresponses occur in a sample. Let us
reorder the sample observations so that the responses come first and the nonre-
sponses last. The log likelihood is then
d dln(1 — 7, ) eo . a
oF ap leer! ot
eww“ - tS —_— ._ => —-T7: 3
Thus
dln L
a = y TX;
@ i=r+l
Mm
v
M-~ sas en
I i=1
wx = Do ax, (10-75)
al a
The left-hand side is the sum of the x vectors just for the individuals displaying a
response. The right-hand side is nonlinear in B, and an iterative nonlinear
program is required for the estimation of B. The asymptotic standard errors may
be obtained as follows. From Eggs. (10-74)
On,
i=" e*
Os; (1 +e%)
= (1-7)
Thus
Oa SC 7, )X,
428 ECONOMETRIC METHODS
and
2 n n
Ga
dB de
op’ i ae ay ts
2B ~ $3 AG os 7;)X ;X’;
i=] i=]
R(B) = my 7, (1 — 7, )X ;X,
i=]
So far we have implicitly assumed that the X variables have been measured
without error and that the only form of error in the equation has been in the
disturbance term u. The latter has generally been thought of as representing the
influence of various explanatory variables that have not actually been included in
the relation. It could, of course, also have a component representing measurement
error in the dependent variable Y, and the previous results would still be valid.
We now pose the question of what happens if the X variables are subject to
measurement error. We assume that the B vector represents the coefficients of the
correctly measured X variables. Thus the model is assumed to be
y=XBP+u (10-76)
where X is the n X k matrix of the true (but unobserved) values of the explana-
tory variables. The matrix of observed values is
X=X+V (10-77) ©
where V is the n X k matrix of measurement errors. If some variables are
measured without error, the appropriate columns of V are zero vectors. Combin-
ing Eqs. (10-76) and (10-77) gives the following relation between the observed
variables:
y = XB + (u — VB) (10-78)
The OLS estimator of B in Eq. (10-78) is then
1. The
~
measurement errors in X are uncorrelated in the limit with the true values
X. Thus
Aas
plim(—-Xv] =0
and so
1 “Di,
z= plim{—X’X} = plim 1 i
—yX, —LX?
n
1 p
eer ae
where p and o7 denote, respectively, the mean and the variance of X. Further
1 0 0
sate villi) anes 1
Q plim|; Vvv] plim 0 “Zo?
mAliOey.0
DynOiviNor
since there is no error in the dummy variable for the intercept term. Substitution
in Eq. (10-79) gives
i Aili Get 1 —po,B
aly 5 (OMe fe 0,8
430 ECONOMETRIC METHODS
from which
lim(b)= B - =
ORs
Soe
B
po ae o* +0, l +107 /a7
Errors of measurement in X thus bias the estimate of 8 downward. The per-
centage bias is approximately given by the error variance as a percentage of the
variance of the X values. The estimate of the intercept is also inconsistent, and
this result extends to the multivariate case: even if some explanatory variables are
measured correctly, all coefficients will in general be inconsistent.
The measurement error in the X variables thus poses a possibly serious
estimation problem, and alternative estimators are required. There are two main
types of estimator described in the literature. One is based on instrumental
variables of various kinds and the other on ML methods, buttressed with fairly
strong assumptions about the covariance matrix of the measurement errors.
Before describing the estimators it is worth emphasizing the possibility that in
certain circumstances economic agents may react to the measured values rather
than the true values of economic variables. Firms may base investment decisions
on some extrapolation of national income trends and in so doing will use the
latest national income statistics complete with such errors as they contain. If
decision makers respond to measured data, then the measurement error is
irrelevant and our previous techniques will be valid.
where x; and X, denote the means of the values above and below the median and
Y, and Y, the means of the corresponding Y values. The estimator of the slope is
Bo
OO eX
Deal
and the intercept is estimated by
aw = x a bX
This procedure amounts to partitioning the data into two subsets by the median
X value and passing astraight line through the mean points (X,, Y,) and (X3, Y,).
If n is odd, one should omit the central observation before beginning the
computations. This estimator was first proposed by Wald.+ Under fairly general
conditions the Wald estimator is consistent but likely to have a large sampling
variance. Bartlett has shown that the efficiency may be increased by dividing the
X values into approximately three equally sized groups, the first containing
the n/3 smallest X values and the third the 7/3 greatest X values.t Omitting the
central n/3 observation, the slope is estimated by
IV
Yi
xX, ae 1
and the intercept as usual by ayy = Y — bX.
Extension of the grouping methods of Wald and Bartlett to more than one
explanatory variable is cumbersome and tedious. A somewhat different IV
estimator suggested by Durbin does not have this drawback.§ The suggestion is to
rank the X values in ascending order and then define the Z matrix as
TESRba 16gray
Pea yaa yee eter
where the second row indicates the rank values of the X’s.{] Substitution in Eq.
(10-80) then gives the estimate of the slope as
ey
by
= alt ale (
10-81
8 )
+A. Wald, “The Fitting of Straight Lines if Both Variables Are Subject to Error,” Annals of
Mathematical Statistics, vol. 11, 1940, pp. 284-300.
+M. S. Bartlett, “Fitting a Straight Line when Both Variables Are Subject to Error,” Biometrics,
vol. 5, 1949, pp. 207-212. It is easily seen that this is equivalent to making the second row in Z’
consist of equal numbers of zeros and plus and minus ones according to the ranks of the X values.
§ J. M. Durbin, “Errors in Variables,” Review of the International Statistical Institute, vol. 22, 1954,
DD yo aoe
4 With this formulation plim((1/n)Z’Z) would not exist as required for the consistency of the IV
estimator. However, if the second row is replaced by 1/n,2/n,..., 1, the condition will be satisfied
and the same estimates as in Eqs. (10-81) and (10-82) will result.
432 ECONOMETRIC METHODS
Y=a+ BX +u
pte pet a t=l--yn (10-83)
with X, = X, + v,
where X denotes the observed value and X the true unobserved value. The u term
is an amalgam of the conventional disturbance term and any measurement error
in Y. Thus the model might be written equivalently as an exact relation between
two variables, both subject to error, that is,
Y,=a+ BX,
5 : (10-84)
with eee tas and X, = X, + 0,
The errors u, and v, are assumed to follow normal distributions with the following
properties:
Case 10-1. X,, X,,..., X, are a set of given numbers. This case has two possible
interpretations. One is that the set of X’s can be held fixed in repeated sampling.
This situation would be of little interest, even in the experimental sciences, for if
the X’s are truly unobservable, how can the experimenter know that they have
been held constant in repeated trials. The more useful interpretation, especially in
the social sciences, is the one treating the X’s as fixed amounts for making
inferences conditional on the set of X’s underlying the sample observations.
Case 10-2. The X’s are random drawings from a normal distribution with mean p
and variance o*. This is hardly a plausible description of the generating mecha-
nism of most economic variables, but this case leads to the simplest estimating
equations and there are interesting parallels between the estimators in the two
cases.
If the X’s are fixed, then so are the Y’s, and the assumptions already made in
Eq. (10-85) would ensure zero covariances between errors and true values.
A SMORGASBORD OF FURTHER TOPICS 433
Specifically
E(X,u,) = E(X,v,) = E(¥,u,) = E(¥,v,)=0 forall (10-86)
If, however, the assumptions of Case 10-2 apply and the X’s and hence the Y’s
are random variables, the conditions in Eq. (10-86) would constitute an additional
set of assumptions.
Estimation of Case 10-2. Given the assumptions listed above, the observed X, Y
values would come from a bivariate normal distribution which is fully determined
by the following five parameters:
E(X)= E(X)= p
E(Y) = E(Y) =a + Bp
var(X) = 07 + 02 (10-87)
var(Y) = of + of = B’o* + 0,
cov( X, Y) = cov( X,Y) = Bo?
The ML estimates of the parameters on the left-hand side of Eqs. (10-87) are
given by the corresponding sample statistics, and we then hope to solve the
resultant equations for estimates of the parameters of the model. The estimating
equations for a, B,... are
m,, = 6° + 62 (10-88)
where the m’s indicate second-order moments of the sample data, that is,
é= Y—BX (10-90)
2. Knowledge of 62. This is perhaps a less likely situation than prior knowledge
of 0, since «2 incorporates both the measurement error in Y and also the
conventional equation error. If, however, we have a prior estimate s?, the
fourth and fifth equations in Eq. (10-88) yield
p= —— (10-91)
If s;, were zero, this estimate becomes the reciprocal of the slope in the OLS
regression of X on Y.
3. Knowledge of the ratio \ = 62/02. After some manipulation the last three
equations of Eq. (10-88) now give
with roots
The sign of ® must be the same as that of m,,. This will be so only if the
numerator of Eq. (10-93) is positive, and that in turn will be so only if the
positive sign before the square root is taken. Thus the estimator is
Estimation of Case 10-1. We now assume that there is a set of unknown values
X,, X,,..., X, underlying the sample data, and we wish to make inferences
conditional on this set. We still retain assumptions (10-84) and (10- 85). The log
likelihood function is
=
In L = constant — 7n ino, - 7n ino, Be es
X (x, — X,)
=, \2
~, \2
= LS) fea eae (10-95)
20,7 i=]
ioe 2n rv
i mew : at ey™ = 2Bm =r B>m.,,) (10-96)
TV ee
~ t(n
— 2)
vl-r?
2 1/2
(1 = 2)|(m,., — i) + 4m?,|
The corresponding limits for B are the tangents of these angles. The assumptions
required for the development of Eq. (10-97) render this essentially a large sample
method, and, of course, all the above rests on exact knowledge of A, which is not
often likely to be forthcoming. The technique may be extended to a multivariate
regression if the investigator has knowledge of the ratios of all the error variances.
Details are given in the Kendall and Stuart treatise.
+See M. G. Kendall and A. Stuart, The Advanced Theory of Statistics, vol. 2, Griffin, London,
1961, pp. 383 ff.
+ M. G. Kendall and A. Stuart, op. cit., pp. 385-386.
§ M. G. Kendall and A. Stuart, op. cit., pp. 388-391.
436 ECONOMETRIC METHODS
PROBLEMS
10-1 Prove Eq. (10-5) by the method suggested in the text. [Hint: Remember that expressions such as
x’,(X’,_,X,_,)'x, are scalars and may be moved back and forth in matrix formulas, that is,
cAB = AcB = ABc, where c is a scalar and A and B are matrices.]
10-2 Relation (10-5) is a special case of a general result given by Plackett.+ His problem and method
of proof may be stated as follows:
First sample data y}, X,(" X k)
The problem is to find the simplest computational way of updating least-squares statistics from the
first sample to the complete sample.
Method: Define
2 =i ,
R, = X,(X,X,) X45
R,R=R,-R
and hence that
CyeRat R) = I,,
Then show that
(XX)
, zi
"= (XX)
, ik
— (KX)
, al
XS [L, + Ry] "X2(XX,)
, = , re
Finally show that this result yields Eq. (10-5) when X, is just a row vector of observations on one
additional sample point.
10-3 For the recursive residuals defined in Sec. 10-1, prove
RSS, = RSS,_, + w
[Hint: Express y, — X,b, as y, — X,b,_, — X,(b, — b,_,). Partition
18 Yr-1 = De 1
yn [75| and x= ("|
Applying the partitioning again and using Eq. (10-5) gives the desired result.]
10-4 Take a simple time series and verify that the restricted estimation of Eq. (10-17) yields the
same
point estimates of the a and £ parameters as those derived from the estimated coefficients of the spline
function (10-14).
7+ R. L. Plackett, “Some Theorems in Least Squares,” Biometrika, vol. 37, 1950, pp. 149-157.
A SMORGASBORD OF FURTHER TOPICS 437
10-5 For the disturbance term in Eq. (10-21) make the following assumptions:
E(u? = 6;
B=J,@I,,
G,
2 Oo
2
CeO 22 ra 2
ania 2 eee
ee eeew N
i = 1,2,..., n (panel members), j = 1,2,..., ¢ (time periods), and the X’s are exogenous variables.
The ¢;; are assumed to be normally and independently distributed with zero mean and constant
variance for all i, /.
(a) If X3,, is not observed and an investigator regresses Y,; on just X,;; and X,,; with a constant
term in the regression, what is the bias in the least-squares estimate of a5? If the algebraic sign of the
simple correlation coefficient for X3;; and X3,; were known, is this sufficient information to determine
the algebraic sign of the bias? If not, explain what information is required to determine the algebraic
sign of the bias.
(b) If the unobserved independent variable X%; ; is assumed to satisfy X3;; = X3; for all j and is
assumed to be nonstochastic, explain how to obtain estimates of a, and a, and their associated
standard errors.
(University of Chicago, 1977)
438 ECONOMETRIC METHODS
E(yireis) = E(eréis) = 0
E(e) =o? E(e7) =o2 E(en8:2) = p02 E(e€€:2) = p02
E(x#?)=02 = E(x4x%)=p,02 —foralli;t = 1,2;5 = 1,2
Xjz» Vix ate Observed for i = 1,..., N; t= 1,2. Let b be the IV estimate of b from a cross-section
regression using data from the second time period and x,, as the instrument. Let b be the IV estimate
of b using the same cross section hut with y,, as the instrument. Show that if p,, p,, and p, are all
positive, then plim(b) < b < plim(4).
(UL, 1981)
CHAPTER
ELEVEN
SIMULTANEOUS EQUATION SYSTEMS
So far our interest has centered mainly on the inference problems associated with
a single equation, although there was some discussion of groups of equations in
Chap. 8. Economists, of course, often focus on a single equation, such as an
aggregate consumption function, a demand function for gasoline, a wage-change
equation, and so forth. However, economic theory teaches that such equations are
embedded in a system or subset of related equations. Thus one must examine
whether the presence of these related equations has any implications for the
estimation of the focus equation. More importantly, the estimation of a complete
system of equations is often an important practical problem, whether the objec-
tive is to test economic theories about the nature of the system or to use the
complete system to make joint predictions of a set of related variables.
In this section we will consider a few very simplified systems in order to illustrate
the main problems that arise, and then in subsequent sections we will give a more
general and formal treatment.
Consider first an even simpler income determination model than the one
outlined in Chap. 1. This one consists solely of a consumption function and the
national income identity, namely,
C=a+
t
BY, + u, (11-1)
Y= C4
t (11-2)
439
440 ECONOMETRIC METHODS
1. u~ NO, 021)
2. Z and u are independent, which will be satisfied if either Z is a set of fixed
numbers or Z is a random variable distributed independently of u. Z could be
taken as representing autonomous investment and government spending
controlled by some central authority. The model does not discuss the determi-
nants of Z.
Ci a
Tee B Gee
ate (11-3)
a ]
Y-Toptpepot (11-4)
+ As shown in Chap. 1, the reduced form is obtained by solving the model so as to express each
current endogenous variable solely in terms of exogenous variables and lagged endogenous variables.
¥ If necessary, review the discussion of consistency in Sec. 7-2 and illustrations of inconsistency in
the presence of lagged variables in Sec. 9-2 and in the presence of errors of measurement in Sec. 10-6.
§ The range is, of course, infinite for a normally distributed disturbance, but the finite range is a
convenient assumption to keep the diagram simple.
SIMULTANEOUS EQUATION SYSTEMS 441
a+ BY
Figure 11-1
then trace out points in successive periods in the range P, to P, along the Y — Z’
line. If Z never changed from Z’, these would be the only points ever observed for
this economy, no matter how many observations were taken. The estimated
regression of C on Y would coincide with the line Y — Z’, and the estimated
marginal propensity to consume would be unity, no matter what the true B happened
to be. Now suppose that over a large number of time periods Z ranges between Z’
and Z”. Observations on C and Y would then fill in the parallelogram P, P, P; P,.
The least-squares regression of C on Y minimizes the sum of squares of the
residuals measured in the vertical (that is, C) direction. Thus in the limit the OLS
line will tend to pass through the points P,, P,;. The estimated slope will now be
less than unity but will still be greater than the true B, so that the asymptotic bias
is positive.
DoCz
and bry = Dyz (11-6)
where c, y, and z denote deviations from the sample means. From Eqs. (11-3) and
(11-4) we may derive
eee
Yen = 7B
2 + }izo
]
Lyz = Lz? + Lz
=f,
Thus, provided
a OL =
plim(—:20 =0 and plim( 52" =M,,
a finite number,
plim(d,;y) = B
and hence plim(a,y) = a
6+ ae
Sz
variables and not deviations from sample means. We will reserve the letter y for
endogenous variables so that y,, denotes the rth observation on the ith endoge-
nous variable. Likewise x,, will denote the th observation on the / th exogenous
variable. The structural parameters B and y also have two subscripts, the first
indicating the equation and the second the variable to which it is attached.
Model (11-7) would be a conventional demand-and-supply model if y,
denotes price, y, denotes quantity, and we impose the restrictions
y\
S: Bai¥, + Yo + ¥21 = 0
yf H—-————— ——
|
| Diet Bynys
vig
|
||
O | WY
V3 2 Figure 11-2
SIMULTANEOUS EQUATION SYSTEMS 445
where A = | — £,,8,,. The first term on the right-hand side of each equation is a
constant. Thus we may write the reduced form more simply as
(11-9)
Vara Fy TOiy
y21 = by 7 U2,
where
hy =
See inee
ma
yt
f
2
ees
(11-10)
7
OTS
— Mir Bio
A
t
Ome
~ ai
_ B A
t
+Ure
If we postulate that
E(u,)
, Gije 3Cq2
lag 1 ie ss
=
then
E(v,)=0
9 pe
B3101, + Aa 61
— e281ae
a E(v3,) = Eire
var(v,)
and
— B01; — B29 tall 5 B 2B) Ov
cov(v;, v>) as E(v,0,) am A
+ Here there are no lagged endogenous variables and the only exogenous variable is the dummy
variable x,;, = 1 for all t, which is required to take care of the intercept term in the structural
equations.
446 ECONOMETRIC METHODS
E(y,) ="
E(y)
= Mo
var( y,) = var(v,) (11-11)
var( y,) = var(v,)
= BY ty) 2= Uy (11-12)
01}, = 9, = 1 01, = 0.5
Equations (11-7) define a model, and a structure like Eqs. (11-12) is obtained from
a model by assigning specific numerical values to the 8 and y parameters and also
to the variances and the covariance of the u’s. Solving this structure for Eqs.
(11-9) gives
yy =2+0,
Va Aas
u, — 2u
where v, Leen
ee 3u, LS
+u 2
—1.5
cov( y1, ¥2) = cov(0,, 02) = o—
The true structure (11-12) is, of course, known only to the “deity” who sets the
economic system in motion. Now suppose that one of the deity’s vice-presidents
tinkers with the institutions in an attempt to confuse the econometricians of the
world and concocts a new structure by the following rule, where (1) and (2)
SIMULTANEOUS EQUATION SYSTEMS 447
The new structure obeys the same a priori constraints on signs as Eqs. (11-12).
Solving this structure for Eqs. (11-9) gives
y=a2t+vy
yy = 44 v3
where
uy = 945 uy
— 205
LU aeOV Se, CMEep
10g, 2us) Buy,
DoT eNG ITCtie reg Ae Ae
Thus the five parameters of the reduced form E(y,), E(y2), var(y,), var(y2), and
cov(y,, ¥>) are identical for the two different structures and indeed for all
structures derived by taking linear combinations of the original structural equa-
tions.
It is instructive to see what type of further information might help identify
one or both equations of this model. There are three basic possibilities, namely,
(1) restrictions on the 8 and y parameters, (2) restrictions on the 2 matrix, and (3)
respecifications of the model to incorporate additional variables. To illustrate the
first category, suppose the supply function is presumed to go through the origin.
The a priori restriction is thus
Yr,
=9
This reduces the number of structural parameters to six, but the number of
reduced-form parameters is five, as before, so that it is still not clear that any
structural parameters can be identified. However, making the substitution y,, = 0
in Egs. (11-10) gives
isp!
be An
oo Buti
oe A
me
so that be
es By
448 ECONOMETRIC METHODS
= Bin095
var( y,) oi A2
0.
var( y,) = a
bee — B 12%
cov(y;, 2) = = aie
so that
ea — var(y,)
cov( y;, V2)
and thus the slope of the demand function is identified. Taking expectations of
the demand function in Eqs. (11-7) gives
Yin = Shake
and substitution for », and pw, from Eqs. (11-10) verifies that this relation holds.
Thus y,, and £,, can both be expressed in terms of the parameters in Eqs. (11-11),
and the demand equation is identified. This case is pictured in Fig.11-3. The
combination of o,, =0 and o,, + 0 generates a set of observations on the
demand function.
A less extreme version of this case would occur if o,, were “small” as
compared with o,,. The scatter of observations would then tend to be con-
centrated around the demand function rather than lying exactly on it. However,
knowledge about the relative sizes of disturbance variances is not likely to be
generally available, though a possible reason for a large o,, might be the omission
of important explanatory variables from the supply function in Eqs. (11-7). The
appropriate remedy is the respecification of the supply function to include such
variables. In practice the demand function should also be looked at since the
simple two-variable model of Eqs. (11-7) is hardly a realistic specification with
which to commence empirical work.
SIMULTANEOUS EQUATION SYSTEMS 449
where A = 1 — £,,8,, and the v’s are given in Eqs. (11-10). Let us denote the
reduced-form coefficients by 1;, (i= 1,2; j=1,..., 4). It is clear that the
structural coefficients can be obtained from the reduced-form coefficients. For
example,
Bo, = eh)
21 TT
Let us assume a linear model containing G structural relations. The /th relation at
time ¢ may be written
It is plausible to assume that the B matrix is nonsingular since, if it were not, one
+ Notice that for the moment we‘have not normalized the structural equations
by setting any of the
8 coefficients at unity.
SIMULTANEOUS EQUATION SYSTEMS 451
where
E(y,|x,) = Ix,
Thus the mean of the conditional distribution of y,, given x,, depends solely on
the IT matrix. A finite sample of observations (y,,x,; t = 1,..., 1) will yield some
estimate ne which will deviate from the true II due to the fluctuations of random
sampling. Suppose, however, that we dispense with sampling problems by assum-
ing that an infinitely large sample of observations can be made available. In
general the true II may then be determined with any desired degree of precision.
This is all that can be afforded by the sample data. Thus knowledge of B and
can only come from knowledge of II.
To see the same point in a likelihood context, let us assume
u,~ N(O, =)
and also that the u, vectors are serially independent. It then follows from Eqs.
(11-19) that
v,~ N(O, &)
where C= BoISBa) (11-20)
and the y, are serially independent. From the reduced-form equation (11-18)
+ When x, contains lagged y values, this expectation has to be read as conditional on these lagged
endogenous values.
452 ECONOMETRIC METHODS
P(yIx,) = p(u,) eS
= p(u,) - ||BI|
where ||B|| denotes the absolute value of the determinant of B. The likelihood of
the sample y’s conditions on the x’s is then
L =(20) aoe
"°"|BII"|2|
n St
Perel— 3Dw
l
t=1
[SS
1,
0
There may also be linear homogeneous restrictions involving two or more
elements of a,. The specification that, say, the coefficients of y, and y, are equal
would be expressed as
a
eee ee are rea re . =)0)
0
If these were the only a priori restrictions on «,, they may be expressed in the
form
a,® = 0 (11-24)
where
0 1
(ee |
® =) 1 0
00 00
The ® matrix has G + K rows and a column for each a priori restriction on the
first equation.
In addition to the restrictions embodied in Eq. (11-24) there will also be
restrictions on a, arising from the relations between structural and reduced-form
coefficients. From Eqs. (11-19) we may write
BIL+T=0
or AW =0
454 ECONOMETRIC METHODS
where W= LT
The restrictions on the coefficients of the first structural equation are thus
aW =0 (11-25)
Combining Eqs. (11-24) and (11-25) gives
a[W o]=0 (11-26)
There are G + K unknowns in a,. The matrix [W 9] is of order (G + K) X (K
+ R), where R is the number of columns in ®. On the assumption that II is
known all the elements in [W ©] are known. Thus Eq. (11-26) constitutes a set
of K + R equations in G+ K unknowns. Identification of the first equation
requires that the rank of [W ®] be G+ K — 1, for then all solutions to Eq.
(11-26) would lie on a single ray through the origin. This suffices to determine the
coefficients of the first equation uniquely, for in specifying the general model in
Eq. (11-17) a B or y coefficient was attached to each variable in every equation.
Normalizing the first equation by setting one coefficient at unity (say, B,, = 1)
will now give a single point on the solution ray, and this determines a, uniquely.
e[W ®]=G+K-1 (11-27)
is clearly a necessary and sufficient condition for the identifiability of the first
equation. The condition for the identification of the ith structural equation is
e[W ®]=G+K-1
where ®, is the matrix embodying the a priori restrictions on the ith equation. The
basic difficulty with the rank condition, as stated in Eq. (11-27), is that it is not a
convenient one to apply since it requires the construction of the II matrix, which
is complicated even in small models. We will give below an equivalent condition ~
in terms of structural parameters which is easier to apply. However, condition
(11-27) does yield necessary conditions for identification which are very simple to
apply. Since[W ®] has K + R columns, a necessary condition for Eq. (11-27) to
hold is that
IM drdk = Gar ix = Il
or Kea Geael (11-28)
that is,
The number of a priori restrictions should not be less than the number of
equations in the model less 1.
When the restrictions are solely exclusion restrictions, the necessary condition is
restated as:
The number of variables excluded from the equation must be at least as great as
the number of equations in the model less 1.
SIMULTANEOUS EQUATION SYSTEMS 455
+ See F. M. Fisher, The Identification Problem in Econometrics, McGraw-Hill, New York, 1966,
Chap. 2; or for a shorter proof, R. W. Farebrother, “A Short Proof of the Basic Lemma of the Linear
Identification Problem,” International Economic Review, vol. 12, 1971, pp. 515-516.
456 ECONOMETRIC METHODS
structure could yield a new structure which satisfied the same a priori constraints
as the original structure and had identical reduced-form coefficients. Let
A=[B TI]
denote an original set of structural coefficients (that is, with specific numerical
values), and let FA denote a new structure obtained from A by premultiplication
with an arbitrary G X G nonsingular transformation matrix F. The new structure
is said to be admissible, or equivalently F is said to be an admissible transforma-
tion matrix, if FA satisfies all a priori restrictions on A.f Identifiability of the first
equation then requires that the first equation of every admissible structure be
some scalar multiple of the true first equation. The first row of A may be
expressed as
a, =e,A
where e, is a 1 X G row vector with unity in the first position and zero elsewhere.
Thus the a priori restrictions on the first equation may be written
e,(A®) = 0
The first row of coefficients in the transformed structure may be written as f,A,
where f, denotes the first row of F. For an admissible structure this must obey the
same restrictions as «,, and so we must have
f,(A®) = 0
Identifiability requires that f,A be a scalar multiple of e,A, that is, that f, be a
scalar multiple of e,, which gives the condition that p(A®) = G — 1. If all the
equations of a model are identified, the only admissible transformation matrices
are diagonal matrices.
7 The general definition of admissibility also requires that the variance matrix of the transformed
disturbances satisfy all the a priori restrictions on the original variance matrix, but we are restricting
consideration here to the structural coefficients.
SIMULTANEOUS EQUATION SYSTEMS 457
Yi2 0
and A® = ne] = ie
0 1gach
that is,
Bum, + Bim, + W11 = 9
By M2 + ByyM + Y2 = 9
Yi2 = 0
If we normalize by setting, say, 8,, = 1, these give
pis 2
9
ig, 22
and Nitta aah
eaterare
2
which shows explicitly how the parameters of the first equation may be
derived uniquely from those of the reduced form. The parameters of the
second equation may be obtained in a similar fashion.
-—-
O°
oO
458 ECONOMETRIC METHODS
Be i
and A® = a
which has zero rank. Thus the first equation is not identifiable; nor is the
second, for this is the case we alluded to in Example 11-1, where x, appears
in neither equation.
ib at Yi2 = 0 Yx.
=0
This example might be treated in two ways. In one approach we note that the
restrictions y,, = 0 = y>, mean that x, does not appear in the model at all.
Thus the model could be reduced to one with just a single exogenous variable,
in which case the only restriction is y,, = 0, and that suffices to identify the
first equation, but leaves the second unidentified. Alternatively, retaining the
dimensions of the original model, the restrictions on the first equation give
0 O
ATEN ; 205 ad)
® = 1 0 with A® ey |
OF al
Thus p(A®) = 1 = G — 1, and so the first equation is identified. For the
second equation
0
® eer= O0 with A® ae |
]
so that this equation is not identified. Alternatively, for the second equation .
a,[W ®|=0
gives
™m, M2 O
Yi =e ¥o
—0
SIMULTANEOUS EQUATION SYSTEMS 459
= ie a oe |Pte fara
7, 72 A= Yq aay)
and av =|F]
0
In all the above examples readers should check for themselves that the
necessary condition (or order condition, as it is often called) would correctly
indicate the presence or absence of identification. This need not always be the
case. For example, if 8,, in Example 11-5 were zero, the rank condition would fail
even though there is one restriction on the second equation.
Treatment of Identities
Identities themselves do not raise any identification problems since in general the
coefficients are known and indeed are usually unity. The general model
By, + Ix, =u,
may, however, be formulated in two alternative fashions. In one version all
identities appear explicitly in the model. In the alternative version the identities
may be substituted in other structural equations, thus effectively reducing the size
of the model. The identification rules may be applied to either version. Solving .
out the identities will not change any conclusions about the identifiability of any
behavioral or other structural equation whether in its original or revised form.
As an illustration consider the simple supply-and-demand model
qP=at+aptu
Gq = By + Bip. + Baw tu
q?=¢°
where q? = quantity demanded
q° = quantity supplied
P price
w = an index of weather conditions
This is a model containing three endogenous variables q”, g°, and p (G = 3) and
two exogenous variables w and z (a dummy variable) set at unity to take care of
the intercept term in the first two equations. Rearranging the model in more
SIMULTANEOUS EQUATION SYSTEMS 461
D
1 OS aeO. 0 — Qo ei uy
0 losweBin Aa! 1 Bo i! seg
ee | 0 0 LO) w 0
Z
For the first equation
0 0
A® = ae Bo
al 0
and p(A®) = 2 = G — | so that the equation is identified. Notice that when we
have exclusion restrictions, the A® matrix can be written down directly by taking
the columns of the A matrix which contain zeros in the row corresponding to the
equation under study. For the second equation
1
A® = |0
1
which only has rank unity, and so the second equation is not identified.
If we rewrite the model without the identity, it becomes a two-equation model
in two endogenous variables q and p,
T= Ay 1 aptr uy
g=B)+B\p+
Bw u,
where now G = 2, and the first equation is again just identified because it has one
restriction on its coefficients while the second equation is not identified because
there are no restrictions on its coefficients.
it can be written as
Bitty vein ©
plus the normalization rule B,, = 1. Thus inhomogeneous [Link] be
recast in homogeneous form before normalization and the previous procedures
still apply.
462 ECONOMETRIC METHODS
where
+ The argument to the contrary in G. S. Maddala, Econometrics, McGraw-Hill, New York, 1977, p.
230, is incorrect. Maddala investigates identifiability via transformation matrices. However, he
aia
essentially postulates a transformation matrix
and then finds that the restriction implies A = 0, which leads him to conclude that both equations are
identified. But F has already assumed that the second equation is identified, which is an invalid
assumption. The identifiability of both equations has to be considered jointly.
SIMULTANEOUS EQUATION SYSTEMS 463
ie ye
GRE a
the transformed structure is
by + fo Boi) V1 + for
Yo + (firn + foyYo1) X = uy
The requirement that the transformed structure satisfies the same a priori
constraints as the original structure, namely, that y, does not appear in the
first equation, gives
fio = 9
The normalized transformed structure is then
Veo Yi te
(2 + frBo havin + fora x ia
|
Vt Io = iS)
fo fo
If we now impose the cross-equation constraint y,, + y2, = 0 on the original
structure, the same condition on the transformed structure gives
ae fallFa =
or fay = 9
giving
fa, = 9
so that all admissible transformation matrices are diagonal and both equa-
tions are identified.
Alternatively the reduced form of the model is
Vi re
Vola (Bayi a Yo1) x 105. ¥11 (Ba am 1)x, + Uv,
The parameter y,, can be obtained from the first reduced-form coefficient and
8, can be derived from the second reduced-form coefficient, so that both
equations are identified.
So far the only explicit assumption about the disturbances has been that of serial
independence, but we have made no explicit assumptions about contemporaneous
correlations between disturbances in different structural equations. Let
z= E(u’)
> is then a G X G matrix, the terms on the principal diagonal indicating the
464 ECONOMETRIC METHODS
fir + fi2Ba = 1
fiz = 9
giving f;, = 1 and f,, = 0. The only restriction on the second equation is the
normalization condition, which is held in abeyance. Thus admissible transforma-
tion matrices are given by
] 0
he B al
showing that the first equation is identified and the second not.
Suppose we can now postulate
O11 0
> —
0) 05>
~ wale Sf
f,2f5,= 0
that is,
fo), = 0
which gives
fy, = 9
The value of f,, is then settled by the normalization condition that the coefficient
of y, in the second equation must be unity. The coefficients of the transformed
giving the coefficient of y, in the second equation as f,,. Thus f,, = 1, and the
rls
only admissible transformation matrix is
+ It is convenient algebraically, but not necessary, to impose the normalization condition on the
coefficients of y, and y, in the first and second equations, respectively, of the transformed structure.
The absence of y, from the first equation gives |, = 0. The zero covariance term then gives f,, = 0
and so the class of admissible transformation matrices is
ea
r-(4 |
which secures the identification of both equations.
466 ECONOMETRIC METHODS
a, 0 0 hry
[1 0° O} 0) <o55% 0 1 | =,0;,=0
0 Ona to
so that
hay = 0
and in a similar fashion the second and third conditions gives f;,; = 0 and f;, = 0.
Thus the only admissible transformation matrix is
Ee0a 0
F = \0ipele0
Oe Oe
and all three equations are identified.
The above model has two special features, namely, a triangular B matrix and
a diagonal = matrix. The presence of these two features defines a recursive system.
All the equations of the recursive system are identified and, as we shall see below,
simple estimation procedures are available for this model.
Zero covariances can aid identification and not necessarily just in recursive
systems. For example, in
Yt Byy2 =
Bo¥, + Yo + Yn1%1 = U2
the first equation is identified and the second is not. However, the additional
specification o,, = 0 would serve to identify the second equation as readers can
easily prove for themselves. There is no simple necessary and sufficient condition
for the zero covariance case as there was for restrictions on the B and y
parameters, so each case must be examined from first principles.
SIMULTANEOUS EQUATION SYSTEMS 467
The discussion has dealt only with models which are linear in variables and
parameters. Many realistic models, however, may be nonlinear in variables
and/or a priori restrictions. Identification theory for such models is difficult and
has only been partially developed. Owing to the unsatisfactory state of the theory
it will not be summarized here. Interested readers should consult Fisher.
Recursive Systems
As we have seen already, the two crucial features of a recursive system are a
triangular B matrix and a diagonal 2 matrix. As an illustration consider the
model
Viet WX = Ue
Boi Vis + Yar + Yai, = Ure
with the specification
o,, O
E(uu’) = = =
0 05>
To explore the connection between the y’s and the u’s we look at the reduced-form
equations which are
Vie = TVM1% + U1
(11-30)
When Bis lower triangular, then so is B~'. Thus Eq. (11-30) gives
Vir =f (uy)
Yor =f (Us Uo,)
Vou = F (tags Uays M34)
0
Yor = f(Uygs ays <+s MGs)
The assumption of a diagonal = matrix then ensures that y,, is uncorrelated with .
u5,, that y,, is uncorrelated with u,,, and so forth. Thus the second structural
equation may be estimated consistently by an OLS regression with y, as the
dependent variable, the third with y, as the dependent variable, and so on.
It is also easy to show that if the u’s are normally distributed, OLS yields ML
estimates. As was shown in the previous section, the likelihood of the sample y’s,
conditional on the x’s, for the model
By, + Ix, =u,
is given by
aids ut 1 n
For recursive systems |B| is unity and = and ="! are both diagonal. Thus finding
the B and f to minimize L is equivalent to finding the B and f to minimize
n
Lu, 2 'u,
t=1
SIMULTANEOUS EQUATION SYSTEMS 469
]
91)
n : 1
‘ Ur, Tee us
= — $+ SH
Re oolioe . -927-, 9 O33
Thus the partial derivatives of In L with respect to the coefficients of the ith
structural equation are simply the partial derivatives of
2
se
ga Oi
Setting these partial derivatives to zero gives the OLS equations for the ith
structural equation. Thus under the special assumptions of the recursive model
the OLS estimators of the structural equations will have the desirable properties
of consistency, asymptotic normality, and efficiency. They will also have the usual
small sample properties.
Vit X11
y21 X91
Y= Ne ands ky =
YGt X kt
+ For a proof that the usual small sample inference procedures apply see E. Malinvaud, Statistical
Methods of Econometrics, 2nd edition, North-Holland, Amsterdam, 1970, pp. 679-681.
470 ECONOMETRIC METHODS
OY ees aaa
/ /
er kop a =) he ae
Y= ; xs ;
/ /
7 Yn a i xX, oot
| -B =|5
0 (11-37)
0
J J y
KX Ge Goo) Kx 1
Substituting in this from Eq. (11-35) gives the ILS coefficients as the vectors b and
c obtained by solving
: c
(XX) “X’Y =p} = i (11-38)
0
The crucial question is whether there are unique solution vectors b and c.
Rewriting Eq. (11-38) as
l c
(XX) 'X’[y Y, Y,] Fi |0|
0
gives
will have the same number of columns as Yj. This suggests using
[X, X,]
as the set of instruments for [Y, X,]. The resultant IV estimates are given by
which are identical with Eqs. (11-40) and (11-41). Notice that the ordering of the
instrumental variables is unimportant. We can just as well take
x= [X, X,]
z=
1
(Yee
1
xsl] and ||
the IV estimator of 8 is
b z,
diy = | I= (X’Z,) 'X’y (11-42)
Civ
which is easily seen to be identical to the ILS estimator defined in Eqs. (11-40)
and (11-41).
= Y/X(X’X) 'X’Y,
and VX, = (ViANV,)X,
= YX,
Thus the equations for the 2SLS estimator can now be written
The equivalence between Eqs. (11-45) and (11-46) may be proved by the reader as
an exercise.
matrix is
—_
X’X =
SY >
nT
oo
Se)"=) oO
ohowore
]
and so
WVX(X°K)
©XY) = [1 te eet |
2
Y’X(X’X) 'X’y =[0.1 0 05 0.5] ; OG
1
The 2SLS equations are then
1:65, 1 Ocha. 2
Li 910% 0) ine raeaen ee
0 0 © Sre5 3
with solution
bi 1.6667
Crp P= 11020333
Cp 0.6000
Example 11-9 For a model
Vie = Bio aetna
Yor = Bar Wie + Yo2%or + 23X31 + Ur,
SIMULTANEOUS EQUATION SYSTEMS 475
Ps el-[4
The 2SLS equations are thus
with solution
loaBey
The second equation is just identified and thus may be estimated by
2SLS or ILS. For the 2SLS approach
| | ae |
y=|%2 Y= ("1 X,=|*%2 %3} X,=|*1
+In this and the previous example the X’X matrices are assumed to be diagonal to keep the
arithmetic simple. In realistic situations orthogonal variables are, of course, very rare.
476 ECONOMETRIC METHODS
Thus
5 40 10
XVA="40 V5 ie X’y = | 20
20 30
7 4a 20 ele 4
A= ba ag |0 10
lowe 0 5
Y/X(X’X) 'X’Y, =[5 40 20]]}0 0.05 0 || 40
0 0 0.1} 20
5
=[5 2 | 0|- 5
20
10
Y/X(X’X)'X’y =[5 2 2]] 20] = 150
30
The 2SLS equations are thus
145 40 20}] 5), 150
20 QO: 10 |Hhes 30
with solution
by, D
> |= \'—3
P’ = (X’X) ‘X’Y
LO 0 ar ah!
= Obs OOS eer 40 20
OnRO OF 120> 30
Selo
=r 2 1
ee
SIMULTANEOUS EQUATION SYSTEMS 477
giving
Viper ekap eka Oe
and Vor LU take, teks es,
The reduced-form matrix may also be estimated by substituting the estimated
structural coefficients B and f in Eq. (11-19),
i= -B-'f
However, care must be taken in making this substitution since Eq. (11-19)
was derived from the structural equations specified as By, + Tx, = u,, whereas
the equations of this model have been specified with just a single endogenous
variable on the left-hand side of each equation. The 2SLS estimates of the
structure are
Valea 1) 21me 4tx1, + uy,
Vp = Vi, — 3g, — Xz, + Ud,
Rearranging with all variables on the left-hand side gives
Xt
| eee Bole are On| ee =e
—2 1 y20 0 Se ery Uy,
3t
Thus
fe Lt eet “N44 0 0
ae 1 Ore ar ei
A SPss as
10> %35
These are somewhat different than the OLS estimates. The reason is that the
OLS estimates are unrestricted and thus fail to satisfy the restrictions placed
on the reduced-form parameters by the overidentification in the system. With
two endogenous and three predetermined variables there are six reduced-form
coefficients, which are functions of just five structural coefficients. The true
reduced-form matrix is
1 Yu Bir¥x. Bi z¥x3
II =
which is the difficulty with Z, in Eq. (11-47). Provided a matrix W can be found
such that
S plim(—-W'u] =0
the IV estimator
Le Tila)
so that Y, is the set of instruments for Y,. The IV estimator defined in Eq. (11-48)
is then
VV eye x! oy i Yiy
xy 11-50
(11-50)
xiY, Xx XxX, Civ
but we have already seen that Y/Y, = Y/Y, and Y/X, = Y/X,. Thus Eqs. (11-50)
and (11-44) are identical, so that 2SLS is in fact an IV estimator with Y, as the
instruments for Y,.
The consistency of the 2SLS (IV) estimator requires the three conditions on
W, stated above, to be fulfilled. We will assume that
ale!
plim|Ww) and plim|,W72}
SIMULTANEOUS EQUATION SYSTEMS 479
1 plim(+ Yu]
plim(—-W'] = ¢ =0
ek ore le,
plim| Xiu}
=0
since the first two terms are finite and the last is the zero vector.
It was also shown in Sec. 9-2 that the IV estimators are asymptotically
normally distributed with an asymptotic variance matrix estimated by Eq. (11-49).
Substituting for W and Z and using the fact that Y/Y, = Y;Y, and Y;X, = Y;X,
gives
b NY UN |
asy var =a in
c KYA DKK)
Se WX(XX)
/ , =A
XV Gn YX |
r , =
nn
XY, xX,
where
2 = = Vib — X,0)(y — Yb = Xo) (11-52)
n
which is a consistent estimator of «7. Some authors prefer to use the number of
degrees of freedom n — g — k + 1 as the divisor in s” rather than n. This is also a
consistent estimator of 0,7. The 2SLS estimators are thus consistent and asymptot-
ically normally distributed with estimated variance matrix given in Eq. (11-51).
A problem sometimes arises in the application of 2SLS to equations in
medium-size or large-size econometric models. The difficulty is that the number of
predetermined variables in such a model may become large in relation to the
number of observation points. Suppose, to consider a special case, that the
number of predetermined variables becomes as great as the number of observa-
tions, K = n. The X matrix is then square and, in the absence of any exact linear
+ The detailed conditions for this to be true are set out in H. Theil, Principles of Econometrics,
Wiley, New York, 1971, pp. 484-488.
480 ECONOMETRIC METHODS
and 2SLS is equivalent to OLS. The 2SLS estimates would, of course, no longer
be consistent, since the matrix of instrumental variables is now W = [Y, X,] and
ial
plim(—-Yiu +0 so that plim(--W') +0
(X’X)p, = X’y,
(X’X)p, = X’y,
X’(Xp — y,) = 0
J J
KXn nx1
Since X’ has rank n (< K), the only solution vector is Xp — y, = 0 so that
¥, = y,. The same result will hold for each variable in Y, so that once again
Y, = Y,, and 2SLS would be equivalent to OLS.
Various suggestions have been made for dealing with the problem of an
excess of predetermined variables. Kloek and Mennes suggested replacing X, in
the first-stage regressions by a smaller number of principal components.} Let F
denote the n X / matrix of the / chosen principal components and then define
Z=[X, F]
This Z matrix takes the place of the X matrix in Eq. (11-45), and the 2SLS
estimates based on the principal components approach would then be given by
from these two estimators with OLS and FIML, Klein found 2SLS based on just
four principal components to give the smallest absolute percentage error followed
by the other 2SLS estimator, OLS, and FIML in that order.+
An alternative approach based on instrumental variables has been suggested
by Brundy and Jorgenson to bypass the substantial computation involved in
calculating the reduced-form coefficients required for Y,.¢ Let
E(Y,) = XII,
where II, is the K X (g — 1) submatrix of reduced-form coefficients relevant to
the variables in Y,. The Brundy-Jorgenson suggestion is as follows.
The regular 2SLS estimator satisfies these conditions, for ne = (X’X) 'X’Y,
is a consistent estimator of II, and W, then becomes [Y, X,]. The novelty of the
Brundy-Jorgenson approach is to avoid computing reduced-form coefficients and
to derive an appropriate I, by first obtaining B and fas consistent estimators of
B and [ and then using
bat
from which the relevant submatrix II, can be extracted and XII, computed for
insertion in Eq. (11-55). Thus even if one is interested in Just a single structural
equation, this approach requires the initial computation of consistent estimators
of all structural coefficients. On the other hand, if one is estimating all
the
equations of a model, the single [I matrix is used to provide all relevant IT,
submatrices.
Several suggestions are offered for initial consistent estimation of the B and
r
matrices, all of them essentially IV estimators. Considering Eq. (11-47) again,
the
matrix of right-hand side variables is
Z, = [Y, X,]
where Y, is n X (g — 1) and X, isn X k. Define
Wi = [xt X,]
7 L. R. Klein, “Estimation of Interdependent Systems in Macroeconometrics,” Econometrica, vol.
STNO6SS ppl 92»
¢J. M. Brundy and D. W. Jorgenson, “Efficient Estima
tion of Simultaneous Equations by
Instrumental Variables,” Review of Economics and Statistics,
vol. 53, 1971, pp. 207-224.
SIMULTANEOUS EQUATION SYSTEMS 483
= [F, X,]
where F, is a subset of g — 1 principal components of X. This differs, of course,
from the Kloek and Mennes procedure, where the principal components were
used in quasireduced-form estimation to compute Y,. Here the principal compo-
nents are used as instrumental variables ina first-round estimation of structural
coefficients. The Brundy-Jorgenson estimator is known as the limited-information
instrumental variables efficient (LIVE) estimator. The asymptotic variance-covari-
ance matrix for d is estimated by
2 — LY = ZA)'y ~ Za)
n
The LIVE estimates can thus be computed even where the 2SLS estimates cannot,
but the actual point estimates will, of course, vary with the variables chosen as
instruments.
Y,By - Xv =u (11-58)
where
Let us suppose that the endogenous variables have been so numbered that Y,
constitutes the first g such variables and likewise that X, refers to the first k
predetermined variables. The likelihood function for the endogenous variables in
Y, will involve the parameters in the first g rows of the reduced-form matrix II.
Let these rows be partitioned into the two submatrices [II,, II,,] which are of
order g X k and g X (K — k), respectively. We know that
BIT
= -T
The first row of each side of this equation may be written
[Bx 0,JII=[-y' 0]
where 0, indicates a row vector of G — g zeros and 0, a row vector of K — k
zeros. Using the partitioning of II then gives
Bl ante (11-60)
BxIT,, = 0, (11-61)
Eq. (11-61) constitutes K — k homogeneous equations in the g elements of By.
However, one of the 8’s has been set at unity so that we merely need to determine
the ratios of the elements in B,. This can be done uniquely if the rank of II,> is
g — |. Even in the overidentified case where K — k > g — 1 and II,, thus has g
rows and at least g columns, the rank of II,, cannot exceed g —) a eats ts
obvious intuitively since Eq. (11-61) is just a subset of equations from BII = —T,
which gives the relations between the true structural coefficients and the true
reduced-form coefficients. However, the true II,, is unknown, and when it is
replaced in Eq. (11-61) by, say, the ML estimate Ihe this matrix in the
overidentified case will almost certainly have rank g so that one cannot solve
for
nonzero B,, except by arbitrarily dropping one of the equations.
The limited-information maximum likelihood (LIML) approach is to maxi-
mize the likelihood function for the g endogenous variables in Y, subject
to the
restriction that e(II,5) = g — 1. This approach was developed by Anderson
and
Rubin.¢ The application of the method requires one to know, in additio
n to the
specification of the equation being estimated, merely the predetermined
variables
appearing in the other equations of the model, as in 2SLS. The
mathematical
development of the LIML estimator is complicated and lengthy,
but it may be
shown that it reduces to the choice of the elements of B, to minimize
jo B’Wars By
Zi Br Wa By (1 #62)
where Wy‘, and W,, are certain matrices of residuals.+ The explanation of these
residuals is given in the following account of least variance ratio (LVR) estima-
tors.
Rewrite Eq. (11-58) as
z=X,ytu
where z= YP
so that the z vector is a linear combination of the endogenous variables appearing
in the equation, the coefficients of the combination being the unknown B
parameters. If z is regressed on X,, the residual sum of squares is
l
_ BiWesBs
BxWa Ba
which is the same criterion as that for the LIML estimator. Differentiating / with
respect to B, and setting the result equal to the zero vector gives
(Wiis — 7Wys)By = 0 (11-65)
This set of equations will only have a nontrivial solution if the determinantal
equation
|Waa a LWy al =)
is satisfied. This gives a polynomial in /, which must be solved for the smallest
root /. This root is substituted back on Eq. (11-65) and the estimator B, obtained
+ T. W. Anderson and H. Rubin, op. cit.; see also W. C. Hood and T. C. Koopmans, op. cit.,
Chap. 6. Hood and Koopmans arrive at Eq. (11-62) by a different method from the original approach
of Anderson and Rubin, who maximized the likelihood function subject to appropriate constraints by
using Lagrange multipliers. Hood and Koopmans start with the likelihood function for the complete
model of G equations for all G endogenous variables and then, by a series of stepwise maximizations,
eliminate from the likelihood function all parameters other than those of the equation to be estimated.
Finally, even y is eliminated and the concentrated likelihood function expressed in term of By.
486 ECONOMETRIC METHODS
from
a= Y, By
and regressing Z on X, gives
w= Wd+ Vv (11-79)
where the definition of the symbols in Eq. (11-79) is obvious from the comparison
with Eq. (11-78). The variance matrix for the v vector is
Ol opt --- ol
V= E(w’) =| 1 oyI +--+ gl |= Tel (11-80)
ani a Jo nee ay
The variance terms in Eq. (11-80) follow directly from Eq. (11-76). The typical
covariance term is
Si
a7a,
ae
ie
for alli, j
giving
V=Se1
The 3SLS estimator of 8 is then
dssts = (WV
'W) (WV! w (11-81)
SIMULTANEOUS EQUATION SYSTEMS 489
G
wel ZeX(X’X) Xy J
vt
P matrix defined in Eq. (11-73) is also square of order K and nonsingular. Thus
W, = P’X’Z,
is K X K and nonsingular. This result holds for all i = 1,..., G. Thus the
block-diagonal W matrix defined in Eqs. (11-78) and (11-79) is nonsingular and
each component submatrix is nonsingular. The 3SLS estimator defined in Eq.
(11-81) may then be written
d, = W, l ‘Ww,
Thus the collection of 2SLS estimators for the complete system may be written
“t
d,
[wrt 0 Daal Mb)
dosts = <1 a 0 wW,' 0 : catia
dq 0 0 ars WwW, Wo
with BU)SSOF er Sa Ee
E(u) ==
If it is assumed that the G disturbances follow a multivariate normal distribution
>
we may write
any
Gey
1
P| 7l ue
"y-1
u,]
Assuming, in addition, that the u vectors are serially uncorrelated, the likelihood
for the n vectors u,,U,,..., u, is then
If we write
l
|
vz A'S'Az, = — 5tr(ZA’E- 147’)
7—
— 5tr(B>'AZ/ZA’)
where
yx)
Z=(Y XJ=/y x,
yc xy,
is the n X (G+ K) matrix of observations on all the endogenous and prede-
termined variables. Defining
M = loz
n
tr(=~'AZ’ZA’) = ntr(=~'AMA’)
492 ECONOMETRIC METHODS
Total Number of
number of stochastic Estimation
Country Data} equations equations method
Australia Q 82 42 OLS
Austria A 128 54 OLS
Belgium Q 25 19 OLS
Canada A 183 44 OLS
Finland Q 144 60 OLS
France A 32 19 OLS
West Germany A 137 51 FIML
Italy Q 104 53 OLS
Japan Q 78 43 OLS
Netherlands A 87 13 LIML and 2SLS
Sweden A 133 75 OLS
United Kingdom Q 226 106 OLS
United States Q 207 70 OLS
Developing America A 12 11 OLS
Developing South and East Asia A 14 13 OLS
Developing Middle East and Libya A 10 9 OLS
Developing Africa less Libya A 11 10 . OLS
PROBLEMS
11-1 For the model defined by Eqs. (11-1) and (11-2) show that
11-2 Prove the equivalence between the alternative expressions for the 2SLS estimator in Egs. (11-45)
and (11-46).
11-3 The structure of the Klein model is
T=
By + Bl + B,T_, + B3K_, + uy
Y=C+/+G
i= Y= W,—7
KG Kee tel
The six endogenous variables are Y (output), C (consumption), / (net investment), W, (private wages),
II (profits), and K (capital stock at year-end). The four exogenous variables are G (government
nonwage expenditure), W, (public wages), T (business taxes), and f (time).
Examine the rank condition for the identifiability of the consumption function.
11-4 Tintner’s model of the U.S. meat market is specified as follows:
Vir + BiaYar + WX ie = Me
the y’s are endogenous, the x’s exogenous, and u/ = [u,v ,] is a vector of serially independent
normal random disturbances with mean zero vector and the same nonsingular covariance matrix for
494 ECONOMETRIC METHODS
yy 1) xy xX X3
yy 10 0 l 0 al
V2 0 10 =i = 0
xy ] sail! ] 0 0
X5 0 cal 0 l 0
x3 =I 0 0 0 l
(a) Estimate the parameters a, and a, by 2SLS and test the hypothesis
a, = 0 against the
alternative a, + 0.
(6) Repeat part (a) using IV estimates of a, and a5 obtained with x5,
as an instrument for Yo,
and x,, as its own instrument.
(Yale University, 1980)
SIMULTANEOUS EQUATION SYSTEMS 495
11-8 (a) Assess the identification of the parameters of the following five-equation system:
Calculate:
(a) Least-squares estimates of the unrestricted reduced-form parameters
(6) ILS estimates of the parameters of Eq. (1)
(c) 2SLS estimates of the parameters of Eq. (2)
(d) The restricted reduced form derived from parts (6) and (c)
(e) A consistent estimate of E(€)7&,) = 0)
(WEF 1973)
11-10 Let the model be
vy -|
80.0 -4.0 rox 2.0 10 -3.0 7
-40 50 -0.5 15 05 -1.0
3.0520 0 0
0 DD 10 0
X’X =
0 0 12 OREO)
0 0 0 0.5
find the 2SLS estimates of the coefficients of the first equation and their standard errors.
(UL, 1970)
11-12 In the following market model
Supply = =9,=ByP,
+ Yio + 41,
Demand = Q, = ByP,+ Y20 + Y21Z11 + Y22Z2¢ + U2,
quantity Q, and price P, are endogenous, while income Z,, and the price of some other good Z,, are
exogenous. If the supply function is estimated directly by least squares, will the resulting estimate of
8, be biased? If so, in which direction will the bias occur?
(UL, 1972)
11-13 If
7 0 oan
wee 10 en)
ce Son Se 4
l 0 bet
Only the first of these exogenous variables has a nonzero coefficient in a structural equation to be -
estimated by 2SLS. This equation includes two endogenous variables, and the least-squares estimates
of the reduced-form coefficients for these two variables are
iP pas |
Lie eal el
Taking the first endogenous variable as the dependent variable, state and solve the equation for the
2SLS estimates.
11-15 For the model
Vie = BizYar + YX + Uy
Yay = Bar Vir + Yo2%21 + ¥23X3, + U,
you are given the following information:
SO Al
KD MO Ss
2. The estimates of variance of the errors of the coefficients in the first reduced-form equation
are |,
Ws), Ofte
3. The corresponding covariances are estimated to be all zero,
4. The estimated variance of the error on the first reduced-form equation is 2.0.
SIMULTANEOUS EQUATION SYSTEMS 497
Use this information to reconstruct the 2SLS equations for the estimates
of the coefficients of the
first structural equation, and compute these estimates.
(UL, 1969)
11-16
Vit = Bia Yar + Big ¥3e + Yx1, + uy,
is One equation in a three-equation model which contains three other exogenou
s variables x5,, x3,, and
X4,- Observations give the following matrices:
20 15 ae) 2 2 4 2) ; : ;
YY =] 15 COP = 457 ¥ X= 0 4 Wh 5)|| 2,05 O10 eee
4S 70 ON 2 ee 10 0 OF Os
Obtain 2SLS estimates of the parameters of the equation and estimate their standard
errors (on the
assumption that the sample consisted of 30 observation points).
(UL, 1968)
CHAPTER
TWELVE
ECONOMETRICS IN PRACTICE:
PROBLEMS AND PERSPECTIVES
A careful study of the material covered in the previous eleven chapters would not,
unfortunately, equip the reader to conduct a successful piece of applied econo-
metric research, since that involves many more problems than those already
discussed. We will tentatively explore some of these issues in the present chapter,
but the reality should be faced at the outset that it is not feasible to write a.
comprehensive manual that would prepare applied econometricians for all the
problems that can arise in a wide variety of research projects. Successful econo-
metric modeling is not a collection of mechanistic and routine procedures but
more of an art requiring wide-ranging knowledge and judgment. Such an art is
best learned by practice, hopefully with talented supervisors and colleagues, and
by study of “best practice” examples. It is, however, not always easy to find the
latter. Indeed a very instructive book might be written under the title, How NOT
to Do Econometrics, with every chapter illustrated by one or more published
articles. The author of such a book would have to time its publication carefully in
relation to his own impending demise or retirement from contact with his
professional colleagues: he might also face the difficult problem of choosing some
of his own previous work for inclusion.
There is a widespread view that econometrics has in some sense not lived up
to its early promise, and there is much scepticism about the value of the plethora
of empirical results embedded in the literature. This state of affairs should not be
too surprising. There is, after all, a sound proposition in economics that the use of
a good or service tends to expand to the point at which price and marginal utility
498
ECONOMETRICS IN PRACTICE: PROBLEMS AND PERSPECTIVES 499
are equated. If computers are essentially free goods, if “researchers” can plug into
a data bank without any understanding of where the series come from or how
they were constructed, if they can press buttons to implement computer programs
whose contents they dimly comprehend, it then follows, as night follows the day,
that some work of zero worth will emerge. Indeed, given the uncertainty inherent
in the research process, compounded by the fallibility of the researcher, some
outputs may err on the wrong side of zero and be positively dangerous. Our twin
defenses against further encroachments by a flow of dubious work lie in improv-
ing still further the quality of the editorial screening process and raising also the
quality of the training given to would-be practitioners.
require the modeling of the refining decision and of the relationship between the
price of crude and the prices of refined products. In this decision process it is of
great importance to have as much knowledge as possible of what may be called
the “institutional realities” of the situation, specifically in this case such things as
the nature of the refining process and the constraints on the refining decision, the
quantitative importance of various groups of consumers, and the crucial factors in
their decision processes. An econometrician coming cold to the study would run
the risk of very slow progress with much searching through inappropriate formu-
lations. In my own experience collaboration with an experienced oil specialist
greatly improved the research efficiency.
Knowledge of the “institutional realities” is, of course, valuable in all areas.
In a study of cost-output relationships in coal mining this author felt it necessary
to don a safety helmet and get to the coal face in the narrow and twisting seams
of the Lancashire coal field in order to see at first hand the nature of the
production process before sitting down to peruse the statistics at the regional
headquarters of the National Coal Board. Similarly in studies of scale, costs, and
profitability in road passenger transport and of cost-output variations in a
multiple-product firm the author spent time at each firm talking to accountants
and managers to study their accounting and decision processes before extracting
the relevant data by hand from the firm’s records.} To take a final data problem,
monetary theory postulates the demand for money to be positively related to
income and negatively related to the rate of interest. Each of the three nouns in
this proposition raises formidable problems of definition and measurement. There
are numerous definitions of money and almost continual evolution of payments
technology, there are many interest rates, and even income is not unambiguous.
When appropriate data series have been identified, the next decision in
time-series contexts is what data period (hourly, weekly, monthly, quarterly,
annual, or whatever) to use. Again if we had institutional information about
decision processes (who decides when about what) we could make the appropriate’
choice. If, for example, production decisions are revised at the start of each
month, a model of the production decision employing monthly data would have
the best chance of capturing the essential features of the process. Quarterly or
annual data would in this case involve an inappropriate aggregation over time,
thus making it difficult, if not impossible, to determine the lag structure. Often,
however, there is little firm information about decision procedures, and the main
choice between quarterly and annual data is based largely on a mixture of
empirical considerations and the objectives of the modeling process. As Table
11-1 shows, the macroeconometric models for the developed economies are split
roughly evenly between those based on quarterly and those based on annual data.
By far the most difficult problem of all is the initial specification of the
model, be it a single equation or a set of equations. By specification we mean the
+ For these and other studies see J. Johnston, Statistical Cost Analysis, McGraw-H
ill, New York
1960.
ECONOMETRICS IN PRACTICE: PROBLEMS AND PERSPECTIVES 501
following:
+ Our brief discussion cannot hope to do adequate justice to this topic. The interested reader will
find much nourishment in the elegant, entertaining, and enlightening E. E. Leamer, Specification
Searches, Wiley, New York, 1978.
502 ECONOMETRIC METHODS
at a time. Thus he computes all '°C, = 120 possible multiple regressions and the
attendant F statistics for the overall fit. The true value of all 120 population F
statistics is of course zero, but the reader would not be surprised to find that our
researcher discovers some significant sample regressions.+ His theoretical and
institutional knowledge enables him to write a plausible commentary on these
regressions and perhaps select one as the seemingly best theory for the explana-
tion of y. Sending the write-up to an editor, who likes to publish “significant”
results, guarantees another “scientific” paper and a further small step by the
author up the academic ladder.
Another variant of Example 12-1 is a theory that only identifies the three
candidate variables, none of which, in fact, has any relevance to y. A series of
investigators drawing different sets of sample data from y, x,, x», x; fail to find a
significant regression, consigning their computer printout to the waste paper
basket or filing cabinet, according to temperament. In either case their profes-
sional colleagues are unaware of this accumulation of “negative” results, and so
testing of the theory continues. Working at any conventional level of significance,
it is only a matter of time until a set of sample data is drawn that yields a
“significant” result, which will, of course, have a good chance of being published.
The moral of Example 12-1 is clear. In an area where theory is poor and
provides little guidance on specification to the researcher, data mining is a highly
dangerous activity. Combined with the propensity of editors to publish only
significant results, it can in extreme cases result in the publication of falsehoods
and the suppression of truth. However, take heart, faint reader, the above surely
cannot be a description of economics, the queen of the social sciences, richly
endowed with well articulated theory. Consider then Example 12-2.
Hy: B= B = By
Working at the 5 percent level of significance, the probability of accepting the null hypothesis
for any
specific model is 0.95. Assuming independence of the models, the probability of accepting
the null
hypothesis for all the models considered is (0.95)'7° = 0.0021. Thus the chance that the researcher
finds ar least one “significant” regression is 0.998. Working at the more stringent
| percent level of
significance, the probability of finding at least one significant regression is still as
high as 0.70. The
models will not all be independent of each other because of overlapping explanator
y variables, so
these startling probabilities need not be taken too seriously, but they do indicate
the nature of the
potential problem associated with data mining.
+A small but constructive step toward addressing this problem was taken a few
years ago by the
editors of the Journal of Political Economy, who initiated a section for the publication
of “confirma-
tions and contradictions.”
ECONOMETRICS IN PRACTICE: PROBLEMS AND PERSPECTIVES 503
+ The development of econometric theory has been heavily influenced by the early work of the
Cowles Commission, which emphasized problems of equation error to the almost total exclusion of
problems of measurement error. Little is known about significance levels or the relative properties of
different estimators when these problems jointly coexist, as indeed they do in practice.
504 ECONOMETRIC METHODS
Residual variance (R”) criterion. Most of the operational criteria have been
developed in the context of a single equation model. The first is the residual
variance, or R*, criterion. Suppose there are just two competing models for the
explanation of y, namely,
y=X,B, + u, and y = X,B,
+ u,
where X, is nonstochastic, of order n X k;, and of full column rank. Suppose that,
in fact, the first model is correct. If the second model is fitted, the vector of OLS
residuals is
e, = Moy
= M,(X,B, + u,)
where M, =1-—X,(XX,)
'X
Thus the residual sum of squares is
E(s}) = of
where
Criteria for individual coefficients. There are two important criteria under this
heading. Economic theory is rich in qualitative predictions about the direction of
various effects. Thus one looks for agreement between a priori expectation and
the signs of estimated coefficients. Second, one looks for correctly signed coeffi-
cients which have reasonable statistical significance. The latter criterion should
not be applied too stringently since we have seen, for example, that collinearity
among the regressors can inflate estimated standard errors. The R? criterion also
has implications for the significance level of individual coefficients. As shown in
Problem 5-12, R* only increases with the addition of an extra regressor if the F or,
equivalently, the ¢ statistic for that variable exceeds unity, which corresponds to
the use of a significance level of about 30 percent rather than the conventional 5
or | percent level.
The previous remark is in the context of a fixed sample size. However, any
substantial increase in sample size has implications for significance levels. As seen
in Chap. 5, the test of the hypothesis that a subvector of g elements in B is the
zero vector is given by
_ (€xex — e’e)/q if
Ae e'e/(n — k) Aga ie)
where e’e is the residual sum of squares from the unrestricted model and eye,
that from the restricted model, the relevant g variables having been omitted. This
statistic is written equivalently as
Re SoRA Usk
1a Re q
Thus even though R* — R% may be very small, the test statistic can become
arbitrarily large with increasing sample size. Using a given significance level, the
null hypothesis is more and more likely to be rejected as n increases. This point
has been emphasized by Leamer, who, along with others, argues that the signifi-
cance level for this kind of test should be adjusted downward for larger samples.+
with ie On (12-5)
If the restriction were valid, estimation of Eqs. (12-5) would involve just three
parameters, namely 8,, yy, and o,, whereas estimation of Eq. (12-3) involves four
parameters. However, Eq. (12-3) has, in fact, to be estimated to test the restric-
tion. The payoff is improved statistical efficiency of the parameter estimates if the
restriction is upheld. Comparing Eqs. (12-3) and (12-4) the restriction implies that
{y,} and {x,} have a common factor with root B,.t There may be no economic
rationale for the restriction or common factor. If so, it is likely to be rejected and
the “general” equation (12-3) cannot then legitimately be reduced to the “simpler”
form in Eq. (12-5). Sargan’s COMFAC program tests for the existence of
Pe yea ay(Lyx e,
=,»,
5(L)u
(12-7)
which involves considerably fewer parameters than Eq. (12-6). Suppose, for
example, that a relationship
Vy as Biy-| ee By y,—> a Yo*X, at YixX1-1 a Y2X 1-2 a ¥3%;-3 ntsv, (12-8)
was estimated and a common polynomial 6(L) = (1 — L)(1 — pL) found. The
relation (12-8) may then be written
(1—L)(1 — pL)y, = (1 — L)(1 = pL) (vs + IL)x, + ©,
which only involves three parameters instead of six. For estimation purposes it
may be put in the form
DV = Ve AX tay (NXeae ta
(12-9)
t
+J. D. Sargan and J. D. Sylwestrowicz, “COMFAC: Algorithm for Wald Tests of Common
Factors in Lag Polynomials,” User’s Manual, London School of Economics, London, 1976.
508 ECONOMETRIC METHODS
and it is assumed, as usual, that u ~ N(0, o7I,,). Now suppose a new set of m
(< k) observations on these same variables becomes available. On the assumption
that the original model still holds, the new observations may be characterized by
Yo= XoB+ uy
where E(u,u,) = o7I,,. The m observations are insufficient to allow reestimation
of the model, but one may forecast the yy vector by
Jo = Xob
The vector of forecast errors is
Since e’e/o* has an independent x?(n — k) distribution, it follows that under the
hypothesis of parameter constancy
/ / il / =,
7G. C. Chow, “Tests of Equality between Sets of Coefficients in Two Linear Regressions,”
Econometrica, vol. 28, 1960, pp. 591-605.
+D. W. Jorgenson, J. Hunter, and M. I. Nadiri, “A Comparison of Alternative Econometric
Models of Quarterly Investment Behavior,” and “The Predictive Performa
nce of Econometric Models
of Quarterly Investment Behavior,” Econometrica, vol. 38, 1970, pp. 187-224.
§D. F. Hendry, “Predictive Failure and Econometric Modelling
in Macro-Economics: The
Transactions Demand for Money,” London School of Economics, London,
September 1978.
ECONOMETRICS IN PRACTICE: PROBLEMS AND PERSPECTIVES 509
where
nee
er eay
The test of forecast errors in Eq. (12-10) may be extended to deal with joint
forecasts from the reduced form of a simultaneous equation model.}
An aspect of prediction which is frequently ignored, and unjustly so, is the
longer-term implications of the dynamic regression that has been estimated. For
example, return again to our hypothetical demand function for oil, and ask
questions such as:
The model’s answers to questions such as these have to be put up against the
intuition and good sense of the researchers themselves and, more importantly, the
intuition and good sense of informed critics. This may seem very “unscientific”
and perhaps it is, but it is nonetheless very important and in the next two sections
we present a brief discussion of some ways in which it is attempted.
We have suggested that there are various aspects of any specification which are
important, namely,
+ See P. H. Dhrymes et al., “Criteria for Evaluation of Econometric Models,” Annals of Economic
and Social Measurement, vol. 1, 1972, pp. 307-308.
510 ECONOMETRIC METHODS
This point was brought home to me forcefully and convincingly a few years
ago when I was working on an energy research project. Each month a report on
the econometric activity had to be presented to a steering committee in London,
presided over by Sir Alex Cairncross.¢ Cairncross had (and still has) a healthy
scepticism of econometrics, no doubt partly due to his days at the U.K. Treasury
when, to quote, “the young men might present me with thirty different equations
to ‘explain’ British imports, so that, at the end of the day, neither they nor I knew
what determined British imports.” Each month the econometric output was
subjected to his shrewd, informed, and penetrating scrutiny. Eventually, however,
there came a monthly report which secured the approbation, “I wouldn’t mind
getting on a plane and taking this to Riyadh.” Presumably I had been engaged in
some successful data mining or, perhaps, had been “learning by doing,” so I
suggested to him jokingly that the Cairncross test would appear in the next
edition of Econometric Methods. The two-step Cairncross test is thus as follows.
¥ Sir Alex Cairncross, a very distinguished British economist, was for many
years economic advisor
to Her Majesty’s Government and subsequently Master of an Oxford College.
I hesitate to give his
present address lest he be deluged with manuscripts from aspiring econometricians,
but I suspect he is
mostly to be found in his Scottish retreat north of the Solway Firth, enjoying
the Scotsman’s favorite
view “looking down upon England.”
+ The Bayesian approach requires a book of its own. The premier
references are A. Zellner, An
Introduction to Bayesian Inference in Econometrics, Wiley,
New York, 1971; and E. E. Leamer,
Specification Searches, Wiley, New York, 1978.
ECONOMETRICS IN PRACTICE: PROBLEMS AND PERSPECTIVES 511
p(y) = (2003)
'””en] 3 ian? — 1) i
where o, is known, but the mean p is unknown. This gives the vector
This is the likelihood for the sample observations, conditional on the parameters pu.
and o,, but the latter has been omitted from the left-hand side since it is assumed
known.
The first crucial element in the Bayesian approach is to postulate the
existence of prior information about p. This may come from theoretical sources,
from previous empirical studies, hunch, judgment, or what have you. Such
information cannot be exact, so it is formulated in a stochastic fashion. It is
theoretically convenient to model this information in a way that is compatible
with the likelihood in Eq. (12-12). This leads to the concept of the conjugate prior.
The prior pdf for p is thus taken to be normal and written
+ We are using p(-) to indicate a pdf for sample data and 7(-) to indicate a pdf for parameters.
This practice was suggested in K. M. Gaver and M. S. Geisel, “Discriminating Among Alternative
Models: Bayesian and non-Bayesian Methods,” in P. Zarembka, Ed., Frontiers in Econometrics,
Academic Press, New York, 1974, Chap. 2.
512 ECONOMETRIC METHODS
from which
PE ry
Pr( B) - Pr( A|B)
Pr( B| A) = ———_——— (12-14)
12-14
7(uIy) =
mu): p(y|e) (12-15)
P(y)
In Eq. (12-15) the expression 7(u|y) represents the posterior pdf for m, and
comparison with 7() indicates the change in the researcher’s beliefs about p
brought about by the sample information in y. The denominator in Eq. (12-15) is
given by
P(y) = Jp(y|u)7(H) dp
For given y, m, oj, and o” this reduces to a constant. Thus Eq. (12-15) can be
rewritten as
(ply) x e9|
-1) ee
Pineal
al |
m(ply) & o|- |
207o3/n Ore}. oo/n
where ji = Ly,/n. Thus the posterior pdf for p is also normal with mean
E(p) =
fi(ag/n)a' + m(o?)
a
| (12-16)
(o5/n) + (07)
and
1
var(p) =
(og /n)' + (0?)
Formula (12-16) shows that the posterior mean is a weighted average
of the
sample mean and the prior mean, the weights being the reciprocals
of the
respective variances. Strong prior information (low 0°) gives the prior mean
a
large role to play in determining the posterior mean, and conversely,
strong
sample information (large n and/or low 06.) gives the sample mean a dominat
ing
role. The importance of the posterior mean rests on a basic result in
Bayesian
statistics that if one assumes a quadratic loss function for errors in estimating p,
the estimate which minimizes the expected loss is the posterior mean.+
A parallel result holds in the linear regression case when the form of the
model is assumed known, but a multivariate Bayesian prior distribution is
specified for the B vector.t Assuming a multivariate normal prior, this involves
specifying the mean vector by and all the elements in the variance-covariance
matrix, indicated, say, by o7N, '. Assuming the usual linear model
y=XB+u_ with u~ N(0,071)
the mean of the posterior distribution is§
have no simple treatment for the nonnested case, where Bayesian procedures
admit, in principle, of a very simple solution.
For any model the marginal density of the observations M,, sometimes
referred to as the predictive pdf, is given byt
In Eq. (12-18) p(y|B;, M;) is the likelihood for the sample observations, condi-
tional on the model M, and its parameters B,, and 7(B,|M,) is the prior density for
the parameters, given that the model is M,. Equation (12-18) says that if M, is the
true model, the marginal pdf for the sample observations is found by taking a
weighted average of the sample likelihoods, where the weights are the elements of
the prior distribution for the parameters, given the model. Now suppose that
associated with each model there is a nonnegative fraction P(M,) indicating the
prior subjective probability that M, is the true model and, in this case, such that
P(M,) + P(M,) =1
The unconditional pdf for the sample observations is then
P(M,) p(y|M,)
P(M\ly) = (12-19)
P(y)
In comparing two models there are just two possible losses, one if M, is chosen
when M;j is the true model and the other if M, is chosen when M , applies. If these
losses were equal, the decision rule that minimizes the posterior expected loss is as ©
follows. Choose M, if
P(M,|ly)
P(My\y)
= P(M,) p(y|M,)
P(M,) p(y|M,)
(12-20)
is greater than 1. Equation (12-20) defines the posterior odds ratio, which is seen to
be equal to the prior odds ratio multiplied by the ratio of the marginal densities
(weighted likelihood functions). If there are more than two models, posterior odds
as defined in Eq. (12-20) can be computed for any pair.
It is clear from Eq. (12-18) that the formidable task in the computation of
Eq.
(12-20) is the evaluation of the marginal pdf’s. If a multivariate normal prior
is
assumed for B,, given M,, that is,
7(b;|M,) 1s N(b;+,Njx')
then Leamer has shown that p(y|M,) varies inversely with a quadratic form Q,
+ We are assuming unrealistically, but for simplicity, that there is no uncertainty about
the
disturbance variances.
ECONOMETRICS IN PRACTICE: PROBLEMS AND PERSPECTIVES 515
Or
where N, = X/X, and b, = (X,X,)~'X‘y,. The first term in Eq. (12-21) is the
residual sum of squares from the OLS fit of model M,. This has to be increased by
a factor depending on the discrepancy between the OLS vector and the prior
mean vector. Alternatively, the first term in Eq. (12-22) is the error sum of squares
if the coefficient vector were set equal to the prior mean vector. This is adjusted
downward by a term which is again dependent on the discrepancy between the
sample and the prior coefficient vectors. Thus apart from the prior odds ratio, the
choice between models would depend on these adjusted sums of squares, which
are a mixture of sample and prior information.
Readers must judge for themselves whether a criterion such as Eq. (12-20),
for all its elegance and simplicity, is a valid guide for choice. Suppose that just
two crude and simple models are being compared. M,, say, is a “Keynesian”
reduced-form equation relating GNP to “exogenous expenditures,” while M, is a
“Friedmanian” equation relating GNP to “money.” If the Ghost of Keynes could
be contacted, he would presumably offer a prior-odds ratio P(M,)/P(M,),
dramatically different from that forthcoming from Professor Friedman. How can
the protagonist of one theory begin to specify the prior pdf’s for the parameters
of the opposing theory, which he basically regards as false? Must then a Bayesian
researcher be certified ideologically pure and unbiased before being allowed to
specify prior odds and prior densities for model parameters?
A partial resolution to the problem of excessive dependence on priors, which
are, perhaps, spuriously precise, idiosyncratic, or just personal to one investigator,
is provided by some recent work by Chamberlin and Leamer.{ It is assumed that
in a single equation there are one or more “focus” variables, whose coefficients
are of crucial interest. The equation may also contain other “doubtful” variables.
The investigator specifies a prior zero mean vector for the doubtful variables.
However, he does not have to specify the elements of the prior variance matrix,
merely that it belongs to the class of positive definite or semidefinite matrices.
Leamer’s SEARCH program computes bounds on the focus coefficients, so that
the researcher can study the robustness of these coefficients under a variety of
specifications. This approach is appealing and seems likely to be developed and
considerably extended. It adds yet another dimension to the array of information
that we can obtain on any specific problem. How to weigh and interpret the
jigsaw of computation and information will still depend on the vital spark of
human imagination and powers of judgment.
The position can best be summarized by a quotation from the late Jacob
Bronowski.+ Though writing of the physical world, his comments are very
apposite to the economic and social world that we study.
The world is not a fixed, solid array of objects, out there, for it cannot be
fully separated from our perception of it. It shifts under our gaze, it interacts
with us, and the knowledge that it yields has to be interpreted by us. There is
no way of exchanging information that does not demand an act of judgment.
Science is a very human form of knowledge. We are always at the brink of the
known, we always feel forward for what is to be hoped. Every judgment in
science stands on the edge of error, and is personal.
A
MATHEMATICAL
AND STATISTICAL APPENDICES
The purpose of this section is merely to remind the reader of various notational
conventions for functions and derivatives. It is not intended to review the basic
rules of differentiation.} If y is a function of x, the relationship may be denoted
variously as
The derivative measures the slope of the function at a specific point and is, in
general, a function of x, as is emphasized by the f’(x) notation. Thus it may itself
+ For a lucid introduction to the calculus and other mathematical topics of special relevance to
economists see A. C. Chiang, Fundamental Methods of Mathematical Economics, 2d edition, McGraw-
Hill, 1974.
517
518 ECONOMETRIC METHODS
Vesa OX ore |
where the x’s are capable of moving independently of one another, then one may
study the change in y in response to the change in any one of the independent
variables (or arguments) of the function, the other independent variables being
held constant at any arbitrary set of values. This gives rise to the partial
derivatives, denoted by
In this notation f,, for example, would indicate the rate of change of y with
respect to x,. Once again further partial differentiation may be carried out,
yielding the second-order partial derivatives
d*y a ae =
Ox;0x, Ox,0x, “4
Alternatively, if the independent variables are given separate labels, as in
Vi flues)
the partial derivatives may be denoted by
Ces Cvs Oy
One Ap ie
Onde ae
though, even here, one may see f, used for f, and Jo lOEp.
y= b* b>0O (A-1)
This is called an exponential function since the variable x appears as the exponent
of the constant, or base, b. We rule out negative values for b, since if x were, say,
one-half, » would be the square root of a negative number, which is imaginary. If
x denoted time ¢ measured at equal intervals, then
y, = b' and Vt =b
Veet
(a)
x = log, b y A-2
This is the inverse of the exponential function. The first expresses y as a function
of x and the second expresses x as a function of y. Typical graphs for b > 1 are
shown in Fig. A-1. If the graph in Fig. A-15 were superimposed on Fig. A-la with
the y axis on the y axis and the x axis on the x axis, the curves would coincide.
Numerical calculations are facilitated by the tables of common logarithms, which
are taken to the base 10. Thus, for example, log,,100 = 2 since 100 = (10)?. In
practice the subscript 10 is rarely shown explicitly. For mathematical purposes it
is usually much more convenient to work with natural logarithms, which are taken
to base e. This is the mathematical constant defined byt
e= lim ( +2 = 2.41828
then
dx dx?
that is, all derivatives are equal to the original function. The function is written in
alternative forms as
y=e* or y= exp{x}
and the inverse logarithmic function is written ast
x = log, y or x=Iny
The general exponential function is written as
y = Ae™ or y= Aexp{cx}
which has the effect of stretching or contracting the typical exponential shape in
Fig. A-la vertically and horizontally.
If the inverse function exists, as it does when y = f(x) is monotonic (that is,
to each value of x there corresponds a unique value of y and vice versa), then
dx 1
dy dy/dx
Ify = e*, then dy/dx = e* = y, and so for the inverse function, x = In y, dx /dy
= 1/y. Thus we have the two standard forms:
Wiaes
yH=e ay meeahs
Be
e
y=Inx Deo
ei
d(logx) 1
ie ie log e (A-3)
Finally we may note a frequently used connection between logarithms and
elasticities. Ify = f(x) and a change Ax is imposed leading to a change Ay, then
BE ee ge
perl iy Ax yy
measures the proportionate change in y per unit proportionate change in
x. The
elasticity of y with respect to x is defined as the limiting value of this ratio
as
Ax — 0, that is,
oe usin ON eae
ven 5 dw dw/dx — %
= elasticity of y with respect to x
The second part of the identity shows that the same relation holds if logarithms
are taken to base 10, since there is a proportionate relationship between loga-
rithms to the two bases.
It follows from Eq. (A-4) that a functional form which implies a linear
relation between the logs of the variables is a constant elasticity function. For
instance,
y = Ax*
gives
log y = log A + a(log x)
so that a is the elasticity of y with respect to x. A simple way to fix the meaning of
an elasticity is that it measures the percentage change in y produced by a / percent
change in x.
The elasticity concept extends to functions of several variables. Thus
y = AxFz¥
is a constant elasticity function, where a, 8, and y are the partial elasticities with
respect to the arguments x, v, and z.
De Ng
Mae cick X,)
i=l
This sum is variously denoted by
(LX
i=n n n
Seat
i=1
Oeee ima ys orjust!
i=] 1
Dea dian
j= i=
x=
UX; t
n
It follows directly from the definition that
2 (Xin Xia Xp
i=]
so that the algebraic sum of deviations around an arithmetic mean is zero. The
sum of squared deviations from the arithmetic mean is
E(=
Y
i=]
¥(aes
-¥ axea yy i=]
I DEX DX SE ae
-Dx- (Ex)
or alternatively,
n 2
SoKay
= aeee
i=]1
= OGY) as
l
a)
Suppose a variable has two subscripts, say,
| Xi, X2 Xin,
2 X91, X225-++5 Man,
Class :
As an example, X might measure personal income and the sample data consist of
n, observations from social group 1, n, observations from social group 2, and so
forth. Total income in the sample is defined by
peak
» » xX, or, more simply, ae
Hes | Tey
ae 2, Xi
n
The mean income for the ith group, or class, is
nj
af ie j
x; n;
The sum of squared deviations about the overall mean is
=(%,-
my %) +0(¥-X)
Al: +20(%,-
; ¥)(¥-%)
The last term may be written
P
since the factor (X, — X) does not involve the j subscript and so may be moved in
front of the summation over /. But
524 ECONOMETRIC METHODS
for each 7, and so the whole term vanishes. The middle term may be written
since (X, — x )* is a constant for each element in the ith group, so that the sum
overj is simply n,;(X; — X). Thus
(x,-¥) =L(%,-
us
RY +Da(%-¥) Ey) I
EXE) = DE xp,
t=]
} For the statistical paragraphs in this appendix, two of the most lucid
texts at an introductory
level are P. J. Hoel, Introduction to: Mathematical Statistics, 4th
edition, Wiley, New York, 1971, and
L. D. Taylor, Probability and Mathematical Statistics, Harper and Row, New
York, 1974.
MATHEMATICAL AND STATISTICAL APPENDICES 525
or expected squared deviation about the mean. This is usually denoted by o”. Thus
2 : | 2
BCX)= Cas By
i=l
= BCX) Eee)
This result may also be obtained by squaring the expression in Eq. (A-7) and
applying the expectation operator to each term in turn. Thus
[f(x) ax =]
The mean and the variance are defined as before, but integrals now replace
summation signs.
Figure A-2
526 ECONOMETRIC METHODS
[{fC y) dx dy = ]
and
Given the joint density, a marginal density is obtained for each variable by
integrating over the range of the other variable. Thus
and
Tey)
(
EO)
{(x)
aes (A-8)
and similarly, a conditional pdf for X, given Y, is defined as
Ae)
(
Lip =
/0)
Two variables are said to be statistically independent, or independently distrib-
uted, if the marginal and conditional densities are the same. Thus the joint
density can be written as the product of the marginal densities
Be = E(X) = ffxfy),
dxdy= fxf(x) ax
oy = var(X) = f(x — w,)’f(x) ax
and similarly for the mean and the variance of Y. A new statistic for the bivariat
e
case is the covariance. It is defined as
f(x) = 1
ep|- 1
(x-0) 2 (A-10)
This defines a two-parameter family of distributions, the parameters being the
mean p and the variance o*. The bell-shaped curve reaches its maximum at x =
and is symmetrical about that point. A special member of the family is the
standard normal distribution, which has zero mean and unit variance. An area
under any specific normal distribution may be expressed as an equivalent area
under the standard distribution by defining
eee eg
ea
Clearly, E(z) = 0 and var(z) = 1, so that
l 2
f(z)z)= as eRe A-11
(A-11)
Then [fo dx = [°f(z) a
where z; = (x; — »)/o. The areas under Eq. (A-11) are tabulated in App. B-1.
Three very important results about the normal distribution are as follows.
+ See L. D. Taylor, Probability and Mathematical Statistics, Harper and Row, New York, 1974, pp.
154-160.
528 ECONOMETRIC METHODS
m <n, then
y ~ N(Dp, DID’)
2. Central limit theorem.} If (qv,= )15-a phere: a independent random
variables with means (j1,, “,... ) and variances (0/7, o7,... ), then
ys (ea ile)
lim
n— oo
Pr Ae
= 3
Ss
= a ee
exe" dz) eC Ae)
0;
i=]
Notice first of all that nothing is assumed about the specific forms of the
various pdf’s other than the existence of means and variances. The remark-
able result embodied in Eq. (A-12) is that the limiting or asymptotic distribu-
tion of the quantity U(x;— p;)/\Zo? iis the standard normal distribution. A
special case of the result may help to make its meaning clearer. Suppose the
means and variances are all identical. The statistic in Eq. (A-12) then reduces
to
n _ _*¥-p
Date np o/Vvn
1=1
f(x») =mae l
2770,0,¥ 1 — p°
1
||x — p,\?
|
/ pe) Ox
a r a | l e o e l a l a = ) | |
“2lir
where p = 0,,/0,0,. When the covariance d,, 18 Zero, the joint pdf simplifi
es
7S. S. Wilks, Mathematical Statistics, Wiley, New York, 1962, pp. 257-258.
MATHEMATICAL AND STATISTICAL APPENDICES 529
OTE eal
dz
=x 422-10=0
The third equation ensures that the constraint is satisfied. Eliminating A from the
first two gives 2x = z, which on substitution in the third gives x = 2(z = 4) and,
as before,
Pmin a Yimin = 20
530 ECONOMETRIC METHODS
ae (A-13)
vo
has Student’s ¢ distribution with v degrees of freedom. The ¢ distribution, like x’,
is a one-parameter family. It is symmetrical about zero and tends asymptotically
to the standard normal distribution. Its critical values are given in App. B-2.
The F distribution is defined in terms of two independent x? variables. Let u
and v be independently distributed x? variables with vy, and pv, degrees of
freedom, respectively. Then the statistic
u/V,
Fete
4s, 2
(A-14)
has the F distribution with (v,, v,) degrees of freedom. Critical values are given in
App. B-4. In using the table note carefully that v, refers to the degrees of freedom
attaching to the expression in the numerator and v, to the expression in the
denominator.
MATHEMATICAL AND STATISTICAL APPENDICES 531
t
Oe 27
v/v
where z? eine the square of a standard normal variable, has the x?(1) distribu-
tion. Tus t? = F(1,v), that is, the square of a ¢ variable with v degrees of
freedom i Syan F variable with (1, v) degrees of freedom.
The x* variable was formed from the sum of squares of a standard normal
variable. Suppose, however, that the z variables are still independent but distrib-
uted as
Z; ia N(u,;, 1)
The statistic z> + z3 + --- + z? now has the noncentral x? distribution with n
degrees of freedom. The previous distribution is sometimes referred to as the
central x* distribution. Corresponding to a noncentral x? distribution, there are
noncentral ¢ and F distributions, the former arising when the v variable in Eq.
(A-13) is noncentral and the latter when u in Eq. (A-14) is noncentral, but v is
central.+
Let X and Y be two variables with a bivariate pdf denoted by f(x, y). Let g(x, y)
be some function of the variables. The problem is to evaluate E{g(x, y)}. By
definition
+ For references to some tables for noncentral distributions see B. W. Lindgren, Statistical Theory,
2d edition, Macmillan, New York, 1968, p. 383.
+ We exclude functions g(-) which may have some values undefined such as x/0 or 0/0.
532 ECONOMETRIC METHODS
Let g(x, y) = x/). The straightforward application of Eq. (A-15) would then
give
Var Pp tas
where the u’s are well-behaved, the x’s are stochastic and distributed indepen-
dently of the u’s so that, in particular,
E(x
=,u,
E(x,)E(u)
,)=0 forallz
The OLS estimator of B is
hale pee
xs DE
MATHEMATICAL AND STATISTICAL APPENDICES 533
Sat] {Bue S|
xu Dx.
Be a Ee Ne)
so that 5 is still unbiased when x is stochastic, provided it is independent of wu.
The sampling avariance is given by
var(b) Sep
= E{(b — B)’} heeeies oH|aaa
sh
Now
ie yay : a os
Sepang? Dn,
Thus
]
var(b) = ott
t
which is the one-dimensional version of the general result given in Eq. (7-26).
Now consider the case where x is stochastic but no longer independent of u.
Suppose, for example, that x, u follow a bivariate normal distribution
f(x, u) =
-2»(22#\(4)+ (#)]
tomcat HET
with marginal densities
and
f(u).= ee
—exe|5)
f(ulx) me=
an eat eeae
‘ zs i= Ze ep shat 262(1 amexa pea
0”) | 0, ( )
(A-17)
Thus
E(xu|x) = xE(u|x) =x-
po,
- (otSic)
534 ECONOMETRIC METHODS
since Eq. (A-17) shows f(u|x) to be normal about mean po,(x — p,)/o,. It then
follows that
go ee ue al
exe Bele eee
— PO,
l ee —
= z.|6, oe Dexa K)|
oP On,, E De em eo
0. Ling |
On dividing the top and bottom by n, the term in brackets is approximately the
ratio of the sample variance of the x observations to their sum of squares. This
expectation does not vanish, and so b is a biased estimator, both for finite samples
and also asymptotically. This example is a legitimate application of Eq. (A-16)
since f(x, uw) is a well-defined bivariate distribution.
Now consider the model
Be = EBay | = 0
ye fal e ye
However, the OLS 5is well known to be biased in finite samples.} The source of
the error is that (y,_,,u,) does not have a well-defined bivariate pdf, which
renders the application of Eq. (A-16) invalid. Given some starting value Yo, Once a
u vector is drawn, the y vector is exactly determined by Eq. (A-18). The stochastic
behavior of y is completely determined by u, so there is, in effect, only one
stochastic variable. Thus E(y,_ ,u,/Xy/_,} has to be evaluated solely over the u
distribution, and the two-step procedure of Eq. (A-16) does not apply.
A more complicated version of the same error can arise in IV estimation.
Consider
+ J. S. White, “Asymptotic Expansions for the Mean and Variance of the Serial Correlation
Coefficient,” Biometrika, vol. 48, 1961, pp. 85-94.
MATHEMATICAL AND STATISTICAL APPENDICES 535
¥,—;. The W’Z matrix of Sec. 9-2 will then include terms in Lx, y,.,/and
2x,-1);—-1. These are stochastic simply because wu is stochastic and, as in the
simple example given above, a two-stage evaluation via E.,,, followed by E,, is
invalid.
The basic idea may be simply illustrated for the univariate case. Suppose u is a
random variable with density function p(w), and suppose that a new variable y is
defined by the relation y = f(u). The y variable must then also have a density
function for y in terms of the density function for u and the relation y =f(u).
Suppose the relation between y and u is monotonically increasing, as shown in
Fig. A-3. Whenever u lies in the interval Au, y will be in the corresponding
interval A y. Thus
Pr{y lies in Ay} = Pr{u lies in Au}
or p(y’) Ay
= p(w’) Au
where u’ and y’ denote appropriate values of u and y in the intervals Au and Ay,
and p(y) indicates the postulated density function for y. Taking limits as Au goes
to zero gives
7 o Figure A-3
read
du
P(y) =p(u)- dy (A-20)
If y = f(u) were not a monotonic function, Eq. (A-20) would require amendment,
but we are only concerned here with monotonic functions.
In the multivariate case u and y now indicate vectors of, say, n variables each.
Under suitable conditions a result similar to Eq. (A-20) still holds, namely,
ou
p(y) = p(u) dy (A-21)
where |du/dy| indicates the absolute value of the determinant formed from the
matrix of partial derivatives
du, ee
| du, eedu,
dtr, tk ws du,
ORO
i ae
eeieon
er ae
where the observations have been expressed as deviations from the sample means,
for we are concerned with studying the variation in the data.
The nature of principal components may be approached in a number of ways.
One is to ask how many dimensions there are or how much independence there
really is in the set of k variables. More explicitly we consider the transformation
of the X’s to a new set of variables which will be pairwise uncorrelated and of
which the first will have the maximum possible variance, the second the maximum
possible variance among those uncorrelated with the first, and so forth. Let
21p— yy, ¥ OaiXae Mee Gye, belt
denote the first new variable. In matrix form
z, = Xa, (A-22)
where z, is an n-element vector and a, a k-element vector. The sum of squares of
z, 1S
Z\Z, = a',X'Xa, (A-23)
MATHEMATICAL AND STATISTICAL APPENDICES 537
The problem now is to maximize Eq. (A-23) subject Eq. (A-24). Define
o = a X’'Xa, — A,(aia, — 1)
where A, is a Lagrange multiplier. Thus
a = 2X’'Xa, — 2,2,
Setting
Op
dan
gives
(X’X)a, =A,a, (A-25)
Thus a, is an eigenvector of X’X corresponding to the root A,. From Eqs. (A-23)
and (A-25) we see that
Z\z, = A,a\a, =A,
and so we must choose A, as the largest eigenvalue of X’X. The X’X matrix, in the
absence of perfect collinearity, will be positive definite and thus have positive
eigenvalues. The first principal component of X is then z,.
Now define z, = Xa,. We wish to choose a, to maximize a’, X’Xa, subject to
a’,a, = | and aja, = 0. The reason for the second condition is that z, is to be
uncorrelated with z,. The covariation between them is given by
a) X’Xa, = A, aa,
=0 if and only if aja, = 0
Define
@ = a,X’Ka, — A,(aa, — 1) — w(a\az)
where A, and p are Lagrange multipliers.
SOON
da,
a Eh as = ia 0
Premultiply by a’
2a) X’Xa, — p = 0
But from
(X’X)a, = A,a,
a’,(X’X)a, = A,a5a, = 0
Thus
w=0
538 ECONOMETRIC METHODS
and we have
and A, should obviously be chosen as the second largest latent root of X’X.
We can proceed in this way for each of the k roots of X’K and assemble the
resultant vectors in the orthogonal matrix
Na) 0
ZZ=AXXA=A=]0 A, <*> 0 (A-29)
geet .
showing that the principal components are indeed pairwise uncorrelated and that
their variances are given by
z.z,=A, ies lee (A-30)
If the rank of X were r < k, k — r eigenvalues would be zero and the variation in
the X’s could be completely expressed in terms of r independent variables. Even
if X has full column rank, some of the A’s may be fairly close to zero so that a
small number of principal components account for a substantial proportion of the
variance of the X’s. The total variation in the X’s is given by
tr(A’X’KA) = tr(X’XAA’)
tr(X’X)
since AA’ = I, and so from Eq. (A-29)
k n k
DY Lo xf = (XX)= VA,
=212, +--+ + 2,2,
i=1 t=1 i=]
Thus
Ale
AN ean
represent the proportionate contributions of each principal component to the
total variation of the X’s, and since the components are orthogonal, these
contributions sum to unity.
It is sometimes difficult to attach a concrete meaning to specific principal
components. Occasionally a suggestion may be found in the correlations of a
component with various X’s. To find the correlation between, say, the first
MATHEMATICAL AND STATISTICAL APPENDICES 539
principal component and the X variables, we proceed as follows. The vector X’Z,
gives the cross products between z, and each X variable. But
X’z, = X’Xa, = X,a,
Thus the correlation between X, and z, is
an Aiaiy
hms a
2
Ay Di
t=1
aA,
eea ee Le re ke (A-31)
2
ys Xir
Gal
where a;, is the ith element in the vector a,. In general, the correlation between_X,
and z;, is
a;, d,
rj =———= — ii, j=l,...,k (A-32)
| LuXit; t
These correlation coefficients may also be used to show how the variations in
each X variable may be decomposed into the contribution due to each compo-
nent. From
Z=XA
we have
L — XG
and X’ = AZ’
since A is orthogonal.
So X’X = AZ’ZA’
= AAA’
from Eq. (A-29), and so
n k
xe Dag A Lele i (A-33)
t=1 j=l
Dividing both sides of Eq. (A-33) by X,x7, gives
ie aid . 4i2rdr ee did 4 (A-34)
Lixe DXi DXi
where the terms on the right-hand side are the squares of the correlation
coefficients defined in Eq. (A-32). Thus the proportions of the variation in X;
associated with the various principal components are given by
22 aed, 2
Vito Vid.+ ++ Tik
540 ECONOMETRIC METHODS
and since the components are uncorrelated, these proportions sum to unity, as is
shown by Eq. (A-34).
A note of warning should be inserted here. The development so far has
proceeded on the implicit assumption that the X variables are all measured in the
same units. If not, it is difficult to attach a meaning to concepts such as the total
variation of the X’s and the partitioning of that total variation into the contribu-
tion due to each component. It is still, of course, possible to compute the
eigenvalues and eigenvectors of X’X even if the dimensions of the variables are
not all the same and the correlations in Eq. (A-32) and the partitioning in Eq.
(A-34) would still be meaningful even though the partitioning of the total
variation in the X’s would not. As an alternative, analyses are sometimes carried
out after all the X variables have been standardized, that is, each deviation from
the sample mean is divided by Vn times the sample standard deviation of that
variable. X’X is now the matrix of zero-order correlation coefficients of the X
variables. The analysis can proceed from X’X as before. Now tr(X’X) = k, and
from the development following Eq. (A-29),
ED Neenem,
The eigenvalues and eigenvectors will in general be different from those yielded
by unstandardized variables. We leave it as an exercise for the reader to establish
whether the correlation coefficients in Eq. (A-32) are affected by the standardi-
zation of the X variables.
Empirically, then, one may compute the principal components for a given X
matrix and see how much of the variation of the X’s is accounted for by various
components. Frequently the intercorrelation of economic and social data means
that a small number of components will account for a large proportion of the
total variation, and it is desirable to have a test for Judging the number of
components to retain for further analysis. Suppose that we have computed the:
roots A,,A,,..., A, and that the first r roots Aj, Ag,---,A, (7 < k) seem
both
sufficiently large and sufficiently different to be retained. The question then
is
whether the remaining k — r roots and their associated vectors and the compo-
nents are sufficiently alike for one to conclude that the true values are equal. A
very approximate test is based on
Z = XA
and hence
xX = ZA’ (A-36)
Equations (A-36) express the X’s as exact linear combinations of the components
with coefficients given by the elements of A. If, however, we retain less than k
principal components, Eqs. (A-36) would have to be replaced by
X = Z*A*’ + U (A-37)
where Z* and A* denote the submatrices of Z and A giving the retained
components and the corresponding eigenvectors, and U is a matrix of errors.
Principal components is obviously a possible estimation method in factor analy-
sis, but slight modifications are required to the A* coefficients to conform to the
imposed assumption that the factors should have unit variance. Without addi-
tional restrictions z;z; = A,, as we have seen in Eq. (A-29). When the A*
coefficients have been adjusted, they are referred to as factor loadings. However,
several other estimation methods are used in factor analysis, and we do not
propose to discuss them here. An interesting application of factor analysis is
given by Adelman and Morris, who find that 66 percent of the variance of the
GNP per capita in 74 underdeveloped countries associated with just four factors,
which have in turn been based on a complex of more than 20 social and political
variables.t
Table A-1 shows another example in which a small number of components
effectively account for the variation in a set of data. The basic data are 11 series
of average quarterly interest rates in the United Kingdom from the first quarter of
1963 to the first quarter of 1969. They include various national and local
government rates as well as commercial rates, such as those on Building Society
deposits. The series were standardized and the second row of the table gives the
values of A,/LA for the first four principal components. The first principal
component, which turned out to be effectively a simple arithmetic average of the
standardized series, accounts for over 83 percent of the total variance and the first
three components account for almost 97 percent. The last seven components
account for less than 2 percent of the total variation.
+ See J. T. Scott, Jr., “Factor Analysis and Regression,” Econometrica, vol. 34, 1966, pp. 552-562;
M. G. Kendall and A. Stuart, op. cit., pp. 306-311; and H. H. Hyman, Modern Factor Analysis,
University of Chicago Press, Chicago, 1960.
+ Adelman and C. T. Morris, “Factor Analysis of the Interrelationship between Social and
Political Variables and Per Capita Gross National Product,” Quarterly Journal of Economics, vol. 79,
1965, pp. 555-578.
542 ECONOMETRIC METHODS
Component l 2 3 4
7a) | ee ee
ae |e ans
and for A, = 0,
ey Ree
PEE AGE aS
The second principal component does not exist, for
Z,= poe x, =0
eee
since x, = 2x,. However, the first component does exist, for
l 2
a Rear + —x, = 75x,
v5 v5
and so the coefficient of z, in the regression with Y as the dependent variable can
be computed as
= — = =
oy
wat A, (oma
Y=5,z, +e
544 ECONOMETRIC METHODS
APPENDIX
—_—_--
Oreeooooo
— o—!_—O
STATISTICAL TABLES
545
ba ieee
9 we tiie so
ihtapdl , re ee
sf"
ee } ray wr 7
7 , . oy fae ie tr
is
; 1s iy
CA
A i
a+ j
a) es
i
/!
STATISTICAL TABLES 547
-05
ooo°z<e
L1Z°97
8LS°0¢
607°¢¢
So08°7e
086°77
682°00
7£6°8E
999°LE
Z209°SY
£96°97
picyy
Z268°0S
T61°9¢
980°SI
LLO*<ET
Z18°9T
SLv°8I
060°02
ZL0°SZ 889°LZ
saaibap
£L8°97 TVT°6Z
889°67
8L2°80
10°0 so9°9
O12°6
Jo
BE9
Ive°tl 99917
602°€Z
It
7S0°7Z
L89°¢Ce
6S2°8Z
$66°0¢
££9°6Z
Oz20°S¢
IETS
778°L
899°TT
8Be°el 89T°8T
6L9°61 £7O°9S
896°8¢
6S9°LE 99S9S8°77
OvI°7y
£6990
Jaquunu
SL9°61 819°2Z
yo
Z17°S L¢8°6 17
ZZ9°91 T9T°IZ
£¢0°ST OLZ°0” 617°St
796°LY
962°9Z
1Z 970°
ie OW
166°S OLO°TTL90°7T
L0S°SI CS9°LE Leeiy
LS9°20
ayj
C9IE*7Z
S$89°¢Z
966°7Z
L8S°LZ
pyt°Os
698°8Z
“O-D ‘OUy
178°¢ ST8°L
8817°6 69°71 616°91
LO<"81
U st
1L9°Z¢
CLI“SE
URTV] SUTYsTqng
NZ6°CE
SsT7°9¢
see"se
e107 ellen
aJaym
S79°O1
LIO°ZI
L86°ST
789°
162°9
‘aouelseA
SLO°LIZ18°61
697S°8T
790°1Z 69L°0C S19°6ZLO0°Z¢78S" L80°6¢
90L°Z
9£2°6
s09°)
686°S 6LL°L
O<O" 79S"
Ee
Z02°71 1
LO¢*2Z
CUS°ST686°SZ
702°LZ
Z17°8Z <18°0¢961°SE£ISSE
TvL°g9e
9I6°LE9S2°07
T
c7°¢
c79°7
1 279°
IL1°9Z sL9°0¢Z16°ZE
Lz0°v<
UQIM y1UN
S86°91
Z18°ST
$917°0Z
006°¢Z
8<0°SZ
IsT’st
612°¢
682°L
£08°6
8SS°8
6217°82
TO0<°Lz
TT
£SS°62S6L°Ie 6£1°SE
osz°9¢
a3eIAap
T¢e9°t
TI<*61
S19
IZ
O9L°72
“YIOX
YLO°I 790°9 £82°8 9S59°01 668°Z1
Tee°k 61ST cZe°LI
c22°91
T10°vT 8I7°st 689°1Z
ogs"e¢
JeuJOU
s99°¢
807°2 8L8°7 02S°618Z°IT T1S°6I
109°0ZSLL°7Z
MON
8S8°¢Z
810°9Z
6£6°7Z
ZLI°8Z
612°O¢
960°LZ
907°6Z
19%°Z<
l6c"l¢
6£2°CT
6£o°VT
B<e"8T
Le¢"6l
Ove*Z
Beers
Yip] “po
pasnse
8ee°9l
8eerLT
17°01
ssv°0992°C
Looe
Icey
B1E"S
900°9
perl 70S°6
Ove
YoADasay ‘siay4Os4
L¢ee°0z
LESES
7£0°6 TT
LEerlz
Leese9EE°ST 9EE°6Z
99Z°91
ZS°ST
1Z8°OT
Aewaq
926°6
7T Ory"
nen"! o00°¢
£1ZL°0 £98°0Z Ly9°<Z
£6°61 LLS°0Z
nc9°Z1
Tesrel
IZL°11
n9E"¢T
Z00°Z1
L0<°01
est°ll
076°81
LS8°Z1
0Z8°6I
SL7°2Z
s77°st
PIS°9l
85°71
290°81
FSET
9£9°8
7XZ/\
uolssaidxe
$00"|
679°1
OLore
O8<°S
6L1°9
L907°6
¢<0L°02
L8°LI
686°9
2£08°L
[0I11SNDISspoyja=
4of
£9o°7
cZ8°¢
21790°0
9007°0
"69°07
$80°01
s98°01
e772
819°S
Z70°L
8s10°0
O19"
790°1 89I°y OVZ°<T878°7T
170°71 £Ly°9T
6S9°ST VIT°SI
Z6Z°L1 89L°6I
6£6°81
L0S°8
O6L°L
70s°9
ZIL"6
1s9°TT
06°0 78S°0
1120 £oB°e
702°Z 060"¢ S98"
66S°0Z
ueY] ‘O¢ ay3
8ee°Z1
826°91
67°81
—
£6¢00°0
tI 119°
6LE°ST
QLS°ST 802°LT
“IaYSi
ISst°9T
16S°TT
160°<T
Wopaady Jaqea1H
S6°0 Sv"
<01°O 1120 eel°e Ov6"¢ SLO? Z68°S cL9°8
1L49°9 C96°L 06£°6 1S8°OT
009°01
90¢°9T
266°1T
ScI°vl
69°21
£62°1T
600°<
S16°6
cSL°0 799°
MET"T BLT" 892°S SSo-L
906°L
Wor “WY
98°
86°0 829000°0
¢8t°0
7070°0620°0 CES°Z
c£0°7 6S0°¢ 609°¢ S9L°Y S86°S
v19°9 L9S°8
LE2°6
1
poyuudey
£s1000°0
saaibap
092°8
Z18°S
622°S
6£2°1
yo
SII°O 979°T
880°2 961°OI
20S°6 029° 6L8°Z1
9S8°01 952°
S9S°¢T
807°9
CBOE:
<s0°¢
STOL
LOI”
099°
IZs°¢
wopaady
TT
7SS°0
10Z0°0462°0 2L8°0 8S9°2 L68°8 1
861°ZI £S6°71
saaibaq
4 JO
ANN
TN OM ONO
So 92
-X
uoHNgLysip
qe
¢-g jo Iz cz SZ 92 Le 82 62 O¢
a gd
148,p-¢7 uONNgLysIp
550
¢ yUI0I0d URWIOY)
(adA} pue
| yusoIadore) (ad) syutod
Joy ay} UoRNGMIST
JO 4 p
saaibaq
jo
wopaaiy
JO} saaibag
jo wopaaiy
10} s0ye1awnu
(17)
JoJeulwousp ae
Fe ae +
(am) I é € 0 S 9 L 8 6 ol Il | PA val 91 0z 02 o¢ Ov os SL oor 002 00s oo
I 191 002 912 S2Z Oe nEZ Lez 6272 19Z CZ £92 902 SZ 90Z 802 6nZ= OSZ 1SZ 752 £SZ £62 SZ
ZS0P 666% E€0PS SZ9S 0SZ 967
POLS 6585 Bz6S T86S zZz09 9509 72809 90T9 ZbI9 6919 80€9 >PEZI BSz9 9879 ZOE9 EZE9 HEED ZSE9 I9E9 99€9
z 15°81 OO'6I YI°6I SZ°6I OF6I <°6I 95°61 LEI BEI 6S°6I Ov6l I7'6I ZH'6Il £761 7°61 sv6l 9n6I Ly6l LH’6I BEI
6P°86 TO°66 LT°66 SZ°66 6H°6I 6Y°EI OS°6I OS*61
OF66 CECE PEEE 9E°66 BE°66 OF°66 Tb*66 7Zh°66 Eb°66 bb°66 SV°66 96°66 LP*6E6 8F°66 8h°66 6F°66 6h°66 6°66 OS°66 OS*66
€ <I°OI SS 826 ZI°6 [0° 76°8 88°8 48° I8°8 8L°8 9/L°8 %2°8 IZL°8 698 99°8 99°8 z9°g 09°g 86°8 1S6°8
ZI*PE TB*OE 9F°6Z TL°8Z 95°8 S°8 75°8 £S°8
Prez TELE LO*LZ 6P°LE HELE ES*LZ ET°LE SOLE 26°9Z EB°9% 69°9c O9°9c 05°92 TH*9t OE9Z Lz*9% EL°9Z BI-9Z PI-9Z cT°9f
Y TY SE) SRE) (SEE) CYA SII) YSU VED aR) SYS SARIS YSIS TEKS) KS (UES (HESS isc Tig.
Oz*Iz OO°8T SOLS) 99° 599°C S975 THIS, 9S
69°9T 86°ST ZS°ST T2°ST 86°PT O8*PT 99°PT SPT SH°PT LET H2*PT ST°PT ZOOL E6°ET E€8°ET PLTET 69°ET I9°ET LS*ET ZS*ET BP°ET OP°ET
S 19°9 6/6 Ih¢ 6IS GOIS G6h 88°) Zev 8Li7 Wis OLth 89°h 49° 09°F 95°7 £67 OS°7 97°%
9z°9T LZ*ET 90°ZT GE*IT thr Zhh On BSD LED 92°77
LOOT L9°OT SP°OT Lz*OT SOTOT:«<ST*
= §=—96*E
OT «6686 §=LL°E 6896 =6SS°6 47°6 «=CLT°G6
«CETTE)—Z*EC*E
=6—LO"E sEG
«=bOE cO°6
9 §=66°S DI°S OLY Sth 6ER 82h 127 Sih Ol” 90°7 <0°7 HOt
| 968G) Z6rGe BIG. VOI LIEGE
«LET ZE*OT
= =—BL*6 KG PAE SU RS GES MSS
ST*E) =6SL°B Lee 92°38 «COTS «6B86L LBL 6L°L ZL*L O9L CSL 6&L TEL SPT*LSCE*L
GOL ZO*L 669. 6°9 06°9 88°9
L §65°S Dl’? Geer HnZL oC CumO°mnoC OGL
Sie Com
9: ESOS O91 EL Re GIS Tae 1RS MUEIS WSIS
Sz7T $5°6 SPB TOG BGT BCs Gren CAS CISION
SB*L 9F°L 61°L O00°L B8°9 «z9r9ss—TZ"
«PS*9
9 LBRO SE°9: LZ*9 ST*9 L0°9 86°S OBS SEs 82:5) SLES O£5 £95 SHS
8 ZE°S
= «98° LO°7? HBS 69°C BSS OSE HE BES MEE ISS B8ce repeats
| (aPAYe AGI ASS RS US
92°IT $9°8 65°L TOL <Ore OOS 862 96°% 62 £6°C
£999 LE*9 619 €0°9 T6e°S zB =—L9°S)—BL*
«69G*S
S «BRS OES =82°S Of°S IT'S 90°S 00°S 96% T6% 88> 98>
6 =ZI°S 92°7 9RtG -<o"G ahie. EG cccs
6 ECC less
G See Shi OLS LOS ZO°S 862 S6°% O62 «(98°
95°0T 20°8 CBZ Q*Z,
en 127 maf OSC CAROL. CLe MOS
66°99 Zr9 90°9 08S Z9°S LS SES 92°5 Bs ITS O00°S ZE*P O08 ELD vS> 96°p Ise Shr Tre OF €€b T&>
or 96°47 Oh =f Bye 6S 2226) ISG. LOLS COLwe G:CNL WG:C CeNG DErCM ZOtz. LEZ) aUlsc. NOLS
pOrOT 95°2 LDC, WIC 1927 EGG Ce MIGCemGGeC Ro
S5°9 66°S P9°S 6ES TZS 90°S S6*b S8b BLE TLR OOF ZS Tre E€&b Sob 4Tb ZIP Sor TOP 96°F EGE T6°E
09° os*Z 12°2 L8°7 26°LS°Z 88°T 78°1 18°TIES Toye: IL IZT
0n°Z 9E°E oie £12oo°e L0°Z 10°ZSLE 96°1S97 67°e cre 8L°I Clare StayTez £re
ERE
BEE
c2°Z
COPE
1¢°Z
80°Z
LLS
17°Z
ere
912
68°72
Z0°7
7L°7 92°7 I~ oz 66°1 Ly°T 78°T 6L°TcEZ 9L"T oL*T
18°T
c7°7
E99° THe T&E€ 90°E 26°7 70°Z0a°z Olt S6°1coz 16°oS°7 L8°1 cre LEZ Lez EEE
LEZ
08*T
EES
LOE EERE, LO°Z 06°1
Ly°z
98°T
z8°T
cre
S77OLE So°z90°F 90°72 612 VeLE°7 9L°7 86°1897 76°09°z
98°7 ZO°Z ESS 23° |(YES62°Z
9¢°7 82°2Of°E SsV°Z 68°7 TL°Z OSC TS*Z THz o8*T
LZOLE 6r°E 122bre oo°e 60°Z 90°Z6L°2 00°Z 96x£9°~ z6"1 68°1 "19n°C 78°1 z8*19E°% CES
0n7°Z 22°Z 92°Z 81°2 <1°Z 80°Z 70°Z8L°7 o0°zOL°e 9671 £6°18S°~ 16°I€S°e bre
os°2ose aoe. Lone T2°€ LO°E 96°7 98°7 ESS 88°TBre 98°1 78°1one
n2°7CHE LOS 12°Zere 91°72
TO°E 11°Z76°7 LO0°Z Z0°Z 96°1 161€S°Z 68°1
£S°798°E 7°72T9°E 92°E E€8°Z 9L°7 667169°C €9°% £6°18S°e 6r°z L8°1Sz
OL°E 1<°Z sv% 12T6°Z o8°~ 70°Z L9H 76°1
L9°2v6°E 97°7 8<°ZTS°E bee S220z°E 02°2ore oore L0°2 LL°Z 00°ZcL°7 86° 9671c9°z 8S°e z6"1SZ
cO°o
BL°E
62°7
92°7
80°€
0o°€
26°S
19°%
S0°Z
80°Z
coe
SZ
os°2
6S°E
Z7°7
6E°S
8T°E
SoZ
ERE
98°C
08°z
99°72
612
£0°Z
8671
00°2
OL°e
96°1
SLE
IZ
Bbr°e
77°Z
LOE
6£°2Z
ere
19°2
SO°E
66°C
S12
£12
Ora
125%
90°2
Ta°z
Te*p
OL°Z
09°Z
BL°E
cOrE
£o°7
LEE
62°7
6re
S27
8I°Z
b6°~
68°7
60°2
SB°~
HL°Z OL°E LETS SOC 0z°2 ZO°E L6°7 11°Z
62°D 99°ZSO°D S9°ZSBE 87°72 £7°Z9S°E SHE £o°7SEE 62°ZLEE 9C°%6T°E ETE LO°E 81°Z VIZ £V°Z£6°% 68°7
6L°Z 09°2 oa°e 87°Z 7E°7 OE cre 02°72 IVS
Opp 69°7
booT 96°F £S°Z LYE cv°zSS°E 82°ZSh°E LEE Lez 82°7EZ°E S22LTE £2°7 LOPE 8I°ZEO°E 66°S
ELSE
cao
s7°Z
60°E
O&E
SO°E
2°72
02°2
IEC
172
9b°>D
Z8°Z
cO°n
ZL°Z
£9°7
99°7
98°E
cS°e
7¢°7
1o°Z
oo°E
82°7
THE
ore
LES
92°7
GEE
T<°z
Ble
ore
Of>
orp
96°7
9L°%
09°2
69°E
S92
67°7
S77
TS°€
8¢°~
92°E
SEZ
ones
os°z
92°C
72°CETE
oS°d
D6°E
L9°%
oa°e
6S°E
17°Z
ERE
LEE
TEE
ofS
Toe
82°Z
S8°Z
S62
CES
O2z"¢ 69°D 06°2
96°C 18°ZDED 99S66°E 99°T
bere 29°Z06°E 09°Z
I<90°S c0"¢98°D 9S°b S8°Zbro LECSZ°D VL°Z40°? WEOr’e 89°Cb0°D 98°E
9°¢ NERS 8I°e 8S°D Os*> 78°C Z8°Z 9L°7
LES Tres 0z°s Ile£0°S 90°¢68°R To"eLLP 96°7L9°b £6°S 06°2 2°72ern LED Tee 08°Z
90° 8L°Zcoe 8T°o
6S°¢f2°9 67°¢S6°S bL°S 62°¢crs 62°S O2*¢ 9I°<60°S c8°p To"<cL°p
Ie nee9S°S 92° 8r°s cI¢TO°S ors6°? LO*¢L8°R sO°¢ <0°¢9L°9 66°289°F
Of*Z 08*< Ci9, To°9 Sas 8L°S CLS
86°¢ €6°9
88°< 0L°9 yL'¢qTs°9 89°¢9E°9 £9°€EZ59) 69°¢ SS°E cs"¢E6°S 60°¢ Lv¢ ne cre99-9 onsoS BereLS°S:
08°17 €E°6 40°6 98°32 0S" €s°8 ov? 17782°38 88°Z 92°17
$9°6 sly L9°7 09°) 89°8 67°17 sty 8T°8 sonors conzo°s os"bErL 82°07
8<°7 L£2°2
ca°L 92°07
él
II
<1 aI SI 91 LI 6l 02
81
IZ ce £2 ne ‘Go
551
148.p-@ (panunuoD)
LT
2a)w
¢ JUDdI0d UeWIOY)
(adAj pue
| yusdIadoye!) (adAy syutod
Joy 94} UORNQUIsIp
JO 4
saaibaq
jo
wopaad}
JO} saasb6aq
yo Wopaady
10} s0je1ausnNU
(Ia)
Jojeulwousp
(22) I Z £ ” S 9 i ie 6 I él val gT Oz 92 0c On Os SL O01 002
8
00s oo
Or
2@ mmecosy eco CueL
GR aCe 6G 2e CML
CSi LZOZ
0 <ZZs20 BI5Zs Iscn eG sce OGulCSGu
mm SOs]limeO6ul
en COnt aun
ulOL
RGO;CMOl Nanos
cL Olsen CPs
ELIZ «ESS OH PTR BE =6S°E ChE LTE 60°F COE 96°C «98+ «772% 997% 8S*S. OST THT. GEST. Boe,
c£°7
Soe, GIS, SIC) oie,
62°E
LE Cat,
CGOG SSC LSC NEC DSC Ose ES cdOG Sher Cama ipa kd TU Tesi) GERI ah TIRKIC Cyl Mie WEN GERI
B9°Z 6RS O9°F TI*R 6L°E 9S°E 6E°E PTE 90°F 86°F £6°C EGC PL*Z E997 GG8Ca LACE BESTE EELS SCHC Caceat DEC CTsCeMN OTRO
O¢°292°E
- 8c OGSY: ComSS SOcCun LC IGS NUNC FELCH 72°E 612 Gime vAGe chips
= aia SYSit Rit [GP TGR GYR) Gif iil GEA IE SEA
8S9*Z SPSS LS*P LO°P) 9ZIE) VESTER DEE ITE EO°9E S6S 06° BT Lr O9% EE HHS SEZ OFS fe BT
62°72
ETS 60°F 90°C
ECE
6z OT" OEE CEBTZ «(OLS 69S £72 Sot Cee Ce
lex
8 CeWN O1eZ <O%z. 0072) svGel [Link] SOrI OSs at
/ eh Eg alee eeena eae
09°L 25°S PS*R POR ELE OS*E EEE BOE S—C«OO"EE-s« ZO"ESC LL°ELB*E
89° LS: «6FT THT TET LEZ 61% SIZ, ONC. IO:Ce- ECS
82°7OE
o< TAY) RASS (ASA CRG SESS LAI UESCE crc cc,we Chad MOO DOC G6| -<6al) 6Bcl) Vealh
ee nGlall Ola Coen Gone eee ale Cee
9672= (6ESS TS¢h. CORE OL{Em LELEN ORES 90°E 86°F OB PE? PLT 99° SEZ LHT BET OC HES OTF ETS LOS EOS TOS
LCseLEE:
ce ST OSS «(COGS CLO IG (ORS CSC lec econame OleC Oram COpCm
ee EGal NSAI SiR rah CNR TH (ERE TER UERIE LER TUN
OS*Z PESGS OPP ZELEN (IFES CHER, SCIEN TOE FEZ 98°F 08° OL zZerZ Ist Ze PET Sez OFZ ZT BOS ZO"? 86°T 96°T
LALAere
nE Iv BE BB S97 G6H% 88 Os Lee Clete ZeeOO: OOsCGOrCe
Gu Gaal Gslen st 0Gst leery
Salma Oem Sela GSalime ES
PPTL 6F°S CH*H EGE TOE BEE To*E LENT 6B CBS QL 99°C BS% LHe BEC
£2°7
OF ez STZ BO°X OT 86°T POT T6T
9¢ Tce 9ZIE 9852 $952 BN SESS 82o ec cme :Cum90;C
oO OUilame
aeOl 6a Gel Oct G/alan Clale 69u SIs CIN OS Sale Seale
GEL StS BER 68°F BSE SEE 8T°E ENT OBE BLT «CLT CHS OSS ER SET
80°e 12°72
9% LIZ 2I% POTS O07 ET O6T 48°T
8¢ Oley GCse SBC CIC CONC EGS DCC HI'Z 60°% <O°% 20% 961 C61 S81 GBI CYA (VR Cie WS SURE VASA SRLS
SEZ FS PER 98°F; PSE CEE STE
(SEN
652 CBS SEZ 69°C) Ce NGS ESSER
bO°E 61%
OF zET eee PIe BOS 00% LET O6T 98T Pe°T
core
Ov BD CEZ"E HBTZ
«OO OI9Z «SHS SS S22 ZI*z LOZ ¥0°2Z 00:2: S61 0671) VOT GET “HLL. 69°T (99 19%L) 65:1. SS eGo
“TESL STS) TESDY PESSE TSIE. 6CLES CIFES
Rast
(8822 OBS ELST) 995ZE 9S:Z) 6RiC) LEZ bez ore IT? Sot LOT PET B88T PST T8°T
81°66°C
cy CLO OCZZE BZ 6SZ (MS 28S "oo II°Z 90% ZO°Z 66° él 68°l ZBI BLT Sil 89° 7921) O9Te LST Poel
ZEAL STSS] 6CSP OBER (6526= SC°ES ODSES* TER Cita
Wiha MACE EE ATE ME SEZ 927 LIZ Bor Zor POT Té6T SBT O8T SLT
ny «(90H «CIZE «7B Bc°Z EHS 4I1€% <o2 ~=—CoT°*z «<0 «1O"%. 86" Z6zl 885i 18 ILA /et
zc 9977 S977 OG 9Sct leecoo beerOSt
PEL CET*S OZ*H
= «BLE =—ORE =—PE°E LOE Syn
«GL*ZOvB*T
BOT EIT CSS bee
LUZ96°C Dee
CES peer SI*z 90% 00% 726I 88T Z8°T 8£°T SLT
b6°S
90 SsOr OE 18% LOZ cyt Of coe OZ «=00TZ_—SsswOTZ_~=s
«=L6I:«G «6=T6"I =L8°T (O8l SZ°T) TILT Corl COTE Ge Ssheee
TzZ OTS PER OLE PRE 22°E SO°E a Siete
Pee
COS
| VELS. IPT ORC OSC PZ
912
OFZ eer ETZ PO'Z B86°T O6T 98T O8°T 9LT CLT
f6°7
87 “HORY, 6IeS. (ORFs ‘9SiZe ined)a OSC Se VIC 8024 SOi%s 665) PSIG GRO6 FBT GL VEST OLS VOSS SEIT SEISI SSN OSes WEVA GOA
ETL 8802S ©2S°R FLOE) ERE 0CSE. PO°E ((06°S; OBS, TLE «PSE A8S*S BRE OWE BE OC Tics. “COC 96T A8EE PET S8L5T* TELE OLT-
Os Ony, Olson eG/aCer OSC© CamOlcG mcr OCC. elie LO eCOrce —“86ply meSGrll meOGnl| SOrl aie
8 Lala
7 Osan Ose OOetm MSGR SEN MOC
Oy OVA im
ZTSZ -90%S OSB. UEZLSE. TEE’ BITE COPE BEE)
6 «BLE «COL «68COE SE OG OPS GES OES 8T5o) (OTSe -OO2S PET 9ST (CET OLE TET OST.
09 ‘SUGEOOat,
“Isc VESrCw GOCE, LC Oi VOR 66ylee S6ule ec6rl)”
= 9Beee Bal Lelie:mS OLae GI GSolen OGateee OSme Oi7 alieny yale 6Fnv
BOLLar (86:7en CISD =SOtE PEE SacIne SCS COC CLS ESSN (9GTE GOG:S ONC CEC. [Link] CILES EO;CN 86.0) OE LE CLT TOLER OIA NESE OFT,
OL OG.S Clee aULic OSG. MESCee ConG Com iy LOccemme (Orc <L6ulen -SOnle mvanlarGGnl.
laGL Ziel Sule Cosme OG eGo meU eeYl eG maSoa
LO-La o6<ha 80th ~O9Sfs WGCee, LOEa 160 LLG LOC 65:0 TS: aese.e. Secwe eee SEC LOC lanS86 SE
| =eCOue Lal mee TODD OSESESSE
08 9656 Uileeae CEC Bye SSe crc Ce cI SOle) 66ull S6ule N6aly eSSel COnlae VL Oa Soot JOSTes 47Set Gal) Svolaeem Sole Sa CONS
—~96°9) BaF - Oth 95°F §SELE (POLE LBZ PLS, HOF SSS- BPS THCA CES vere? TEE E€O°% FET, F8°T 82° VOLT SST LSE) CSE 6U-T
scl Uicuicone
ame O9:cwo oie mGGeCu Ll mO0\Cmm Occ Sonlamel OGrlen Sue cOells
em ol S9dleurcaleane
aanOS GSelen Mi OV MGV Ore) sl Cal Cale GCL
88s9) BLP B68 (LESEP LICE SG6e, GL SIC OSC) “(L92Cm (OPIS SEEESC: ECLS: c AST COC P6rTs Set Ta (SL OS OSit SEB be OV OVE LET
“H
Aq
Aq
AQ
“MA
YT
pue
0861
wWoIy
PMO]
a1P1S
We
IBIOaH
WUIAag
"WMO]
553
spoyjay
‘UONIPA
‘Ssoig
‘SoUTY
JOIIpsug
‘ueIyIOD
poiuUdsay
jpI11S1NHIg
UOISstUIad
AVISIOATUA)
JGeT.C-g A-UIGING
UOS}E INSHL}IS M-UIABS)
9914 (so]qu} -UTGING
A uOs}e :onsneis
| yuso10d souvoyruais
sjutod
jo Tp pue 2fP
1=4 c=
554
oA =A S=At 91 l= 8=1 6=4
u Tp Np Tp p Tp Np Tp Mp Tp Np Tp Np Tp Op Tp p Tp Op
O1=.4
9 O6¢"0 ZvVI°l AS See re Oe eae es ——— a oe aaa aon =e ad ---- “see
L S£v"0 9<0°I 762°0 9L9°T eee ee ——— ee ee a ere oe —— —_—_
8 L6t°0 <00°I Sve"O 687°l 622°0 ZOI°Z ere a oe oe anne —<<= ee —_-
6 SS°0 866°0 807°O 68¢°1 622°0 SL8°l ¢BI°O se7°Z ee = cso cacao ee anos a oor
ol 09°0 100°! -—-
990°0 ece"l OvE'O <cL°l O€Z°0 €61°Z2 OSI°O 069°2 aea oe te,
i Abou© cana s =k)
Il ¢£s9°0 OIo’l 61S5°0 L6Z°1_ 96¢°0 079°T 98Z2°0 O<0°Z <61°O £s7°z ~Z1°0 268°2 n= Ps, Sree ee. aie
Zl 1L69°0 £20"! 695°0 HLZ°1 6707°0 SLS°Il 6<<°O €16°l 7Z°0 082°Z 79I°0 S99°2 SOI°O €S0°€ ees! pn ee ee
<I 8£Z°0 8<0°l 919°0 19z°T 667°0 92S°1 16¢°O 978°I 762°0 OSI°Z 11Z°0 067°2 OVID 8<8°Z 060°0 Z8I°< oeoe eae
val 9LL°O 7sO0°l 099°0 #S2°1 LS°0 064° I%77°0 LSL°1 £€H<°0 60”0°2 LSZ°0 9So°% ¢8I°0 L99°% ZZ1°0 186°Z 820°0 L82°¢
SI =118°0 oLo’t O0ZL°O
+ ZS2°T 165°0 797°T 88°0 7OL°l 16¢°0 L£96°1T ~<0E"0 97Z°% 92Z2°0 0<s°Z I91°O L18°Z LOO Tore 890°0 ples
91 78°0 980°! CSCHLCLSO£€£9°0 9n7°T Z€S°0 £991 L¢é"0 006° 675°O £SI*Z 692°0 917°Z 002°0 189°Z ZHI"O 9H6°Z %60°0 10z*€
<I ~=%728°0 coll ZLL°0 SSz°l Z29°0 zev"l 7LS°0 O<9°T O8t°0 Lv8°I <£6£°0 820°2 <€1¢°0 61<°2 12°0 99S°Z 6Z1°O T18°Z LZ1°0 <s0°e
81 206°0 8II°l $08°0 6SZ2°1 ~=80L°0 ZZ7°I ¢€19°0 709° Z2Zs°0 <08°T S¢v°O S10°2 ~SS¢°0 8E72°Z 78Z°0 L9”°Z 91Z°0 L69°Z O91°0 $26°Z
61 826°0 ZEIT S¢8°0 s9z°1l Z72°0 SIv*l 0S9°0 78S°T 1995°0 L9L°T 9Ln°0 £96°1 96¢°0 691°Z 72Z€°0 1g<°Z $SZ°0 L6S°Z 961°0 18°C
02 2S66°0 LYI°l £98°0 TZZ2*1 ¢Z2°0 Tiv*t $89°0 L9S°1 86S°0 LEL*l SIs°O 816°l 9¢°0 OII°Z 79¢°0 g0¢°Z %62°0 OIS°2 2€Z°0 HIL°Z
1Z SZ6°0 T9T*T 068°0 LLZ°T <£08°0 807°T 8IZL°O 7SS°I €¢9°O
+ ZILT ZSS°0 188° 7L7°0 6S0°2 O00°0 90Z°Z 1€€°O ven"z 89Z°0 SZ9°Z
2 L66°0 HLT 16°0 "8Z°1 1£€8°0 LOv°l 8Z°0 £9S°T 299°0 169°I L8S°0 678° OIS°0 S10°2 LEv"O 88I°Z 89¢°0 L9€°Z 70E"0 87S°2
£2 8Iovl Z8I°l 8¢6°0 162°T 8sS8°0 LOv°T LLL°O nes°t 869°0 £L9°T 0z9°O 1Z8°l $S°0 LL6"1 ¢Lv°O OVI°Z 07°D g0¢°Z OVE'O 6L7°Z
02 LEO 66I1°T 096°0 862°I Z88°0 LOv*l S08°0 8ZS°I 8Z2°0 8S9°1T 7S9°0 L6L°T 8LS°0 976°1 L0S°0 460°Z 6£7°0 SS2°Z SL¢°0 LI”°Z
SZ S$S0°T T1121 186°0 sO¢e*T 906°0 607°T 1¢€8°0 €ZS°I_ 9SL°0 S79°T 789°0 99L°T O19°0 S16°I 0”S°0 6S0°2 ¢Lv°0 602°Z 60°0
+ Z9€°Z
92 COCMNCLOM
100° 21g" 826°0 TI?"l SS8°0 8IS°T €8L°0 scl TIL°O 6SL°1 079°0 688°I ZLS°0 970°Z S0S°0 89T°Z I%7"0 e1¢°Z
Le 680°1 €e2°1 610°I 61g*l 676°0 <17°l 828°0 SIS*I 808°0 979°1 BELO €vL°l 699°0 L98°T Z09°0 L£66°1 9€S5°0 TE1°Z €Ln°O 69772
82 vOr'l
= nZ°1 e207 S207 696°0 SI?°T 006°0 €1S*t z<8°O 819°T 79L°0 6ZL°1 969°0 L”8°T 0£9°0 OL6"1 995°0 860°Z +0S°0 622°Z
62 6II°T "S21 S01 cee"l 886°0 817°l 126°0 ZIS°I $S8°0 TI9°T 88Z°0 8IZ°T ¢ZL°0 O<e"l 8S59°0 L”6°1 S6S°0 890°Z £€S°0
O¢ tial Culses
SO OLO'T 6€<°1 900°
£61°2
TZ? 176°0 IIS*T LL8°0 909°T ZI8°0 LOL*I 87Z°0 VI8°T 789°0 Sz6°I_ 72Z9°0 170°Z 29S°0 O91°Z
l¢ Lytl €LZ2°1 S80 Stel ¢Z0°1 Scv"l 096°0 OIS*T 268°0 109°T 7¢8°0 869°Il ZZL°0 aos"! OIZL°O 906°1 69°0 L10°Z
ce ~O9I"T c8Z°1 OOI'T 2Se°T
685°0 T<1°Z
Ov0°T 8Z7°T 626°0 OIS*I LI6°0 L6S°T 958°0 069°I 76Z°0 88Z°I 7EL°0 688°I %719°0 S66°I S19°0 HOI°Z
ce ZLI°I 162° PITT 8ST SSO°T ZE°T 966°0 OIS"T 9£€6°0 76S°1 928°0 £89°l LetO PELE LSL°O HL8°1 869°0 SL6"Il
ne PBI" 662° BZI'T 9S" OLO'I
1179°0 080°2
sev" ZIO°I TIS*] 756°0 16S°I 968°0 LL9°l LEB°O 99L*l 6ZL°0
Se 098°1 ZZL°0 LS6°I S99°0 Ls0°Z
lal LOSSES OVI'T OLE*T S80°l 6e7°1l 8ZO'l ZIS*T 126°0 68S°T 716°0 1Z9°Il LS8°0 LSL°T 008°0 Z£?8°I 77L°O Ov6°T 689°0
9¢ 90271 STE*T €ST°l 9LE°T 86071 cyl £70°T
2¢0°Z
€1S*l 886°0 88S"I 7Z£6°0 999°1 LL8°0 60L°1 128°0 9¢8°1
le SLA SCOT 99L°0 SZ6°I TIL°0 810°2
~SIT*T CBE" ZIT 9071 8S0°I vIS*T 700°T 98S°I 0S6°0 Z99°I S68°0 ZvL*l 178°0 sz8°I 28Z°0 T16*l ¢¢L°0 100°Z
8¢ L221 O<e"l IZLII 88e"l PZI°l 6071 ZLO°I SIS°I 610°T S8S°I 996°0 8S9°l ¢€16°0 S¢L*l 098°0 918°T Z08°0
6£ Lez Leet LBITT 668°1 S2°0 S86°I
£6E°1 LETT ESV S80°l LIS°T EOI 78S°T Z86°0 Ss9°I 0£€6°0 6ZL°T 8Z8°0 208°! 928°0 L88°I PLL°O OL6°T
Ov 92°71 Nye] Bé6I"T 86o°I 8vI°l Losv°l 860°I 8I1S°T 80°I 78S°T L66°0 ZS9°I 976°0 HZL°l S68°0 66L°1 778°0
7] 8821 9LE°T SHz°l 9L8°l 68Z°0 9S6°1
€Zvl 102°1 glyl 9ST" 8ZS°I =TIT'l "8ST S90°T £79°T 6I10°T POL*I 26°0 892°I LZ6°0 7E8°l 188°0 Z06"I
Os 7Z<C°l <Onl S8Z°1 9071 S77! T6n"l SO0Z°T 8ES°I P9T"l
= L8S°T €ZI°l 629° 180°1 Z69°1 6£<0°1 872° 266°0
SS 9S¢°T Levl OZE"l 997°T s08°! $S6°0 798°1
8Z°T 90S°I LaZ°I BYS°T 602° Z6S°1 ZLIT 8e9°T PET°l $89°l S60°T MEL°l LSO°T S8Z°T BI0°I L¢8°l
o9 <BErl
l 607 OSE"l perl LIET Ozs*l €82°1 8SS°T 61271 86S°I 1 VIZ 6E9°T 6LI°T 289° DPT] 97L°1 S8OI°l TZZ°*T
S9 LOnT 8907°T LL¢°T oos*T 9NE°T NEST
Z20°1 Z18°1
SIE] 89S°T <8Z°1 709° =1S2°1 f09°T 8Iz°l 089°T
OL 98I°T OZL°I €ST°T T9L°T =OZI*I ZOB*T
627°1 SBn"l T OOF SIS*T ZLET 97S°T €HE"I BLS°Il Sele LISS €82°1 S79°l <¢S7°1 O89°T | ERLE CHEN Z6I°T 9SL°l Z9T°T
SL 8t7h"T
= 10s*T 220° 62S°T S6E°T LSS°T 89" L8S°T Z6L°1
OVE"T LI9*T CIE" 979°T =782°1 789°l 962°1 9JIL*T Leeal SUE
08 990° SIS*T 17" T9s°t) 9It°l 89ST 06¢°I
66I°T SBL°T
S6S°1 79E°1 HZ9°T BEE" £s9°T ZIT €89°T S8Z°1 MILT 6S2°1
S8 Z8t°T 8ZS°1 8S9°T £6S°T SEnT SPL°l CCEA PEPE
BLS°T TIT <O9°T 98E°T Os9°T 29st LOOT Leet SB9°l
06 96n°T O7S°1 ZI< MILT L8Z°T Sol" Ck SEERCI
PLP°l £9S°T 2S" L8S°I 624°1 T1191 900°T 9E9°T ¢€8E"T 199°T
S6 OIS*T O9¢"T L89°T 9¢E°T HILT Z1S*T Tel? 88z°! 69L°1
2SS°T 689°1 €LS°l 89n°T 96S°T 991°T BINT SZ CONT ¢€0°T 999°T 18¢°T 069°T 8S¢°I SIL°T 9E¢"T Tell
Oot 22S°T Z9S°T <0S°T £8S°T Z8t°T 709°I 72991
Eel LOLT
SZ9°1 TT Leo") 12771 OL9T O01 £69°T SZ¢rl LILI
osI ATS LEIA 86S°T 1S9°T 78S"T LS¢"I Tell SEErl S9L*Il
S99°l 1LS°1 649° LSS*T £69°T €9S°T 80L°T O€G"l ZL"
002 799° 789°l SIS*T LEL’l 10S°T ZSL°T 98h°T L9L°T
£€S9°T £691 €79°T HOLT ¢€¢9°T SILT €Z9°T SZL°l €19°T SEL] €09°T 9PL°T Z6S°1 LSL°T Z8S°1 892°1 ILS°1 6LL°1
21981,S- (panuyuoy)
pu
Cleat 7 tis Tp Mm 27, 15
| 090°0 907°< ket ae haan y E+. Opes i ene dea eas a Wake} mK: ea aa oe “eee ang oe aa
LI 180°0 982°¢ ¢sS0°0 90s°¢ eas ow eae ca) ee Medes Tae ae ee re ce ee eer sar eee)
81 “SUTCO SPIE $/0°0 see L0°0 £ss°¢€ Cher Re @==s> See ae ey betes See SSeS ste sane seem oe
61 ~SHI"O <20°e =ZOl°O Lé2°< 190°0 Ozv"¢ ¢t70°0 Ta9°< Soe eetSa Toe. ae ee es sa sern
a sre A
02 B8ZT°O v16°2 T<1°O 60I°< +260°0 L6Z°€ 190°0 pLy"< 8<0°0 6£9°€ prs SSe50 ests S254 weses are ses Ss ae
1Z RCO ATSIC Z9IT°O 700"< 611°0 s8i°< 1780°0 gc<e"¢ $S0°0 1Zs*€ S¢0°O 1L9°€ tare Sy emer ara Aas Se ewan care
cc 992°0 62L°% 761°0 606°2 87I°O ”80°¢ ~601°O zsz*< LL0°0 ZIv7°< 0S0°0 z9S°¢€ Z<£0°0 OOL*< Seas Sees seer, 2 ea eS
<Z 182°0 169°Z ~LZ2°0 228° 8ZI°0 166°2 9¢T°O SST°€ OOO Tl¢"¢ OL0°0 6S7°¢ 970°0 L46S°€ 6Z20°0 ScL°€ ease Sas Sar Soe
92 SI¢°O Ogs*z 09Z°0 WHL°Z 602°0 906°% S9T°O s90°< SZI°O 81Z°€ 760°0 £9¢°€ $90°0 10s°¢ £€70°0 629°€ LZ0°0 Lyl°< eo sooo
SZ 8v<°0 LIs°2 762°0 HL9°Z OVZ°0 628°2 761°0 786°Z 72S1°O Tete 9IT°O Hle< sS80°0 OIv7°’< 090°0 8<s°¢ 6£€0°0 Ls9°€ $Z0°0 99L°€
92 18¢°0 o9n°z ~HZ<E"0 OI9°Z ZLZ°0 8S/°2 2Z2°0 906°2 O8t*O aso’s I71°0 Té6te LOT°O sze"e 6L0°0 Zs" $S0°0 ZLS°€ 9€0°0 Z89°¢
LZ <17°0 607°2 96¢°0 CSS°Z ¢€0E°0 769°% €SZ°0 9€8°Z 802°0 9L6°Z LITO elle TET" snee OOO IZe"¢ ¢L0°0 067°< 160°0 zo9°¢
82 HHO £9¢°% L8C°0 667°2 EEE°0 Se9°% €82°0 CLL°Z L¢z°0 L06°2 761°0 O70°< 9ST°O 691°< ~ZZ1°0 "62°E €60°0 ZIv°< +890°0 9ZS°€
62 L070 12e°z LIv°O Isv°Z £9¢°0 c8S°Z <I¢eO ¢IL°% 992°0 £VR°Z 722°0 ZL6°Z Z8I°O 860°¢ 97T°O Oz2°< VITO sees L80°0 Ose
o¢ £0S°0 £82°2 Lth°0 Lov°2 £6¢°0 ££S°% ZHE°O 6S9°2 62°0 S8L°2 6%72°0 606°2 80Z°0 z<0°< {LTO cst°e LET°O L9z°< LOTTO 6Le"¢
I¢ 1£S°0 872°2 SLv°O L£9¢°% ZZ07°0 L87°Z IL¢°O 609°2 Z22¢°0 O£L°2 LL2°0 168°2 EZ°0 OL6°2 961°O L80°¢ O9T°O 10z*< 8ZI°O Ile"
ce 8SS°0 9122 €0S°0 Oge"Z OS7°0 9072 66¢°0 £9S°2 OS<"0 089°Z 70¢°0 L6L°% 192°0 Z16°Z 122°0 920°€ 81°0 Lets STO ICE
£¢ S8S°0 L81°Z 0£S°0 962°2 LLY°O 80”°zZ 927°0 O2S°Z SELEO ESIC EEO 9PL°% L8Z°0 8S8°2 97Z°0 696°% 602°0 8Z0°< L1°O mete
¢ O19°0 o9t°z 965°0 99Z°2 £0S°0 eL¢°% ZS7°0 18#°2 0%°0 06S5°2 LS¢°0 669°Z <I<°0 s08°z ZLZ°0 S16°2 ~£EZ°0 ZZ0°s L6T°O 9z1°<
Se 729°0 9<1°% =18S°0 L¢z°z 62S5°0 Ove"Z 8L7°0 vonZ O€7°O 0Sss°2 ¢8E"O SS9°Z 6£€£°0 19L°% 1L62°0 $98°% LSZ°0 696° 122°0 TLO°€
9¢ 8S9°0 ¢1I°Z S09°0 O1Z°Z 769°0 O1<*2 0S°0 O12 SSO ZIS°Z 601°0 V19°Z 9<°0 LIL°2 ZZ<°0 818°Z 7282°0 616°Z 772°0 610°
Le 089°0 Z60°Z 82z9°0 981°Z ~8LS°0 Z82°2 82S5°0 6L¢°Z O8%°0 LLv°Z ~HE7°0 §=9LS*% 682°0) SL9°Z Leo PLL°Z 90¢°0
+ ZL8°Z 89Z2°0 696°Z
8¢ ZOL°O ¢L0°Z 159°0 991°Z 10950 ISC2 7Z6S5°0 Os¢°Z 70S°0 sz 8S7°0 Ovs*z I1v"0 L¢9°2 LEO eel°Z O<E"0 8zZ8°Z 16Z2°0 £26°Z
6¢ €ZL°0 Ss0°Z €19°0 spr°c 7£z°Z€79°0
SLS°O £éo°% 8ZS°0 VI17°Z 28”°0 L0s°Z. 8¢°0 009°2 S6¢°O 769°% S¢°0 L8L°2 SIE°O 6L8°2
oO” 77L°0 6£0°2 69°0 £212 S9°0 OIZ°2 L6S°0 L6Z°Z 16S5°0 98¢°Z S0S°0 9L7°% 19%°0 995°2 8I7°0 £69°2 LLE"O 80L°2 BEe"0 8e8°Z
7] S¢8°0 cL6"1 O6L°0 ”70°Z 7L°0 8I1°Z OOL"O €61°Z_:
= S690 692° 719°0 9967 ~OLS°0 727°Z 82S°0 £0S°Z 88°0
+ Z8S°Z 8t707°0 199°Z
Os £160 SZ6"I IZ28°0 £86°1 628°0 1s0°Z L8L°0 911°Z ~97L°0 Z8I°Z SOL°0 0S2°Z $99°0 8I1g°Z $29°0 L8¢°2 985°0 9S7°Z 81S°0 92S°7
SS 626°0 168°T 07670 Sv6°l Z06°0 cO0°Z £98°0 §=©6S0°2 S28°O LIT ~98L°0 =ILI°Z BLO LEz°Z IIL°0 862°2 7L9°0 6S€°Z L¢9°O 127°Z
09 L¢eOrl s98°T 100°T v16"I S96°0 996°! 626°0 S10°Z £68°0 L£90°2 LS8°0 OZ1°Z 7278°0 £L1°Z -98L°0 L222 I1SL°0 €82°Z 9IL°O 8ee°2
S9 80°! S78" €S0°I 688° OZO"T 7€6°1 986°0 086°I £€56°0 L£ZO°Z 616°0 SLO°Z 988°0 €Z1°Z 72S8°0 ZLI°Z 618°0 1ZZ°Z 98L°0 CLEC
OL TEI°l T<8°T 660°1 OL8°T 890°T TT6°l LEO"l £S6°1 SO0°T S66°T 17L6°0 8<£0°2 £76°0
+ Z80°2 116°0 L212 088°0 ZL1°Z 6178°0 L12°Z
SL OLI"I 618°1 ItvI'l 9S68°T TIT'T 68° Z80°1 1g6°T ZSO°T OL6°T €ZO°l 600°2 £66°0 670°2 96°0 060°Z 7€6°0 Tet°% $06°0 ZLI°Z
08 s0Z°T O18T LLL*T p7s8"t OSI"l 828°T cele SlGylie 760°T 676° 990°T 786°1 ~6€0°1 220° I10°l £g0°2 £€86°0 L60°Z $S66°0 SE1°Z
S8 9€Z°I £08°l O12 HBT vBI°l 998°T 8cI°l 868°T cole RSG 9OI°T S96°1 O80"! 666°1 €S0°T €<0°2 LZ0°I 890°Z O00"! YOI°Z
06 79271 86L°1 O21 Lé8°1 SIZ7T 9S8°T I6I°l 988°T 99T"T LI6°T IIT 876° 9II"I 6L6°T =T60°I c10°% 990°1 770°Z 1470°1 LLO°%
S6 06271 £6L°1 LOceT ZEA 7Z°T 878°l 12271 9LB°T L6I'l S06°T LIT HE6°l OST°l £961 9ZI°T £66°1 ZOI"l ¢Z0°2 6L0°1 9S0°2
Ol VIET O6L°I 26271 918°T OLZ°1 Tv8°l 81Z°I 898° SZZ°T S68°T £0Z°1 f26°1 I8I"l 676° 8SI°l LL6°T 9<I°I 900°Z ENT DEO°S
OST <lV°l £8L°1 8Sh"I 66L°1 7t7h"T HIBT 620°1 O<el PIn'T LO8°l O0"T £98°1 s8¢"l ose’l OLE"! L68°T Scerl €16°l OvE"l Te6°l
002 196°T T6Z°T OSs°T TO8*T 6£S°1 £18°l 8ZS°T H28°T BIS"T 928° LOS*I Z£98°t sSén"T O98"! 78t7°T TZ8°T L7°T £88°l Z9%°1 968°1
,% si ay} Jaquinu
jo siossaibai burpntaoxe
ay] *ydaasaqur
555
71421S- (panuuod)
556
Z09°T Zell 6LS°I SSL°T LSS°T BLL°T OZ7"T 606"1
O0T 7S9°IT 769° SEs" cO8*l ZIS*T LeB°l 68h°T ZS8°l S9n"T
pE9"T SILT E€I19°T 9EL°T Z6S°I 8SL°I IZS°T LLB" ZHI £06°T
OST “OZL*T 9HL*T O8Z°T 0SS°I £08°T 8ZS°T 9Z8°l 90S°T OS8°T
= 9OLT ~—O9L"T C691 MLL= 6L9I BBLTT= 18h"T PLB] 971 868"T
«BSL"I_~BLLTT_- S991 ZOB*T= S91 LIST= Le"1 zeg*T 2291 §=LHBT
—o0z= -BHL*T EBL*T
= BELT 66L"T= B2L"l 18s] SLE OZER= BOT Z9BI= "6ST CLB°l
LO/e1 L691 vOut
") eH ec ee DT
7198],S-q (panuuoy)
,% si ayy Jaquinu
yo siossaibas Bulpnjaxa
ay} *ydaosajur
payudsy
Kq UoIsstULIAd
Wo ‘vaLJaWoU0rg
"JOA ‘Sp ‘OU ‘8 ‘1161“dd '$661-2661
557
558 ECONOMETRIC METHODS
6 -J01 .386 = .163 3.413 3.881 4.233 36 1.452 1.241 1.022 2.551 2.767 2.994
a -790 .464 .228 3.299 3.731 4.095 37 1.460 1.251 1.034 2.544 2.757 2.982
8 861 537 .285 3.206 3.618 3.973 38 1.467 1.261 1.045 2.536 2.747 2.969
a 22 601 2539 3.131 3.524 3.871 39 1.474 1.270 1.057 2.529 2.738 2.957
10 iD: %657, 590 3.069 3.445 3.784 40 1.480 1.279 1.067 2.522 2.729 2.946
11 1.020 .708 .438 3.016 3.378 3.710 41 1.487 1.287 1.078 2.516 2.720 2.935
12 1.060 .753 .482 2.970 3.319 3.645 42 1.493 1.295 1.088 2.510 2.711 2.925
13 V096 .795 .523 2.930 3.268 3.587 43 1.499 1.303 1.097 2.504 2.703 2.914
14 1.128 .832 .561 2.8959 5.222) 3.955 44 1.504 1.311 1.107 2.498 2.695 2.904
15 1157, 6866-597 2.863 3.181 3.488 45 1.510 1.318 1.116 2.492 2.687 2.895
16 1.183 .898 .630 2.835 3.144 3.445 46 1.515 1.325 1.125 2.487 2.680 2.885
17 1.207 .927 .661 2.809 3.110 3.406 47 1.520 1.332 1.133 2.482 2.673 2.876
18 12228°) 954.6911 2.785 3.079 3.370 48 1.525 1.339 1.142 2.477 2.666 2.868
19 15249 979 718 2.764 3.051 3.337 49 1.530 1.346 1.150 2.472 2.659 2.859
20 1.267 1.003 .744 2.744 3.025 3.306 50 1.535 1.352 1.158 2.467 2.653 2.851
21 1.285 1.024 .769 2.725 3.000 3.277 51 1.540 1.358 1.165 2.462 2.646 2.843
22 1.301 1.045 .792 2.708 2.978 3.250 52 1.544 1.364 1.173 2.458 2.640 2.835
23 1.316 1.064 .814 2.692 2.957 3.225 53 1.548 1.370 1.180 2.453 2.634 2.828
24 1.330 1.082 .834 2.677 2.937 3.201 54 1.552 1.376 1.187 2.449 2.628 2.820
25 1.344 1.100 .854 2.663 2.918 3.179 55 1.557 1.381 1.194 2.445 2.623 2.813
26 1.356 1.116 .873 2.650 2.901 3.157 56 1.561 1.387 1.201 2.441 2.617 2.806
27 1.368 1.131 .891 2.638 2.884 3.137 oy) 1.564 1.392 1.207 2.437 2.612 2.799
28 1.380 1.146 .908 2.626 2.868 3.118 58 1.568 1.397 1.214 2.433 2.606 2.793
29 155909 PV6O) 9.925 2.615 2.854 3.100 59 1.572 1.402 1.220 2.429 2.601 2.786
30 1.400 1.173 .940 2.605 2.839 3.083 60 1.575 1.407 1.226 2.426 2.596 2.780
ee Sa
Reprinted by permission of S. J. Press and R. B. Brooks from Report No. 6911, Center for
Mathematical Studies in Business and Economics, University of Chicago, Chicago, 1969.
560 ECONOMETRIC METHODS
561
562 INDEX
Ee
ee
eae
De
EIS EL
EASA ea
z 7 Pere
Eee
aE
ie:
LIE,
ee
aa
ee
EE
Soa LE
es
LY
Zep LL
Ze
Ee
eeeEEO EE
SEES
Les ie
Heteroskedasticity, which refers to the presence of non-constant variance in errors across observations, significantly affects econometric models as it violates the assumptions of homoscedasticity necessary for best linear unbiased estimators (BLUE). It complicates the prediction and estimation outputs, leading to inefficient estimators and invalid hypothesis tests if not corrected. Often seen in regression analyses involving cross-sectional data, addressing heteroskedasticity involves using generalized least squares (GLS) or robust standard errors to ensure reliable interpretations .
Transposition in matrix operations alternates the arrangement of elements from rows to columns (and vice versa) without changing the actual values. This process allows for multiplication operations that require compatible dimensions, such as turning a column vector into a row vector for the inner product calculation. Transposition thus facilitates various mathematical manipulations necessary in linear algebra and ensures operations adhere to required structural formats .
Matrix algebra simplifies the solution of least squares problems by facilitating the manipulation of entire datasets in a structured way. Using matrix operations, problems are reduced to solving normal equations like (X'X)b = X'y, where the vector b is computed using inverses of matrix products. This algebraic approach enables efficient calculations and generalization across different datasets and models within econometrics .
In matrix algebra, vectors and matrices are integral in defining the inner product. The inner product of two vectors is calculated by transposing one vector and multiplying it by another vector, which results in a scalar value. This involves row vectors from one matrix and column vectors from another, highlighting the role of vector alignment and dimensions in determining the matrix product .
Using restrictions in simultaneous equation systems aids in the model identification process. Restrictions, such as exclusion restrictions or linear homogeneous restrictions, guide the assignation of variables to specific equations, simplifying the complexity of multi-equation models. They ensure unique and meaningful solutions by allowing the determination of unknowns in structurally defined relationships, addressing identification challenges like multicollinearity and simultaneity biases .
The F-statistic in one-way ANOVA assesses the variation among class means by comparing it to the variation within classes. It effectively tests the homogeneity of class means by evaluating the sum of squares due to the difference between the class means and the overall mean, relative to the sum of squares due to variation within the classes .
Logistic regression plays a crucial role in addressing the inherent limitations of linear regression types, notably when the dependent variable is categorical. By employing the logistic function, it maps predicted probabilities to a range between 0 and 1, mitigating issues like heteroscedasticity and ensuring meaningful, bounded predictions. It successfully models binary, and sometimes multinomial outcomes, providing more robust, interpretable models under conditions where traditional linear regression assumptions do not hold .
The identifiability condition ensures that econometric models yield unique solutions to their parameters by setting necessary constraints, often through rank conditions and restrictions on variables. It prevents overspecification and ensures that each equation in a model can be isolated and estimated accurately. This condition is vital in complex multi-equation or structural models, where failure to meet identifiability can lead to indefinite parameter estimation, introducing ambiguity and bias .
Sums of squares are utilized in regression models to quantify different sources of variability within the data. Total sum of squares (TSS) is partitioned into explained sum of squares (ESS) and residual sum of squares (RSS). ESS measures the explained variability due to the regression model, while RSS measures the unexplained variability. This partitioning enables tests of significance, such as F-tests, to determine the contribution of certain variables or the overall fit of the model .
The concept of linear independence is crucial for the existence of an inverse matrix because only if the columns (and equivalently, rows) of a matrix are linearly independent can it span the full space necessary to form a basis. This ensures the matrix is non-singular and invertible, allowing for the unique solution of equations, as seen in the case of the matrix X'X in linear regression .









