Econometric Methods

100% found this document useful (1 vote)
717 views584 pages

Uploaded by

hasniayoub528

PANN

Ae
Rg
a

OT
Loe

i ile

abit alt
AUVea,Rei
faite
o,

ONIN
LPP de
RAte
NEr
danni
Johnston, J.(John),

imui 2931 0135033


DUE

O at
Te

[AO
| 2 as a
> | co
by Mm
Pp|< 7
mt :
V
jo e CN) eel =
ECONOMETRIC
METHODS
Third Edition

J. Johnston
University of California, Irvine

McGraw-Hill Book Company


New York [Link] San Francisco Auckland Bogota Hamburg
Johannesburg London Madrid Mexico Montreal New Delhi
Panama Paris Sao Paulo Singapore Sydney Tokyo Toronto

197964
IN MEMORY OF
B. and J.

This book was set in Times Roman by Science Typographers, Inc.


The editors were Patricia A. Mitchell and Scott Amerman;
the cover was designed by Nadja Furlan;
the production supervisor was Phil Galea.
The drawings were done by Burmar.
Halliday Lithograph Corporation was printer and binder.

ECONOMETRIC METHODS

Copyright © 1984, 1972, 1963 by McGraw-Hill, Inc. All rights reserved.


Printed in the United States of America. Except as permitted under the
United States Copyright Act of 1976, no part of this publication may be
reproduced or distributed in any form or by any means, or stored in a data
base or retrieval system, without the prior written permission of the
publisher.

1234567890 HALHAL 8987654

ESBNeG-O?-03eb65-1

Library of Congress Cataloging in Publication Data


Johnston, J. (John), date
Econometric methods.

Includes bibliographical references and index.


1. Econometrics. Ieelitie:
HB139.J65 1984 330’.028 83-14899
ISBN 0-07-032685-1
CONTENTS

Preface Vil

Chapter 1 The Nature of Econometrics


il Economic Model Building
1-2 A National Income Model
1-3 Unanswered Questions
1-4 Role of Econometrics
1-5 Structural and Reduced Forms
1-6 Multipliers and Dynamic Properties BRN
CONN

Chapter 2 The Two-Variable Linear Model 12


ek The Linear Specification 12
39 Least-Squares Estimators 16
2-3 The Correlation Coefficient
2-4 Properties of the Least-Squares Estimators 25
Ons Inference in the Least-Squares Model 34
2-6 Analysis of Variance in Least-Squares Regression 39
7 Prediction in the Least-Squares Model

Chapter 3 Extensions of the Two-Variable Linear Model 48


3-1 Repeated Observations and a Test of Linearity 48
a) Nonlinear Relations
3-3 Transformations of Variables 61
3-4 Three-Variable Regression 74

Chapter 4 Elements of Matrix Algebra 89


4-1 Operations on Vectors and Matrices 92
4-2 Matrix Formulation of the Least-Squares Problem 101
4-3 Geometric Interpretation of Least Squares 104

iii
iv CONTENTS

4-4 Solutions of Sets of Equations 113


4-5 The Eigenvalue Problem 141
4-6 Quadratic Forms and Positive Definite Matrices 150
4-7 Maximum and Minimum Values 153

Chapter 5 The k-Variable Linear Model 161


5-1 Preliminary Statistical Results 161
5 Assumptions of the Linear Model 168
53 Ordinary Least-Squares (OLS) Estimators 171
5-4 Inference in the OLS Model 181

Chapter 6 Further Topics in the k-Variable Linear Model 204


6-1 Estimation Subject to Linear Restrictions 204
6-2 Tests of Structural Change 207
6-3 Dummy Variables Deo
6-4 Seasonal Adjustment 234
6-5 Multicollinearity 239
6-6 Specification Error 259

Chapter 7 Maximum Likelihood Estimators and


Asymptotic Distributions 267
Tel Review and Preview 267
1p Some Remarks on Asymptotic Theory 268
73 Maximum Likelihood Estimators 274
7-4 Some Asymptotic Results for the k-Variable Linear Model 279

Chapter 8 Generalized Least Squares 287


8-1 Sources of Nonspherical Disturbances 287
8-2 Properties of OLS Estimators under Nonspherical Disturbances 290
8-3 The Generalized Least-Squares Estimator 291
8-4 Heteroscedasticity 293
8-5 Autocorrelation 304
8-6 Sets of Equations 330

Chapter 9 Lagged Variables 343


9-1 Sources of Lagged Variables 343
9-2 Estimation Methods 352
9-3 Time-Series Methods 37)

Chapter 10 A Smorgasbord of Further Topics 384


10-1 Recursive Residuals 384
10-2 Spline Functions 392
10-3 Pooling of Time-Series and Cross-Section Data 396
10-4 Variable-Parameter Models 407
10-5 Qualitative Dependent Variables 419
10-6 Errors in Variables 428
CONTENTS V

Chapter 11 Simultaneous Equation Systems 439


11-1 Some Illustrative Simultaneous Systems 439
11-2 The Identification Problem 450
11-3 Estimation of Simultaneous Equation Models 467

Chapter 12 Econometrics in Practice:


Problems and Perspectives 498

Appendix A— Mathematical and


Statistical Appendixes S17
A-1 Functions and Derivatives SAF
A-2 Exponential and Logarithmic Functions 518
A-3 Operations with Summation Signs 521
Random Variables and Probability Distributions 524
A-5 Normal Probability Distribution 527
A-6 Lagrange Multipliers and Constrained Optimization 529
A-7 Relations between the Normal, x7, ¢, and F Distributions 530
A-8 Expectations in Bivariate Distributions 531
A-9 Change of Variables in Density Functions 535
A-10 Principal Components 536

Appendix B— Statistical Tables 545


B-1 Areas of a Standard Normal Distribution 547
Student’s ¢ Distribution 548
B-3 x? Distribution 549
B-4 F Distribution 550
B-5 Durbin-Watson Statistic (Savin- White Tables) 554
B-6 Wallis Statistic for Fourth-Order Autocorrelation 558
B-7 The Modified von Neumann Ratio 559
B-8 Significance Values for cy in the Cusum of Squares Test 560

Index 561
ae vachefate
eeie: Mt end ig Mi
ie “whe aah A
Oe
‘WAySee iér ne sil a gue at
«MC NA Mien ines
mi
ole
SS ’ To ; ais OThiiig Bes
ae t ond ’ eat dheott Cyee pee ae ;

aa ae
PREFACE

This edition has been completely rewritten. The main features of the new edition
are the following:

1. The mathematical, statistical threshold has been lowered in order to make the
material more accessible to students with only an elementary prior knowledge
of statistics. This has resulted in a somewhat larger proportion of words to
symbols in the early chapters than would otherwise have been the case. A
series of paragraphs on mathematical, statistical topics has also been provided
in Appendix A. These are keyed into the early chapters to ease the transition
into the heart of the book. For the same reason the chapter on matrix algebra
has been retained and, indeed, expanded to include a geometric as well as an
algebraic treatment of some topics.
2. All the inference procedures for the general linear model have been derived as
special cases of a single basic procedure, namely, the testing of a set of linear
restrictions on the parameters of the model (Chapter 5). This leads in turn to
an exhaustive treatment of tests for structural change (Chapter 6). Chapter 6
also contains extended treatments of the use of dummy variables and of
multicollinearity among the regressors.
3. Every effort has been made to cover both new and old topics on which
substantial work has been done in recent years and which are thought to be
significant and enduring rather than passing fancies. Such topics include the
estimation of sets of equations with special reference to transcendental
logarithmic approximations and applications in energy economics (Chapter 8),
autocorrelated error terms (Chapter 8), time series techniques (Chapter 9) and,
in A Smorgasbord of Further Topics (Chapter 10), the following “menu”:
recursive residuals, spline functions, pooling of time-series and cross-section
data, variable-parameter models, qualitative dependent variables, and errors
in variables. The author has also granted himself the indulgence. of some
personal comments on the present state of econometrics (Chapter 12).

vii
Vili PREFACE

4. The problem sets have been extended to become truly Anglo-American, with
offerings from the Royal Statistical Society plus the universities of Cambridge,
London, Manchester, and Oxford on one side of the pond, and Chicago,
Michigan, Yale, and Washington on the other side. Grateful acknowledgment
is made to various anonymous authorities for the first set, and to Arnold
Zellner, Jan Kmenta, Peter Phillips, and Charles Nelson for the second.
Appendix B also contains an extensive set of statistical, econometric tables,
and grateful acknowledgment is made to the appropriate sources for permis-
sion to publish them.

My debts to many individuals can be warmly acknowledged but never fully


recompensed. Craig Riddell (University of British Columbia) read the entire
manuscript and contributed many valuable comments. I am very grateful to both
of them. My thanks also go to Ian McAvinchey (University of Aberdeen) and to
my colleagues Ken Chomitz, Max Fry, and Charles Lave (University of Cali-
fornia, Irvine) fora similar service. Ken Chomitz has also produced a solutions
manual for comments on various chapters. This in some places is almost a
supplementary text in that extended solutions have been written for various
problems, outlining particular issues that could not be dealt with in the main text.
Copies of the solutions manual are available to instructors on application to the
publishers. Kathy Alberti and Barbara Sawyer did a magnificent job on a difficult
manuscript. In addition, Barbara Sawyer did the preliminary artwork, prepared
the tables in Appendix B, and proofread the entire manuscript. Finally, I
gratefully acknowledge the questions and suggestions from teachers and students
in many parts of the world during the two decades since the first publication of
this book. I can only hope that this third edition will inspire a similar response.

J. Johnston |
CHAPTER

ONE
THE NATURE OF ECONOMETRICS

Before asking the question, “What is econometrics?,” one must pose the prior
question, “What is economics?” The answer to the second question will indicate
the role that econometrics can play in the development of economics. Although
the focus of the exposition in this chapter will be on economic models, the
methods that have been developed in econometrics can and do play an important
role in other social sciences, where there is a concern with building and estimating
models of the interconnections between various sets of variables in a predomi-
nantly nonexperimental situation.

1-1 ECONOMIC MODEL BUILDING

Economists seek to understand the nature and functioning of economic systems.


Their concerns may relate to global aggregates, or macro quantities, such as the
value of the gross national product (GNP), the level of employment, or the
current level of the consumer price index. Alternatively, the focus of attention
may be some sector or area of the economy, such as production and employment
in the automobile industry or the price and volume of the peanut crop in Georgia.
One objective of such an understanding is to be able to make conditional
predictions of the likely future development of the system and hopefully enable
economic agents, whether government, business, or consumers, to take action to
control to some degree the evolution of the system. Another important objective
is to test economic theories about the system.
2 ECONOMETRIC METHODS

The first step in seeking to understand the functioning of a system is to build


a theoretical model. All models are inevitably simplifications of reality, and the
model builder seeks to capture the fundamental features of the system being
studied. The performance of an economy, or a sector of an economy, at any point
in time will depend upon the decisions of various economic agents, taken in the
context of the existing state of technology with given stocks of capital, labor, and
other limited productive resources. Thus theoretical models typically contain
behavioral relations, which describe the forces thought to determine the behavior
of various groups of economic agents, and technological relations, which describe
the restrictions imposed by the current technology and endowments of the system.
Often technological relations, such as the production function, describing the
maximum output achievable with various inputs of capital, labor, and other
productive resources, may not appear explicitly in the model, but will have been
used in the derivation of behavioral relations, such as the demand function for
labor, and so on. In addition to behavioral and technological relations, economic
models typically contain identities or definitional relations.

1-2 A NATIONAL INCOME MODEL

As an example of the model building process let us consider one of the simplest
forms of the national income model, which is used as a pedagogic device in most
elementary textbooks on economics. Such models begin with the national income
identity. For a closed economy with no foreign trade, this identity in any period is
Y =cChi ae (1-1)
where y= gross national product (GNP)
c=consumption expenditure
i= investment expenditure
g=government expenditure
all expenditure flows being measured in real terms. The construction of the model
proceeds with the formulation of hypotheses about the determinants of the
expenditure components of GNP.
Consumption expenditure might be hypothesized as dependent on disposable
income, net of tax, and the rate of interest. Thus we write}

c= f((1-7)y,r) (1-2)
where T= tax rate (assumed constant across the economy)
r=rate of interest

The theoretical expectations about this relation are


Orefr Sule se0 (1-3)
where f; indicates the partial derivative of the function with respect to the ith
7 See App. A-1, Functions and Derivatives.
THE NATURE OF ECONOMETRICS 3

argument. The first assumption in Eq. (1-3) is that the marginal propensity to
consume out of disposable income is a positive fraction less than unity. The
second assumption is that a rise in the rate of interest will have a depressing effect
on consumption since it raises the return on savings, increases the cost. of
financing «consumer durables, and also reduces the nominal value of bonds, which
are a part of wealth, which in turn might appear as an argument of the
consumption function but has been omitted from Eq. (I--2) on grounds of
simplicity.
The investment function may be specified as
i= f(Ay,r) (1-4)
with
f a 0, h <0 (1-5)
The term A y indicates the change in GNP. Investment is positively influenced by
profit expectations, and the crude assumption here is that observed changes in
real GNP serve as a proxy for these profit expectations. The rate of interest is
again expected to be negatively related to this form of expenditure.
Collecting results, so far we have a three-equation model, namely,
Scale ge

Cm (Vy f)
Aya)
supplemented by the expected signs on derivatives expressed in Eqs. (1-3) and
(1-5). This model then constitutes a theory about the joint determination, or
“explanation,” of the three variables c, i, and y. Such an explanation is obviously
conditional on the values assumed for g, r, and r. The model builder now faces a
decision on how to treat these remaining variables. Should one formulate theories
to explain the determination of government expenditure, the rate of interest, and
the tax rate, thus expanding the system to one of six equations? If one does, the
new equations will almost certainly contain some explanatory variables on the
right-hand side that have not previously appeared in the system, and these, in
turn, raise the question of how they are to be treated. It might seem that economic
models must become infinitely large, but there is not, of course, an infinite
number of variables to be explained. In any case the behavior of model builders is
very pragmatic. Everything is relative: all depends on the problem at hand. For
some purposes a small model is sufficient and some variables, which in larger
models would have explanatory equations, may be left “unexplained.” In the
present instance we make no pretense at economic realism, but only require a
model for illustrative purposes, so we will restrict it to the three equations already
specified.
The model contains only two behavioral relations, one for consumption and
the other for investment. Economic theory has done two things. First, it has
specified the list of explanatory variables on the right-hand side of each equation,
and second, it has indicated the expected signs on the partial derivatives. This is
usually as far as theory per se can go, but it still leaves a series of important
questions unanswered.
4 ECONOMETRIC METHODS

1-3 UNANSWERED QUESTIONS

Functional form Theoretical considerations alone cannot usually specify the


functional form connecting the variables in a relationship. Many functional forms
are consistent with a priori signs on derivatives. Letting
z=(l-T)y
denote disposable income and omitting the rate of interest variable, the following
functional forms all give c as a monotonically increasing function of z and, with
appropriate restrictions on parameters, could satisfy the condition that the
marginal propensity to consume is a positive fraction:
C= Op OZ
c Az"
=1

These functions, however, have different qualitative implications. In the first, an


extra $100 of income always produces the same absolute increase in consumption
expenditure. The second and third functions both exhibit a declining marginal
propensity to consume as income rises. However, the second function implies that
consumption rises indefinitely with income, while the third shows consumption
approaching a saturation or asymptotic level a) as income becomes very large.
This is a typical example of the fact that the qualitative restrictions deriving from
economic theory do not serve to delimit functional forms very closely.

Data definition and measurement Theory is sometimes precise and sometimes


sloppy in the matter of definitions. In this model, for instance, should consump-
tion be taken to mean expenditure, including actual expenditure on consumer
durables, or should consumption of durables be treated as an implicit flow -
measured by the value of services from the existing stock of consumer durables?
If the second definition is taken, is this consistent with the definition of the same
variable in the national income identity? What is meant by income? Should it be
adjusted for purely seasonal fluctuations or not? Is it to be taken as some recently
observed level, or should it be interpreted as some kind of “permanent” or “long
run” income? There are many different rates of interest. Should we select a
“representative” rate or some combination of rates, and should this variable be
treated the same way in both consumption and investment functions?

Lag structure Somewhat allied with problems of data definition are problems of
lag structure. Should investment be specified as responding to the current interest
rate or to some set of previous interest rates in view of the inevitable time lags
involved in making and implementing investment decisions? Again, by the nature
of things, economic theory cannot be specific about appropriate lag structures.
Moreover, much of economic theorizing has necessarily been about equilibrium
positions, as, for example, the equilibrium rate of consumption corresponding to
some level of income, which has, in theory, remained constant long enough for
consumers to become fully adjusted to it. In practice, the world is always
THE NATURE OF ECONOMETRICS 5

staggering from one disequilibrium position to another, so actual data reflect


adjustment processes rather than equilibrium positions. Equilibrium theory, by
definition, says nothing about adjustment processes, and theories of adaptation
and adjustment are still in a fairly primitive state.

Qualitative versus quantitative implications The theoretical model does yield


unambiguous qualitative implications, such as that a rise in the rate of interest
will depress GNP anda rise in government expenditures will increase it. In more
complicated models qualitative conditions on the various equations may not lead
to unambiguous predictions about the overall behavior of the model. If our
simple model asserted that the rate of interest had a positive effect on consump-
tion and a negative effect on investment, the direction of the rate of interest effect
on GNP could not then be known without quantitative knowledge of the two
separate effects and the magnitudes of consumption and investment. In practice,
of course, policymakers are vitally concerned with the likely magnitude and timing
of the effects of changes in the rate of interest, tax rates, or government
expenditure. The expected signs of partial derivatives cannot provide this kind of
information.

Choice between theories So far, in discussing the previous four problems, we have
implicitly assumed that our theoretical model is “correct,” but how can we tell
whether a theory is sufficiently correct to be used as a valid tool of analysis?
Perhaps there are as many theories as there are theorists. There is, in practice, a
very important and very difficult problem involved in attempting to discriminate
between competing theories. Some theoretical models differ in degree but not in
kind. They might be regarded as variations on a theme. For example, another
theorist might accept the general form of our consumption and investment
functions but wish to add wealth as an additional explanatory variable to the first
equation and capital stock to the second. At the other end of the spectrum would
be a theorist who rejected the Keynesian flavor of our model and advanced
instead a supply-determined theory of output or a model in which the fundamen-
tal driving force was the money supply.

1-4 ROLE OF ECONOMETRICS

Econometrics tackles all five questions. Its basic task is to put empirical flesh and
blood on theoretical structures. This involves several crucial steps. First of all, the
‘theory or model must be specified in explicit functional form. The econometrician
does not have any special insights in this area that are denied to the economic
theorist, so one usually starts with the simplest functional forms that are con-
sistent with the a priori specifications. At the same time one makes an initial
specification of the lag structure. As an example we might specify the three-equa-
6 ECONOMETRIC METHODS

tion national income model as


C,=
a) + a,(1—7)y,
+ ayy, (1-6)

i, = By + Bi(y-1 = yo) pray (1-7)

Ve Cp Bg, (1-8)
with a priori expectations
USar<st, a, < 0, B; > 0; B, <0

The subscripts on the variables refer to time periods. The unit time period can be
anything considered relevant by the econometrician, provided there exist ap-
propriate data in terms of that unit. However, it is typically a quarter or a year,
and the model is in discrete, not continuous, time.
The second task of the econometrician is to decide on the appropriate data
definitions and assemble the relevant data series for the variables which enter the
model. The third task is to perform a “marriage” of theory and data by means of
statistical methods. The “offspring” of the marriage are various sets of statistics,
which shed crucial light on the validity of the theoretical model that has been
specified. The most important set consists of the numerical estimates of the
parameters of the structural form. The Greek letters of Eqs. (1-6) and (1-7) are
now replaced by numbers. There are further statistics which enable one to assess
the reliability or precision with which these parameters have been estimated,
which in turn helps us to check whether the model conforms to the theoretical
expectations about signs of derivatives. There are still further statistics and
diagnostic tests that help one to assess the performance of the model and decide
whether or not to proceed sequentially by modifying the specification in certain
directions and testing out the new variant of the model against the data.
Most of this book will be concerned with the statistical methods used by
econometricians in estimating, testing, and evaluating economic models. Histori-
cally, econometrics started with the corpus of methods inherited from classical
Statistics. These methods, however, were mainly developed in the context of the
experimental sciences. Special problems of statistical inference arise in economics,
where the possibility of controlled experiments is the exception, not the rule, and
these will be described in the chapters to follow. All that remains to be done in
this introductory chapter is to indicate some of the possible applications of an
econometric model, once it has been estimated. This will again be done with the
simple model outlined above.

1-5 STRUCTURAL AND REDUCED FORMS

Equations (1-6) to (1-8) constitute the structural form of the model. The structural
form may be regarded as a theoretical explanation, or hypothesis, about the
determination of the three variables y,, c,, and i,, conditional on the values
currently assumed by g, and r, and also on the recent history of the system as
represented by y,_,, y,5, and r_,. This enables us to make the following
THE NATURE OF ECONOMETRICS 7

classification of the variables in the system:


Current endogenous variables: Corley,
Lagged endogenous variables: Wor vis
Current exogenous variables: GAL,
Lagged exogenous variables: hey

The crucial distinction is between endogenous and exogenous variables. The


former are those variables whose current values are, in theory, explained by
the functioning of the model. The model, however, has nothing to say about the
determination of the exogenous variables. A second important distinction is that
between the current time period ¢ and previous periods, such as ¢ — 1, t — 2, and
so on. When we come to study the functioning of the model in period 1, all lagged
values, whether of endogenous or exogenous variables, are already given and
cannot now assume new values. Once values are also fed in for the current
exogenous variables g, and r,, the model then delivers the values of the current
endogenous variables c,, i,, and y,. This point may be expressed formally by
recasting Eqs. (1-6) to (1-8) in an alternative form. Substituting Eqs. (1-6) and
(1-7) in Eq. (1-8) and rearranging gives

Y, = (a + By)6 + 067, + BiS(y%-1 — Y-2) + Bodn%_, + 8g, (1-9)


where
1
. 1—a(1 —+1)

The important point about Eq. (1-9) is that only one current endogenous variable
appears in the equation, namely, y, on the left-hand side. The right-hand-side
variables are a mixture of current exogenous variables and lagged variables,
whether endogenous or exogenous. This collection of three sets of variables is
labeled the class of predetermined variables, since, from the viewpoint of the
model. in period ¢, their values either are determined -by_the past history of the
system or_are [Link] in the current period. The investment equation
already has nothing but predetermined variables on the right-hand side, so we
repeat it here:
i, = Bo + B\(Y-1 — Y-2) + Boni (1-10)
Finally, substituting Eq. (1-9) in the consumption function gives

c, = [ay + a,(1 — 7)(ap + By)


d] + [e, + a,(1 — 7)a, |r,

SOL eh) BO oY,


<2)ita lam 8,07—1 a,(1 — 7)dg,
(1-11)

The three Eqs. (1-9), (1-10), and (1-11) constitute the reduced form of the
model. Each equation of the reduced form expresses a current endogenous
variable as a function only of predetermined variables. The reduced form may be
8 ECONOMETRIC METHODS

Inputs Time period ¢ Output

Exogenous variables
current and lagged ~~_,
Predetermined Current endogenous
variables variables
Lagged endogenous oat
is variables

Figure 1-1

written compactly as

Vz, = Mo + M18, + Mh FM3%—1 + M4W-1 + MsV-2 (1-12)

Cp = My F M918, F Mh FM 3%—1 + MaM—1 + M5 Vi-2 (1-13)


i, = To +733%1 + M41 + 35Y,-2 (1-14)
where the 7’s are the functions of the structural parameters indicated in Eqs.
(1-9) to (1-11). Schematically, the reduced form is indicated in Fig. 1-1.
The reduced form also indicates that there is one-way causation in the model
in the sense that the exogenous variables influence the current endogenous
variables, but there is no feedback in the opposite direction: current endogenous
variables do not influence the exogenous variables.

1-6 MULTIPLIERS AND DYNAMIC PROPERTIES

The 7’s of the reduced-form equations are economically very important parame-
ters. They measure the impact in the current period on each endogenous variable
of a unit change in any predetermined variable. Consider, for example, a unit
increase in the level of g,. From Eq. (1-8) of the structural form there would be a
simultaneous increase of one unit in GNP. But from the consumption function
(1-6), increases in GNP will induce increases in consumption, which in turn, from
Eq. (1-8), will induce further increases in GNP. The reduced-form coefficient
dy, 1
oe eras \

shows the end result of this process in period t. This is the national income
multiplier of simple Keynesian theory. For example, if + = 0.25 and a, = 0.8,
7, = 2.5, so that a unit increase in government expenditure, with tax rates and
all other parameters unchanged, would raise national income in the same period
by 2.5 units. Similarly, an inspection of

To = (a + By)
shows that a unit increase (upward shift) in the intercept of either the consump-
tion or the investment function would have equal multiplier effects on GNP. All
THE NATURE OF ECONOMETRICS 9

the 7’s are multipliers, and they are termed impact multipliers, because they show
the effect in the current period of changes in predetermined variables. Estimates
of the structural coefficients can yield estimates of the reduced-form coefficients,
and so these impact multipliers can be evaluated. Alternatively, the reduced-form
equations may be estimated directly. These topics will be discussed in the chapter
on simultaneous equation estimation later in the book.
The impact effects in period ¢ are not the end of the story. Let us write Eq.
(1-12) in first difference form,
Ay, = m Ag, + m2Ar, + m3An_) + MAAN + MsAY—2 (1-15)
where
Ay, See ile Vet aie:
Let us suppose that g and r have been held constant sufficiently long for y to settle
down at some constant equilibrium level. This involves the implicit assumption
that equilibrium values exist and that the system is stable, and we will return to
this point below. Equilibrium thus implies
Ag,_.=°°° =0
Ag, = Ag,_, =
= 3... = (0
= Ar,_,
Ar, Say ay
= hy ai
Ay = Aye

Now suppose that the level of government expenditure in period ¢ + 1 is raised


by an amount d and then held constant at the new level indefinitely, that is,
Age Go, Soi O85 =. 0
From Eq. (1-15) the impact effect on national income in period ¢ + 1 is
AY41 = m4
In period ¢ + 2, Eq. (1-15) reads
AYi42 = MAN = M414
In period ¢ + 3, the equation reads
AV43 = MAN. + MAM41
= m4 4 + m571d
Thus the one-step change in g sets off a sequence of changes in y because of the
lags in the system. There is thus a whole series of lagged multipliers, namely,

OV eats2 ™) zero lag, or impact multipliers


98441

IYi42 = 11471 one-period lag


08,41

0 é
Weise (iy + m5)
7 two-period lag
08441

The estimated reduced form can be applied sequentially to trace out the dynamic
10 ECONOMETRIC METHODS

effects of a postulated change in any exogenous variable. The multipliers at


various lags are called interim multipliers, and if the impact and interim multi-
pliers are summed over an infinite time horizon, assuming that the sum converges,
we have a ¢otal multiplier giving the final effect on the equilibrium value of an
endogenous variable of a one-step change in an exogenous variable.
Finally, we may look briefly at a third way of expressing the system, which
casts light on the question of stability. Equation (1-12) may be rearranged as
Ve May1 — M52 = Moet Mikey MiG Miaka (1-16)
This is a second-order nonhomogeneous difference equation in y.} It is a fluke of
this simple model that this reduced-form equation did not contain lagged values
of any other endogenous variable, but it is always possible to derive a difference
equation for each endogenous variable, which contains only exogenous variables
on the right-hand side and does not contain any other endogenous variables,
current or lagged. Equation (1-16) may be expressed more simply for present
purposes as
Ve Teas ss, en eee (1-17)
A remarkable feature of linear dynamic models is that each endogenous variable
in the system can be described by a difference equation of the same order and
with identical coefficients, but differing only in the linear combination of exoge-
nous variables appearing on the right-hand side. To demonstrate this result in the
present model, we take the investment function
1s Bot Bi yen ey) een
Lag it one period and multiply by 7,4, lag it two periods and multiply by 7, 5, and
subtract both equations from the current equation. This gives

bie Ugh poets Fist; ee ee yay ae T15Y,-3)


= Bil Veo Taya 75,4)
+ Bo(1 — m4 — m5)
+ By(7,_, — M4h—2 — M53)
The first two terms in parentheses on the right-hand We are seen from Eq. (1-17)
to be functions only of exogenous variables. Thus
iy — Magi, — T5l;-2 = h(g, 1) (1-18)
Finally, applying the same treatment to the national income identity gives
Ce > Malaya M50,9 — VT Vag 5 Y-2)
aie Ti) — T5i,-2)
WG 48,1 — ™58,-2)
and, using Eqs. (1-17) and (1-18), we have

6, — Ma Fist, => = k( g.7) (1-19)


+ For an introduction to difference equations, see A. C. Chiang, Fundamental Methods of Mathe-
matical Economics, McGraw-Hill, New York, 1984.
THE NATURE OF ECONOMETRICS 11

Thus all three endogenous variables are characterized by a second-order difference


equation with the same coefficients. The endogenous variables will therefore
display the same dynamic behavior. The clue to this behavior comes from the
roots of the common characteristic equation

The structural and reduced-form equations show


: B,
CS BNEs 0
te ea
Denoting this parameter by a, the characteristic equation becomes
NV -ac\t+a=0
with roots

a + ja(a — 4)
UMguagar 2
If a < 4, the roots are complex and the economic structure is inherently cyclical.
The product of the roots is also a, and so if a < 1, the cycles are damped, but if
1 < a < 4, the cycles are explosive. We see that a depends on

B,, the acceleration coefficient of the investment equation


a,, the marginal propensity to consume
7, the tax rate

Thus, once again, empirical estimates of these parameters shed crucial light on the
nature of the economic structure. The above model has been highly simplified for
expository purposes, but these methods of analysis can be and are applied to large
systems. We have attempted to illustrate the importance of econometric estima-
tion and testing by reference only to a simplified aggregate system. Other varied
illustrations of the power and range of econometrics will be given in the course of
the book.
CHAPTER

rwoO
THE TWO-VARIABLE LINEAR MODEL

The national income model of Chap. 1 has two complications that we do not wish
to tackle right away. First of all, it is a simultaneous equation model with three
equations to explain the determination of three endogenous variables. Second,
each behavioral equation contains more than two variables. We will, however,
begin our exposition of econometric methods by concentrating upon a single.
equation with just two variables. It is not claimed that a single two-variable
equation is an adequate model of any economic process, but starting with it has
the double advantage that certain fundamental ideas can be introduced in the
simplest of all settings and that the tools and concepts developed for the
two-variable model are essential building blocks for the more complicated cases
which are treated in the rest of the book.

2-1 THE LINEAR SPECIFICATION

The relevant theory is now assumed to postulate

Y =/(X) (2-1)
where Y indicates the dependent (explained) variable and X the independent
(explanatory) variable. We may have theoretical expectations about the sign of
f’(X) or about the range of. values in which it lies. In this chapter we will deal
only with linear specifications.
12
THE TWO-VARIABLE LINEAR MODEL 13

A linear specification means that Y, or some transformation of Y,9 can be


expressed as a linear function of X, or some transformation of X. In this sense,

Y=a+ BX (2-2)
Y = aX? | (2-3)
and Yee exp a+ p>} (2-4)
are all linear specifications. The first is already linear in Y and X. The second, on
taking logarithms of both sides of the equation, may be written as
log Y = loga + Blog X (2-5)
which is linear in log Y and log X. The third is

logY=a+t Bo (2-6)
which is linear in the logarithm of Y and the reciprocal of X. The function
Y=a+ BX + yx?
is linear in Y, X, and X?, but it is not a two-variable linear function, and so its
treatment will be postponed to Chap. 3. The function
6
Y=a+
Ap
however, where a, 8, and 6 are unknown parameters, cannot be reduced to a
linear function of some transformations of Y and X, and so cannot be treated by
the methods of this chapter.
The first step in the econometric investigation of the relationship between Y
and X is to obtain a sample of n pairs of observations on the two variables. The
sample data are thus indicated by
X,, Y,ql b= lezen
Next we must make a choice between specifications such as Eqs. (2-2), (2-3), and
(2-4). At this stage the choice is made by plotting the raw data or various
transformations of them on two-dimensional scatter diagrams to see which, if
any, yields an approximately linear scatter. Examples of various typical shapes
and appropriate linearizing transformations will be given in Chap. 3. Here we will
assume that Y and X denote appropriately transformed data, and so we postulate
the linear relationship
Y=a+t+ BX
where a indicates the intercept made by the line on the vertical, Y, axis and B
indicates the slope of the line.
The econometrician now faces the task of using the sample data to obtain
numerical estimates of the unknown parameters a and B. If the postulated
relationship were really true, one would have no problems at all; one would need

+ See App. A-2, Exponential and Logarithmic Functions.


14 ECONOMETRIC METHODS

just two sample points and a ruler to join them. Further sample points would lie
on the same straight line and would convey no additional information. However,
exact functional relationships such as Eq. (2-2) are inadequate descriptions of
economic behavior. Scatter diagrams do not yield points which all lie on a single
straight line. Thus the specification of the linear relationship is expanded to
Y=at+BX+u (2-7)
where u denotes a stochastic variable with some specified probability distribution.
The purpose of the u term is to characterize the discrepancies that emerge
between the actual, observed values of Y and the values that would be given by an
exact functional relationship. mi
To fix these ideas let us suppose that we have data from a budget survey with
X representing household disposable income and Y household consumption
expenditure. Clearly, household expenditure will depend on some crucial factors
in addition to income, such as household size and composition, so let us suppose
that we are looking at the relationship between Y and X within a subset of
households of a given size and composition. Nonetheless it would still be
unrealistic to expect all households with a given income X, to display exactly the
same expenditure a + £X;. First of all, even among households of the same size
and composition and with the same income, there will be variations in the precise
ages of the parents and children, in the number of years since marriage, in
whether the husband is a golfer, drinker, poker player, or bird-watcher, in
whether the wife is addicted to spring hats, Paris fashions, swimming pools, or
foreign sports cars, in whether the household income has been increasing or
decreasing, in whether the parents are themselves the children of thrifty, cautious
folks or carefree spendthrifts, and so forth. This list might be extended ad
infinitum. Many factors may not even be quantifiable, and even if they are, it is
not usually possible to obtain data on all of them. Even if it were, the number of
variables would almost certainly exceed the feasible number of observations, so
that no statistical means exist for estimating their influence. Moreover, many
variables may have very slight effects so that, even with substantial quantities of
data, the statistical estimation of their influence will be difficult and uncertain. We
thus let the ner effect of all these possible influences be represented by a single
stochastic variable uw.
A second reason for the addition of the stochastic term is that there may be a
basic and unpredictable element of randomness in human responses. For pur-
poses of practical statistics the distinction between these two reasons for variabil-
ity does not matter since, for reasons of both theory and data, we hardly ever
claim to have included all distinguishable and relevant factors in any relationship,
so that the insertion of a stochastic term is required on the first count, and the
second merely adds to its variance. Finally, we note that if there were measure-
ment errors in Y so that the recorded values did not accurately reflect the values
given by the theory, this would also be a component of the stochastic term and
add to its variance.

+ See App. A-4, Random Variables and Probability Distributions.


THE TWO-VARIABLE LINEAR MODEL 15

The variable u is often referred to as the disturbance term in the equation, or


as the equation error. We cannot predict the specific value of u that will emerge in
any single observation, but we can make propositions about the main features of
its probability distribution. First of all, it is clear that the u’s may take on positive
or negative values, since the net effect of the many omitted and unmeasurable
variables may push Y up or down from the value it would otherwise have had.
However, there is usually no reason to expect a bias one way or the other, so the
first assumption about u is that its average or expected value will be zero, that is,
E(u) =0
Second, since u is the algebraic sum of many different positive and negative
effects, we expect numerically small values of u to be much more frequent than
very large values, so that the distribution will be unimodal around some fairly
small value of u. If we add the assumption of symmetry, then the modal value will
coincide with the expected value of zero. Third, we will often assume a specific
form for the probability distribution, and an appeal to the central limit theorem
suggests assuming a normal probability distribution for u.} Finally, we postulate
that the various values of u will be distributed independently of one another. In
terms of the budget data, that amounts to saying that if one household displays a
positive disturbance, this does not make a positive (or a negative) disturbance
more likely for neighboring or any other households. Each disturbance is con-
ceived as drawn independently from some normal distribution

N(0,02)
The assumption that the values of u are drawn independently from a normal
distribution with zero mean and variance o? is written compactly as
u ~ NID(0, o2)
where the symbol ~ means “is distributed,” and NID stands for “normally and
independently distributed.”
The specified model is illustrated graphically in Fig. 2-1. For a household
with income X, the average or expected expenditure is given by a + BX;. The
actual expenditure will be a + BX; + u,, where u,; is a random drawing from
N(0, o2). The complete mathematical specification of the model is
Y,=a+ BX, + u; i= 1,2,...,n (2-82)
E(u;) = 6 for alli (2-8b)
\WN ~ uj= ie
E(u,u,)
0 ij, alli, j
parecalbi ei :
(2-8c)

p(u;) = N(0, 02) for alli (2-84)


Assumptions (2-8b), (2-8c), and (2-8d) are a more extensive way of stating
u ~ NID(0, 02)
+ See App. A-5, Normal Probability Distribution.
16 ECONOMETRIC METHODS

Pu)

Figure 2-1

The reason for splitting them is that some of our subsequent derivations only
require assumptions (2-85) and (2-8c) and not the assumption of normality. The
first part of assumption (2-8c) states that all possible covariances of the u’s are
zero, and the second part states that the variances of the u distributions at each
point in Fig. 2-1 are the same.t
The three unknown parameters of the model are a, 8, and o2. We now turn
to methods by which these parameters may be estimated.

2-2 LEAST-SQUARES ESTIMATORS

An estimator is defined_as a formula or method of estimating some unknown .


parameter, and an estimate as the numerical value resulting from the application
of the formula to a specific set of sample data. We start with the n sample
observations and plot them on a scatter diagram as in Fig. 2-2.
We typically plot scatter diagrams in the positive quadrant, but this is purely
a matter of convenience since many economic variables, such as inventory
investment, the balance of trade, the real rate of interest, and so forth, can take
on both positive and negative values. Any straight line drawn through this scatter
of points may be regarded as an estimate of the hypothesized relationship
Y=a+ BX + u. A straight line is indicated by
Y =a bx (2-9)
where Y indicates the height of the line at any given value of X. Once the
numerical values a and b have been set, the line is determined, and one such
line
is shown in Fig. 2-2. If the line has been drawn through the scatter, some
data

} If two variables are independently distributed, their covariance will be zero,


but the reverse does
not necessarily hold, except for normally distributed variables. See App. A-5,
Normal Probability
Distribution.
THE TWO-VARIABLE LINEAR MODEL 17

Figure 2-2

points will lie above the line and some below. We define the residuals from the
line by
ESV
eeYn a DX OT Sl en (2-10)
A

L L

so that any line generates a set of n sample residuals.


It would now seem sensible to choose the line, that is, to choose the values of
a and b, to make the residuals “small.” A possible criterion might be
n ,
ee: VY
Select a, b to make )) e, = 0 Se ye ete eeeseo
i=l
Using Eq. (2-10), this criterion gives}
wep (Y; — a —-bx, ) =0
which, on dividing through by n, givest

Y=a+bx (2-11)
Equation (2-11) merely gives the condition that a and b should be chosen to make
the line go through the point ofmeans (X, Y). Thus we could pass a line with any
slope whatsoever through (X, Y), and it would make the algebraic sum of the
residuals zero. The criterion is thus inadequate to determine a specific line.
The least-squares criterion is stronger. If each residual is squared, negative
signs disappear, and the sum of squared residuals is a nonnegative quantity. The
least-squares principle is

Select a, b to minimize Le?

+ Where there is no ambiguity about the range of summation, we will use Le;, or sometimes just
Le, instead of the more cumbersome L7_ e;, and similarly for other expressions.
+See App. A-3, Operations with Summation Signs.
18 ECONOMETRIC METHODS

From Eq. (2-10),


Le? = L(Y, — a — bx,) (2-12)
Thus
Le* = f(a, db)
since the sample data are given, so that passing different lines through the scatter
(that is, choosing different values for a and b) produces variation in the residual
sum of squares. The necessary conditions for a stationary value are

a(ze?)
da
_ abe?)
0b
_ :

Applying these conditions to Eq. (2-12) gives}

LY = na + bX
(2-13)
AY ay Ap
ae
These are termed the normal equations for the straight line, for reasons that will
become clear when we discuss the geometry of least squares in Chap. 4.
To estimate the line implied by Eq. (2-13) we first compute five quantities
from the sample data, namely,

n, DEX, Lye XY. and eke


Substitution of the resultant numbers in Eq. (2-13) gives two simultaneous
equations, which can then be solved for the two unknowns a and b. This gives the

sueda =I(Y
—aaa yma 2a
gives LY= na + bYX, and
2

ales es = NY Sn) aero


gives LXY = aLX + bY xX’. In obtaining the derivatives we leave the summation sign where it is,
differentiate the typical term (Y, — a — bX,)? with respect to a and b in turn, and simply observe the
rule that any constant can be moved in front of the summation sign but anything which varies from
one sample point to another, such as X; and Y,, must be kept to the right of the summation sign.
Finally, we have dropped the subscripts to leave the equations uncluttered, since there is no ambiguity
about the range of summation. Strictly speaking, one should distinguish between the a and b which
appear in the expression for the residual sum of squares

Ye? = L(Y—a-— bx)


and the solution values for a and b obtained by solving the equations

Ozes egal een ae é


Jang a Ap ee
but again no ambiguity is involved, and we have kept the expressions as simple as possible.
THE TWO-VARIABLE LINEAR MODEL 19

Table 2-1

Y XY Xe Y e=Y-Y

2) 4 8 4 4.50 —0.50
3 7 21 9 6.25 0.75
1 3 3 1 DUIS 0.25
5 9 45 25 9.75 = 075
9 147 153 81 16.75 0.25

Sums 20 40 230 120 40.00 0

least-squares regression of Y on X, namely,


Y=a+bXx
or Y=Yt+e=a+bX+e
' The following example illustrates the application of these techniques to the X, Y
data in Table 2-1.

Example 2-1 The normal equations are


40 = 5a + 206
230 = 20a + 120b
with solution
a=1, b = 1.75
The regression of Y on X is
Y=141.75X
Substituting each sample value of X in the regression equation gives the ¥
and e values shown in the last two columns of Table 2-1. Note that the i
values sum to the same total as the sample Y values, and the residuals, of
course, sum to zero.

The linear regression has a number of important properties.

x 1. The regression line passes through the point of means X, Y (i.e., the sum of the
residuals is zero).

This follows directly from the first equation in Eqs. (2-13), which, on division
by n, gives
Y=a+bX
and it is also shown in the footnote on page 18.

+ 2. The residuals have zero covariance with the sample X values and also with the
predicted Y values.
20 ECONOMETRIC METHODS

This also follows from the footnote on page 18 where d(Le”)/db = 0 gives
Xe = 0. The sample covariance between X and e, by definition, is
] o :
cov( X,e) = Ax = X \(e-@)

=23(X- ¥)e since é = 0

= ee _ ee
n n

= “[Link] since Le = 0

Since Yis a linear function of X it follows directly that cov(Y, e)=0.

3. The regression coefficients may be computed sequentially from

il
Daya a
ey (2-2140
14 )
and
a= Y-bxX (2-14b)
where x and y denote derivations from sample means,
x=X-X, y=Y-Y

Equation (2-145) is merely a rearrangement of the first equation in Eqs.


(2-13), and Eq. (2-14a@) follows from substituting Eq. (2-14) into the second
equation of Egs. (2-13) to get

UXY = (Y —bX )EX + bY.X?


giving
1 1
p|zx? - = (DK): = LXY - 7 (2X (LY )
or Dee = xy
Alternatively, since the least-squares line passes through the point of means,
we may take (X, Y) as a new origin, as in Fig. 2-3. Consider a point P with
coordinates (X;, Y,). The first coordinate can also be expressed as x;, its distance
from X. The second coordinate can similarly be expressed as y, and split into two
components, namely,
Y= I, + e;
where j=Y— 7
which is also a proper deviation since Y and Y have the same mean value. The
regression equation can now be written
y= bx
THE TWO-VARIABLE LINEAR MODEL 21

O i > X
>s| >

Figure 2-3

and the sum of squared residuals expressed as

Ze? = E(y — yy’


= L(y — bx)
= Ly? — 2bYxy + b?Lx?
which is only a function of b. Setting the derivative equal to zero gives

[b= bxy |

The sum of squared residuals is seen to be a quadratic in b. The coefficient of b? is


x, which is necessarily positive (unless all X values were identical). The
quadratic is thus U-shaped, and the stationary point must give the minimum sum
of squares.

4. Decomposition of sum of squares

The total variation in Y may be expressed as the sum of just two components,
the variation “explained”’ by the linear regression and the variation “unexplained”
by the regression. From property 3 we have
Ve ae UX tec,
Squaring and summing over all n observations gives
Ly? = Ly? + Le? + Wye = b?Lx? + Le? + 2rYxe
or Ly? = Lp? + Le? = b?7Lx? + Le? (2-15)
22 ECONOMETRIC METHODS

since y and x each have zero covariance with e. The crucial quantities in Eq.
(2-15) are

Ly?=total sum of squares in the dependent variable, measured about its mean
(TSS)
Le*=residual or unexplained sum of squares (RSS)
~y*=explained sum of squares (ESS)

It follows from Eq. (2-15) that the explained sum of squares may be
expressed in several alternative ways,
2
ESS = Dj? = B°Ex? = bExy = ‘ey
xX

using Eq. (2-14a).

Example 2-2 The data of Table 2-1 may be expressed in deviation form, as in
Table 2-2. Thus
neylO
oa = 7 = 1-75

and a= Y —bX
= 8 — 1.75(4)
=1
as before. The explained sum of squares may be calculated as
ESS = bY xy = 1.75(70) = 122.5
and the residual sum of squares may be obtained by subtraction as
RSS = TSS — ESS = 124 — 122.5 = 1.5
The proportion of the Y variation explained by the linear regression is
ESS@ 21275
Tss 124 0-788

The u values underlying the sample data are unknown and unobservable, for
we could only measure them if we actually knew the true values « and 8. Thus the
variance of the disturbance distribution o7 cannot be estimated from a sample of

Table 2-2

5 y xy ee y?
peer SS SE Se SO FBS eee ee ee
—2 —4 8 4 16
—] —] 1 1 1
ao —5 15 9 DS
1 1 1 1 l
5 9 45 25 81

Sums 0 0 70 40 124
THE TWO-VARIABLE LINEAR MODEL 23

u values. We do, however, have the regression residuals e,,e,,...,¢,, and it is


plausible to base an estimate of the disturbance variance on them. Two alterna-
tive estimators are
Ye? Sex
—— or
n n—-2

Both are in use, but for reasons to be explained later in this chapter we typically
use

so= (2-16)

2-3 THE CORRELATION COEFFICIENT

The regression estimated from Eq. (2-13) fixes a line which passes through the
sample scatter of points in the X, Y space. The correlation coefficient indicates the
“closeness” of the scatter about the fitted regression line. A visual inspection of
the scatter cannot indicate the degree of closeness, since changes in the units of
measurement for X and Y can stretch or contract scatters to give very different
impressions of the relationship.
The correlation coefficient is defined as

Ee
nS,S,
(2-17)

where x and y denote deviations from sample means and s,, 5, are the sample
standard deviations,

2x? _
oes n y n

This is known as the Pearsonian (after the distinguished statistician Karl Pearson),
or product moment, coefficient of correlation. Its rationale may be explained as
follows. Referring to Fig. 2-3, the perpendiculars erected at X and Y divide the
diagram into four quadrants. We pay particular attention to the sign of the
product x;y, in each quadrant.
NE quadrant xy positive
NW quadrant xy negative
SW quadrant xy positive
SE quadrant xy negative

Thus if we have a positive relationship, with sample points lying mainly in the NE
and SW quadrants, Lxy tends to be a positive number. Conversely, a negative
24 ECONOMETRIC METHODS

relationship will generate points mainly in the NW and SE quadrants, with Lxy
tending to be a negative number. If there is little, if any, relationship between the
two variables, sample points will be scattered in all four quadrants, and Lxy will
tend to zero. Lxy, however, has two defects as a measure of association between X
and Y. The first is that its numerical value may be increased by simply adding
further observations. This is corrected by dividing by the sample size to give the
sample covariance

cov( X,Y) = ze
n
Second, the covariance depends on the units in which X and Y are measured.
Shifting from dollars to cents for each variable would increase the covariance by a
factor of ten thousand. The covariance is standardized by dividing each deviation
by the sample standard deviation of the variable in question. Defining

x
variance of X = var( X) = s? = ——
n
and so on, and by some algebraic manipulation, we have a variety of ways of
looking at and computing the correlation coefficient:

oN cov( X, Y) (2-18)
/var( X) yvar(Y)

s =o (2-185)

any (2-18¢)
se
nuXY —(ZX)(XY)
= 5 (2-18d)
\n=X? — (ZX) pay? — (cYy
Looking at Eq. (2-18c) and rearranging gives

ks(==) ae2 f
x y7

= b=Sy
from Eq. (2-14a)
Sy

or b=r—
i Pe
a (2-19)
-
x

which shows the relationship between the regression slope and the correlation
THE TWO-VARIABLE LINEAR MODEL 25

coefficient. Squaring Eq. (2-18c) gives

ae ee (Zxy)°
(Zx°)(Ly’)
-= bExy
Fy?

_ ESS
TSS
= ]— RSS
TSS
Bee 2
al (2-20)
Ly*/n
Thus r? measures the proportion of the total sum of squares explained by the
regression. The last two expressions in Eq. (2-20) show that the Jimits of r are +1.
The residual sum of squares is nonnegative. It is only equal to zero if each and
every residual e, is zero, that is, if all the scatter points lie exactly on a straight
line. A value of unity for r? thus corresponds to all points lying on the regression
line. The sign of r depends upon the sign of the regression slope, that is, on the
sign of the covariance term. The relationships in Eq. (2-20) also indicate why the
correlation coefficient may be taken as a measure of the degree to which the
scatter points lie close to the regression line. Note, finally, that r is a measure of
the Jinear relation between X and Y; it is an inappropriate and misleading statistic
if the relationship is nonlinear. Suppose, for example, that X indicates a firm’s
rate of output and Y the average variable cost per unit of output. Traditional
theory postulates a U-shaped curve, which would generate points in all four
quadrants of Fig. 2-3, with an r? tending toward zero.

2-4 PROPERTIES OF THE LEAST-SQUARES ESTIMATORS

Least squares is just one possible method of estimation. Other estimators may
easily be defined. For instance, we might order the sample data by increasing size
of X and pass a line through the first and last points. Or one might average the
lowest two points and the highest two points and pass a line through these
averages. Applying the first principle to the data of Table 2-1, we have

xX Mu

Lowest point 1 3
Highest point g 17

giving an estimated slope of 14/8 = 1.75, which happens to coincide with the
26 ECONOMETRIC METHODS

least-squares slope. To make the line pass through the lowest point we have
3=a+1.75
= a= 1.25
This also ensures that the line passes through the upper point. The second
principle gives

X vg

Average of two lowest points 1.5 3.5


Average of two highest points 7.0 13.0

giving an estimated slope of 9.5/5.5 = 1.727 and an intercept of

Ga ote 7(1.727) = 0.911

Recapitulating the least-squares regression, we now have three possible equations,


namely,
Y = lol 5
Y = 1.25 + 1.75X
Y= 0,911 + 1 27x
In this example the scatter is almost perfectly linear, and so there is little
difference between the equations yielded by the three methods. The more disper-
sed the scatter, the greater are likely to be the differences between the results.
The crucial question now is how to choose between equations or, equiva-
lently, how to choose between estimating principles. The answer given in classical
Statistics is to choose on the basis of certain important properties of the various
estimators. These properties refer to the behavior of the estimators in repeated
sampling. At first sight this may seem a strange criterion. We usually have just one
set of sample data, and we wish to do the best we can with that. Bayesian
inference techniques, which are briefly discussed in Chap. 12, focus directly on
that concern, but in this chapter we will outline the classical approach.
To fix ideas, let us return to the data of Table 2-1. We assume these data to
have been generated by the model Y = a + BX + u, and we now perform the
conceptual experiment of imagining that repeated samples of five observations are
drawn, but with the X values fixed from sample to sample. This is not to imply that
economic variables are subject to experimental control and can actually be held
constant as further samples are drawn. It is, rather, an assumption that substan-
tially simplifies the derivation of the properties of the estimators, and it can be
relaxed at a later stage. With the X’s fixed, the only source of variation from
sample to sample is in the u’s, which in turn is reflected in the Y’s. Suppose that,
say, 10,000 samples were drawn, each consisting of five pairs of X, Y values. The
application of the least-squares principle would thus yield 10,000 pairs of a, b
values. These could be arranged in a bivariate frequency distribution. As the
number of samples increases indefinitely, this distribution would tend to some
THE TWO-VARIABLE LINEAR MODEL 27

smooth continuous function

f(a, b)
This is defined as the joint sampling distribution of a and b.
As with any bivariate distribution, we can integrate to obtain the marginal
distributions f(a) and f(b), which are the sampling distributions of a and b,
respectively. Concentrating on f(b), we can imagine some distribution such as
that shown in Fig. 2-4. This is a picture of the various sample values of b that
would be obtained by the repeated application of the least-squares method to
successive samples of n observations. The true parameter, B, that is being
estimated is unknown, but it is reasonable to expect some sample estimates to be
above 6 and others to be below.
There are three features of a sampling distribution that are crucial to the
assessment of an estimator. These are the mean, the variance, and the mean-
squared error. The mean value

E(b) = [of(b) db
indicates the average value that would be yielded by the estimator in repeated
applications. The bias of the estimator is defined as
bias(b) = E(b) — B (2-21)
If the bias is zero, the estimator is said to be unbiased. If the bias is nonzero, the
estimator is said to be biased. The variance of the distribution

var(b) = 0? = fle — E(b)]° f(b) db (2-22)


measures the spread of the distribution about its mean value. The standard
deviation of the distribution o, is often referred to as the standard error of b. The

S(0)

Figure 2-4
28 ECONOMETRIC METHODS

smaller the variance of the sampling distribution, the greater is the precision of the
estimator, that is, the greater is the chance of a sample estimate lying within some
specified interval about the true value. If we are comparing two estimators which
are both unbiased but have different variances, one would naturally prefer the
estimator with the smaller variance. If we consider the class of all unbiased
estimators and can find one with a smaller variance than any other, it is said to be
a best unbiased estimator.
A more difficult choice problem arises in comparing two estimators if both
are biased and also have different variances. If one estimator has a larger bias but
a smaller sampling variance than the other it is intuitively plausible to consider a
tradeoff between the two characteristics. This notion is given formal expression in
the mean-squared error,

MSE(b) = E{(b — B)’) = f(b — BY f(b) ab


Using Eqs. (2-21) and (2-22),

MSE(5) = E{[[b — E(b)] + [E(b) — B]]*}


= E{[b - E(b)]°} + E{[E(b) — B}’}
+2£{[b — E(b)][E(b) — B]}
= var(b) + [bias (b)]? (2-23)
as the cross-product term vanishes.f The mean-squared error measures the spread
of the estimates around the true value 8. Equation (2-23) shows that on the
mean-squared-error criterion a biased estimator may be preferred to one with a
smaller or zero bias if its variance is sufficiently small to offset the larger bias.
With this introduction we can return to the least-squares estimator and use
assumptions (2-8a), (2-8b), and (2-8c) to derive the means and variances of their _
sampling distributions. Looking first at the least-squares slope, b may be ex-
pressed as
hee Lx;y; _ Lx,Y,
satan el = rw, (2-24)

where x;
W, =re (2-25)

It follows from Eq. (2-25) that

“w,=0, LwX,=1 and w= ee (2-26)

tfIn evaluating E({b — E(b)}[E(b) — B]) the rules are the same as for summations. Any factor
which is a constant may be moved to the left and put as a multiplier in front of the expectation sign.
The expectation of a constant is that same constant. Thus

E{[b — E(b))[E(4) — B]) = [E(b) — B] - E{b — E(6))


= [E(b)
—B]- [E()
-E(5)]
=
THE TWO-VARIABLE LINEAR MODEL 29

To establish the properties of b in repeated sampling we need to express b in


terms of the underlying stochastic variable u. Combining Eqs. (2-8a) and (2-24)
gives
b = Lw,(a + BX, + u;)
=B+Ywu, using Eq. (2-26)
Taking expectations,
E(b)=8 (2-27)
since E(u;) = 0 for all i. Thus the least-squares slope is an unbiased estimator of
the true slope. Its variance may be expressed as

var(b) = E{(b — B)}


= E{(Lw,u, .
But

(Xw,u;)” ae “wu; aye > Ww uu;


i<j
Taking expectations and using Eq. (2-8c),
var(b) = 622w? (2-28)
which, on using Eq. (2-26), gives

(0) == 25 (2-28)
2
9,
var(b) 2-29

Turning to the intercept,


1 OX
=a+BX+iu— bX
=a-—(b-B)X+a
Since E(b) = B and E(u) = 0, it then follows that
E(a)=a (2-30)
The variance of a is

var(a) = E{(a - a) }
= X°E{(b — B)) + E(w) — 2X E{(b — B)a)
LG 4 02
=e
exe act

= 6, i + ee (2-31)
Meee
In this derivation E{a?} = 62/n since w@ is the mean of a random sample of n
30 ECONOMETRIC METHODS

drawings from the wu distribution which has zero mean and variance o,;. The
. . . . . . os

cross-product term vanishes since

E{(b — B)u) = E{(Sma, (-.=4,}|

= B|5 me + Yi (we+ wun]


i+]

=0
The covariance between a and b is
cov(a, b) = E{(a — a)(b — B)}

= E{|a (2B) |(Gark)}


~XE{(b — B)’) since E{(b— B)u)=0

Satan (2-32)
Formulas (2-29), (2-31), and (2-32) all involve the unknown oa”. To make the
formulas operational, this is replaced by its estimated value s* = Le?/(n — 2).

Example 2-3 From the data in Table 2-2,


eras Wer seo. and n= 5
Thus ;
= 1.5/3 = 0.5 “Z¢
and :
0.5 1
var(b) = “40 = 80 — 0.0125

leak
var(a) e— 2 ee =
—— ——

] 16 12
-os|z+H|-2 0.3

We might use the same data to estimate the sampling variance of the
slope estimated by passing a line through the lowest and highest points. Table
2-1 shows that the smallest X value occurs at the third observation and the
largest at the fifth observation. Thus the alternative slope estimator is

ji we sas

NS iN
]
er ee)

gla +98 + us) — (a+ 6 +u)]


om
B+ g (us — us)
THE TWO-VARIABLE LINEAR MODEL 31

Hence E(b’) = B
so that the alternative estimator is also unbiased. Its variance is

var(b’).= E{(b’ — B)’)


1
EB)aq + uz - 2usus)}

2
0,

ap
Replacing a? by the same estimate as in the least-squares case, the estimated
sampling variance is
var(b’) = + = 0.0156

which is about 25 percent greater than the variance of the least-squares


estimate. If we form the ratio
var(b) — 0.0125
=
aha ROO =i). a
sbest
we have the efficiency of b’ relative to b so that, by the least-squares criterion,
b’ is 80 percent efficient.

This last example is an illustration of an important theorem, the Gauss-


Markov theorem, that the least-squares estimators have minimum variance in the
class of linear unbiased estimators. We will now prove this theorem just for the
regression slope and give a proof of the general case in Chap. 5. We already know
that b is an unbiased estimator. It is also said to be a linear estimator since its
definition in Eq. (2-24) shows it to be a linear combination of the Y values, and
hence it is also expressible as a linear combination of the stochastic u variables.
Define a general linear estimator of B as
by = Lc,Y, (2-33)
where the c; (i = 1,2,..., 2) are some set of weights. Substituting for Y, from Eq.
(2-8a) gives
Dee obec + BicxX. Leu,

+ The alternative estimator does not provide any means of estimating 07, and we have only been
able to estimate var (b’) above by using the s” based on the least-squares residuals. However, this does
not affect the calculation of the efficiency of b’ since the proper definition is
var (b)
Efficiency of b’ = are)

AE
02/32
32
AQT
32 ECONOMETRIC METHODS

with E(by) = adc; + B&c;X;


We require the weights to be such as to make b, an unbiased estimator of B. This
imposes the conditions
Ye; =0 and == Le,X, = Lex; = 1 (2-34)
Under these conditions the variance of b, is
Var( by)'=07c7 (2-35)
To compare this variance with that of the ordinary least squares (OLS)b, write
c, = w, + (¢; — w,)
Thus
Xe? = Lw? + X(c, - w,) + 22w,(c; — w,) (2-36)
But

Lw;c; = Bee
Ex
] :
cao using Eq. (2-34) (2-37)

and, as already shown in Eq. (2-26),

LW? = els
ye
Thus
Lw,(c; — w,) = 0 (2-38)
and so

025c? = o2dw;? + 022(c, — w,)°


which, using Eq. (2-28), gives

var(b,) = var(b) + s2X(c, — w,)° (2-39)


Since L(c, — w,)* > 0, unless c, = w, for all i, this establishes that the OLS
estimator has minimum variance in the class of linear unbiased estimators, and
we write

b is a best linear unbiased estimator (b.].u.e.) of B

The proof of the similar result for the intercept is left as an exercise for the
reader.
This minimum variance property of least squares is the main reason for the
widespread use of the technique. It rests on the assumption that the X’s are fixed
in repeated sampling, that the relationship has been correctly specified in Eq.
(2-8a), and that the disturbances have zero mean, constant variance, and zero
covariances. Notice that assumption (2-8d) on the normality of the u distribution
THE TWO-VARIABLE LINEAR MODEL 33

has not been used so far. If we bring this assumption into play, the sampling
distributions of a and b are fully determined. Since a and }are linear functions of
the u’s, they in turn are normally distributed variables. Once the mean and the
variance are known, a normal distribution is completely specified. Thus
‘y2
o~ Mavei|tae | (2-40)

and
oe
bn 8.28 (2-41)

It still remains to justify the expression given in Eq. (2-16) for the estimator
of the disturbance variance. The ith residual is
A

Cea
=ait BX us; — a — bx;
Sita len) (Dia8 )XG
From the expression for a just before Eq. (2-30),
a-a=iu-—(b—B)X
Thus
e, = (u,— &) — (b- B)x;
Notice that z, being a sample mean, cannot be set at zero, even though E(u;) = 0.
Squaring and summing,

Le? = L(u, — #) + (b — B) Lx? — 2(b — B)L(u,; — a) x;


= Yu? — na + (b — B) Ux? — 2(b — B)Lu,x;
We note that
E(u?) = no?
N

E(a*) = var(a@) = =&


2
E(b — B)° = var(b) = = 2
i

and E((b ag B)(Lu;x;)} a E{(Zw,u;)(Lu;x;)}


= 07)w;x;
— 0,

Thus
E(Xe?) = (n — 2)o?
So
E| Le 2 )=02
n—2 u
34 ECONOMETRIC METHODS

and the estimator proposed in Eq. (2-16) is unbiased. The distribution of this
estimator will be established for the general case in Chap. 5.

2-5 INFERENCE IN THE LEAST-SQUARES MODEL

In the previous section we established the sampling distributions of a and b under


assumptions (2-8a) to (2-8d). Results (2-40) and (2-41), however, involve the
unknown o,, and so they are not operational as they stand. To derive the
sampling distributions of a and b when o, is replaced by its estimator s*, we
merely state here two results which will be proven later in Chap. 5, namely, that
under the assumptions made so far,
2
ee) (2-42)
Gy

and
Le? is distributed independently of f(a, b)
Concentrating first of all on inferences about B and recalling that the ¢
distribution is given by the ratio of a standard normal variable to the square root
of a x” variable divided by its degrees of freedom,
biap ~ N(0,1)
——£_
G/N Bee
from Eq. (2-41) so

(B= BEX? Ee? /o,


0, (n oa 2)

giving

item t(n — 2) (2-43)


s/\Ex2

Replacing 07 by its estimator s? shifts us from a normal distribution to a ¢


distribution. As the sample size becomes very large, the ¢ distribution tends
toward the standard normal distribution, so for sample sizes in excess of 30 or so
no great harm is done in treating (b — B)/Zx2 /S as if it were a standard normal
variable.
The standard inference procedures based on Eq. (2-43) are then as follows. A
95 percent confidence interval for B is given by

DEES fase Sy ece (2-44)


where yx? is computed from the sample data, b from Eq. (2-14a), s from Eq.
+ See App. A-7 on the x”, t, and F distributions.
THE TWO-VARIABLE LINEAR MODEL 35

(2-16), and ty 9); is read off as the 25 percent point of the ¢ distribution with n — 2
degrees of freedom. In general a 100(1 — e) percent confidence interval for B is
given by
B+ t,95/(Ex? (2-45)
To test the null hypothesis that 8 has some Rpecined value Bp, that is,

Hy: B= Bo
against the alternative hypothesis that B has some value other than Bo, that is,
HH: BB
we insert By in Eq. (2-43) and then have the conditional statement. If the null
hypothesis is true,
DSBS
~ t(n
— 2)
s/\Ex2

This gives the sampling distribution of b under the null hypothesis, as shown in
Fig. 2-5. If the null hypothesis were true, 95 percent of all sample values of b
would lie within fp9), standard errors of Bp, that is, inside the symmetrical region
about 8) shown in the figure. If our sample b is found in either tail of the
distribution, then either

1. The null hypothesis is true and an unlikely event has occurred


or
2. The null hypothesis is not true

In such a case we deliberately choose the second interpretation and thus follow
this procedure. Reject H, at the 5 percent level of significance if

b— By
> lo025
s/\Zx2

Bo — to.o2s8/V Ex? Bo Bo + to,258/V Ex?

Figure 2-5 Sampling distribution of b under Ho: B = Bo.


36 ECONOMETRIC METHODS

Accept Ho at the 5 percent level of significance if

ORO a ee
s/\ix*

In general the procedure is as follows. Reject Hp at the 100e percent level of


significance if

Accept H, at the 100e percent level of significance if

cba ae
b =

s/\Xx?
The null hypothesis most frequently tested is
H,: B=0
This is referred to as testing the significance of X. If the hypothesis is true, the X
variable plays no role in the determination of Y. Nonetheless when sample values
of X and Y are drawn from such a population and the least-squares formula is
applied, we will usually find nonzero values for b. These can arise from the
fluctuations of random sampling even though the underlying £ is truly zero.
The appropriate significance test follows directly from replacing 8, by zero in the
results above. Thus reject Hp: B = 0 at the 100e percent level of significance if

b
> t, 2
Sex i
The test statistic is now simply the ratio of b to its estimated standard error. Most
computer programs for regression analysis print out the value of b and either its
estimated standard error or else the ratio of b to its estimated standard error, and
this is often referred to as the sample ¢statistic.+ If the sample f¢ statistic is
numerically greater than the preselected critical value of t, we accept the alterna-
tive hypothesis and conclude that X plays a significant role in the determination
of Y. The presentation of sample ¢statistics is thus directly useful for making
significance tests. If, however, one wishes to test a null hypothesis that B has some
value other than zero, one requires the standard error (s.e.) for substitution in the
appropriate test statistic. This can be obtained from the computer printout as
b
carpe ant
7 In the rest of the text we will normally drop the distinction between the true standard error and
the estimated standard error, as it is usually clear from the context which concept is implied.
THE TWO-VARIABLE LINEAR MODEL 37

By a similar development, tests on the intercept are based on the ¢ distribu-


tion:
oe
UD pee

V(r]
2) (2-46)
piel Ae
Ss een

Thus a 100(1 — e) percent confidence interval for « is given by

leek:
at va +
eee | (2-47)
Rey Ne
and the hypothesis
Hy: a= a
would be rejected at the 100e percent level of significance if

SS tay)

Ss aliens.
a Ba

Tests on 0; may be derived from the result stated in Eq. (2-42). Using that
result one may, for example, write
n — 2)s?
PrxBox = ( =e = x30s| = 0.95 (2-48)
u

which merely states that 95 percent of the values of a x? variable will lie between
the values that cut off 25 percent in each tail of the distribution. This is illustrated
in Fig. 2-6. The critical values are read off from the x? distribution with n — 2

P(x?)

a iz,
Xo .005 X0.975

Figure 2-6
38 ECONOMETRIC METHODS

degrees of freedom. The only unknown in Eq. (2-48) is «7, and the contents of the
probability statement may be rearranged to give a 95 percent confidence interval
for 62 as
(n — 2)s? - (n — 2)s?
2 2
X0.975 X0.025

Example 2-4 From the data of Tables 2-1 and 2-2 we have already computed
a=] var(a) = 0.3
b = 1.75 var(b) = 0.0125
Thus
$.€. (a) = v0.3 05471
s.e.(b) = ¥0.0125= 0.1118
Since n = 5, from the ¢ distribution with 3 degrees of freedom,
tooos = 3.182
Thus a 95 percent confidence interval for a is
1 + 3.182(0.5477)
that is,
—0.74 to 2.74
and a 95 percent confidence interval for B is
1.75 + 3.182(0.1118)
that is,
1.39 to 204
The intercept is not significantly different from zero since
a
s.e.(a) 0.5477 = 1.826 < 3.182
while the slope is strongly significant since
b Lea
s.e.(b) = 0.1118 = 15.653 > 3.182

Once confidence intervals have been computed, there is no need to actually


compute the significance tests, since a confidence interval which includes zero
is equivalent to accepting the hypothesis that the true value of the parameter
is zero, and an interval which does not embrace zero is equivalent to rejecting
the null hypothesis.
From the x? distribution with 3 degrees of freedom Xoos= U2ileand
Xd.075 = 9.35. We also have Le? = (n — 2)s?= 1.5. Thus a 95 percent
confidence interval for o? is
15 165
935 “96
THE TWO-VARIABLE LINEAR MODEL 39

that is,
0.16 to 6.94

2-6 ANALYSIS OF VARIANCE IN LEAST-SQUARES REGRESSION

The test for the significance of X (Hy: = 0), derived in the previous section,
may also be set out in an analysis of variance framework, and this alternative
approach will be especially helpful when we treat problems of multiple regression
in Chap. 5.
From Eq. (2-41) we have the result

eDiciB-’ ss N(O, 1)
0,/ {Ex
From the definition of the x? variable in App. A-7 we then have

(6 a B) oe x7 (1)
0,/LX*
and since
we?
Faget 2)
independently of b

ge b — B) Xx’? aD) (2-49)


Le?/(n — 2)
recalling that F is the ratio of two independent x? variables, each divided by the
number of its degrees of freedom. If B = 0,
e bea
AF (in —2) (2-50)
be2/ (2)
Referring to the decomposition of the sum of squares in Eq. (2-15), the F statistic
in Eq. (2-50) is seen to be
Tess!
(2-51)
a RSS/(n — 2)
Following this approach, the data are set out in an analysis of variance (ANOVA)
table (Table 2-3). The entries in col. (ii) and (iii) of the table are additive. The
mean squares in the final column are obtained by dividing the sum of squares in
each row by the corresponding number of degrees of freedom. An intuitive
explanation of the degrees of freedom concept is that it is equal to the number of
values that may be set arbitrarily. Thus we may set n — 1 values of y at will, but
the nth is then determined by the condition that Uy = 0. Likewise, we may set
n — 2 values of e at will, but the least-squares fit imposes two conditions on e,
40 ECONOMETRIC METHODS

Table 2-3 ANOVA for two-variable regression

Source of Degrees of
variation Sum of squares freedom Mean square
(i) (ii) (iit) (iv)
x ESS = Ly? = b*Dx7 l ESS/1
= bY xy
Residual RSS = Le? n-2 RSS/(n — 2)
Total TSS = Ly? n-1

namely, Le = Lxe = 0, and finally there is only 1 degree of freedom attached to


the explained sum of squares since that depends only on a single parameter, B.
The F statistic in Eq. (2-51) is seen to be the ratio of the mean square due to
X to the residual mean square. The latter may be regarded as a measure of the
“noise” in the system, and thus an X effect is only detected if it is greater than
the inherent noise level. The significance of X is thus tested by examining whether
the sample F exceeds the appropriate critical value of F taken from the upper tail
of the F distribution. Thus the test procedure is as follows. Reject Hj»: B = 0 at
the 5 percent level of significance if
= 2 ESS!
RSS/(n — 2) > Fogs(1, n — 2)
where F),; indicates the value of F such that just 5 percent of the distribution lies
to the right of the ordinate at Fy,;. Other levels of significance may be handled in
a similar fashion.

Example 2-5 Table 2-4 shows the analysis of variance for the data of Table.
2-2.
The sample F statistic is

Sample F = = = 245.0

and Fo 45(1,3) = 10.1. Thus we reject Hy: 6B = 0.

The analysis of variance test is merely the significance test on B in another


guise, but it is useful to have introduced the procedure here, as it will be applied

Table 2-4

Source of Source of Degrees of


variation squares freedom Mean square

XxX 122 ] 122.5


Residual 1.5 3 0.5
Total 124 4
—_—.
eESSSSSSSSSSSSSSSSSSSSSSSSMSsssFFsesessFFseFee
THE TWO-VARIABLE LINEAR MODEL 41

extensively in later work. To establish the equivalence, recall that the ¢ test for the
significance of B is as follows. Reject Hj: $= 0 at the 5 percent level of
significance if
“b
> tonos(#'— 2)
s/\Lx?
From Eq. (2-50) the F test procedure is as follows. Reject Hy: 6 = 0 at the 5
percent level of significance if

bx
=a
Ye2/(n BS 2) os
0.95 ( )

The second test statistic is seen to be the square of the first. Thus

SampleF = (sampler)
It is also shown in App. A-7 that the F variable with (1, r) degrees of freedom is
the square of a ¢ variable with r degrees of freedom. Thus

Critical F = (critical 1)”


and the two tests are completely equivalent.
Finally, there is another way of looking at the F test that will also be helpful
later. It was shown in Eq. (2-20) that

ESS\= -770S9\= 1297


and

RSS = (1 — r”)TSS = (1 -— r?)¥y?


Thus the sample F statistic may be writtenf

wearin Leer
"0-2 te
Again, exploiting the relation between the ¢ and F distributions, Eq. (2-52) gives

fli es2
ry(n — 2)
(253)
a=")
and this statistic may be referred to as the ¢ distribution with n — 2 degrees of
freedom to test the significance of the relationship between Y and X. From
Example 2-2 we have
r
ISLESS 6. 9122.5
TRS TORN es
+ It is, of course, superfluous to insert unity as the divisor of the numerator in both Eqs. (2-51) and
(2-52), but it maintains a correspondence between these expressions in models where there is only one
explanatory variable to later models with several explanatory variables.
42 ECONOMETRIC METHODS

giving r = 0.9939. Substituting in Eq. (2-53) gives

2 OR
(1 — 0.9879)

From the data given in Examples 2-2 and 2-3

oe Le 15.65
~ s.e.(b) y0.0125
and the square root of the sample F statistic in Example 2-5 is
t= VF = 245.0 = 15.65
Thus all three tests are simply three versions of a single test. The test involving r
may be regarded as a test of the significance of the correlation coefficient, that
based on b as a test of the significance of the regression slope, and the analysis of
variance formulation tests the significance of the explained sum of squares, but
they are just three different ways of essentially asking the same question.

2-7 PREDICTION IN THE LEAST-SQUARES MODEL

Suppose we have a set of sample observations X,, Y, (i = 1,..., n) to which we


have applied the least-squares techniques of the previous sections. In addition we
suppose that our interest now focuses on some specific value of the independent
variable Xo, and we are required to forecast, or predict, the value Y, likely to be
associated with X). For instance, if Y denoted the consumption of gasoline and X
the price of gasoline, we might be interested in predicting the demand for gasoline
at some higher future price. The value of X, may lie within the range of sample X .
values, or, more frequently, we may be concerned with predicting Y for a value of
X outside the sample observations. In either case the prediction involves the
assumption that the relationship presumed to have generated the sample data still
holds for the new observation, whether it relates to a future time period or to a
unit that was not included in a sample cross section. Alternatively, we may have a
new observation (Xo, Y)), and the question arises whether this observation may
be presumed to have come from the same population that generated the sample
data. For example, an appeal might be made to motorists on grounds of
patriotism to reduce their consumption of gasoline, and we may use the new
observation to test whether the appeal is having any effect. Prediction theory
enables us to perform both tasks. We may make two kinds of predictions, a point
prediction or an interval prediction, in just the same way as we can give a point
estimate or a confidence interval estimate of a parameter B. But in practice, a
point estimate is of little use without some indication of its precision, so one
should always provide an estimate of the prediction error. The point prediction is
given by the regression value corresponding to Xp, that is,
A
THE TWO-VARIABLE LINEAR MODEL 43

The true value of Y in the prediction period is given by

where uy indicates the value that would be drawn from the disturbance distribu-
tion in the prediction period. The prediction error may then be defined as

6 = ig Y
=U
—(a— a) — (b- B)X, (2-55)
Taking expectations
Ee) 20
since E(u,) = 0 and a and b are unbiased estimators of a and B. Thus the
least-squares predictor, Eq. (2-54), is an unbiased predictor. The variance of the
prediction error is then found by squaring Eq. (2-55) and taking expectations:
var(e,) = E(e)
= var(u,) + var(a) + X$ var(b) + 2X,cov(a, b)
since the other two covariances vanish.} On substitution from Eqs. (2-28), (2-31),
and (2-32) this gives

var(e,) = 62

(2-56)
The variance of the prediction error is thus at its minimum value when X, = X
and increases nonlinearly as Xj departs from X. From Eq. (2-55) ey is seen to be a
linear function of normal variables and so is itself distributed normally. Thus
e
- ~ N(0, 1)
0,u
| 1
bee (Xo ao y
n x

Replacing the unknown o, by its estimate s = /Le*/(n — 2) then gives

a ein) (2-57)
: ppale. a
n yx?

Everything in Eq. (2-57) is known except Yo, and so, in the usual way, we derive a

+ By assumption wy is independent of u,, u2,..., u,, and thus it has zero covariance with (a — a)
and (b — B), since these are each linear functions of u), uz,..., Uy.
44 ECONOMETRIC METHODS

95 percent confidence interval for Yo as

(a + bX,) + tooo55 pPea1 RE


(Xo)
TEN (2-58)

Sometimes interest centers on predicting the mean value of Yo, that is,
E(Y,) =a + BX,
rather than Yp itself, since there is, of course, no way of predicting the value of a
single drawing from p(u). The prediction error is now

ey = E (Yi Y
= ai Gao ie NO eB) Xe
which gives
1 2 (4) 2)= \2
var(eé,)
( o) = 9; 2+ ——_———
x2

and so a 95 percent confidence interval for E(Y) is

iy CX)= \2

The width of the confidence interval in Eqs. (2-58) and (2-59) is seen to increase
symmetrically the further X, is from the sample mean X, as shown in Fig. 2-7.

a+ bX

ou
>| Xo
Figure 2-7
THE TWO-VARIABLE LINEAR MODEL 45

Example 2-6 Assembling relevant results for the data of Table 2-1,
A
Y=1+4+1.75X
n=5
X=4 :
Ber es
S 2= EAD =e 3 0.5

Lx?
= 40
Suppose we require a 95 percent confidence interval for Y given X = 10.
Applying Eq. (2-58) gives
een se
1 + 1.75(10) + 3.182V0.5 /( A : Wess) |
40
that is,
18 Deere 3:26
or 15.24 to 21.76
The 95 percent interval for E(Y|X = 10) is

Teen (Or 34 ys
18.5 + 3.182y0.5 5 is “Frag ae

that is,
18.5 = 2.36
or 16.14 to 20.86
To test whether a new observation (Xp, Yo) may be thought to come from the
structure generating the sample data, one merely contrasts the observation
with the confidence interval for Y). For example, the point (10, 25) gives a Y
value which lies outside the interval 15.24 to 21.76, and one would conclude
that it was unlikely to have been generated by the same structure as the
sample data.

PROBLEMS

2-1 The least-squares estimate of a in Y= a+ BX +uisa=L(l1/n — Xw,)Y,, where w, = x,/Xx?


with x, = X;— X and

var(a) =,
ries,a
stn
Lx?
i=]

Show that no other linear unbiased estimate of a can be constructed with a smaller variance.
2-2 Show that if z, are independent quantities from the same population, with variance o”, then the
sampling variance of
n

b= ye a;Z;
f=)
46 ECONOMETRIC METHODS

is o?L"_,a?. Observations Y, are related to fixed quantities X, and the quantities z, above by the
t

relations ¥, = a + BX, + z; (i = 1,..., n). If the values of X; are


AG OG BG AR BE) 2G
] 23. See Ae 45 e'G
an alternative estimate of is

a(¥%o Ys =)
Deduce the sampling variance of this estimate and compare with it with the sampling variance of the
least-squares estimate.
(Oxford University, 1958)
2-3 From a sample of 200 pairs of observations the following quantities were calculated:

LX = 11:34;. . LY = 20.72, “EX? = SY 84.960 eas


Estimate the regressions Y = a + BX and X¥=y+ dy.
(R.S.S. Certificate, 1956)
2-4 Show that if r is the correlation coefficient between n pairs of values (X;, Y;), then the correlation
coefficient between the n pairs (aX; +\b, cY,; + d), where a, b, c, and d are constants, is also r.
(R.S.S. Certificate, 1956)
2-5 The percentage of fat X and the percentage of nonfat solids Y are measured on milk samples of a
number of dairy cows in two herds. A summary of the data is set out below. Calculate the linear
regression of Y on X for each herd, and test whether the two lines differ in slope.
Herd A, number of cows = 16:

EX = 51.13, “LY =mo1l725, x2 = 127, Sy Asse) yell 64


Herd B, number of cows = 10:

DX i= 37.20, 9LY 78-15," Lx = 03) ya eye 10


F : P A(R.S.S. Certificate, 1956)
[Note: If B, is N(B,, 07/Lx?) and B, is N(B,, 0 /Lx4), where B, and B are independent, then
B, — By is N(B, — By, 07/Lx? + of/Xx3). If 6? and o? are unknown, then a shift to the f
distribution can be made if we assume of = 07 = 0? and pool the sum of squared residuals from each
regression so that (Le? + Le})/o7 has a x? distribution with n, + n, — 4 degrees of freedom]
2-6 Data on aggregate income Y and consumption C yield the following regressions, expressed in
deviation form:

If Y = C + Z (where Zis savings), compute the correlation between Y and Z, the correlation between
C and Z, and the ratio of the standard deviations of Z and Y.
(R.S.S. Certificate, 1948)
2-7 The table below gives the means and the standard deviations of two variables ¥ and Y and the
correlation between them for each of two samples.

Number
Sample in sample x uy Se Sy ies
] 600 5 12 2 3) 0.6
2 400 W 10 3 4 0.7

Calculate the correlation between X and Y for the composite sample consisting of the two samples
THE TWO-VARIABLE LINEAR MODEL 47

taken together. Comment on the fact that this correlation is lower than either of the two original
values.
(R.S.S. Certificate, 1955)
2-8 An investigator is interested in the two following series:

1935: B36 375 938" 8395 {404s 14056435 Ae 4546

X, deaths of children
under
| year, thousands 60 G2 Ole) Sier 5) OOM NO3E Sut 2a AGerao) AS
Y, consumption of beer,
bulk barrels 23 2S 20 Ooo 0 es Os ao Sueno!

(a) Calculate the coefficient of correlation between X and Y.


(b) A linear time trend may be fitted to X (or Y) by calculating an OLS regression of X (or Y) on
time r. This requires choosing an origin and a unit of measurement for time. For example, if the origin
is set at mid-1935 and the unit of measurement is | year, then the year 1942 corresponds to ¢ = 7. If
the origin is set at end-1940 (beginning of 1941) and the unit of measurement is 6 months, then 1937
corresponds to t = —7. Show that any computed trend value X, = a + bt is unaffected by the choice
of origin and the unit of measurement.
(c) Let X be X with any time trend removed; that is, X,= Oto X.. Calculate the correlation
between X and Y, and between X and Y. Compare these values with that obtained in part (a), and
comment on the difference.
(R.S.S. Certificate, 1954)
2-9 A sample of 20 observations corresponding to the regression model
Y=a+BX+e
where ¢ is normal with zero mean and unknown variance a”, gave the following data:

YY=219,. “(Y-Y¥)=86.9, “(x-X)(vY-Y) = 1064,


EX = 1962 CK ek ) = 2154
Estimate « and and calculate estimates of variance of your estimates. Estimate the (conditional)
mean value of Y corresponding to a value of X fixed at X = 10 and find a 95 percent confidence
interval for this (conditional) mean.
(UL, 1958)
2-10 Consider the regression without an intercept Y,= BX; + u; (i = 1,..., n) for which all standard
assumptions hold. Suppose we want to predict Yo. Find a predictor ¥2 sucht that E Gn N= Eee):
(University of Michigan, 1981)
2-11 If the sample values of X in the linear model
Oop Xena

have zero mean, show that the covariance of the least-squares estimates of « and £ is zero. Hence, or
otherwise, prove that an unbiased estimator of 8 can be derived by estimating the equation

, Y, = BX,

which is constrained to pass through the origin. What is the variance of this estimator of 8?
(UL, 1973)
CHAPTER

THREE
EXTENSIONS OF THE
TWO-VARIABLE LINEAR MODEL

The obvious limitations of the two-variable linear model are that it is linear and
embraces only two variables. These restrictions limit the variety of statistical
phenomena for which it provides an adequate description. In this chapter we
describe some of the more important ways of extending the range of the model:
First of all we discuss the case of replicated observations for various XY values,
which enables us to test the adequacy of the linear representation against the
alternative hypothesis that the relationship of Y to X is nonlinear. Then we discuss
various types of nonlinearity and the ways in which they may be handled. In
some cases suitable transformations of the variables return the problem toa linear
framework, in which case the simple techniques of Chap. 2 may be applied. In
others, nonlinear relations have to be fitted directly, but the discussion of
nonlinear estimation is beyond the scope of this book. Finally, an introduction to
three-variable regression is provided, preparing the way for a general treatment of
multiple regression in Chap 5.

3-1 REPEATED OBSERVATIONS AND A TEST OF LINEARITY

Suppose now that we have sample data which could be arranged in the form
shown schematically in Table 3-1.
Here we have p distinct values of X and m observations on Y corresponding
to each X observation, giving n = mp sample observations altogether.
This
48
EXTENSIONS OF THE TWO-VARIABLE LINEAR MODEL 49

Table 3-1 Sample data with replication

x x Mean F¥,

x) Far * ats Y,
Xy YY. sie Tay 0)

X; Yi Yi2 Yim ¥

x, Yo%p2 °° Yom Y,

situation is common in experimental design work: X, for example, might indicate


the amount of fertilizer applied to standard sized plots and Y the yield per plot of
some crop; or X might indicate drug dosage and Y the response of a patient.
Since no two plots (or patients) are completely identical, they may be expected to
display varied responses to any given level of X, and so repeated observations on
Y are desirable. We have assumed in the layout of the table that we have the same
number of replications m for each value of X. This is a simplification to keep the
formulas as uncomplicated as possible. The various formulas can easily be
amended to deal with the general case where there are m, observations on Y
corresponding to X, and a total sample size of

The first step in the analysis of replicated data is to compute the row means
1 m

SI = iADa Yi; (3-1)


j=l

where Y,, denotes the jth observation in the ith row or class. The scatter diagram
might look like Fig. 3-1 for the case of m= 4 and p= 6, with the circles
indicating individual observations and the black squares the sample mean values
of Y. Clearly, we can fit a least-squares regression to this scatter of 24 observa-
tions by the methods of Chap. 2. In this case, however, that turns out to be
identical to the regression fitted to the six mean points shown on the scatter.t To
see that this is so, consider the formula for the regression slope

b => by
xs

The deviations in this formula are measured from the overall sample means,

+ This result does not hold when there are unequal numbers of observations in each class. See Eq.
(3-6).
50 ECONOMETRIC METHODS

>

Figure 3-1

which we may now define as

y- ODSiya
rae y ey
ai
or, using Eq. (3-1),
ees | Sees
——— Yy. 3-2
p oo
It is convenient to denote the X observations by X;;, with the proviso that
X, =X.= °°: = X,,= X; fori = 1,2,..., p. Thus
aol y > ls
E> iy =
TAD oatae eeore
Then
Pom fees
Ext= i=1 j=1xX)
=m¥i=1ze (X,-¥)= \2

ae — —

and Y=) ei Xeeaed eae ate)

-¥(%-
8)E(¥,-¥) 63)
remembering that anything which is constant with respect to a particular summa-
tion sign can be moved leftward in front of that summation sign. But

(Yjids a ae (estan eee meee


m

a pas
EXTENSIONS OF THE TWO-VARIABLE LINEAR MODEL 51

since the first term vanishes and the second is a constant with respect to
summation over j. Thus
P
Pye SX (KF ) (3-4)
i=1
Putting Eqs. (3-3) and (3-4) together, the regression slope is

b= pl Se) ee) (3-5)


DE Ae)
but this is also the regression slope fitted directly to the X,, Y, points. The reason,
of course, is that the m factor in Eqs. (3-3) and (3-4) cancels out and the class
means are, in effect, each given a weight of unity in determining the regression
slope. In the more general case of unequal numbers of observations, the regres-
sion slope would be given by

b= eee = a ae ee Xe (3-6)
Sr xx) Ro OKp XO)”
which can be regarded as a weighted regression applied to the class means, the
weights being equal to the number of observations in each class.
In scatter diagrams such as that in Fig. 3-1, the class means will usually not
lie exactly on the regression line, but will be spread around it in the same way as,
but to a lesser degree than, the sample points are spread around the regression
line. We are thus led to pose two related questions.

1. Is there any relationship between Y and X?


2. If so, is it a linear or a nonlinear relationship?

The problem may be modeled as follows. Set up the hypothesis


Yj,
I
= bi te, for v= 1,2. co) psy = 1, 2325 5m (3-7)
where the ¢,,’s are assumed to be independent normal variables with zero mean
and variance o*. Thus
ECS |X.) =H ~for tal ea D (3-8)
The hypothesis embodied in Eq. (3-7) states that the m observations on Y in the
ith class (corresponding to X = X;) are random drawings from a normal popula-
tion with mean p, and variance o”, and that a similar statement holds for the
observations on Y in each of the p classes, the only systematic variation between
classes being in the underlying mean values
by, Pas-++5 Ly

Note carefully the distinction between p, and Y,, the former being the true but
unknown mean of the distribution from which the Y, observations are drawn and
the latter being the actual mean of those sample observations. .
The variation hypothesized for the m’s allows a very flexible and general
relationship between Y and X. In this context the null hypothesis of no relation-
52 ECONOMETRIC METHODS

ship between Y and X could be set up as

Die i hoe (3-9)


A test of H, can be based on a decomposition of the sum of squares in Y. We
have
(Ved es(peer
I l oe) (3-10)
Squaring and summing over all sample observations givest

SS
ivy
(ii Talia
i,j
ys i (3-11)
The decomposition of the sum of squares in Eq. (3-11) is similar to that given for
the linear regression model in Chap. 2. The left-hand side again represents the
sum of the squared deviations of Y, measured from the overall sample mean. The
first term on the right-hand side is the sum of the squared deviations of the Y’s,
measured now about the relevant class means, and the second sum of squares is
that due to the variation of the class means about the overall mean. We can thus
describe Eq. (3-11) in the form
TSS = RSS oe ESS
residual or error sum of squares “explained” sum of squares due
{Total Se ot Saree y} “unexplained” by the class means to the class means

In conventional analysis of variance treatments the right-hand-side terms are


described respectively as the within class and the between class sums of squares.
The assumption of normality for the Y’s now enables us to derive the
following test of Hj, based on RSS and ESS. From the results on the x?
distribution in App. A-7,
1 2
7 Lh, - %) ~ X?(m-— 1) for each i = 1,2,..., p
J
Further, since the sum of independent x? variables is also distributed as x7,
1 ae
= L(t,— %) =x°(pm=2p) (3-12)
ij
Since the mean of a x” distribution is equal to its degrees of freedom, we have

pean Y,) ee
Oo

+ We now write the double summation L?_,L7_, simply as ©, j- Equation (3-11) is derived by
noting that the cross-product term vanishes, that is,

U(%- %)%- %)-L(K- P)L(v,- F) =0


since L(Y; — Y,) = 0 for each i = [Link]
EXTENSIONS OF THE TWO-VARIABLE LINEAR MODEL 53

or

2
eee) ns
=O, (3-13)
Under the null hypothesis the Y, are independent normal variables with mean be
and variance o*/m. Thus

DEBE eh
<I
oO
(3-14)
and so

pte IT
a =o? (3-15)

We state, without proof, that under the null hypothesis the sums of squares in
Eggs. (3-12) and (3-14) are independently distributed. Thus since the F distribution
is given by the ratio of two independent x7 quantities, each divided by its degrees
of freedom, under H)

F= _mE(¥-Y¥)Ap-1)
_= ESs/(p~1) | Fp sal poet)
Bey, = ¥,)'/p(m i 1) RSS/p(m a 1)

(3-16)
The rationale of this test is easily seen from Eqs. (3-13) and (3-15). Under the null
hypothesis the numerator and the denominator of F are independent estimates of
o*, and thus we may expect F to vary randomly about unity. If, however, the null
hypothesis is not true, the between class sum of squares in the numerator of F will
reflect more than just the random variation of the class means about a common
mean, and F will rise in value. The null hypothesis is then rejected if the
computed value of F exceeds a preselected critical F value from the upper tail of
the distribution.
The sums of squares in this F statistic may also be used to define the
correlation ratio n as follows:

n? = = =
Eee me
Ye
my; =1(Y, eee] eal
ot yom
ty uy
|: (3-17)

Lp Ay Y) aC J Y)
This is analogous to the definition of r? in Chap. 2, with the exception |that the
explained sum of squares is based on the variation of the sample means Y,, rather
than the variation of the regression values Y,. The F statistic of Eq. (3-16) may
then be stated equivalently
fee pew t)
(1 — 9°)/p(m — 1)
which has the same structure as the expression involving r* in Eq. (2-52).
54 ECONOMETRIC METHODS

Table 3-2 One-way analysis of variance (ANOVA)

Source of Degrees of
variation Sum of squares freedom Mean square

P _— —

Between classes m yy GY Y)* = ESS Dial ESS/ (pial)


i=l
Within classes > (%, — FH)? = RSS p(m-— 1) RSS/p(m — 1)
J
— (oe ee ——————
Total - (%,-
¥) =Tss —1
mp
inj

The test defined in Eq. (3-16) is the standard one-way (or one factor) analysis
of variance. It is based solely on the Y observations and is a test of the
homogeneity of the p class means. The only role for the X variable has been to
classify the Y’s into classes associated with a common X value. Thus the X
variables could have been qualitative variables, such as socioeconomic status or
educational level. The data for the test are normally set out in an ANOVA table,
such as Table 3-2.
The second and third columns of the table are additive, as usual. Thus, once
any two of the sums of squares have been calculated, the third follows by using
TSS = ESS + RSS.
We can now carry this analysis one step further and derive a test for the
linearity of the relationship between Y and X. Figure 3-2 shows just one point
(X;, ¥;;) from the X,Y scatter, and Y = a + bX indicates the linear regression
fitted to the data.

Figure 3-2
EXTENSIONS OF THE TWO-VARIABLE LINEAR MODEL 55

The deviation of Y,, from the overall mean Y may be decomposed as


follows};

ia a) Opal ay. f(y oe) (3-18)


This decomposition embraces both the linear analysis of Chap. 2 and the analysis
of class means in this section, as shown by the following groupings.
Linear analysis
= lita= Gia is lee
(%j- Y)=(%,-H+(%-H+(%- 7
Class mean analysis
Squaring Eq. (3-18) and summing over all the sample observations gives

Bile
inj
a= iJ(8 yh (Y= 7) + my es
(3-19)
In words, Eq. (3-19) states
sum of squares
Total sum sum of squares due to variation sum of squares
of squares } = ( about class + ( of class means } +( due to linear
in Y means about linear regression
regression

We notice from Eq. (3-19) that r* is the ratio of the third term on the right-hand
side tod, ,(%; — Y)’*, and 7 is the ratio of the sum of the second and third terms
to the same total sum of squares in Y. Thus

yw>r
The equality would only hold if all class means fell on the linear regression line.
The middle term in Eq. (3-19) is thus proportional to the excess of n” over r* and
is the basis of the linearity test set out in Table 3-3.
The sums of squares may be calculated in various ways. If r* and 7* have
already been calculated, the required quantities follow directly. If not, S, can be
calculated first; S, is then the explained sum of squares due to the regression, to
be calculated by the methods of Chap. 2; S, can be calculated directly; and S,
then follows from the fact that

If the class means deviate significantly from the linear regression, S, will tend to
be large in relation to S,, when appropriate allowance has been made for degrees

+ In Fig. 3-2 we have shown, for simplicity, a case where Y;; exceeds Y,, which in turn exceeds Y,,
which again exceeds Y, so that all the deviations on the right-hand side of Eq. (3-18) are positive. In
general, of course, these deviations vary in sign.
+ All three cross-product-terms vanish, Those involving (Y;; — Y,) vanish since Li — ¥,) is zero
for all i. The term involving Y; (Y, - ¥,)( a= Y) is simply properdonal to the covarianceabetaeen the
regression values Y, and the residuals Y,— Y, about the regression line, and thus also vanishes.
56 ECONOMETRIC METHODS

Table 3-3 Test of linearity

Source of Degrees of Mean


variation Sum of squares freedom square

Regression m>(Y, = Y)o = rs, =$, 1


i
Class means
about 7 2 2
a 2 San = 82
regression mee, = Ta = nay, ap = So/(p —ae 2)

Within classes Ee, Y,)? =(1-77)S, =S; p(m—1) S3/p(m-1)

Total rior) =S, mp—T


as

of freedom. The test for linearity is thus based on

Sb S,/(p eo2) i
way og Be
and the linear hypothesis is rejected in favor of a nonlinear alternative if the
computed F exceeds a preselected critical value from F( p — 2, p(m — 1)).
Example 3-1 We have eight classes (p = 8), each class being defined by a
particular value of X, and five observations (m = 5) in each class, giving 40
observed points on an X, Y scatter. The class means are shown in the final
column of Table 3-4 and display a negative relationship with X, but the class
means do not lie exactly on astraight line.
We now wish to build up the numerical equivalent of Table 3-3 for these
data. The first step is to fit the linear regression of Y, on X,. The eight pairs of
values yield the following numbers:

> X, = 250 >)X7 ="9,500

DEY= 16569 PY, = 0166 oe ey — 8s

Table 3-4 Replicated data

47.2
37.0
25.4
19.6
14.8
10.4
8.8
st DUN
eee
St
Nee
DT ON
PWN 2.4
EXTENSIONS OF THE TWO-VARIABLE LINEAR MODEL 57

from which
ss = 9,500
(x, - X¥) - 250)°
aE = 1,687.50
oe- Y) = 3,585-
YG - ¥)(¥, 250(165.6
SS — 1,590.00
a
X(¥,- ¥)= 5,036.56 — os = 1,608.64
SO

1590
Bea cea ete ab Aee

and a= Aes +0.9422( | = 50.144


giving the regression equation

Y = 50.144 — 0.9422.X
Turning now to Table 3-3, we compute first of all the overall sum of squares:

S4 = L(Y; - ras

=D (Et)
= 25,620 — ay (828) = 8.480.4
The sum of squares within classes, S,, is obtained by calculating the sum of
squared deviations within each class about the class mean and aggregating
the result over all classes. For the ith class we have

D(%,- FY EM
J
-(EY a
For the first row in Table 3-4, this gives

(48) + (51)? + --- + (52)’ — 3(48 + 51 + --- + 52)? = 86.8


Carrying out the same calculations for the remaining rows and aggregating
gives
Si 96.8 4 82.0 + 31.2.4 85.2 + 102:8:+ 1.2 + 20.8 + 27.2 = 437.2
The square of the correlation ratio may now be computed from Eq. (3-17) as
MRM T22
= 0.948446
lee 8,480.4
58 ECONOMETRIC METHODS

Table 3-5 Sums of squares from the data of Table 3-4

Sum of Degrees of Mean


Source of variation squares freedom square

Regression S; = 7,490.67 1
Class means about regression S,= 552.53 6 92.09
Within classes S3 = 437.20 32 13.66
Total S, = 8,480.40 39

The next simplest quantity to compute is


2 = \2
S| = m>, (¥; ert )

From the regression of ¥, on X, the explained sum of squares is

sea oha (NOU) eg


Satie = “1687.50” = 1,498.133

Thus
Ss, = 5(1,498.133) = 7,490.67

The remaining sum of squares, S,, can now be obtained by subtraction, and
Table 3-5 is prepared.}

} There is a subtle point concerning the interpretation of r? in this example. We have seen that in
the special case, where there is a constant number of observations per class, the regression slope may
be calculated by considering Ft: the 40 observations (X;,, Xi) Y; or the eight observations (X,, Y,).
There are, however, two distinct r?.One relates to the ee of Y, on X; and the other to the
regression of Y;; on X;,. It is intuitively clear that the former r? must exceed the latter since the class
means will lie closer to the regression line than the raw data. In Table 3-3 S, is expressed as r2S,, and
this r? is the one relating to the raw data. By definition, it is

e..(%,- Fy], -¥)]


Using Egs. (3-3) and (3-4),

¥ (%, — ¥) = 5(1,687.50)
=8,437.50
i

Lo(%;- X)(%,- ¥) = 5(- 1,590) = -7,950

and oy Y)* has already been calculated as 8,480.40. Thus

r> = 0.883292
Then

Sy = (1 — r7)S4 = (0.948446 — 0.883292)8,480.4 = 552.53


which agrees with the value obtained by subtraction in Table 3-5.
EXTENSIONS OF THE TWO-VARIABLE LINEAR MODEL 59

The test statistic for the linear regression is

EEA ar eeAo0eT
LGR Shyse Sees ae vO
The 5 percent critical value from F(1, 40) is 4.08, and the 1 percent value is
7.31, so this is a highly significant sample statistic, and the data would lead us
to reject decisively the hypothesis of a zero coefficient for X. The test statistic

OO

40

Y = 50.144 — 0.9422x

OO
800—

20

5 10 15 20 25 30 35 40 45 50
60 ECONOMETRIC METHODS

for the linearity of the relation is


_, 8/6 92.09. _
ac 53/3201 66 ola

and the | percent critical value from F(6, 30) is 3.47. This sample statistic is
also significant, and we conclude that the true relationship is probably
nonlinear. The scatter, regression line, and class means are shown in Fig. 3-3.

3-2 NONLINEAR RELATIONS

The relation
Y=at+BpxXt+u (3-21)

is Jinear in the parameters a and £ and also in the variables X and Y, and, as we
have seen, application of the least-squares principle results in two simultaneous
equations, which are linear in the estimates a and b and thus easy to solve. The
relation
logY=a+BX+u (3-22)
may be written
Z=a+BX+u (3-23)
where
Z = logY (3-24)

The scatter from Eq. (3-22) in the X, Y plane will be nonlinear. However, the
transformation defined in Eq. (3-24) will yield a linear scatter in the X, Z plane,
and so the techniques of Chap. 2 may be applied directly to Eq. (3-23). This is an
example of a transformation of a variable, changing a relationship which is
nonlinear in the original variables into one which is linear in the transformed
variables. As will be seen below, simple transformations of one or both variables —
can deal with a wide variety of nonlinear relations.
Another approach is to introduce additional terms in X on the right-hand
side of the relation. If Y denoted this average variable cost per unit of output and
X the rate of output, the U-shaped cost curve of economic theory might be
depicted as

Y=at+BX+yX*+u (3-25)

The nonlinearity between Y and X here requires the introduction of an additional


explanatory variable, X*, and we are now involved in a three-variable regression
problem, which will be taken up in Sec. 3-4. Relation (3-25), however, is still
linear in the parameters a, B, and y, and the least-squares principle can easily be
extended to deal with three (and more) variables.
EXTENSIONS OF THE TWO-VARIABLE LINEAR MODEL 61

A more difficult relation is one such as


y=at Bx’ +u (3-26)
If we attempted to apply least squares directly, we would have to choose values of
a, b, and c to minimize
n n

Le? = L (¥,- a - boxe)


i=1 i=]
The resultant normal equations are nonlinear in the unknowns and cannot be
solved analytically. If, however, we consider a somewhat different relation,
Y = BX ™u (3-27)
taking logarithms of both sides gives
log Y = log B + ylog X + log u (3-28)
There are two crucial differences between Eqs. (3-26) and (3-27): in the latter the
intercept a is assumed to be zero and the disturbance term has also been entered
multiplicatively rather than additively. The consequence of these assumptions is
the linear relation in Eq. (3-28). If b and c denote the intercept and the slope,
respectively, of the least-squares regression of log Y on log X, we obtain estimates
of B and y as follows:
Estimate of y = c
Estimate of 8 = antilog b
The quality of these estimates and the justification for the application of least
squares to Eq. (3-28) depends, however, on log u having the properties postulated
for the disturbance term in Chap. 2. The contrast between the treatment of the
disturbance term in Eq. (3-26) and (3-27) illustrates one of the fundamental
problems of the econometrician. Economic theory cannot yield precise proposi-
tions on the nature of the disturbance term, and so the econometrician proceeds
somewhat pragmatically by first of all assuming the most convenient form for the
disturbance term in any specific application and then attempting to test these
assumptions as far as possible by the methods described in subsequent chapters.f
In principle the two main approaches to the fitting of nonlinear relations are
either to seek transformations of some or all of the variables to reduce the
problem to a linear form or else to fit the nonlinear form directly. The first
approach is analyzed in more detail in the next section.

3-3 TRANSFORMATIONS OF VARIABLES

In very rare cases, economic theory may indicate the appropriate transformation
of variables. As Zarembka has pointed out, the constant elasticity of substitution

+ The late Sir Julian Huxley once defined God as a “personified symbol for man’s residual
ignorance.” In a similar vein the disturbance term might be regarded as the econometrician’s stochastic
symbol for his residual ignorance, and then, as is sometimes done with God, the inscrutable and
unknowable may be ascribed the properties most convenient for the current problem.
62 ECONOMETRIC METHODS

(CES) production function}

Y = (a,K® +a,L°)””
gives
ye’? = a, K® + a,L° (3-29)
Thus each observation on output should be raised to the power p/v, and each
observation on capital and labor inputs should be raised to the power p. This is
an example of power transformations of the variables, and although Eq. (3-29) is
linear in the transformed variables, it still poses difficult estimation problems.
Some special cases of power transformations, however, yield simple estimating
procedures, and we now turn to these.
Returning to two-variable relations, let us denote a transformation of the Y
variable by Y“”. This symbolism indicates that the transformation depends only
on a single parameter A,. Likewise, let us indicate the transformation of the X
variable by X2). A very general form of transformation has been proposed by
Box and Cox, namely,+

y"-1
aad hr
Sa
oees A, + 0
a (3-30)
In Y A, =0
and similarly,

xX — ]
cael Xr 0
XO2) = Ay Dh (3-31)
In X A, =0

At first sight these seem needlessly complicated transformations, and one might
well ask, why not use the simple power transformation Y*', X*2. An examination
of Fig. 3-4 indicates the rationale behind the Box-Cox transformation. Fig. 3-4a
shows the simple power transformation Y* for two illustrative values of Y,
namely, 10 and e = 2.84128. The transformed variables cross at the (0, 1) point,
and to the left and right of that point their ordering is reversed. The simple power
transformation is thus unsatisfactory since different values of X would not
preserve the ordering of the data. Fig. 3-4b shows the graphs of Y*/A for the
same two values of Y, and now the ordering is the same for all values of A, but a

7 Chap. 3, “Transformation of Variables in Econometrics,” in P. Zarembka (Ed.), Frontiers in


Econometrics, Academic Press, New York and London, 1974. The constant elasticity of substitution
(CES) production function was introduced by K. J. Arrow, H. B. Chenery, B. S. Minhas, and R. M.
Solow, “Capital-Labor Substitution and Economic Efficiency,” Review of Economics and Statistics, 43,
1961, pp. 225-250. The elasticity of substitution is given by 1/(1 — p) with p < 1, and vy denotes
returns to scale.
¢ G. E. P. Box and D. R. Cox, “An Analysis of Transformations,” Journal of the Royal Statistical
Society, ser. B, 1964, pp. 211-252. From now on the symbols log and In denote logarithms to base 10
and to base e (natural logarithms), respectively.
(2?) (9) (2)
64 ECONOMETRIC METHODS

discontinuity occurs at A = 0. Finally, Fig. 3-4c shows the graph of the general
transformation
y* = 1
rv
and now the ordering is the same for all values of A, and there is no discontinuity
at A = 0. If we substitute A, = 0 in Eq. (3-30), we obtain Y“” = 0/0, which is
indeterminate. However, the application of L’H6pital’s rule shows thatt
lim YOU = InY
4,70

Suppose now that the transformed variables fit the linear model, that is,
YOU = ay + BXO” + u (3-32)
This model has five basic parameters, namely, aj, B, A,;, Az, and o/. In this
section we will consider only some special cases corresponding to particular
values of A, and A).

Case 3-1: A, = 1 = A, (Linear model). Combining Eqs. (3.30), (3-31), and (3-32)
now gives
Y=a+BX+u
where
a=l+t+a,—8
This is the simple linear model of Chap. 2.

} For a definition of L’Hopital’s rule see, for example, A. C. Chiang, Fundamental Methods for
Mathematical Economics, 2d ed., McGraw-Hill, New York, 1974, p. 420. The application of the rule to’
Y® states that

By e eda) (laa)
jor! ai Loli Cae OR)
= lim
ae (Y*InY )

=InY
This development uses the result that (d/dd,)(Y*') = Y™'In Y. To derive this from first principles,
consider a general function

where a is some constant, and define

z=Iny=xlna

Then dz_dzady_\lw
—= =
Ore Gly Obs yn abe
dz
Ge
But
u 7 Ina

Th us aes aS
ye Inag=a*lna
EXTENSIONS OF THE TWO-VARIABLE LINEAR MODEL 65

ORaIBre1

Figure 3-5 The log-log model.

Case 3-2: , = 0 = A, (Log-log model). Combining Eqs. (3-30), (3-31), and (3-32)
now gives
nY=a,+PlnX+u (3-33)
Again, all the techniques of Chap. 2 may be applied, once the original data have
been transformed to logarithmic form. Notice that although Eg. (3-33) is specified
in terms of logarithms to base e, one may take logarithms to base 10 in carrying
out the empirical work. The estimate of a, will be affected by the choice of base,
but that of B will not.
Ignoring the disturbance term in Eq. (3-33), the relationship between Y and X
is
Yor Anxe (3-34)
where
In Ay = ao
From Eq. (3-34)
dY
a
AX BA, X B-1

so that, if B is positive, the slope is always positive and Y tends to infinity as X


tends to infinity. If 8 exceeds unity, the slope increases continually as X increases,
while if 0 < B < 1, the slope decreases continually, though always remaining
positive. When £ is negative, the slope is always negative. This gives the shapes
pictured in Fig. 3-5. The relationship only exists for positive values of the
variables.
The double logarithmic relationship has a very important characteristic. It is
a constant elasticity function, and that elasticity is given by 8.t Thus when Eq.

+If In Y =N, then Y =e’,


logoY = N logioe = In Ylog joe
Thus Eg. (3-33) may be written
logo¥ = (a@ologye) + BlogioX+ (ulogioe)
so that the intercept term being estimated is a log joe.
+ See App. A-2, Exponential and Logarithmic Functions.
66 ECONOMETRIC METHODS

va x
4

ex

y = ela + BX) y = e(a + BX)

e&

B>0 B<0

O Se @ >X

Figure 3-6

(3-33) is fitted to the data the regression slope is a point estimate of the elasticity.
If 8B= —1, Eq. (3-34) gives
XY = Ay (3-35)
which is a rectangular hyperbola. If Y denoted the quantity purchased and X the
price per unit of some commodity, then Eq. (3-35) would represent a demand
curve with constant elasticity of —1 and a constant total expenditure on the
commodity, whatever its price.

Case 3-3: 4, = 0,A, = 1 (Semilog model). This combination of \ values gives


nY=a+BX+u (3-36)
where
OSG aa
From Eq. (3-36),

oanB (3-37)
Thus the proportionate rate of change in Y per unit change in X is a constant and
equal to B. The function is only defined for positive values of Y. Ignoring the
disturbance in Eq. (3-36), we may rewrite the function as
Y= extBhx (3-38)

and its general shape is shown in Fig. 3-6. The intercept is given by e%, and the
slope is positive or negative, depending on the sign of B.
+ Equation (3-36) is a widely used specification in human capital models, where Y denotes earnings
and X years of schooling. The specific functional form is derived from theoretical considerations by J.
Mincer, School, Experience and Earnings, Columbia University Press, New York, 1974.
EXTENSIONS OF THE TWO-VARIABLE LINEAR MODEL 67

Table 3-6 Bituminous coal output in the United States, 1841-1910

Average annual
output
(1,000 net tons)
Decade ie Log Y Xt

1841-1850 1,837 3.2641 3


1851-1860 4,868 3.6873 ord
1861-1870 12,411 4.0937 il
1871-1880 32,617 4.5135 0
1881-1890 82,770 4.9179 l
1891-1900 148,457 5.1718 2
1901-1910 322,958 5.5092 3

A special case of Eq. (3-38) occurs when X denotes time and the function
then describes a variable Y which displays a constant proportionate rate of
growth (8 > 0) or decay (B < 0).

Example 3-2 The first step in any empirical application of Eq. (3-36) is to
check visually on whether the constant growth assumption seems warranted.
This may be done either by plotting log Y against time or, equivalently, by
plotting Y against time on commercial semilog paper. The first procedure
applied to the data of Table 3-6 gives the scatter diagram shown in Fig. 3-7,
and it is clear that the relationship is approximately linear.

log y
A

Sle

5.0 Sane
logy = 4.4510 + 0.3760¢

4.5

4.07
68 ECONOMETRIC METHODS

In fitting the constant growth curve


Y, = et +8 (3-39)
to the data of Table 3-6, we have to be careful about the treatment of the
time variable, about the base to which logarithms are taken, and about the
way in which growth rates are expressed and measured. Equation (3-39) is a
continuous time formulation. It may be expressed as
Y= We" (3-40)
where Y) = value of Y att = 0
and
led ie ;
Sie instantaneous rate of growth of Y at time f

When time is measured in discrete intervals, such as quarters or years, a


constant growth series would be expressed as

Vat iUeage (3-41)


where g = proportionate rate of growth in Y per unit of time
Taking logarithms of Eq. (3-41) to base 10 gives
log Y, = log Y, + [log(1 + g)]z (3-42)
This is the equation estimated with actual data. Thus
Intercept = estimate of log Yo
Slope = estimate of log(1 + g)
and so an estimate of g can be obtained. Taking logarithms of Eq. (3-40), also
to base 10, gives
log Y, = log Y) + (Blog e)t
Comparison with Eq. (3-42) shows that
Bloge = log(1 + g)
or
B=In(1 + g) (3-43)
which can be used to provide an estimate of B corresponding to any
estimated g. The interpretation is that 8 is the rate which with continuous
compounding would give the same result as a single increment at rate g.t We

¥ Suppose your friendly local bank manager offers interest of 10 percent per year on time deposits
compounded annually. Then $100 deposited now grows in successive years to $110, $121, $133.1, and
so on. In general, interest at 1007 percent per year on an initial investment of $100 will give a sum of
$100(1 + r)” after n years. Suppose further that the manager agrees to your suggestion that, instead
of adding interest of 10 percent once a year, 5 percent should be added twice a year. After 2 years of
graduate school your $100 would now grow to $100(1.05)* = $121.55, so that your perspicacity
has
paid off to the tune of 55 cents. You now become greedy and suggest that | percent be added 10 times
EXTENSIONS OF THE TWO-VARIABLE LINEAR MODEL 69

can select the origin and the units of measurement for time at our conve-
nience. In the final column of Table 3-6 we have chosen the midpoint of the
1871-1880 decade as the origin and measured time in units of 10 years. The
normal equations for fitting a linear regression of log Y on ¢ are

Llog
Y = na + byUt
Ltlog Y = aXt + byt?
Since the t’s have been chosen so that Lt = 0, these equations simplify to
se LYlog Y
n
ie Lrlog Y
vr?
For the data in the table, these equations give
a 31.1575/7 = 4.4511
b 10.5285/28 = 0.3760
Thus the regression estimate of Eq. (3-42) is
ae
log Y = 4.4511 + 0.3760t
The r? is 0.9945, which confirms the story of the scatter diagram that the
constant growth curve fits this series very well. To find the estimated growth
rate,
log(1 + ¢) = 0.3760
1 4¢'= 2.3768
Thus
$= 1.3768
Since ¢ was measured in units of 10 years, this gives the estimated rate of
growth per decade as 137.7 percent. The annual growth rate is found from
(1 + r)'° = 2.3768

a year, or perhaps } percent 20 times a year. You have, in fact, discovered the principle that
2 3) r\”
G+y<(14+5) <(14+ 5) <-> <(142) <---

Unfortunately it is no magic device for unlimited increases in your wealth. The sequence has a limit,
namely,

lim (1+) =e’


noo n

where

e= lim (1oF | = 2.41828


n> 0o n
For a 10 percent growth rate, r = 0.10 and e’ = 1.1052. Thus continuous compounding within a period
at a rate of 10 percent is equivalent to a growth rate of just over 10} percent over the period.
70 ECONOMETRIC METHODS

giving
r = 0.0904
or just over 9 percent per annum. From Eq. (3-43) the corresponding
continuous rate 6 is 0.0866.

Case 3-4: A, = 1, A, = —1 (Reciprocal model). These values give the relation

¥=a+p(y)+u (3-44)
The slope is dY/dX = —B/X?. Thus if B is positive, the slope is everywhere
negative, and conversely it is positive when B is negative. Since 1/X — 0 as
X — o, a denotes an asymptotic value for Y. The shape of this function is
indicated in Fig. 3-8.
Fig. 3-8a indicates a typical shape of a Phillips curve, and many Phillips
curves have been estimated by regressing the rate of wage change on the
reciprocal of the unemployment rate. The estimate of a indicates the asymptotic
floor for wage change; if it turns out negative, the Phillips curve cuts the
unemployment axis. Fig. 3-8b has often been used to represent expenditure
functions, where Y denotes expenditure on some specified commodity or service
and X denotes total expenditure (or income), the data typically coming from a
cross section of households. This particular application only makes sense when a
is positive and £is negative, for a indicates the asymptotic level of expenditure. It
follows that income or total expenditure has to reach some critical level — B/a
before anything is spent on this commodity. Thus Eq. (3-44) cannot serve as a
general model for all types of consumption since we need to allow for finite
consumption of some commodities at values of X close to zero. This particular
difficulty can be overcome by the choice of yet another pair of A values.

B>0 Sa—0

(a) (b)

Figure 3-8 The reciprocal model.


EXTENSIONS OF THE TWO-VARIABLE LINEAR MODEL 71

Case 3-5: 4, = 0, A, = — 1 (Logarithmic reciprocal model). This choice of param-


eters gives
1
iny=a- p(y) +u (3-45)

Ignoring the disturbance term, we may write Eq. (3-45) as


Yrare rhe (3-46)
Y is not defined for X = 0, but Y — 0 as X — 0, so we can define Y(0) as zero,
and we then have a function which is right-hand continuous at the origin. From
Eq. (3-46)


AK of@ (
8) aye
2 Je (3-47)

Thus the slope is positive for positive 8. The second derivative is

aAXtee
(B28 eer
New, aXe
Hence there is a point of inflection where X = 8/2. To the left of this point the
slope increases with X; to the right of it the slope diminishes. As X —> 0,
Y > e*. Substituting X = 8/2 in Eq. (3-46) gives the value of Y at the point of
inflection as 0.135e% or, in other words, about 13 percent of the asymptotic value
of Y. The general shape of this function is shown in Fig. 3-9.
If X represents time, Eq. (3-46) then pictures a growth curve which starts at
zero and approaches an asymptotic level. Rewriting Eq. (3-47) in an equivalent
form gives
Adie SBe
VodXee =¥2
that is, the rate of growth in Y per unit change in X is inversely proportional to
the square of X. Thus the rate of growth falls off sharply after low values of X, as

Figure 3-9 The logarithmic-reciprocal curve.


72 ECONOMETRIC METHODS

shown in Fig. 3-9. This can be a disadvantage of the curve, which may more than
offset the ease of fitting.
A similar curve, which also has an upper asymptote at some finite level and a
lower asymptote at zero, but has a more symmetrical shape between the two, is
the /ogistic. This curve cannot be derived from the general linear relation (3-32) by
a suitable choice of values for A, and A,. Nonetheless it has been widely used in
fitting growth trends, and so a brief account of it is given here. The logistic
equation is

Yall (3-48)
Leiber =
where a, b, and k are parameters to be determined. We have written Y as a
function of time ¢, as this is by far the most common practice, but in some
applications it is quite feasible to replace t by some independent variable X.
From Eq. (3-48) it is clear that

Yok as t— ©
and Y-0 as t> -o

so that k is the upper asymptote and zero the lower asymptote. The first
derivative of Eq. (3-48) is
Giana
—=— Gee Y(k -- Y) (3-49)
-4

Thus the rate of change of Y with respect to ¢ is proportional to the current level
Y and also to the distance still to travel to reach the saturation level k. The first
derivative is positive for all values of t. The second derivative may be written

e a dY
(3-50)
Setting this to zero gives a point of inflection at

ee ee
a
Thus when Y < k/2, the “large” value of k — Y dominates the “small” value of
Y in Eq. (3-49) and causes dY/dt to increase. As Y increases toward kKy2=the
relative balance of the two forces changes so that dY/dt reaches a maximum
value when Y= k/2 and thereafter declines steadily as Y rises toward the
saturation level k. The typical shape of the logistic curve is shown in Fig. 3-10,
the main contrast with Fig. 3-9 being that the point of inflection occurs at half the
saturation value rather than at a much lower level. The logistic curve is frequently
used as a plausible approximation to the growth of any “population,” whether
bacterial, animal, human, or economic, where growth is thought to be positively
related to the size of the existing population and negatively related to the current
distance from a saturation level.
EXTENSIONS OF THE TWO-VARIABLE LINEAR MODEL 73

Figure 3-10 The logistic curve.

Estimation of all the parameters a, b, and k cannot be achieved by the simple


methods that we have described so far. From Eq. (3-49) we can write

idys a
yar 72 (3)¥
If time is measured in constant units, the left-hand side of this equation is
approximately AY/Y, the proportionate rate of growth of Y. Thus one might fit
the linear regression
Tie
z Nyt=a-(f)¥
P a
+e, (3-51)
which yields point estimates of a and k. To obtain an estimate of b, we note that
Eq. (3-48) can be rearranged to give

hay
b= 7 ec. | (3-52)

Thus a value of 6 can be computed for each Y,, given a and k. An estimate b may
then be obtained by averaging some or all of these computed Db values, or
alternatively by substituting Y and ¢ in Eq. (3-52).
There are two main difficulties with this simple procedure for estimating the
logistic function.+ First of all we have point estimates of the parameters, but
inference procedures are difficult, especially for b and k. Second, there is evidence
that the procedure is unsatisfactory compared with a direct estimation by
nonlinear methods.
+ See F. R. Oliver, “Methods of Estimating the Logistic Growth Function,” Applied Statistics,
1964, pp. 57-66. This paper describes an iterative program for computing the (nonlinear) least-squares
estimates. A further paper, F. R. Oliver, “Notes on the Logistic Curve for Human Populations,”
Journal of the Royal Statistical Society, vol. 145, 1982, pp. 359-363, gives formulas for the asymptotic
standard errors of the least-squares estimates.
74 ECONOMETRIC METHODS

Returning to the model


YOD = ay + BX +u
we have outlined five special cases, depending on particular choices of 1, 0, or — 1
for the A parameters. In each case a simple linear relation holds between the
transformed variables, and the techniques of Chap. 2 may be applied, provided
the assumptions about the disturbance term are fulfilled. However, the question
immediately arises: Why restrict consideration to just three possible values for A?
If the A’s are free to take on any values, the approximation to linearity may well
be improved, but we then have five parameters to be estimated, namely, ay, B, Aj,
A,, and a If the A’s can be estimated and if we can test hypotheses about them,
we may be able to discriminate between functional forms. This more general
approach involves nonlinear estimation, which is beyond the scope of this book.

3-4 THREE-VARIABLE REGRESSION


As indicated in Sec. 3-2, some nonlinear relations may be represented by the use
of several explanatory variables on the right-hand side of the relation. For
example, the conventional total cost function of economic theory might be
represented as
TC =a+ 80+ yO? + 6Q? (3-53)
where TC = total cost per period
Q =output per period

and a, 8, y, and 6 are parameters. From Eq. (3-53), the average cost (AC) and the
marginal cost (MC) are obtained, respectively, as

AC =—TC = 4+(B + yO + 8Q?) (3-54)


Ot eo
J L
Average Average
fixed variable
cost cost

and
d(TC
MC = a Js B + 2yO + 38Q? (3-55)
With certain restrictions on the parameters, these functions will have the conven-
tional shapes shown in Fig. 3-11.
Relation (3-54) shows AC asalinear function of the variables

Q, aQ > aide Oe
EXTENSIONS OF THE TWO-VARIABLE LINEAR MODEL 75

Le AC,MC

O = iO ane O) >Q

Figure 3-11

More generally the additional explanatory variables need not be restricted to


transformations of a single variable but will be different variables. For example,
an expectations-augmented Phillips curve may be written as
1 iv
Wet (5 + 8( py)
Uu,

where w,= rate of wage change


u,= unemployment rate
p; =expected inflation rate

The statistical theory of relationships such as these may be more conveniently


handled by using the following general notation:

This states that the dependent variable Y is determined by a linear combination


of k —1 explanatory variables X,, X;,..., X, and a disturbance term. The
subscripts on the X’s indicate different explanatory variables. As the illustrations
above show, some of these X’s may well be transformations of other X’s. The
relation (3-56) is assumed to hold at each sample point. Thus we may write
Y= Bit Bo Xo, t Bika + Bp ta, Pe Lee 657)
or, more compactly,
k
Y=) Ib Xa hai, ean (3-58)
i=l
where X,,= 1 for all t
and the interpretation of the double subscript on X is that X;, denotes the jth
76 ECONOMETRIC METHODS

observation on the variable X,. X, is an example of a dummy variable, that is, a


variable which takes on a priori specified values (in this case unity) at the sample
points.
Equation (3-57) defines the linear model in k variables, and it will be treated
fully in Chap. 5. As an introduction we deal explicitly in this section with a
three-variable model, but the concepts developed will facilitate our understanding
of the k-variable model.
The three-variable model is specified as
Y, = B, + BX, + BX3, + u, =e en (3-59)
Replacing the unknown £’s in Eq. (3-59) by any arbitrary set of numerical
coefficients b,, b,, and b, gives
Y, = b, + b,X,, + b,X3, + e, (3-60)
The residual sum of squares in Eq. (3-60) is then
n n 2

RSS= es G7 a Ss (Y, = Dire OXa) = b; X3,)


t=] t=1

= f(b,, by, bs) (3-61)


The least-squares principle states that b,, b,, and b; should be chosen to minimize
the residual sum of squares defined in Eq. (3-61). The necessary condition is that
a(RSS) _ A(RSS) _ (RSS) _ 0
db, db, doh
This gives the normal equations
EY = ND et OeNoeta eke
UXZY = b UX, + bE XF + b,LX, X, (3-62) .
LUX3Y = bX, + B,D X,X,. + b,x?
All the summations merely involve the sample values of X,, X;, and Y, so Eq.
(3-62) defines a set of three simultaneous equations which can be solved for the
least-squares estimates b,, b,, and b,.{ These equations parallel those for the
two-variable case in Eq. (2-13), and the derivation is the same as that shown in
the footnote below Eq. (2-13). The resultant least-squares regression indicates a
plane in three-dimensional space, and is written as
A
VD dy Xt Diy) Et
1 s
ears (3-63)
or, equivalently,
Yotere | or= 2.2.n (3-64)
Dividing through the first equation in Eqs. (3-62) by the sample size n gives
Y=5b,+6,X,+ bX, (3-65)

+ Again, to keep the notation simple, we are not distinguishing between b; as a variable parameter,
as in Eq. (3-60), and 5; as the least-squares estimator, defined in Eq. (3-62).
EXTENSIONS OF THE TWO-VARIABLE LINEAR MODEL 77

Thus the regression plane passes through the point of means. Summing Eqs.
(3-63) over the sample observations and dividing by n gives

Y = b, + bX, + bX,
Comparison with Eq. (3-65) shows directly that
A
ye=ay,
and so from Eq. (3-64),

e@=0 (3-66)
Thus the mean of the regression values of Y is equal to the mean of the actual
sample values or, in other words, the sum of the least-squares residuals is
identically zero.
Result (3-66) may also be obtained directly from the condition
d(RSS) _ 0
Ob am,
for
0(RSS) : =
ab, =e we b, X,,) = —2 0 e,=0
t=] =A
The equality of the other two partial derivatives to zero gives the important result
that the least-squares residuals are uncorrelated with X,, X,, and Y. For
a(RSS ee
“= ies —2)° X,,¢,=0 (3-67)
2 t=1
and
a(RSS
(RSS)
ab, ie —20X,,e, = 0 (3-68)
Further

Le,Y, = b,Le, + b,DX,,e, + b,XX3,e, = 0 (3-69)


since each term on the right-hand side is zero.
Returning to Eq. (3-64) we can write it as
y,=y, +e, (3-70)
where y = Y— Y and y = Y — Y. Squaring and summing over all observations
gives
Ly? = Ly? + Le? (3-71)
since from Eq. (3-69)
LeY = Le(y + Y)=Lep =0 |
Relation (3-71) is the familiar decomposition of the total sum of squares
TSS = ESS + RSS
78 ECONOMETRIC METHODS

and it suggests the definition of the multiple correlation coefficient R as


2 2
2_ ESS
_ryt _| _der (3-72)
TSS ee
This coefficient is sometimes written as R,>;, indicating that it relates to a
regression where X, and X, are the explanatory variables. Similarly, for explana-
tory variables X,, X;,..., X, it would be written R, 5, ,, but in most cases there
is no ambiguity and the subscripts are not inserted.
It is illuminating to note that R* is also equal to the square of the simple
correlation coefficient between Y and Y. From Eq. (2-20) the latter coefficient
may be written

rz, ea (Zyp)°
eye ae)
But

EVV = Ly (Pre) = Dy
Thus

r2, = ry" = R?
Wy: yy?

The normal equations in Eqs. (3-62) may also be simplified if expressed in


deviation form. Substituting Y = b, + bX, + 6;X; in the second and third
equations gives
Ex) = DyLxs + By xox5 (3-73)
VX,V =Ds ky ks Ds ae
The calculation of ESS, RSS, and R* is then easily made in terms of deviations: -
ESS = Dy?
= LP(b.x2 + bx)
= b,Ux>9 + b,=x,)
= b,x,y + b,Ux3y (3-74)
Thus once b, and b, have been obtained from Eqs. (3-73), ESS is easily computed
from Eq. (3-74), and finally
RSS =e?
= 2) (bE, vad, 2x) (3-75)
and

R2 = zy?
Ly?
The exposition so far has been slanted toward the calculation of estimates by
hand or on a desk calculator. This may seem uncalled for in an era when there is
EXTENSIONS OF THE TWO-VARIABLE LINEAR MODEL 79

w (%)

ib
O
Cle Oo
oO
ae oO O
sere
oO
4 Oo

& O
3 —

2
00

O dA Uu
4.0 5.0 6.0 |‘ Figure 3-12

almost universal access to large-scale computers, but there is no one more


dangerous than the unthinking user of a computer, who has no real understanding
of the nature of the computations being processed inside the machine. The one
sure way to that understanding is to process some fairly simple examples
thoroughly and carefully.

Example 3-3 Figure 3-12 shows a scatter diagram of w, versus u, for the first
16 quarters of the data, that is, from 1954: 2 to 1958: 1. It is suggestive of a
nonlinear negative relationship.f Fig. 3-13 shows the scatter of w, plotted
against the reciprocal of unemployment. The slope is now positive and the
scatter approximately linear. The simple regression fitted to these data gives

Ww, = —9.7530 + 62.9490|1 with r? = 0.8229


t

Computing the same regression for all 26 quarters up to 1960: 3 gives

Ww, = — 1.4486 + 26.2924 with r? = 0.4235


t

+ The nonlinearity is, in fact, very slight. A linear regression of w, on u, has an r? = 0.8192, which
is negligibly smaller than the r? of 0.8229 for the linear regression of w, on 1 /u,.
80 ECONOMETRIC METHODS

w (%)
7

olga
0.15
— !
0.20
ae u,
0.24 Figure 3-13

We see that the coefficients change substantially and the fit to the longer
period is much worse than that to the shorter period. This is just the first
example of how regressions can often change markedly when sample data are
extended, and in Chap. 6 we will outline methods of analyzing such changes »
and testing for structural changes in the hypothesized relationship.
Perry introduced the lagged inflation rate as an additional explanatory
variable, assuming it to be a good proxy for the expected rate of inflation.
The resultant multiple regression for the early period 1954: 2 to 1958: 1 is

w, = —9.7396 + 62.9205 — 0.0054p,_, with R? = 0.8230


t

and for the complete sample period it is

w, = —1.7220 + 25.3698(1] + 0.3069p,_, with R* = 0.5197


t

We see that for the early period the addition of the lagged inflation rate gives
no improvement in the regression: the square of the multiple correlation
coefficient is identical to the third decimal place with the square of the simple
correlation coefficient, and the coefficient of the inflation rate is negligible in
size and perversely signed. For the complete period, the addition of the
EXTENSIONS OF THE TWO-VARIABLE LINEAR MODEL 81

inflation rate yields some improvement in an initially rather unsatisfactory


equation, and the coefficient of the lagged inflation rate is now correctly
signed and no longer numerically negligible. We are not yet in a position to
apply inference procedures to a multiple regression equation, but this exam-
ple should alert us to the possibility of substantial variations in estimated
relationships as a data base is expanded or contracted.

The b’s are point estimates of the hypothetical 8’s, and all the usual inference
questions arise, such as those discussed in Chap. 2. There is little point in working
out these questions explicitly for the three-variable case, for in Chap. 5 we will
develop all the required inference procedures for the general case of k variables.
However, before leaving the three-variable case there are some additional alge-
braic relations to develop, which will also contribute to our understanding of
more complicated relationships later.
With three interrelated variables Y, X,, and X, there are three simple
correlation coefficients denoted by
Tio. 713, and rr,
where the subscript | refers to the Y variable. The techniques of Chap. 2 would
also enable us to compute various regression slopes such as
b, = slope of regression of Y on X,
b,, = slope of regression of Y on X,
by3 slope of regression of X, on X,
and there are, of course, three further regression slopes where the order of the
subscripts on b is reversed.} The first question to explore is the connection
between b, and b,, the slopes of the regression plane, and the simple b’s, the
slopes of the various two-variable scatters. Solving Eq. (3-73) for b, gives
PLA tae ay oe
2
Ex2Ex?2 — (Lx>x5)°
Dividing top and bottom by Lx3Lx; then gives$
eins by3b3. _ M2 = M1373 51 (3-76)
b
7 1 = by by l-rA 52
Similarly,
ie bi3 — By2bo3 _ M13 Miaha3 Si (3-77)
: 1 — by3b3y l-r, 53

+ In the regression X, on X; the residuals are measured in the X, direction, while in the regression
of X; on X, the residuals are measured in the X3 direction.
¢ This follows directly by applying Eqs. (2-14a) and (2-19), that is,
Ux34X a Ux2X3 Ss]
Lyx Dia12 = Tosa
Dia = Ex? ’ b32 ? 23 > 12
ty)
?
“x3 xe
82 ECONOMETRIC METHODS

Thus the coefficient of X, in the multiple regression can be regarded as being


based on the coefficient in the simple regression of Y on X,, subject to a
correction for the presence of X;. If it should happen that X, and xX, are
uncorrelated in the sample, that is, Lx,x, = 0, then b,; = b3, = 0, and the
correction factor is zero, so that
b, = Dj, b; = bi,
The explanatory variables in this case are said to be orthogonal, and the multiple
regression coefficients coincide with the simple regression coefficients. An older,
but expressive notation for b, and 5, is b,,, and 6,35, respectively. A coefficient
like b,, is then described as a zero-order coefficient and b,,, as a first-order
coefficient, the number after the decimal point in the subscript indicating the one
other explanatory variable present in this regression. Equations (3-76) and (3-77)
thus indicate the connection between zero-order and first-order regression coeffi-
cients. It is natural to inquire if there is a similar hierarchy of correlation
coefficients. The simple coefficients r,,, 73,... are defined as zero-order correlation
coefficients. What then is the definition of a first-order correlation coefficient, and
what meaning attaches to it?
This question may be approached in the following way. Suppose X, influences
Y, but X, also influences Y. The simple correlation coefficient between Y and X,
is by definition,
s LyX.
© Sy? Ex3
Part of the variation inY is really due to X;, and so r,, will not correctly measure
the correlation attributable to X,. If one ran a linear regression of Y on X;, the
residuals are given by
y — bi3X3
This series represents the variation in Y left over after the linear effect of X, has
been removed. Correlating these residuals with X, gives
=e eee
X“(y = B33) Xo

yUCy =a bi3x3)° yux5

If X, and X, were uncorrelated in the sample, this coefficient simplifies to


iy ie e eee DX
ry
VE(y = by3x3)° ~x3
which will be numerically greater than the simple coefficient r,.. However, X, and
X; are in general correlated, and the proper correction procedure is to remove the
linear effect of X, from both X, and Y and then to correlate the resulting
residuals. This gives

la eS ame
X(y oe b13X3)(x, - bo3X3)
(3-78)
VE(y im bisxs)s E(x, = bas
EXTENSIONS OF THE TWO-VARIABLE LINEAR MODEL 83

Equation (3-78) defines a first-order correlation coefficient. It measures the


correlation remaining between Y and X, after the linear effect of X, has been
removed from each. It is more generally known as a partial correlation coefficient.
Equation (3-78) simplifies tof .

ease Ne = N3hs (3-79)


12.3 *
Ve rizV1 ac 7
Similarly,
r SS eee
igummeaolos (3-80)
13.2 [ager oOBs
Lie =

This latter expression may also be derived from first principles by finding the
simple correlation between

(y ri bX) and (x3 = b3)X>)


or, more simply, by interchanging subscripts 2 and 3 in Eq. (3-79). A partial
correlation coefficient between any two variables thus measures the correlation
between the residuals left in each variable after the linear effect of all other
explanatory variables has been removed.
The multiple regression slopes b, and b,, defined in Eq. (3-73), may also be
interpreted in terms of these residuals. Regressing (y — b,3x;) on (x, — b,3x3)
gives a regression slope of

= 0 =D
Ely —bisxs)2 ~bass) =EW(%2 ~bass) since x, is uncorrelated with
X(x2 — by3x3) L(x, — by3x3) the residuals (x, — b,3x3)

a LyX ,hxz — Lyx,LxX,


LxzLx;z — (ee),

+ The explicit derivation is as follows. Using Eqs. (2-14a) and (2-20),


numerator of 7,53 = Lyx7 — by,Lx7x3 — by3Lyx3 + by3by,Lx3
5] 52
= S152) His elt 232253 i= "236, 7138153
3
S182 9
113193 ——2 1183
53

= 13 \82(T2 — 113%23)
where s, denotes the sample standard deviation of the Y values, s, the sample standard deviation of
the X, values, and so on. In the denominator L(y — b,3x3)° is the residual sum of squares in the
regression of Y on X3. From Eq. (2-20) this may be written as ns?(1 — rj3). Likewise L(x — by3x3)*
is the residual sum of squares in the regression of X, on X3 and may be written as ns3(1 — 73). Thus
the denominator of Eq. (3-78) is ns,s2y1 — ri, yl 15, and Eq. (3-79) follows.
84 ECONOMETRIC METHODS

from Eq. (3-73). Likewise, b, is just the simple regression slope obtained when the
residuals (y — b,,x,) are regressed on the residuals (x; — 53x).
Finally, the multiple correlation coefficient R,5; is a second-order correlation
coefficient, and it may also be expressed in various ways in terms of lower-order
coefficients. For example, using Eq. (3-74) gives
. Dix, Vat Dg XGy
Ribs = eee eae
Ly
Substituting for b, and b, from Eqs. (3-76) and (3-77) gives

he ps ria + 13 — 2 ishs
Ri23 = 1 2
(3-81)
mala)
The buildup of the explained sum of squares may also be looked at in sequential
fashion, and this is illuminating for the analysis of variance treatment in Chap. 5.
Suppose one first regressed Y on X,. Then
ry? = ESS due to the regression of Y on X,
and
ry? = i) = RSS from the regression of Y on X,
This latter quantity is the sum of squares still to be explained. Regressing the
residuals (y — b,,x,) on the X, residuals, namely, (x, — b3,x,), then gives

rj;.&y?(1 — r2,) = ESS in the regression of the Yresiduals


on the X, residuals
and

(1 — r},.)Xy?(1 — r},) = RSS in the regression of the Y residuals


on the X, residuals

Aggregating the ESS at each stage gives a total explained sum of squares in Y as

Sys (Wipe ae (lee ri») (3-82)


Substituting for r/,, in Eq. (3-82) from Eq. (3-80) and using Eq. (3-81) gives

Ey te + ris2(1 is TD)| = Ly’Rins (3-83)


But Ly*R{,;, by definition, indicates the sum of squares in Y explained by the
multiple regression of Y on X, and X,. Equation (3-83) thus shows that the
multiple regression ESS can be regarded as built up in two steps, first the ESS due
to the simple regression of Y on X, and second the ESS due to the simple
regression of the Y residuals on the X; residuals. The increment in the ESS due to
adding X, to the regression is
Ly? (Riss = ra)
which, from Eq. (3-83), may be written

rdy(I = rp)
EXTENSIONS OF THE TWO-VARIABLE LINEAR MODEL 85

where Yy7(1 — rj) is the RSS after Y has been regressed on els) hes
indicates the proportion of the variation left after the simple regression on i
which is explained by adding X; to the set of explanatory variables. A similar
development and interpretation may be made by starting with the simple regres-
sion of Y on X; and then adding X, to the explanatory variables.

Example 3-4 Denote the variables in Table 3-7 by


l=w,
2=u;!
3 Spe

Table 3-7 Wage change, unemployment, and price change in the United States}

Year Quarter w, oe Dik

1954: 2 3.53 0.2312 1.229


3 1.74 0.1942 0.702
4 1.72 0.1794 0.000
1955: 1 71 0.1827 —0.522
2 Dee 0.1942 — 0.260
3 4.57 0.2128 —= (173
4 4.52 0.2260 0.090
1956: 1 5.06 0.2353 0.699
2 5.56 0.2367 0.263
3 4.37 0.2367 1.048
4 5.95 0.2381 1.997
1957: 1 6.42 0.2395 2.507
2 5.26 0.2424 3.447
3 5.24 0.2410 3.589
4 4.08 0.2286 3.376
1958: 1 3.52 0.2010 3.023
2 3.50 0.1747 3.415
3 3.48 0.1544 3.221
4 2.94 0.1460 2.297
1959: 1 3,88 0.1487 1.966
2 4.35 0.1613 0.814
3 2.88 0.1747 0.404
4 34536) 0.1810 0.967
1960: 1 3.24 0.1878 1.367
2 2.78 0.1869 1.528
3 3.74 0.1869 1.762

+ w= four-quarter percentage change in hourly earnings of production workers in all manufac-


turing
u =unemployment as a percentage of the civilian labor force
p = four-quarter percentage change in the consumer price index (CPI)
Source: G. Perry, Unemployment, Money Wage Rates and Inflation, MIT Press, Cambridge,
MA, 1966.
86 ECONOMETRIC METHODS

For the complete sample period the zero-order correlation coefficients are
ry = 0.6508 43 = 0.3567 —r = 0.0726
r2, = 0.4235 =r, = 0.1272
Thus
|. 0.6508 — 0.3567(0.0726) atieaed
a
[1 — (0.3567)"] \/[1 — (0.0726)"|
0.3567 — 0.6508(0.0726)
hy2 = ee = 0.4087
[1 — (0.6508)?] [1 — (0.0726)"|
and rh, = 0.4498 7r2,,. = 0.1670
and Rj>; = 0.5197. In this example the two explanatory variables are practi-
cally uncorrelated, so there is little difference between the zero-order and the
first-order correlation coefficients. Unemployment alone explains over 42
percent of the variation in wage change, price change alone accounts for
about 13 percent, and the two variables jointly account for about 52 percent.
Unemployment accounts for 45 percent of the variation unexplained by price,
and price accounts for about 17 percent of the variation unexplained by
unemployment.

PROBLEMS

3-1 Given five observations u_,, u_,, Uo, u;, and uy at equally spaced points of time ¢ =
— 2, —1,0,1,2, show how to fit a parabola to the observations by least squares and show that the
value given by the parabola at time t = 0 is

35(—3u_> + 12u_, + 17u9 + 12u, — 3uy)

(R.S.S. Certificate, 1955)


3-2 The “firmness” of cheese depends upon the time allowed for a certain process in the manufacture.
In an experiment on this topic, 18 cheeses were taken, and at each of several times firmness was
determined on samples from three of the cheeses. The results (on an arbitrary scale) are given below:

Time,
h Firmness

4 102 105 115


l 110 120 115
i 126 128 119
2 132 143 139
3 160 149 147
4 164 166 172
EXTENSIONS OF THE TWO-VARIABLE LINEAR MODEL 87

Estimate the parameters in a linear regression of firmness on time. Give standard errors of the
estimates and test the adequacy of a linear regression to describe the results.
(R.S.S. Certificate, 1955)
3-3 Discuss briefly the advantages and disadvantages of the relation
v; = a + Blog vo
as a representation of an Engel curve, where v; is expenditure per person on commodity i and vg is
income per person. Fit such a curve to the following data, and from your results estimate the income
elasticity at an income of £5 per week.

Pounds per week

0; 0.8 He 1.5 1.8 2d 23 2.6 351


Vo Le 27 3.6 4.6 Dail 6.7 8.1 12.0

Does it make any difference to your estimate of the income elasticity if the logarithms of vg are
taken to base 10 or to base e? Explain carefully.
(Manchester University, 1956)
3-4 Response rates at various levels of ratable values.

Range of ratable value A RBG ED EVES G eH Ts:

Assumed central value X, £/annum 3 7 12 #17 #25 #35 45 55 70 120


Response rate Y, percent ek IM Gk) Te ey OY ehh Sil ail 4h

The data relate to a survey recently conducted in England. Estimate the constants in the regression
equation

(Oxford University, 1955)


3-5 X,, X,, and X; are three correlated variables, where s; = 1, sy = 1.3, 53 = 1.9 and 7, = 0.370,
ry3 = —0.641, and r; = —0.736. Compute 7,3. If X, = X, + Xp, obtain ry, 143, and 743. Verify
that the two partial coefficients are equal and explain this result.
(UL, 1952)
3-6 In the regression equation
Y, = Bx, + YX2, + U, t= Wineygit

all variables are expressed as deviations from their sample means. Consider the following alternative
procedures for estimating . ‘
(a) Calculate the estimates B and y ina regression of y on x, and x}.
(b) Regress y on x, and calculate the regression residuals y,*; regress x, on x, and calculate the
regression residuals x*,; regress y* on xf to obtain an estimate b of £.
Show that the two procedures give the same result, that is, B = b.
Show that the regression residuals given by each procedure, that is,

Vice Bx, A YX, and Vp = Oxi


are the same.
(UL, 1969)
88 ECONOMETRIC METHODS

3-7 Outline the properties of the following functions, and sketch their graphs:
(a) y=at Blnx
x
(b) y= oe
ert Bx

Coa eens
Find transformations which linearize functions (b) and (c), that is, for each function find a pair of
transformations f(x) and g(y) such that g() is a linear function of f(x) and the a, 8 parameters
may be estimated.
3-8 Your research assistant reports the following results in several different regression problems. In
which cases could you be certain that an error had been committed? Explain.
(a) R253 = 0.89 and Rj 34 = 0.86
(b) ri = 0.227, r>, = 0.126, and R7,, = 0.701
(c) (Ux?)(Ly?) — (Lxy)? = — 1,732.86
(University of Michigan, 1980)
3-9 Sometimes variables are standardized before the computation of regression and correlation
coefficients. Standardization is achieved by dividing each observation on a variable by its standard
deviation, so that the standard deviation of the transformed variable is unity. If the original relation is,
\
Say,

Y = B, + By)X_ + B3X3 + u
and the corresponding relation between the transformed variables is
Y* = By + By XZ + B3XF + u*
where
Ye = /Sey eX) oe AAS, i= 2,3

what is the relationship between 83, 83, and 8,, B;? Show that the partial correlation coefficients are
unaffected by the transformation.
CHAPTER

FOUR
ELEMENTS OF MATRIX ALGEBRA

It is clear from the last section of Chap. 3 that it would be excessively tedious and
complicated to build up to the general case of k-variable regression in a stepwise
fashion. Fortunately, by the use of matrix algebra we have a compact and
powerful way of treating the problem, and we shall see that the detailed results of
the previous two chapters are merely special cases of a few simple matrix
formulas. The rest of this chapter presents the elements of matrix algebra that are
necessary for following the treatment in the remainder of the book. Most or all of
this chapter may be skipped by those with adequate previous knowledge of
matrices. For those whose knowledge is somewhat rusty it may hopefully serve as
a useful review, but every attempt has been made to make the material accessible
to a student with no prior knowledge of matrices, for the subject is so fundamen-
tal to modern econometrics (and economics) that no serious student can afford to
be without it.
Suppose our theory suggests that a dependent (explained) variable Y is a
linear function of k — 1 independent (explanatory) variables X,, X;,..., X;,.

+ Nonetheless students with no prior knowledge of matrix algebra are likely to get indigestion if
they attempt to work through all the material in this chapter before proceeding with the rest of the
book. The topics are introduced approximately in the order in which they will appear in subsequent
chapters. Thus the students should interact between this chapter and those that follow, learning
enough from Chap. 4 to proceed with Chaps. 5, 6, and so on, and returning to Chap. 4 as necessary.
Summaries of results have also been inserted at various stages in the chapter.

89
90 ECONOMETRIC METHODS

Allowing for a constant term, we would write the function as


Y= B, + BX,
+ BX, + +++ + B,X,
+u
If we have n sample observations, the model gives rise to the following set of n
equations:
Y, = By + By,Xp, + By Xx, + °° + BY +
Y, = By + By Xp. + By Xyq + +++ + By Xpq + Uy
ei © oele (es ©) © ec efhe .olve a (e761 s) fememre: oriemioy @) 6) 6/08) Sie, (Cis ose ees

Y,= 8B + BAo, thBaAgy et = Be Me, te, (4-1)


As afirst step we rewrite these n equations in the form

Y, 1 X, Xx eee Ne 11 Ps uy
Y, & Ag Aa es Be x e (4.2)

¥, 1 yo X3,, AP B, Uu,

or
y=XBp+u (4-3)
where the four boldface symbols correspond to the four sets of elements that have
been enclosed in square brackets in Eq. (4-2). These symbols indicate vectors and
matrices. For example,

is an n-element column vector, with the sample observations on Y arranged in a


specific order in the form of a column. Likewise B is a k-element column vector,
containing the coefficients of the hypothesized relation, and u is an n-element
column vector, containing the n unknown disturbances.

1 X), X3; Xx

Xo eee Kea
ace
1]X, X3,
ee ae
Xu

is a matrix with n rows and k columns. A matrix is simply a rectangular array of


elements, and we say that X is a matrix of order n X k to indicate that the
rectangular array in X has n rows and k columns. In stating the order of a matrix,
the number of rows is always given first and the number of columns second.
Clearly, a column vector is merely a special case of a matrix, namely, a matrix
with only one column. Likewise, a row vector such as

[1 X5, X3) ok. Xl


ELEMENTS OF MATRIX ALGEBRA 9]

is another special case, namely, a matrix with just one row.t We may thus look at
the X matrix in two ways, as an ordered collection of column vectors or as an
ordered collection of row vectors. Each column vector, apart from the first,
denotes the sample observations on a particular explanatory variable. Thus, for
example,

denotes the sample observations on the variable X,. The first column is a
collection of units and, as we will see, is required in order to incorporate the
intercept 8, into the regression. Using this notation for the column vectors, we
could express X as
bat] |
Rex 7) PX ea Xe (4-4)

where each x, is an n-element column vector and x, indicates a column of units.


The rows of X indicate observations on all explanatory variables at a particular
sample point plus a unit in the first position. Thus if the data are in time series
form, the first row indicates the X values in the first time period, the second row
the values in the second time period, and so forth. Thus using s; to indicate the
vector of observations at the ith sample point, the X matrix may also be expressed
ast

X= : (4-5)
—* s —_—

We have inserted vertical and horizontal lines in Eqs. (4-4) and (4-5) to emphasize
that the first is a representation of X in terms of column vectors and the second a
representation in terms of row vectors. In practice one usually writes these
expressions more compactly as
S)
S2
X= [xy yer a) ‘
n

it being clear from the context which are row and which are column vectors.

+ We will adhere to the convention of indicating a matrix by an uppercase boldface letter and a
vector by a lowercase boldface letter. ;
+ We have indicated the row vectors by the letter s since they correspond to sample points. A more
common notation is to let x;. indicate the ith row of the X matrix and x. ; the jth column.
92 ECONOMETRIC METHODS

Equations (4-2) and (4-3) are equivalent ways of stating Eq. (4-1). For this to
be so, operations on matrices must follow certain simple rules, which we will now
describe.

4-1 OPERATIONS ON VECTORS AND MATRICES

The right-hand side of Eq. (4-3) indicates two elementary operations on matrices,
namely, multiplication and addition.

Matrix Multiplication
Matrix multiplication is achieved by repeated applications of vector multiplica-
tion. The multiplication of an n-element row vector into an n-element column
vector is defined as follows:

[a, Cs a,] 5 = a,b, + a,b, + --- + a,b, = Y a,b; (4-6)


i=]

that is, corresponding elements are multiplied together and the results summed.
As a numerical example,

Pave -1 4]= 2(1) + 3(4) + (-1)(5) =9


5
As the definition and the example show, multiplying a 1 x n vector into ann X 1
vector produces a 1 X 1 vector, or a scalar quantity. Notice that the operation is —
not defined if the number of elements in the two vectors is not the same. It is also
clear from the definition that
a, b
a, b,
[OieeDsy eee b\) . |= (4, Ce yeaa

a, b,

Suppose, however, we define the vectors a and b as column vectors, that is,

a, b,
em) b,
ey bee (4-7)

a, b,

The above multiplication definition cannot apply directly to a and b since they are
both column vectors. We thus define an operation of transposition, which turns
ELEMENTS OF MATRIX ALGEBRA 93

column vectors into row vectors, and vice versa. Thus


Transpose ofa=a’=[a,
comet (es
a, --- a,]
Transposition does not change any of the elements of the vector; they are merely
written in the original order, but as a row instead of a column, or vice versa. It is
also clear that transposing a’ gets us back to the original column vector. Thus
(ay =a
Some writers use a superscript T to indicate a transpose, that is, a’ = a’, but we
will use a prime. We will also sometimes find it useful to take the transpose of a
1 X 1 vector, or scalar, and clearly, that leaves the scalar unchanged.

Scalar, Dot, or Inner Product


In general, unless we specifically state the contrary, a vector symbol will indicate a
column vector, as in Eq. (4-7). The multiplication operation already defined in
Eq. (4-6) then enables us to define the scalar, dot, or inner product of the two
vectors a and b in Eq. (4-7) as

a'b = b'a = ) ab, (4-8)


i=l
If we have two matrices A and B, where A is of order m X n and B is of order
n X p, the product AB will be a matrix C of order m xX p, that is,
AB = C (4-9)
(mxn)(nXp) (mxp)

If we indicate the rows of A by a, (i = 1,..., m) and the columns of B by b;


(j = 1,..., p), each (scalar) element in C is the inner product of a row vector
from A and a column vector from B, namely,

aa a
ou ue ey | ajb, a,b, --- ajb,
s beth, 9-2 b=) a,b, aby 9 2) a,b (410)
Re Sorel | a,,b, a,b, a,b,

The basic rule embodied in Eq. (4-10) is


Element in i, jth position in AB = inner product of row i of A and column j of B

Example 4-1

1 2 371 ©} [1G)+20)+3(1) 16) + 2(1)+ 3(1)


E 0 4 ; + 4(1)-—-2(6) + 0(1) + 4(1)
~ |2(1) + 0(0)
eae Al
ap = |4 a
94 ECONOMETRIC METHODS

Example 4-2
[ne 1(1) + 6(2) 1(2) + 6(0) 1(3) + 6(4)
|: IL =| 0(1) +.1(2) 0(2) +10) 0(3)
+1(4)
jem 1(1) +1(2) 12) +10) -1(3)
+1(4)
[gard |
BA = 2~-0 4
oe: 7

As the examples and the definition in Eq. (4-10) make clear, the order in
which matrices are multiplied is of crucial importance:

AB indicates that A is postmultiplied by B or, equivalently, that B is premulti-


plied by A

This operation is only possible if the inner products of rows of A and columns of
B exist, that is, if the number of columns inA is equal to the number of rows in
B. In this event the matrices are said to be conformable. The simplest check is to
write down the order of the two matrices to be multiplied, as in Eq. (4-9), and it is
seen that the common index n disappears to give a product matrix of order
m X p. The product BA would only exist if p = m so that the inner products of
the rows of B and the columns of A could be formed. Note carefully that the
definition of matrix multiplication involves the inner products of the rows of the
first matrix and the columns of the second.
A special case of Eq. (4-10) occurs when one of the matrices is simply a
vector. For example,
—- a — a,b
— a, —||! a,b
: Deas eal oes c (4-11)
| :
— a a,,D
and ¢ is an m X 1 (column) vector.
Returning now to Eq. (4-2), the right-hand side incorporates the multiplica-
tion of a matrix by a vector, and applying the above rules gives

xg = |fit
By
Became et
+ By Xo Bat eee
Xe (4-12)
By HeRG Aon bsg et eee
which is just an n-element column vector.

Matrix Addition
The right-hand side of Eq. (4-2) or Eq. (4-3) is now seen to consist of the addition
of the two vectors XB and u. This addition is simply achieved by adding
corresponding elements. Thus the operation is only defined for vectors with the
ELEMENTS OF MATRIX ALGEBRA 95

same number of elements. In general,


a, b, a, + b,
a, b, a, + by
a+tb=|] |[+] .]/= . (4-13)

a, | b,, Garb.
The definition in Eq. (4-13) is readily extended to the addition of matrices. Two
matrices A and B can be added together only if they are of the same order m X n.
The sum matrix is also of the same order, and each element in it is simply the sum
of the corresponding elements in A and B.
Applying Eq. (4-13), the right-hand side of Eq. (4-3) reduces to an n-element
column vector,

NAW ile ek ey ete ts


Bit Pex, te Bex, 1,
Equality of Matrices
Finally, Eq. (4-3) states an equality between y and XB + u. This simply means
that the first element is y, equal to the first element in XB + u, and so on through
all n elements, that is,
Y, = B, + BX + +++ + BX + uy
Yee FaBo Aone tae PAL, att,
and so we are back to the n equations of Eq. (4-1) and have shown that the simple
rules of matrix addition and multiplication enable us to write the system (4-1) in
the compact form
y=Xp+u

Further Remarks on Matrices


Our primary purpose is to analyze the model y = XB + u by least-squares
techniques, but in order to proceed with that we need to develop some further
properties of matrices.
Transposition has already been defined for vectors. Since a matrix is a
collection of vectors, we can define the transpose of a matrix. Let A be an m X n
matrix, which we write in the form

A’ =
(nXm)
96 ECONOMETRIC METHODS

that is, the first row of A has become the first column of the transpose, the second
row of A the second column of the transpose, and so forth. The definition might
equally well have been stated in terms of the first column of A becoming the first
row of A’, and so on. Clearly, A’ is of order n X m.

Example 4-3

that is,
Ay = a,; fori + j

This property can only hold for square matrices (m = n), since otherwise A’ and
A are not even of the same order.

Example 4-4
Li 4
A=] -1 0 3) =A’
4 5 2

From the definition of a transpose it is immediately obvious that the following


two properties hold:
(A’) =A
that is, the transpose of the transpose equals the original matrix, and
(A + B) =A’+B’
that is, the transpose of a sum is the sum of the transposes. Somewhat less
obvious is the interpretation of (AB)’, the transpose of the product AB. Referring
back to Eq. (4-10), we note again that the i, jth element in C = AB is
c;; = a,b; = inner product of ith row of A into jth column of B

Pest leat; Jil aeeneD


Transposition of C means that the i, jth element in C’ is the J, ith element in C.
Using c;, to denote the i, jth element in C’ gives

CC
le
= a,b;
rah

But
ELEMENTS OF MATRIX ALGEBRA 97

from the definition of an inner product. Thus


c;; = bjaj, = inner product of ith row of B’ and jth column of A’
and so, from the definition [Link] multiplication,
C’ = (AB)’ = B’A’ (4-14)
This rule extends immediately to any number of matrices. Thus
(ABC)’ = C’B’A’ (4-15)
since
(ABC)’ = C’(AB)’
= C’B’A’

by repeated application of Eq. (4-14).


The associative law of addition holds for matrices, that is,
(A+ B)+C=A+(B+C) (4-16)
This result is obvious since matrix addition merely involves adding corresponding
elements, and it does not matter in what order the additions are performed.
The associative law of multiplication also holds, that is,
(AB)C = A(BC) (4-17)
where A, B, and C are assumed to be of the appropriate order for multiplication.
To prove this result, we will show that the i, jth elements on each side of Eq.
(4-17) are the same. The ith row of the product AB is given by

[ajb, a,b, ---]=a,[b, b, ---]=a,B


where a; denotes the ith row of A and b,, b,,... denote the columns of B. Letting
c; denote the jth column of C, the i, jth element of (AB)C is then
a; Be,

Similarly, the jth column of the product BC is


Be;
and so the i, jth element of A(BC) is
a; Be,

The distributive law also holds for matrices, that is,

A(B + C) = AB + AC (4-18)
To see this, let a, denote the ith row of A and b, and ¢; the jth columns of B and
C. The i, jth element on the right-hand side of Eq. (4-18) is then the scalar
a,b; + a,¢

and by the application of the distributive law for scalar algebra this is clearly
equal to the inner product of a; and the vector b; + ¢,, the jth column of B + C,
which gives the i, jth element on the left-hand side of Eq. (4-18).
98 ECONOMETRIC METHODS

If a matrix is multiplied by a scalar, then every element in the matrix is


multiplied by the scalar. For example,t

-2|} 2 ales ie —4 af
ara A =a 0. 5
There are some square matrices of particular importance. First is the unit or
identity matrix of order n X n,

oie) e4. 0) =e) Jeiu oe Beliels) co) ©

with units down the main, or principal, diagonal and zeros everywhere else. As we
shall see, it plays in matrix algebra a role similar to that of unity in scalar algebra.
As one illustration,
IA=AI=A
that is, pre- or postmultiplication by I leaves any matrix unchanged, as may
readily be verified by multiplying out IA and AI. Thus the unit matrix may be
entered or suppressed at will in matrix expressions. For instance,

YoY
= Vie
= (ei
A diagonal matrix is like the identity matrix in that all off-diagonal terms are
zero, but now the diagonal elements are scalar quantities, one of which at least is
nonzero. The diagonal matrix may be written

OO 0
ie OA 5aeeO 0
=a 120Sec Gnomesae aa e
0 Coe)

or, more compactly, A = diag{A, A, -:: A,,}. Examples are

2 0 0
ie i and 3 —4 )
0 0 5
A special case of Eq. (4-19) occurs when the )’s are all equal. This is termed a
scalar matrix and may be written

rA 0 0
0 A 0|}=AI
Galpie ea ham‘

7 The scalar may be placed in front of or behind the matrix.


ELEMENTS OF MATRIX ALGEBRA 99

Another special case of a square matrix is an idempotent matrix. Let A be a


Square symmetric matrix, so that
A’=A
If A is idempotent, then
A=A’=A=.-..
that is, multiplying A by itself, however many times, simply reproduces the
original matrix. As we will see later, idempotent matrices play an important role
in Statistical theory. An example of an idempotent matrix is

ltl; eee nee


A==>|-2 | 4 #2
: Mle l
as the reader can easily verify by multiplication.
Another important matrix is the null matrix 0 whose every element is zero.
Obvious relations are
A+0=A
and

A0=0
Similarly, we may have null row or column vectors.

Partitioned Matrices
Writing the matrix X in the form

Re Xo, ae Xe!
as in Eq. (4-4), is a special example of a partitioned matrix. The elements on the
right-hand side are not scalars but vectors. In general a partitioned matrix
contains submatrices as elements. The submatrices are obtained by partitions of
the rows and columns of the original matrix. For example,

4 0 2!-1 A A
A=| 6 5 Sis, al = ee al (4-20)
us 3 a 0! 5 21 22:

where
AM iOe a
Ay, i Saal Ds | |
(4-21)
AS = lao 20] Ay =5
The dashed lines indicate the partitioning, yielding the four submatrices defined
in Eqs. (4-21).

+A square nonsymmetric matrix is idempotent if it satisfies A? =A, but we will only meet
symmetric idempotent matrices in this book.
100 ECONOMETRIC METHODS

Our previous rules for the addition and multiplication of matrices apply
directly to partitioned matrices provided the submatrices are of appropriate dimen-
sion. For example, if A and B are both written in partitioned form as

AC
ae and B=
i =
a Ay B,, By
then

A,, + By, Ay t+ By
A+B=
A», ztB,, A» a: B,,

provided A andBare of the same overall order (dimension) and each pair A; ,, B;,
is of the same order. As an example of the multiplication of partitioned matrices,

A A
AB mee ie ie ay
B,, B,,
A; A)

A,,B,, + Ay.B, A, By. + A,B,


= | A,,B,, + AB, A,B, + AB,
A;,B,, + A3.B, A3,B,. + A3.B,
For the multiplication to be possible and for these equations to hold, the number
of columns in A must equal the number of rows in B and the same partitioning
must be applied to the columns of A as to the rows of B.

Summary on Matrix Operations


The main results of this section are summarized as follows:

1. The scalar, dot, or inner product of two n-element column vectors a and bis
a’b = bD’a = L?_,a,b,.
2. The typical i, jth element in the product AB, where the matrices are
conformable for multiplication, is D.a,.b
Set Stee S
ic
3. The typical element in A + B is a, 7 Op.
4. A=B means a,, = b,, for all i, /.
5. (AB)’ = B’A’, (ABC) = C’B’A’.
6. (A+ B)+C=A+(B+O).
7. (AB)C = A(BC).
8. A(B + C) = AB + AC.
9. IA= AI=A.
10. The typical element in cA, where cis a scalar, is ca, i
11.A+0=0.
12. AO= 0A = 0.
ELEMENTS OF MATRIX ALGEBRA 101

4-2 MATRIX FORMULATION OF THE LEAST-SQUARES


PROBLEM

Returning again to the linear model y = XB + u, this may be written in the form

B,
B,
Yor X75 Kyat X el ial Gee

B,
that is,

y = Bx, + Bx, +---+ 8x, +u (4-22)


Equation (4-22) states that the observed y vector is the sum of the disturbance
vector u and a linear combination of the columns of X. If we replace the unknown
B’s in Eq. (4-22) by guesses or estimates denoted by 5,, b,,..., b,, then for the
tth observation we have an observed value Y, and a calculated value
DIX, 40, XG et OX,
The difference between these two values defines a residual or error term,
Cel, Oy ee a Xe,
Repeating the procedure for all sample points gives
y = 5.x, + bx, + ---+b,x, +e=Xbt+e (4-23)
where b is a k-element coefficient vector and e is an n-element vector of residuals.
The least-squares principle states that the b’s in Eq. (4-23) should be chosen
to minimize the sum of the squared residuals. This sum of squared residuals is
e}
ey
ee=[e, e, --- @,]| . | =e? tegt+---t+e?
e,
From Eq. (4-23),
e=y— Xb
Hence
e’e = (y — Xb)’(y — Xb)
= (y’ — b’X’)(y — Xb)
= yy — b’X’y — y’Xb + b’X’Xb
= y’y — 2b’X’y + b’X’Xb (4-24)
since y’Xb is a scalar and so is equal to its transpose b’X’y. Once the sample data
have been obtained, y and X consist of known numbers. Thus Eq. (4-24) expresses
e’e as a function of the unknown b vector,
e’e = f(b) (4-25)
102 ECONOMETRIC METHODS

and, treating the elements of b as variables, we have to minimize e’e with respect
to b. This requires some elementary results on matrix differentiation.

Matrix Differentiation
If f(b) contains, say, k different b’s, then we may partially differentiate f(b) with
respect to each 5, in turn, obtaining k partial derivatives. Arranging these partial
derivatives in the form of a column vector gives the general definition

a[ f(b)
ab,
a[ f(b)]
al f(b)]
Ta fe (4-26)

aL f(b)]
These derivatives might equally well have been arranged as a row vector. The
important requirement is consistency of treatment and ensuring that vectors and
matrices of derivatives that have to be added and multiplied are of appropriate
order.
Suppose f(b) is a linear function,
f(b) = a/b
=10 Dias Dye a a,
where the a’s are given constants. Application of Eq. (4-26) then
ay
d(a’b) —A(b’a) ea)
Ab Go. oe alee oes (4-27)
on

Suppose now that f(b) is quadratic in the b’s, that is,


f(b) = b’Ab
This is a quadratic form in b, and we will develop the properties of quadratic
forms later. Without any loss of generality we can suppose A to be a symmetric
matrix denoted byt

41; 4p Ai,
A=]. ay ay,
Gi, Ar, AkK

+ If A were not symmetric, define A* = (A + A’)/2. Then b’Ab = b’Ab/2 + b’Ab/2 =


b/AB/2 +
b’A’b/2, on transposing the second b’Ab. Thus
A+ A’
b’Ab = b 7 b = b’/A*b
ELEMENTS OF MATRIX ALGEBRA 103

Then

+. Gy,
De

Taking partial derivatives,


0(b’Ab
a) = 2(4,,b, + ay.) + +++ + a,,b,) = 2a,b

0(b’Ab
ae ) = 2(a,,b, ta a>,b, Steen iet Gee) = 2a,b

where the a’s indicate the rows of A. Collecting these partial derivatives in a
column vector,
oH a,
,
0(b’Ab) =, a,b
: es a;
ob : :
b = 2Ab (4-28)
ae ay

Equations (4-27) and (4-28) give the standard results on the differentiation of
linear and quadratic forms. Notice the parallel with the differentiation of scalar
functions in that the power of the variable is reduced by | so that it disappears on
differentiation of the linear form and appears linearly on differentiation of the
quadratic form.
These two results may now be applied directly to minimize the residual sum
of squares defined in Eq. (4-24).
aWX'y)
ob
_ yy
using Eq. (4-27), since X’y is just a known k-element vector. Also
0(b’X’Xb) _ ,
masstee 2(X’X)b

using Eq. (4-28). Thus

(ee)
ob
_ _ ayy + 2X'Xb
For a stationary value of the sum of squares all k partial derivatives must be zero,
that is,
d(e’e) _
ae
104 ECONOMETRIC METHODS

and so
(X’X)b = X’y (4-29)
These are the normal equations for the least-squares regression, and include the
equations for the two- and three-variable cases already derived in Chaps. 2 and 3.

Example 4-5 Two-variable regression For the two-variable regression the X


matrix is

eas
5 lex
hes
Thus

oe ee and xy =|
XG Ee DEDE
So Eq. (4-29) gives
nb, + bX = LY
UX + bE XA ee
which are identical with Eqs. (2-13).

Example 4-6 Three-variable regression In this case


n LX, LX,
),CAE) CE ENE NSNEE

and
ai
Xiy = 3| AG
DOXGY
which, on substitution in Eq. (4-29), yield Eq. (3-62).

4-3 GEOMETRIC INTERPRETATION OF LEAST SQUARES

In many problems it is often helpful to have a geometric as well as an algebraic


interpretation. We will thus introduce some basic notions on the geometry of
vectors.
Consider a two-element vector
ELEMENTS OF MATRIX ALGEBRA 105

2d component

1 5 5 ‘i lst component Figure 4-1

This may be pictured as a directed line segment, as shown in Fig. 4-1. The arrow
denoting the segment starts at the origin and ends at the point with coordinates
(2,1). The vector a may also be indicated by the point at which the arrow

--[l
terminates. If we have another vector, say, b,

the geometry of vector addition is conceived as follows. Start with a and then
place the b vector at the terminal point of the a vector. This takes us to the point
P in Fig. 4-1. This point defines the vector ¢ as the sum of vectors a and b, and it
is obviously also reached by starting with the b vector and placing the a vector at
its terminal point. The process is referred to as completing the parallelogram, or
as the parallelogram law for the addition of vectors. Clearly, the coordinates of P

cneve
LFD]-L3
are (3,4), and

so that there is an exact correspondence between the geometric and the algebraic
treatments.

»Ai]-[f
Now consider scalar multiplication of a vector. For example,

gives a vector in exactly the same direction as a, but of twice the length. The
scalar multiplier may also be a negative number. For example,

cae
These two vectors are shown along with a itself in Fig. 4-2. Clearly, all three
terminal points lie on a single line through the origin, that line being uniquely
defined by the vector a.
106 ECONOMETRIC METHODS

2d component

Ist
component

Figure 4-2

Combining the two operations of scalar multiplication and addition of


vectors enables us to represent any two-element vector as a linear combination of
the vectors a and b. If ¢ denotes any arbitrary vector and it is to be represented as
a linear combination of a and b, then we may write
c=A,a+A,b (4-30)
where A, and ), are appropriate scalars. As an illustration of Eq. (4-30) consider
the following examples:

efi
--23ff]-$U
may be expressed as

giving A, = 2; and A, = — §. These values of A, and \, may be obtained


graphically by sketching in the parallelogram whose main diagonal gives the c
vector and then measuring the sides of the parallelogram in relation to a and
b. Alternatively, one may solve the pair of simultaneous equations
2 NE ees - {4
r +X.

ee
1 3 NE OAS

caffe
may be expressed as

giving A, = 3 andA, = 0.
ELEMENTS OF MATRIX ALGEBRA 107

reff
may be expressed as

eo? +22]
giving A, = 2 and A, = 2.

We are now in a position to define a vector space. A vector space is a


collection of vectors with the following properties:

1. If v, and vy, are any two vectors in the space, then v, + vy, is in the space.
2. If v is in the space and Ais a scalar constant, then Av is in the space.

The set of vectors is said to be closed under addition and scalar multiplication, for
these operations do not produce a vector outside the space.
Let us denote the two-dimensional space by the symbol 7. This vector space
consists of all real two-element vectors. Clearly, any vector in the space can be
expressed as a linear combination of the two vectors a and b. Our specification of

a-[] me a[’
a and b, however, was arbitrary. Consider another pair of vectors

These are usually described as unit vectors, and again any vector ¢ in R” may be
expressed as a linear combination of these vectors, only now the determination of
the A’s is particularly simple. The three previous numerical examples in this case
give

[oI 4|1] + 0[°| sodA, = 4,A,=0

2 (5) 6] +3|°| on 6 8

13] 6] +8[°| SO eGAN,


=8
so that the A’s are read off directly as the elements of the ¢ vector.
Each pair of vectors in these examples, that is, a, b, and e,, e,, serves as a
basis for the two-dimensional space ®7. A basis is thus not unique. It is clear
from the geometry that any two vectors can serve as a basis for R? only if they
point in different directions. If a and b point in the same direction, then one is
simply a scalar multiple of the other, as in Fig. 4-2, and only further multiples of
that vector can be expressed as linear combinations of a and b. The condition that
the basis vectors point in different directions may also be expressed by stating
that the vectors should be /inearly independent. Two vectors a and b are said to be
linearly independent if the only solution to
hatA,b=0 (4-31)
108 ECONOMETRIC METHODS

is A, = 0 =A,. If A values can be found, at least one of which is nonzero, to


satisfy Eq. (4-31), then the vectors are said to be /inearly dependent.

EE] aw 9 [f
As an illustration of these definitions, suppose

These vectors lie on the same ray through the origin, and the linear combination
3a — b yields the zero vector. However, if we revert to the original a and b
vectors, namely,
Pie male
a= 1| and b i

it is impossible to find a pair of A values, other than two zeros, such that
A\,at+A,b=0
We can easily find a pair of A values to reduce the first element to zero, but the
same combination will not reduce the second element to zero. A basis for R7 is
thus defined to be any linearly independent pair of two-element vectors: It is clear
from the geometry of the two-dimensional case that the representation of a given
vector in terms ofa given basis is unique, that is, there is one and only one pair of
X,, A, values which satisfy
c=A,a+A,b
Given a basis a,b for ®*, we have seen that any vector c in ®? may be
expressed as a unique linear combination of the basis vectors. Thus the vectors a,
b, and ¢ are linearly dependent, for the equation
A\,at+A,b-—c=0
holds for nonzero A’s. We might ask whether any arbitrary vector v in R” may be
expressed in terms of the expanded set of vectors a, b, and c. The answer is, of ©
course, yes, but the coefficients will not be unique. For example, suppose that the

fl EL Ll
a, b, and ¢ vectors are

and we wish to express


ma
‘ |
as a linear combination of a, b, and c. One such combination is simply

v = 2a + 2b + 0c
but there are infinitely many others. Rewriting the general linear combination
v=A,a+A,b+A,c
in the form
v—A,¢c=A,a+A,b
any arbitrary value can be assigned to A,, and the left-hand side is then some
ELEMENTS OF MATRIX ALGEBRA 109

specific two-element vector, which can be expressed as a linear combination of a


and b. A set of vectors such as a, b, and ¢ is called a spanning set since they span
or generate the space R, that is, any vector in ®* can be expressed as a linear
combination of the spanning vectors. The distinction between a basis and a
spanning set is that the basis consists of linearly independent vectors. A spanning
set may be unnecessarily large, as in the case of the set a, b, and c. This spanning
set can be reduced to a basis by dropping one vector.
Since linearly independent vectors point in different directions, there is then a
nonzero angle between the vectors. This angle may be expressed in terms of the
elements of the vectors.
Reverting to the two-dimensional a,b vectors in Fig. 4-3, let A denote the
angle between the a vector and the horizontal axis and let B denote the angle
between the b vector and the horizontal axis. The angle between the two vectors is
6=B-A
An elementary result in trigonometry states
cos(B — A) = cos B-cosA + sin
B: sin A (4-32)
The length or norm of the vector a is, by Pythagoras’s theorem, ja; + a}. The
length is often denoted by the symbol |lal|, and using the definition of the inner
product in Eq. (4-8), we have
llall = a’a
Likewise,

|[b||* = b’b
Substituting in Eq. (4-32) gives
a,b, a,b,
cos 8§= ——————— + aaee ee

Va’a Vb’b Va’a Vb’b

2d
component

Ist
component
110 ECONOMETRIC METHODS

that is,
a’b
cos 9 = ————— (4-33)
va‘a Vb’b
where it is understood that we take the positive square roots to indicate length.
There are two important special cases of Eq. (4-33). When a and b are
linearly dependent, we may write
b=da
where A is some appropriate scalar. The right-hand side of Eq. (4-33) then reduces
to unity, giving @ = 0°. When a and b are at right angles to each other, 0 = 90°
and cos 6 = 0, giving a’b = 0. Conversely, when a’b = 0, 0 = 90°. Two vectors at
right angles are said to be orthogonal. Thus two vectors are orthogonal if and only
if
a’b = 0

Extensions to Three and Higher Dimensions


If we now consider real three-element vectors, then each vector corresponds to a
point in the three-dimensional space ®*. Any vector v in ®? may then be
expressed as a unique linear combination of an appropriate set of three linearly
independent vectors, which constitute a basis for R’. For instance, choosing
] 0 0
ey =*0 ei es = 10
0 0 ]
as basis vectors, a vector v’ = [3 —2 5] may be written

If we take just two of these vectors, say, e, and e,, then all linear combinations of
e, and e, constitute a vector subspace in ®*, namely, the horizontal plane, since
the third component in each spanning vector is zero. More generally, any two
three-element vectors, say,

1 5
a=|2 and li ==|)
3 1
span or generate a plane surface, as indicated in Fig. 4-4, by the plane containing
Oab.
The set of all real n-element vectors constitutes the space ®”. Each vector in
§” may be expressed as a unique linear combination of some appropriate set of n
linearly independent vectors. To see that the linear combination must be unique,
suppose that a vector v can be expressed as two different linear combinations of
the basis vector v,,¥,,..., v,, namely,
V=Ay,
+A +--- +A nn

and

Vi Piet terse [Link]


ELEMENTS OF MATRIX ALGEBRA 111

3rd
component
4

2d
component

Ist
component

Figure 4-4

Subtracting one equation from the other gives

Oe BV Wo Ge (A eye

But the basis vectors are linearly independent and so

ieepg oe Ba ee A i ne

and the representation is unique. If we take a set of k (< n) linearly independent


n-element vectors, these generate a subspace of ”, which is termed a hyperplane.
The dimension of this subspace is the number of linearly independent vectors
spanning the subspace. The parallelogram law of addition and the cosine law of
Eq. (4-33) apply to the general case of n-element vectors. We can thus conceive of
a set of mutually orthogonal vectors V,,V>,..., V;, if
viv, = 0 for alli, j,i + j

We are now in a position to complete the geometric treatment of the


least-squares problem. The matrix

X = [x, Kgs Xx]

consists of k n-element column vectors, where, by assumption, we have more


observations than variables, so that n > k. The columns of X span a subspace in
&”. The dimension of the subspace cannot exceed k and will only be equal to k if
the columns of X are linearly independent. We refer to this subspace as the
column space of X. It is highly unlikely that the y vector lies in the column space
of X. If it did, y could then be expressed exactly as a linear combination of the x
vectors, giving zero residuals at all sample observations. The general case is
depicted in Fig. 4-5, where y lies outside the column space of X.
112 ECONOMETRIC METHODS

Figure 4-5

Assigning arbitrary b’s to the x vectors gives

Xb [xj x. 2 XE ix boxe

by.
which is then a vector that lies in the column space of X. Choosing different b
vectors gives, in turn, different Xb vectors. To each such Xb vector there
corresponds a vector of residuals e, so that the equation

Y=Xbte
as Fig. 4-5 shows, gives y as the sum of two vectors, of which one, Xb, lies in the
space spanned by the columns of X and the other, e, lies outside that column
space.
We wish to choose the b vector so as to make the point given by the tip of the
Xb vector as close as possible to the tip of the y vector or, in other words, to
minimize the length of the e vector. This is achieved by making the e vector
perpendicular to the hyperplane generated by the columns of X. Thus e must be
orthogonal to any linear combination of the columns of X. We have

e=y
— Xb
and Xc is any arbitrary linear combination of the columns of X. Thus the
orthogonality condition gives

e’X’(y — Xb) = c'(X’y — X’Xb) = 0 (4-34)


Since ¢ is any arbitrary nonnull vector, this condition gives

X’y — X’Xb = 0
ELEMENTS OF MATRIX ALGEBRA 113

or (X’X)b = X’y (4-35)


which are the least-squares normal equations derived algebraically in Eq. (4.29).+

Summary on Vector Geometry

1. A vector space is a collection of vectors with the following properties:


a. If v, and v, are any two vectors in the space, then v, + v, is in the space.
b. If v is in the space andAis a scalar, then Av is in the space.
2. If the only solution to A,;a+A,b=0 is A, =0=A,, then a and b are
linearly independent vectors. Otherwise they are linearly dependent. In general
if the only solution to A,x, + Ax; +--+: +A,x, =OisdA, =A, =--: =A,
= 0, the n vectors are said to be linearly independent. If at least one A is
nonzero, they are linearly dependent.
3. A basis for ®* is any linearly independent pair of two-element vectors.
Likewise, a basis for ®? is any three linearly independent three-element
vectors.
4. Each vector in a space may be expressed as a unique linear combination of a
set of basis vectors, and the minimum number in such a set is the dimension
of the space.
5. The angle @ between vectors a and bis defined by
a’b
cos8 =
Va’a yb’b
6. Two vectors are orthogonal when a’b = 0 (@ = 90°).

4-4 SOLUTION OF SETS OF EQUATIONS

The next problem is how to solve Eq. (4-35) for the desired least-squares
coefficients b. From the original definitions the dimensions of Eq. (4-35) are as
follows: X’X is a square matrix of order k X k, and b and X’y are all k-element
vectors. Thus Eq. (4-35) expresses the X’y vector as a linear combination of the
columns of X’X, and b indicates the coefficients of that linear combination. If the
columns of X’X are linearly independent, they constitute a basis for R*, and any
k-element vector, such as X’y, may then be expressed uniquely in terms of the
basis vectors. In other words, Eq. (4-35) has a unique solution for the b vector.
The solution of Eq. (4-35) for b may be expressed in terms of an inverse
matrix. The meaning of an inverse matrix may be developed as follows. Let A be
a square matrix of order n and let the n columns of A form a linearly independent
set. Does a square matrix B of order n exist such that

AB =I? (4-36)
+ A vector such as e, which is orthogonal to every vector on the hyperplane generated by the
columns of X, is said to be normal to the hyperplane—hence the term normal equations.
114 ECONOMETRIC METHODS

The answer is yes. Letting b, denote the first column in B and equating first
columns on both sides of Eq. (4-36) gives the vector equation
Ab, =e, (4-37)
where e’, =[1 0 0 --- OJ. Since the columns of A are linearly independent,
the vector e, can be expressed as a unique linear combination of those columns.
Thus b, is uniquely determined. By a similar argument each column of B is
uniquely determined, and so there is a matrix B satisfying Eq. (4-36).
We shall see later in this section that if the n columns of A are linearly
independent, then so are the n rows. Then by a similar argument, a square matrix
C of order n can be found such that
CA = I (4-38)
for each row of C is uniquely determined as the coefficients of a linear combina-
tion of the rows of A. Thus Eqs. (4-36) and (4-38) are both true. Postmultiplying
Eq. (4-38) by B gives
CAB = IB=B
But
CAB = CI=C
using Eq. (4-36). Thus
C=B
Thus if the n columns (and rows) of A are linearly independent, a unique square
matrix of order n exists, called the inverse of A, and denoted by A~', such that
AA'=A'A=] (4-39)
If we assume that the k columns of X’X are linearly independent, then the
inverse matrix (X’X)~' exists. Premultiplying both sides of Eq. (4-35) by this
inverse gives

b = (X’X) ‘X’y (4-40)


This expresses the least-squares vector b in terms of the sample data incorporated
in X and y. Equations (4-35) and (4-40) are two equivalent ways of expressing the
vector of least-squares coefficients b. Two distinct issues arise with respect to these
equations. First, there is the numerical, or computational, question of how best to
compute b for given X and y. Second, there is a set of theoretical questions about
inverse matrices like (X’X)~', such as, how are the elements of the inverse defined
and what are the properties of inverse matrices? Our main interest lies with the
second group of questions, though we will give some illustrations of numerical
solution methods. We have defined the inverse (X’X)~' to exist when the columns
of X’X are linearly independent. This concept is intimately related to the concept
of the rank of a matrix, and it is to this topic that we now turn.

Rank of a Matrix
Consider any arbitrary matrix A of order m X n. The columns of A define n
vectors in ®”. Likewise, the rows in A define m vectors in ®”. Let r denote the
ELEMENTS OF MATRIX ALGEBRA 115

maximum number of linearly independent rows in A, so that r < m. When r is


strictly less than m, there may, of course, be more than one subset of row vectors
which are linearly independent. For example, suppose we have a matrix with four
rows (m = 4). It may be that rows 1, 2, and 4 form a linearly independent set and
that rows 1, 3, and 4 also form a linearly independent set, but that all four rows
are linearly dependent. In this case r = 3. Returning to the general matrix A, let
us form a new matrix A by taking any set of r linearly independent rows and
discarding the remaining m — r rows. A is then of order r X n. Let c indicate the
maximum number of linearly independent columns in A. Then c must also
indicate the maximum number of linearly independent columns in A. Each
column in A has r elements. Thus we have immediately that
CES
for any vector in ®’ may be expressed as a linear combination of r linearly
independent vectors.
Reversing this argument we might form a matrix A of order m X c by
retaining a subset of c linearly independent columns of A and discarding the
remaining n — c columns. Since r is defined as the maximum number of linearly
independent rows in A, it also denotes the maximum number of linearly indepen-
dent rows in A. But since each row in A has just c elements, we have
GG
Thus
r=c
that is, for any m X n matrix A the maximum number of linearly independent rows
is equal to the maximum number of linearly independent columns. This number is
defined to be the rank of the matrix, and we will denote the rank of A by the
symbol
p(A)
Example 4-7 Consider
ret o2eer ad
A=/1 0 1 1
Dae) Ase S
By inspection rows 1 and 2 are linearly independent; also rows 1| and 3 are
linearly independent, but row 1 + row 2 — row 3 gives the zero vector, so all
three rows are linearly dependent. Thus r = 2. Let us form a matrix A by
discarding the third row of A. Thus
ee Ag Tel 3 4
ae |ay ay|
Clearly, all pairs of columns of A are linearly independent. Thus c is at least
equal to 2. But it cannot exceed 2, for any column in A can be expressed as a
linear combination of a pair of columns. For example,
col3 = col1 + col2
col4 = col 1 + 1.5 col2
col1 = 3 col3 — 2 col4
116 ECONOMETRIC METHODS

and so on. Thus r = c = 2 = p(A). Alternatively, if we commence with the


columns of A, we cannot finda set of three linearly independent columns for
the relations stated above, for the columns of A also hold for the columns of
the full matrix A, as readers should verify for themselves.

It is obvious that the rank of a matrix cannot exceed the number of columns
or the number of rows, whichever is the smaller. That is,
e(A) < min(m, n) (4-41)
When p(A) = ™, we say that the matrix has full row rank, and when p(A) = n,
that it has full column rank, but, of course, in any specific case, row rank and
column rank are identical, and we speak unambiguously of the rank of the matrix.
Notice that it follows directly from the definition of rank that the rank of the
transpose of A is equal to the rank of A, that is,

p(A’) = p(A) (4-42)


for p(A’) = number of linearly independent columns (rows) in A’
number of linearly independent rows (columns) in A
p(A)
In the special case where A is a square matrix of order n and rank n, then A is
said to be nonsingular, and a unique inverse A~! exists, such that
AA“! =A-A=1],
When the rank ofAis less than n, A is said to be a singular matrix and its inverse
does not exist.
Returning now to the general case of an m X n matrix A, let us suppose
p(A) = r. Thus there is at /east one set of r linearly independent rows and at least
one set of r linearly independent columns. If necessary, rows and columns may be »
interchanged so that the first r rows and the first r columns are linearly
independent. The matrix may then be partitioned by the first r rows and columns:

Ai | Ay } r rows
a eae ee Se

A>, | Ax } m — r rows
ers es
r n-r
columns columns

Thus Aj, is a square nonsingular matrix of order r. Consider now the set of
homogeneous equations

Ax = 0 (4-43)
where x denotes a column vector of n unknowns. The equations are said to be
homogeneous because of the 0 vector on the right-hand side of Eq. (4-43). If the
equations read Ax = b, for b = 0, they are said to be nonhomogeneous. Clearly,

if x, is a solution to Eq. (4-43), then so is cx, for any scalar c,


ELEMENTS OF MATRIX ALGEBRA 117

and

if x, and x, are two distinct solutions to Eq. (4-43), then c,x, + Xx, is also a
solution.

Thus the set of solutions to Eq. (4-43) constitutes a vector space called the
nullspace of A. Our immediate concern is to establish the dimension of this
nullspace (that is, the number of linearly independent vectors which span the
subspace). Let us drop the last m — r rows from A and partition x conformably
with the columns of A. This gives

[Ai Aled =0 (4-44)


where x, contains r elements and x, the remaining n — r elements. This gives a
set of r linearly independent equations in n > r unknowns. Rewriting as
A,X, + Appx, = 0
and premultiplying by A;,', which exists since A,, is nonsingular, gives

x, = —AyApx, (4-45)
The x, subvector is arbitrary or “free” in the sense that we can specify the n — r
elements in x, at will, but for any such specification the subvector x, is
determined by Eq. (4-45). Using Eq. (4-45), the general solution vector to Eq.
(4-44) may be written

AGA
I tPF
-
(4-46)
X9
The matrix in Eq. (4-46) has n rows and n — r columns. The n — r columns are
linearly independent. This fact is guaranteed by the presence of the I,,_, sub-
matrix, whose columns are necessarily linearly independent. Thus Eq. (4-46)
expresses all solutions to Eq. (4-44) as linear combinations of n — r linearly
independent n-element vectors. But any solution to Eq. (4-44) is also a solution to
Eq. (4-43), for the rows that have been discarded from Ato arrive at Eq. (4-44)
are linear combinations of the rows of [A,, Aj]. Any discarded row may thus
be expressed in the form
e[A,, Ay]
where ¢’ is some appropriate row vector of r elements. Postmultiplying by x gives
e[A,, Ay]x=0
since x satisfies Eq. (4-44). Thus each solution x holds for the discarded rows, and
Eq. (4-46) defines the solution vector for Eq. (4-43). Thus the nullspace of A has
dimension n — r. This gives the important result that for an m X n matrix A with
rank r
Number of columns = rank + dimension of nullspace (4-47)
n=r+(n-r)
118 ECONOMETRIC METHODS

The nullspace is sometimes referred to as the kernel of A and its dimension as the
nullity. Thus the result may also be stated as
Number of columns = rank + nullity

Example 4-8 Consider


A x 0

xy
12a emai
Le 2 eee Reer itO (4-48)
2 4 Raa a 4
0
The rank of the matrix is seen to be 2 since rows | and 2 are clearly linearly
independent, as are rows | and 3, but all three rows are not linearly
independent, since
row1 + row2 — row3 = 0
Discarding the third row gives the set
x]
Loe 3 ae sl eee)
| 2a x3 =a eae)
x4
Columns | and 3 are linearly independent, so we rewrite Eq. (4-49) as
|a3 ie 241 X>
(ee ee5 ene E LS
Solving this pair of equations for x, and x, gives
Xp ie het
x3, = eS ONE
and the solution vector to Eq. (4-49) may be expressed as

cane

Veg ie
So | ||| (4-50)
0 ©
YW
vl
=

The matrix in Eq. (4-50) has two linearly independent columns, and so any
solution to Eq. (4-49) may be expressed as a linear combination of two
linearly independent four-element vectors. The solution vectors thus form a
two-dimensional subspace in R*. There are infinitely many solution vectors
since the vector |
.

sl on the right-hand side of Eq. (4-50) is arbitrary. Any


x . . . .

solution defined by Eq. (4-50) is also a solution to the initial set of equations
(4-48). This may be seen by noticing that each column vector in Eq. (4-50)
ELEMENTS OF MATRIX ALGEBRA 119

satisfies the equation discarded from (4-48), that is,


—2
[2 4 4 5] 0eae eo= ()

0
and
if
2

0
[2 4 4 5] as =0
2)

1
Since any solution vector x is a linear combination of these two column
vectors, then x satisfies the third equation in Eqs. (4-48). Since it already
satisfies the first two equations, it is a solution to Eq. (4-48). The nullspace of
A thus has dimension 2, which is equal to the number of columns in A minus
the rank of A. Each vector in the nullspace is orthogonal to each row in A.
There is a seeming element of arbitrariness in the partitioning that we
applied to Eq. (4-49) and also in the choice of the row of A to be discarded.
But this is apparent, not real. For example, suppose we partition Eq. (4-49) as

which gives
ee(| ec
Xg= 2x,+ 4x,
with solution vector
1 0
- 0 ee :
elas 6 ie (oy)
2 4
Equation (4-51) again defines the nullspace of the matrix A in Eq. (4-48). It
has dimension 2, and the columns in the matrix of Eq. (4-51) are linearly
independent. This, however, is the same nullspace as defined by Eq. (4-50),
for each column vector in Eq. (4-51) may be expressed as a linear combina-
tion of the column vectors in Eq. (4-50):

1 23 x

0 =i 1 +) 0 SeAieesOs
3 1 0 2 a8 l Nye
2

2 0 1
and
0 -2 4
1 =i 1 0 =>), ,=1,
a 1 0 +X 2 a l A,=4
2

4 0 1
120 ECONOMETRIC METHODS

Thus the nullspaces are the same. Discarding the first or the second equation
from A in the initial stage would also make no difference to the determination
of the nullspace.

We may note here a particular application of Eq. (4-47), which will be very
useful in the treatment of identification in Chap. 11. If A is m X n and has rank
n — 1, then the dimension of the nullspace of A is 1, that is, all solutions to
Ax = 0
lie on a single ray through the origin. Thus if

Xai Se end
is a solution, then so is
Ex! =" [cx aia eee nes)
for any constant c.
Result (4-47) also yields simple proofs of some important theorems on the
ranks of various matrices. We notice that the crucial matrix for the least-squares
vector in Eqs. (4-35) and (4-40) is X’X. The first important theorem states that

p(X’X) = p(XX’) = p(X) (4-52)


Let X be n Xk with p(X)=r. Then by Eq. (4-47) the nullspace of X has
dimension k — r. If m denotes any vector in this nullspace,
Xm = 0
Premultiplying by X’ gives
X’Xm = 0
Thus m also lies in the nullspace of X’X. Let s be any vector in the nullspace of
X’X. Then
X’Xs = 0
Premultiplying by s’ gives
s’X’Xs = (Xs)’(Xs) = 0
Thus Xs is a vector with zero length and so must be the null vector, that is,
Xs = 0
Thus s lies in the nullspace of X.
We have shown that X and X’X have the same nullspace and hence the same
nullity. Each matrix has k columns. Thus by Eq. (4-47) each matrix has rank r
since
Rank = number of columns — nullity
For the least-squares case X is n X k with k <n. Provided there are no exact
linear relations between the explanatory variables, X has full column rank,
and so
o(X’X) =k
Since X’X is a square matrix of order k, it is then nonsingular and the inverse
(X’X)~! exists.
ELEMENTS OF MATRIX ALGEBRA 121

To prove the rest of theorem (4-52) we merely note that p(X) = p(X’), and
the above proof immediately gives

p(XX’) = p(X’) = p(X)


Notice that XX’ is a square matrix of order n.(> k) so that even if X has full
column rank, XX’ is still singular.
Another important theorem on rank may be stated as follows. If A is any
m X n matrix with rank r, and P and Q are square nonsingular matrices of order
m and n, respectively, then

p(PA) = p(AQ) = p(PAQ) = p(A) (4-53)


that is, pre- or postmultiplication of A by a nonsingular matrix does not change its
rank.
To prove p(PA) = p(A), let m be any vector in the nullspace of A. Then
Am = 0
Thus
PAm = 0
and m also lies in the nullspace of PA. Conversely, let s be any vector in the
nullspace of PA. Then
PAs = 0
Since P is nonsingular, we may premultiply this equation by P~! to obtain
As = 0
Thus salso lies in the nullspace of A, and PA and A have the same nullity and the
same number of columns. Hence the ranks are the same.
To prove p(AQ) = p(A) we note that

p(AQ) = p(Q’A’)
= p(A’) by the above proof
= p(A)
and finally
p(PAQ) = p(A)
follows directly from the previous results.
Both previous theorems involve special cases of the multiplication of one
matrix by another. In Eq. (4-52) a matrix was multiplied by its transpose. In Eq.
(4-53) multiplication was by a nonsingular matrix. Our final theorem on rank
relates to the perfectly general case of the multiplication of one rectangular matrix
by another conformable rectangular matrix. Let A be m X n and let B ben X s.
Then
p(AB) < min[p(A), p(B)] (4-54)
that is, the rank of the product AB is less than or equal to the smaller of the ranks of
the constituent matrices.
122 ECONOMETRIC METHODS

If x denotes any vector in the nullspace of B, then


Bx = 0
and so
ABx = 0
Thus xalso lies in the nullspace of AB. But this time we cannot go in the opposite
direction and prove that
ABy = 0 implies By =0
Thus all we can say is that the nullspace of B is contained in (or is a subspace of)
the nullspace of AB. Therefore, we have
Dimension of nullspace of B < dimension of nullspace of AB
Since B and AB have the same number of columns, it then follows from Eq.
(4-47) that
p(AB) < p(B)
By the usual trick with transposes,
p(AB) = p(B’A’) < p(A’) = p(A)
and so Eq. (4-54) is proved.

Summary on the Rank of an m X n Matrix A

1. The maximum number of linearly independent rows is equal to the maximum


number of linearly independent columns. This number is the rank of the
matrix, denoted by p(A).
2. p(A) < min(m, n).
3. p(A) = p(A’).
4. If p(A) = m = n, then A is nonsingular and a unique inverse A“! exists.
5. n = p(A) + nullity of A where the nullity of A is the dimension of the
subspace containing all vectors x which are solutions to Ax = 0.
6. p(X’X) = p(XX’) = p(X).
7. If P and Q are nonsingular matrices of orders m and n, respectively, then
p(PA) = p(AQ) = p(PAQ) = p(A).
8. p(AB) < min[p(A), p(B)].

The Inverse Matrix


It is now time to return to the topic of matrix inversion and develop some of the
properties of inverse matrices. We will also see how to compute inverse matrices,
though this is a tedious and inefficient procedure for the numerical solution of
equations. The procedure does, however, shed light on the theoretical properties
of the inverse.
Let A denote a square matrix of order n. The condition for the inverse to exist
may be stated in several equivalent ways:

1. A is nonsingular.
2. A has rank n.
ELEMENTS OF MATRIX ALGEBRA 123

3. The n rows of A are linearly independent.


4. The n columns of A are linearly independent.

To study the jo of the inverse matrix let us begin with the 2 x 2 case.
Denote A and A~' as follows:
S| a 1 a "| re a1 a|
42, 422 Oo OD?
So far we have regarded matrices mostly as collections of vectors and paid little
attention to the individual elements. The standard notation is to use the first
subscript of an element to indicate the row in which that element appears and the
second subscript to indicate the column. The definition of the inverse gives the
general equation
AA~! =] (4-55)
Specializing this equation to the 2 X 2 case and taking just the first column from
each side of the equation gives
an eae |e
Qo, 40 || @21 0
Treating the elements of the inverse as unknowns, the solution of this pair of
equations gives
a2
MNS 11422 a ka 12421
aa
— 41
Sel sk 11422
Waa ala
12421
Similarly, equating the second columns in Eq. (4-55) and solving gives
— 412
[2.2 ata 1242)
11422
a)
22 a
a SS —_—————

411492 — 41242)

Thus the inverse has been derived as

ee 1 Eee
aitEtLD Ay | 4-56)
— a (
=
A\14n7 — 4424) | ZI 11

and it may readily be checked that indeed AA! = AA”! =I. Each element in
A~' is a function of the elements in A, and even for the 2 X 2 case certain
important features of A~' are apparent. First, each element in the inverse has a
common divisor, namely, a,,45) — 4,74 ,. This is a function of all the elements in
A. It is a scalar quantity and is defined as the determinant of A. For the 2 Xx 2 case
we thus have .

detA = JA| = 4,142) — 4942, = LY + 1a 42g (4-57)


a,B
124 ECONOMETRIC METHODS

The two expressions on the left of Eq. (4-57) are alternative ways of indicating the
determinant. The final expression on the right means

Ds a Q\q42B 7.
a,B
sum of all possible products of the elements of A, taken two at a time,
with the first subscript in natural order 1,2 and a, B indicating all
possible permutations of 1,2 for the second subscript, each product
term being affixed with a positive (negative) sign as the number of
inversions of the natural order in the second subscript is even (odd).
There are only two possible permutations of 1,2, namely, 1,2 itself and 2, 1.
There is one inversion of the natural order in 2, 1 since 2 comes before 1. Thus the
terms in the expansion are simply
411422 — 4129)
The numerators of the elements in A~' could have been produced by the
following two rules:

1. For each element in A, strike out the row and column containing that element
and write down the remaining element prefixed with a positive or negative
sign in the pattern

This gives the matrix


| 449 |
mor a1)
2. Transpose the matrix obtained in rule | to get
| Q22 ne
— 42 ai

Let us try to apply these rules to the 3 X 3 case. Now we have


Ty 42 G3
A= |421 “422 423
G3, 377 G35
By extension of Eq. (4-57) we define the determinant of A as

a,B,y
There will be 3! = 6 terms in the expansion, since that is the number of possible
permutations of 1, 2, 3. Half will have a positive sign and half a negative sign. The
explicit expression is
JA] = @)14y2433 + 41747343) + G1347)43) — 441473432 — 412471433 — A)349)45,
(4-59)
As a check on the signs we may notice, for instance, that in the third term in Eg.
ELEMENTS OF MATRIX ALGEBRA 125

(4-59) the order of the second subscripts is


Belen
which contains two inversions since 3 comes in front of both 1 and 2. The final
term ;
a 258
yields three inversions (3 before 2 and 1 and 2 before 1).
Expressions (4-58) and (4-59) correctly define the determinant of a third-order
matrix. The expression is already so cumbersome that the generalization to the
nth-order case would be unpleasant. However, we shall derive below a more
tractable expression for the determinant.
The numerator rule for the second-order case, however, does not extend to
the third-order case without modifications. If we strike out the row and column
containing, say, a,,, we are now left with the 2 x 2 submatrix

|
Gy O23
Mann 433
rather than a scalar element. We, in fact, replace a,, with the determinant of this
submatrix, appropriately signed, and similarly for the other elements. The general
rules for determining the elements of A~! in the 3 x 3 case may now be stated.
Let M;,; be the determinant of the 2 x 2 submatrix obtained when row i and
column j are deleted from A. M,, is termed a minor. Further define
Pe Nery
C,, = ( 1) M;;

C,; denotes a cofactor and is simply a signed minor. Thus the sign of M,; does not
change if i + / is an even number and does change if that sum is odd.
The rules then become as follows:

1. Form a matrix in which each element (a;,;) is replaced by the corresponding


cofactor (C;;).
2. Transpose this matrix. The result is sometimes referred to as the adjugate or
adjoint matrix.
3. Divide each element in rule 2 by |A|. The result is A~'.

For the third-order case

A5n G93).
ir ~ |432 agi 447433 ~ 473439

a a
Ces = i ast = — (4,43) — 4)243;)

and so on, and


1 Cr Gi Gi
AU! = IA] Cro Cy Cy (4-60)
126 ECONOMETRIC METHODS

Finally, we may note an alternative expression for the determinant of A. Return-


ing to Eq. (4-59) and collecting terms in the elements of the first row gives
JA] = 4 (492433 — 493432) + G12 (— 491433 + 4343,) + 443 (421432 — 4y243,)
Using the definitions of cofactors just given, we then have
JA] = a Cy + ayCi + 4)3C)3 (4-61)
This defines |A| as a linear combination of the elements in the first row, each
element being multiplied by its cofactor. This definition is clearly not unique. |A|
may be expressed in terms of the elements of any row (or column), provided that
in each case the elements are multiplied by the corresponding cofactors. Readers
should satisfy themselves by direct substitution that any other similar expansion
gives the same result as Eq. (4-59).
These rules for the 3 X 3 case have been rather plucked out of the air. Let us
check that they work for a numerical example before continuing to the nth-order
case.

Example 4-9
1 Sa
A Ne lite? eee
Died ad
Replacing each element by its minor gives the matrix
Fi 4 E | E a
ans m1 5 2 4
E ‘| | ‘ : 3 -|-1 = z
ic 5d OMS hia apeeg Ns he
erst lever plees
Jami Lot fe
Signing the minors gives the matrix of cofactors as
On=3 0
ec 2
— eal
Transposing gives the adjugate matrix
6 Ie 5)
=O eteS 3
0 ot onal
Expressing the determinant of A in terms of the elements in the first row gives
JA] = 4 Cy, + ayCyp + 4)3C\3 = 1(6) + 3(—3) + 4(0) = —3
Thus the inverse matrix is
ELEMENTS OF MATRIX ALGEBRA 127

It is easily checked that AA7' = AA =I.

For the nth-order case the rules for obtaining A~! are essentially those
already stated for the third-order case. The determinant is defined as

|A| = De a Aig rp nt Any (4-62)


EBLE. v
or alternatively
|A| = a,,C,, + aC. + --- + a,,C,, for anyi = l,...,n
or (4-63)
JA] = a, ,C); + 4.,G, +--+ + 4,,C,, for anyj = 1,...,7
The cofactors are now the signed minors of matrices of order n — 1, and the
inverse matrix is
Ci Cn Gui
]
AUl= jAl Cz Gp Ge (4-64)
In 2n Cun

Properties of Determinants
The following properties are stated for the determinants of nth-order matrices,
but they will often be illustrated for the 2 x 2 case. To economize on space,
proofs will not always be given.

1. |A’| = |A| even if A is not symmetric.

: |e
b= ad ZS be Jal =|,
Ge eG

2. If B is obtained from A by interchanging any two rows (or columns) of A,


|B] = —|A|.
|B = Ce ee as as
i d= ob ad
an
: b|= a
Suppose in an n X n matrix we interchange the first two rows. Let a;; denote
the i, jth element in the original matrix and b;, the i, jth element in the new
matrix. A term such as
D447 p93, nt Any

in the expansion of |A| thus becomes


5.51253, ie OS

in the expansion of |B|, where


by;
=4
i} fea ile2e
ek oh
by,= Qa);

b.,=4,, for all j; i = 3,4,...,n


128 ECONOMETRIC METHODS

The numerical values of these two terms are identical; the crucial question is the
sign. To determine the sign of the second term, the first subscripts must be put in
natural order and the number of inversions in the second subscript determined.
This gives
bi pb.4b3y 2a bs
and, compared with the corresponding term in |A|, one inversion has been
introduced or removed, so that this term (and each and every term) changes sign
in |B| as compared with |A|. If we interchange rows i and j, which are separated
by, say, r rows, reordering the b elements in any term to put the first subscripts in
natural order will involve 2r + 1 changes, where each change introduces a new
inversion in the second subscript or removes an existing inversion. This is
illustrated below, where only the first subscripts on the b’s are shown.
ee
15:4 1O;42 aaa 5-15;
r elements
Since 2r + 1 is an odd number, the sign of the term changes, and so |B| = —|A].
Property 1 then ensures that interchanging any two columns will also change the
sign of the determinant.
3. If a matrix has two or more identical rows (or columns), its determinant is zero.
a b
=ab-—ab=a
aan
From property 2, interchanging identical rows would change the sign of the
determinant. But the new matrix is identical with the old, and so its determinant
is unchanged. This gives
|A| = -|A|
so that
|A| =0
4. Expansions in terms of alien cofactors vanish. By this is meant an expression
such as
jn~in

where the elements of row j are multiplied by the cofactors of the elements of
row i.
This is exactly the expression we would obtain for the determinant of a
matrix whose rows i and / are identical. By property 3, that determinant is zero.

5. If B is formed from A by adding a multiple of one row (or column) to another


row (or column), the value of the determinant is unchanged.

|B| = at+Ac b+dd


=(a+Ac)d-(b+dAd)c
c d
ad — be + A(cd
— cd)
a Fl al
ape
ELEMENTS OF MATRIX ALGEBRA 129

Suppose row i of B = rowi of A + ) - row of A. Expanding |B| in terms of


its ith row gives

|B| = (a, zt AG}, )Cy iv (a; ete Nan) Go Sine? ae (4, as Aaj, )C;n

as (aC, TG pO eo ct ZnO.) + ACC, 7 A joCin Rarer inca)


=,(Al
since the coefficient of A is an expansion in terms of alien cofactors, which
vanishes.

6. If the rows (or columns) of A are linearly dependent, |A| = 0, and if they are
linearly independent, |A| + 0. If the rows of

uae eb
a ie a
are linearly dependent, there exist nonzero scalars \,, \, such that
A\,a+A,c=0
\,5+A,d=0
Thus

C= A,
ie and d=-— AAe

and A may be written


Poe b
ie ne 1
where X = —X,/A,, and so |A| = A(ab — ab) = 0.

In the general case if row i is a linear combination of certain other rows,


subtracting that linear combination from row i will produce a zero row. Subtract-
ing the linear combination is merely a repeated application of property 5 and so
leaves the determinant unchanged, but that determinant is zero since the process
has ended with a matrix containing a zero row.
If the rows (columns) of A are linearly independent, there is no way to
produce a zero row (column) and |A| + 0. Thus nonsingular matrices have
nonzero determinants. If this were not so, the inverse matrix defined in Eq. (4-64)
would not exist, since each element is divided by |A|. Conversely, singular
matrices have zero determinants.
This result also provides a means of checking on the rank of low-order
matrices. If A is m X n and has rank r, then there must be at least one square
submatrix of order 7, which is nonsingular and thus has a nonzero determinant,
and all square matrices of order r + 1,7 + 2,..., have zero determinants. Thus
we have the following alternative definition of rank: .
Rank of m X n matrix = order of largest nonvanishing determinant
130 ECONOMETRIC METHODS

Example 4-10
reserskveaye
A=]1 2 ] 1
2 4 -6 -—10
The rank must be at least 2, since although

ee 2aRe
f 3|=0
there are plenty of nonvanishing second-order determinants that can be
formed from the elements of A. For example,

; el
1 3
=-12 Z
4
]
_y|7 -4
and so on. Notice that these are the determinants of second-order sub-
matrices obtained by deleting any one row and any two columns from A.
There are four possible third-order determinants to evaluate. Deleting the
fourth column,

2 3 0 0 2 i 2
] ie el 2 ] = 2) |= 0
2 a>) —6 Z A os 6
The first step in this evaluation has been to subtract row 2 from row 1, which
by property 5 does not alter the value of the determinant. This gives an
expansion, using Eq. (4-63), in terms of the first row, which now contains just
a single term. To evaluate
1 3 4
1 ] 1
7a Comet ()
we might subtract row 1 from row 2 and we also subtract twice row 1 from
row 3 to get
1 3 4
0 outiedks = lias Rl=°
0 -12 -18
The two other third-order determinants may similarly be seen to be ZeTO, SO
that p(A) = 2. Alternatively, we might have spotted that

row 3 = 6- row2 — 4: row l


which establishes that the rank cannot be 3, without a need of evaluating
third-order determinants. This matrix also illustrates another important point.
Looking at the square submatrix formed from the first three columns of A, we
have already shown that its determinant is zero, and inspection of the
second-order determinants within it shows that its rank is 2. Its three rows are
connected by the relationship stated above, but the same relationship does
ELEMENTS OF MATRIX ALGEBRA 131

not hold between the three columns. The linear dependence between the three
columns is expressed by
2-column | — 1 - column2 = 0
or )
2- column| — 1- column2 + 0- column3 = 0
The important point is that a set of vectors is linearly dependent even if some
(but not all) of the coefficients in the linear combination are zero.

7. The determinant of a triangular matrix is equal to the products of the diagonal


elements.
a 0| Gap
0 q|— 4
Gaudi

A triangular matrix may be lower triangular, as in

Zee 0 0
ee ay, Ay O 0
Cae C323 0
Gn) a2 an3 nn

or upper triangular, as in

Gn. 2419 413 Dip


aa 0 O90) 433 22n
0 0 033 43n
Re es

a, 0 0
|A| = 4);|432 433 0
an2 an3 ann

Expanding the new determinant by its first row and repeating the process n times
gives
|A| = 19970 °° * Any

Expanding |A*| successively by the first column similarly gives


|A*| = 4);4y) °°° @ nn
Two special cases of this result follow directly:

«The determinant of a diagonal matrix is simply the product of the diagonal


elements:
132 ECONOMETRIC METHODS

e The determinant of the unit or identity matrix is unity:


a, 0 0
|A] =| 9 eee
Oe ay 0 |= aay Gan

10 0
|AJ =|0 1 0;=1
eceree '

8. Multiplying any row (column) of a matrix by a constant multiplies the


determinant by X. Multiplying every element in a matrix by X multiplies the
determinant by X".

These properties follow directly from the definition of the determinant in Eq.
(4-62), where it is seen that each term in the expansion is the product of n
elements, one and only one from each row and column of the matrix.

9. The determinant of the product of two square matrices is the product of the
determinants.
|AB| = |A| + |B|
This rule is only of interest when A and B are both nonsingular. If either is
singular, AB is singular and both sides of the equation are zero. If A is
nonsingular, repeated applications of property 5, that is, additions of multiples of
rows and columns, can produce a diagonal matrix D, such that |D| = [A].

Example 4-11

el 3742] with |A] = —2


Subtract 3 - row | from row 2 to get

Then add row 2 to row |to get


lo 3
02 a2

Die OT 2
4 with |D| = —2 =
|A|
If these steps are performed on the matrix AB, the result is a matrix DB with
|AB| = |DB| by property 5. This statement in general requires that only row
operations have been performed on A to obtain the diagonal matrix D. This
is always possible. The first step in the example is equivalent to premultiply-
ing A by
ELEMENTS OF MATRIX ALGEBRA 133

and the second step to a further premultiplication by

E, = L. 1
; 6 1|
The sequence of operations is then described by premultiplication by a single
matrix
Saleen |
F-%F,-| 72

so that FA = D, as the reader may verify. By property 8,

|DB| = d,,d,. --- d,,,|B|


= |D| > |B
= |A| > [BI
Thus
|AB| = |A| - |B

Properties of Inverse Matrices

1. (AB)~' = B~'A™! provided A and B are each nonsingular.

The simplest proof is to multiply AB by the suggested inverse and see that the
unit matrix results, since we already know that the inverse matrix is unique:
ABB 'A~! = AJIA~! =
and similarly,
B-'A~'AB =I
This technique is sometimes useful in deriving inverse matrices, namely, guess at a
plausible inverse and check by multiplication to see whether it works. The above
result extends readily to products of three or more matrices. Thus
(ABC) | = C7'B-'A7!
The warning must again be inserted that this result only holds when the
constituent matrices are nonsingular. Students occasionally produce “‘ miraculous”
proofs by applying this theorem to rectangular matrices.

2(At et =A
that is, taking the inverse of the inverse reproduces the original matrix.

From the definition of an inverse,

(A“!)(AT')' =]
Premultiplying by A gives the required result.

3. (A) 1 =A)
that is, the inverse of the transpose equals the transpose of the inverse.
134 ECONOMETRIC METHODS

We have

AAS I]
Transposing,
(A-')'A’ =I]
Postmultiplying by (A’)~!,

(A“TYA(A)' = (A)
Thus

(Anty = (a)
]
4. |A|A7'| | = ——
iA]
that is, the determinant of A“' is the reciprocal of the determinant of A.

This follows directly from properties 7 and 9 of determinants, for

AA t=]
gives

JAI? [AT = 1
5. The inverse of an upper (lower) triangular matrix is also an upper (lower)
triangular matrix.

We merely illustrate this result for a lower triangular 3 x 3 matrix:

a, 90 0
A= la, ay 0
43, 432 33
By inspection it is seen that three cofactors are zero, namely,

&: 0 0 yi lit-Owee 10 a, O
Cy re Az, 33)” C3 = as, OP Cy. rary an 0

Thus

Gowen 0
A=! = |A| Ci Cy 0

C3 G3 Cay

6. The inverse of a partitioned matrix: If

ne A 11
| A |
Ax, Ax
ELEMENTS OF MATRIX ALGEBRA 135

where A,, and A, are square nonsingular matrices,

A-'= Bi, — BA, A>) (4-65)

— Ax A,B, Ay " Ax A> By AA


where B,, = (Ay, — A,2A3)'A>,)~', or alternatively,
A-'= Ai a0 Ay 'ABy A Ai —Aj'A,, By (4-66)

—B yA Ai By
where

B, = (Ay Z Aer Ad AG)

These formulas are frequently used. The first form, Eq. (4-65), is the simpler
if we are interested in an expression that involves just the first row of the inverse.
Conversely, Eq. (4-66) is the simpler for expressions involving the second row.
The derivation of the formulas is straightforward but tedious. Let
Aahes =|
B,, By
where the B,,; submatrices have the same dimensions as the corresponding A,;
submatrices. Postmultiplying A by A”! gives the matrix equations
A,B, + A,B, =1
A, By. + A,B, =0
A,B), + A.B), = 0
A,B, + Ax
B =I
where the unit matrix has been partitioned conformably with A. The third
equation in this set gives
B,, = —Ax A,B, (4-67)
Substituting this in the first and solving for B,, gives
F. =1
B,, = (Ai, a A, Ay Ad) (4-68)
A similar treatment of the second and fourth equations yields
Bi) = —Aj'A,) By (4-69)
=1
and B,, = (Ay 7 A») Ai Ai) (4-70)
These four expressions are seen to constitute, respectively, the first and second
columns in the two alternative formulations of A7'.
To derive the remaining columns in Eqs. (4-65) and (4-66) we multiply out
A~'A = I to obtain
B,,A\, + By,A2, = 1
B,,A,. + B,,A. = 0
B,,A,, + B.A, = 0
B,,A,, + B,A =I
136 ECONOMETRIC METHODS

The second equation in this set gives


B= — BAA (4-71)
Substituting this and Eq. (4-67) in the fourth equation of the first set above and
solving for B,, gives
B,, = An Ds A5'A>,B,,AiAd (4-72)
These two expressions complete the second column in the definition of A~! in Eq.
(4-65). The third equation of this set and the first equation of the previous set
yield
B= —B,A,,Aq' (4-73)
and By, = An’ + Ay'ApBy AyAi (4-74)
which completes the first column in Eq. (4-66).

7. The inverse of a block diagonal matrix: Let A be


ne A m0
Cont ORGEASS
where A, A,,, and A,, are all square matrices. If A is nonsingular, then so are
A,, and A,, since each has linearly independent columns (rows). Then
Ate gee O
A-l=
0 ASS
(4-75)
This is merely a special case of property 6 or, alternatively, it may be seen
directly as the inverse since AA ' is clearly I. A special case of this result is the
inverse of a diagonal matrix. If

a,, 0 0
A=|090 ay 0
jot pee hee

1
— 0 0
a);
‘A 1
Aol= 0 = 0
A)

0 0
a

8. The inverse of a Kronecker product: The direct or Kronecker product of two


matrices A and B is defined as

a4,B 4,B --- a,B


A ® B =, a>,B aB age? 2. a>, B (4-76)

Ani{Bo a7B “Ann B


ELEMENTS OF MATRIX ALGEBRA 137

In this definition A is a general matrix of order m X n, and likewise B can be


a rectangular matrix of any order, say, p X q. In this case A @ B is of order
mp X nq. Suppose A is square of order m and nonsingular and that B is square of
order p and nonsingular. The Kronecker product A @ B is square of order mp
and nonsingular. Its inverse is given by

(A®@B) '=A-'@B"! (4-77)


The proof may be obtained by multiplying out. The right-hand side of Eq. (4-77)
is

C,,Bo' CyB! a Cn Bo!


¥ x l 2 = e
RAREST NEBR ene fa) ois qeGeaPhe
Cy. Be Ci, B* Cube.

Multiplication by Eq. (4-76) yields the identity matrix.

9. Determinants of partitioned matrices: We sometimes need to express the de-


terminant of A in terms of the determinants of submatrices. We begin by noting
that

0 I im |Ai,| (4-78)

for if we evaluate the determinant on the left-hand side by expanding in terms of


the elements of the last row, the only nonzero term is the last one, which is unity
multiplied by a determinant of the same form, except that the order of 1 has been
reduced by |. Proceeding in this way the result follows.

A block diagonal matrix may be expressed as the product of two simpler


block diagonal matrices, namely,
if Al mow An) ¥0,|(Mind
S10 TA OME TO. Ay.
Applying property 9 of determinants and also Eq. (4-78) gives
|A| = |Ai,| * |Adol (4-79)
Now consider

An Ala tan
A A
(4-80)
This follows from the same argument used to establish Eq. (4-78). We can now
find the determinant of a block-triangular matrix.

bts Ay Ai “(4 0 Noe! )


|} 0 A», 0 Ax»|l 0 I
Thus
|A| = |Ay,] * [Az9| (4-81)
138 ECONOMETRIC METHODS

This is a matrix generalization of property 7 of determinants, namely, that the


determinant of a triangular matrix is the product of the diagonal elements. The
final step is to establish the determinant of a general partitioned matrix

a A1
| Aa
Az, Ax)
where A,, and A,, are square and nonsingular. Define

B, =
[Av ae
AGt and B, =
I
be
|
|! I | : ee I
Then
BAR|!A,, — Ap 12442249]
Az JA 0

and since |B,| = |B,| = 1,

|A| = [Ao] * [Ay — AyAnAs,| (4-82)


An alternative expression may be derived in a similar fashion as

|A] = |Ay,| * [Az — A Ay Aj! (4-83)

Cramer’s Rule
This inordinately long section on the solution of equations may be rounded off by
the derivation of Cramer’s rule for the solution of a set of n nonhomogeneous
equations in n unknowns. The set of equations may be written

Ax =b (4-84)
where, by assumption, A is a square known matrix of order n and nonsingular, x
is a vector of n unknowns, and b is a known n-element vector. There is an
unfortunate clash of notation between conventions in algebra and conventions in
Statistics. The normal equations for the least-squares vector are
(X’X)b = X’y
Here (X’X) is a known matrix and X’y a known vector, each depending on the
empirical data in a given problem, and b denotes a vector of unknown coefficients.
It is too late in the day to resolve this conflict; the student must maintain
sufficient intellectual agility to interpret the symbols according to the context.
Returning to Eq. (4-84), the solution vector is written

x=A'b
Substitution for A~! from Eq. (4-64) gives

Cy Cy Cu |} 5
pya | Crp Cy Cro || 2.
TA AM et eee eee
Cin Gy Cin || On
ELEMENTS OF MATRIX ALGEBRA 139

Thus
1
xy = ja 2G WD5C3j oh ve C1)

The expression in parentheses is seen to be the evaluation of the following


determinant by the elements in the first column:’

by ayn a3 Jee yy,


B,Cy, + BoC + +++ + BG, =], a2 493 8° Ay
b,, an2 an3 ann

Similar results hold for each element in x. The ith element is thus the ratio of two
determinants, the denominator being the determinant of A and the numerator the
determinant of the matrix obtained from A by replacing the ith column of A by b
and leaving the other n — 1 columns unchanged.

Example 4-12 Solve the system


BX pr axa = Xa 15
Xj Oko ek SO
6X5 29X54, 1X2 S28
This may be accomplished by several methods.

(a) Cramer’s rule First calculate |A|, and it is helpful to expand in


terms of the elements of a column, say, the first. Thus
2 ae 1
|AJ=|16 -35 2|=
1
2 -3rea 2 |-1]alae2 l [+6]_3
4 -]
v |
= 2(—13)
—9 + 6(5) = —5
Then
15 2 |
~5x,=|-5 —3 2]= 15(—13) + 5(9) + 28(5) = —10
28 5 1
so that x, =2
2ts a a us
ee et rd = -13|) 2|— 5/2 |=28]4 a
6) 28 1
= —15(—11) — 5(8) — 28(5) = —15
so that x, =3
and
2 4 15
=539= 1 23.) = 5 = 15|2 nelle : + 28|7 =
6 5 28
15(23) + 5(=14) + 28(—10) = —5
140 ECONOMETRIC METHODS

giving
X3 = ]

Thus the solution vector is x’ = [2 3 1].

(b) Calculation of A~' From the calculations already completed in (a),

El ek 11 8 =a)
° 23 14 -—10
Thus

:{-13 -9 — 5]f 15 ;{-10] [2


Ke cis|e Miia 8s ae ae |eae ti le
23 14-10) 28 SS veal
Methods (a) and (b) are essentially slightly different ways of laying out the
same set of tedious calculations. A computationally much more efficient
method, not just for small systems but more especially for large systems, is
the elimination method.

(c) Elimination method Lay out the system in matrix form as


Z Saale, 15
ie S 2X nee
6 5 1} | x; 28
In the first step we produce zeros in the second and third positions of the first
column by subtracting one-half the first equation from the second and three
times the first from the third. This gives
2 40 b= ne 15
ee 2 SAXpile [a 129
Oia, “ Xx, =a)
Next we produce a zero in the third position of the second column by
subtracting { times the second equation from the third.

2 4a xy 15
OT 2: Ieee oleae
0 0 0.5 }| *3 0.5
This gives an upper triangular system, which is solved for the x’s by back
substitution. The third equation gives directly
x3, =1
The second equation
= 9X5 4S 22K = DNS or Sky a5
then gives
x, =3
ELEMENTS OF MATRIX ALGEBRA 141

and the first equation

gives
x, =2
In the elimination method the inverse A~' is never calculated at all. The
calculations are fast and simple compared with the first two methods, but it
does not shed light on the theoretical properties of the inverse.

4-5 THE EIGENVALUE PROBLEM

The previous section was concerned with the solution of the set of equations
Ax = Db (4-84)
This section is concerned with solutions of
Ax = Ax (4-85)
where A is a known square matrix of order n, x is an unknown n-element column
vector, and A is an unknown scalar. This problem will arise in a number of places
later in the book. It is known as the eigenvalue problem. In contrast with Eq.
(4-84) there are now two unknowns, a vector and a scalar. Solutions will come in
pairs; to each A there will correspond an x vector. The A’s are known as
eigenvalues, latent roots, or characteristic roots and the x’s as eigenvectors, latent
vectors, or characteristic vectors.
For n = 2, Eq. (4-85), written out in full, becomes
(a,, —A)x, + ax, = 0
yx, + (ay — A)x, = 0
which may be put back in matrix form as
(A —XI)x =0 (4-86)
Equation (4-86) is equivalent to Eq. (4-85) for any n. If the matrix A — AI is
nonsingular, the only solution to Eq. (4-86) is the trivial x = 0. Thus for a
nontrivial solution to exist, the matrix must be singular or, in other words, have a
zero determinant. This condition gives
|A —AI| =0 (4-87)
which is known as the characteristic equation for the matrix A. This gives a
polynomial equation in the unknown A. Each root or eigenvalue A; may be
substituted back into Eq. (4-86) and the corresponding eigenvector x, obtained.
For the 2 X 2 case it is easily seen that the characteristic equation is
= (ay, + ay )A + (441422 — 442431) = 0 (4-88)
with roots
142 ECONOMETRIC METHODS

In the special case of a 2 X 2 symmetric matrix, a, = a), the roots become


]
A= 5|(en tides et Vaan = Gaal a 43 |
and since the content of the square root sign is the sum of two squares, the roots
are necessarily real for a real symmetric matrix. Notice also that the characteristic
equation may be written
(A; ~A)(A, —A) =H’ - (A, +A,)AF rA,A, = 0
Comparison with Eq. (4-88) shows that
Sum of roots = A, + A, = a), + ay)
= trace (sum of diagonal elements of A) (4-89)
Product of roots = A,A, = a,,ay) — a),a>, = |A| (4-90)
These two properties hold true in the general nth-order case, as does the previous
result on real roots for a real symmetric matrix.

Example 4-13
4 x
eel
Thus

(a-an= [45% a
and the characteristic equation is

’-— 5A =0
with roots
A, =5 and A, =0
For A, = 5, substitution in Eq. (4-86) gives

9) =A ce =0=> x, = 2X5

Thus one element in the eigenvector is arbitrary, and so ifx satisfies Eq.
(4-86) for some A, then so does cx, where c is an arbitrary constant.
It is
conventional to normalize the vector by setting its length at unity, that
is,
making

Xp xe =]
which, with x, = 2x,, gives
2

xX, =
v5 corresponding toA, = 5
ee
5
ELEMENTS OF MATRIX ALGEBRA 143

Similarly, it may be shown that

Mo
v5
Xo =

2.
v5
is the eigenvector corresponding to A, = 0.
It is seen that the eigenvectors are orthogonal, xx, = 0. If we assemble
the eigenvectors in a matrix X,

ena Agee
v5 v5
ca ae as
KS — =

v5 v5
and then form X’X, we obtain the result that

X’X = XX’ = i a (4-91)


We will derive this result for the general case below. Forming the matrix
product X’AX gives
2 ] 2 1
5) S| aa es thavoaaa 5 0
ncaa ll 2 |elo ole
ee Gee
The diagonal matrix on the right-hand side of Eq. (4-92) displays the
eigenvalues 5 and 0 on the main diagonal.

Properties of Eigenvectors and Eigenvalues of a Real Symmetric Matrix of


Order n
In statistical applications we are mainly concerned with symmetric real matrices.
Properties with an asterisk apply specifically to real symmetric matrices; those
without an asterisk apply to real nonsymmetric as well as to symmetric matrices.

1.* The eigenvalues are real.

Suppose we have a complex eigenvalue A + ip, where i denotes y—1, and a


corresponding complex eigenvector x + iy. Then
A(x + iy) = (A + ip)(x + iy)
Multiplying out and equating real and imaginary parts gives
Ax = Ax — py
144 ECONOMETRIC METHODS

and Ay = ux + Ay
Premultiplying the first equation by y’ and the second by x’ gives
y’Ax = dx’y — py’y
x’Ay = px’x + Ax’y
When A is symmetric, y‘Ax = x’Ay (a scalar equals its transpose). Subtracting the
first equation from the second then yields

0 = w(x’x + y’y)
Since the eigenvectors must be nontrivial, x’x > 0 or y’y > 0 (or both), so
pw =0
that is, there cannot be a complex eigenvalue. Real eigenvalues in turn generate
real eigenvectors, that is, y = 0.

2.* Eigenvectors corresponding to distinct eigenvalues are pairwise orthogonal.

If x,,x, denote the eigenvectors corresponding to A,, A,, then


Ax, = A,X; => x, Ax, = A,x3x,
and Ax, = A,x, = xj Ax, = A>xix,
The symmetry of A gives
x, Ax, = x,Ax,
Thus
XXX, = A>xix,
If A, + A3, this last equation gives
x)X, = 0

3.* If an eigenvalue has multiplicity k (that is, is repeated k times), there will be
k orthogonal vectors corresponding to this root.+

As an illustration of this result consider the diagonal matrix

ee)
A=
OZ 0
OF Uae
The characteristic equation is

(1-A)(2-A) =0
with roots
A, =1 with multiplicity 2

} For a proof, see G. Hadley, Linear Algebra, Addison Wesley, Reading,


MA, 1961, pp. 243-245.
ELEMENTS OF MATRIX ALGEBRA 145

For A, = 2, (A — ADx = 0 gives


=] 0 O}[ x, 0

0 Ob ts
The multiple root gives
Ord. OLX; 1 0
OF One, =0=>x, =~x,|9| +x, 0
Obed Oi 01h ee 0 ]
The root with multiplicity 2 thus yields two orthogonal eigenvectors e, and e3.

4.* The nth-order symmetric matrix A has eigenvalues X,, X4,..., Aq, possibly not
all distinct.} Properties 2 and 3 then guarantee a set of n orthogonal eigenvec-
tors X,,X,-..,; X,; Such that
x)x, = 0 pe Tet ea een (4-93)

As we have seen, any eigenvector is arbitrary up to a scale factor,


Ax; =A,x; — A(cx,)
= A,(cx;,)
where c is any constant. The arbitrariness may be removed by normalizing the x
vectors, and the most common normalization is to set the length of each vector at
unity, that is,
xx,= 1 pm le) on (4-94)
Conditions (4-93) and (4-94) define an orthonormal set of vectors. The conditions
may be combined in a single statement,
: wil iz]
xix, = 6,,, 5,= ey (4-95)

where 6;, is known as the Kronecker delta. Define X to be an nth-order matrix


whose columns are the vectors x,,X,,..., X,- Condition (4-95) may then be
written in the alternative form
XX =I (4-96)
From the definition and uniqueness of the inverse matrix it then follows that
xX’ = x7! (4-97)
The matrix X is then said to be an orthogonal matrix, that is, a matrix such that its
inverse is simply its transpose. It would be more appropriate to call it an
orthonormal matrix, since Eq. (4-96) requires all the columns to have unit length
as well as being orthogonal, but the former designation is the one established in
the literature. A remarkable property of orthogonal matrices follows immediately
from Eq. (4-97). Since the inverse is unique,
XX’ =I (4-98)
+ See G. Hadley, op. cit., p. 245.
146 ECONOMETRIC METHODS

that is, although X was constructed as a matrix with orthogonal columns, its row
vectors are also orthogonal. Thus an orthogonal matrix is defined by
X’X’ = XX = I (4-99)
Example 4-14
ie 0
A=| 202/72
Oyo Gl
The characteristic equation is then
[aera 2 0
2° 2d 2 |=0
0 v2 .1-A
thatis,
(1 =) )(1 oA) ae ero
with roots
A, =1 A,=-1 A, =4

st
Osriez Oe ex 0 V3
i (A ~T)x=}2) 1 -y2 ||] eo 0 xa] ae
2 ollx] lo 2v3
=
i D> 0 3 0 oe
Ag= ols. CA Dx= 2 > 3. “V2 35 |= 0 =>X,= Vio
Ce ey eiies 0
v5
wee
= 2 O ll x, 0
A; =4: (A — 4I)x = 2 ey | 0|}>x,= V5
0 y2 -3]1% 0 2

v5
ae ees
v1l0.— V5
Thus X= balay
2 iy cians
vl0.— v5
cin eaion
Vee uany15)
ELEMENTS OF MATRIX ALGEBRA 147

The reader can check numerically that the rows of x all have unit length and
are pairwise orthogonal (as, of course, are the columns).

5.* The orthogonal matrix of eigenvectors diagonalizes A, that is,


X’‘AX =A (4-100)
where A = diag{A,, A,..., A,)-

For A j and x j we have

Premultiplying by x’,
x,Ax, = A,x'x, =A,6,, using Eq. (4-95) (4-101)
Equation (4-101) displays the i, jth element in X’AX, and collecting for all i, 7
gives Eq. (4-100). An alternative proof illustrates a useful exercise in matrix
manipulation.

AX=|Ax, A2x, AXn


| | |
A;
hie ul | ne
a2 =
[ass |
a

= XA
Premultiplying by X’ then gives Eq. (4-100). We should not conclude from this
result that only symmetric matrices can be diagonalized. If for any matrix A there
are n linearly independent eigenvectors and we arrange them as the columns of a
matrix X, then
X-'AX=A (4-102)
The contrast with Eq. (4-100) is that the columns of X are not necessarily of unit
length, nor are they necessarily orthogonal.

6. The sum of the eigenvalues is equal to the sum of the diagonal elements (trace)
of A.

This property is true for any matrix, but the proof is particularly simple for
symmetric matrices. Denote the trace of a (square) matrix A by
tr(A) = a,, + a,+-:-+4,,
For two matrices, A of order m X n and B of order n X m,
tr(AB) = tr(BA) (4-103)
148 ECONOMETRIC METHODS

AB is of order m X m. Its ith diagonal element is

Thus

tr(AB) =) 5)s a; ;b;;


iaijail
=]
BA is of order n X n. Its jth diagonal element is

Thus

tr(BA) = x i b,,a;; = tr(AB)


j=li=1
This result extends simply to
tr(ABC) = tr(BCA) = tr(CAB) (4-104)
Turning now to
X’/AX=A
trA = tr(X’AX)
= tr(AXX’) _ using Eq. (4-104)
Thus trA = tr(A) (4-105)
or

A, tA, tes +A, =a, tay+-:-+a

7. The product of the eigenvalues is equal to the determinant of A.

This result is again true for any matrix, but the proof is very simple for
symmetric matrices. We note first that when X is an orthogonal matrix,

Sr ece (4-106)
for
XX = I= |X’| + |X| =1
but |X| = |X’|
Thus |X| = +1
Returning again to
X’AX =A
[X"| + |A] + |X] = |A|
Thus [A] =A,A,-=A, (4-107)
ELEMENTS OF MATRIX ALGEBRA 149

8. The rank of A is equal to the number of nonzero eigenvalues.

We established in Eq. (4-53) that pre- or postmultiplication of any matrix by


nonsingular matrices does not change its rank. Thus, again from Eq. (4-102),
p(A) = p(A) - (4-108)
and the easiest way to establish the rank of A is to determine the order of the
largest nonvanishing determinant that can be formed from its elements. This is
simply equal to the number of nonvanishing eigenvalues.

9. The eigenvalues of A® are the squares of the eigenvalues of A, but the


eigenvectors of both matrices are the same.

Ax = Ax
Premultiplying by A,
A’x = \Ax = 0’x
which establishes the result. We may note, in passing, a very useful application of
this result in analyzing the stability of dynamic systems. Suppose y, denotes a
vector of the values taken by a number of economic variables in time period f,
and suppose y, can be expressed in terms of the previous values by the system of
equations
Y= AY (4-109)
Even if the original specification of the system involves lags of more than one
period, an appropriate definition of new variables can produce a derived system
of the type of Eq. (4-109).+ Successive substitution in Eq. (4-109) gives

y,= A’Yo
where y, denotes initial values of the variables. Provided A has a linearly
independent set of eigenvectors,
xX'AX=A
or A=XAX'
Thus AP = XAXT'XAX b= XNV’-X"!
So At = XA'X™!
and the elements of y, are seen to be linear combinations of the tth powers of the
eigenvalues of A. Thus if the system is to be stable, we need
|A,| < 1, b= licen

10. The eigenvalues of A~' are the reciprocals of the eigenvalues of A, but the
eigenvectors of both matrices are the same.

Ax
= Ax

+ See G. Chow, Analysis and Control of Dynamic Economic Systems, Wiley, New York, 1975, pp.
21-35.
150 ECONOMETRIC METHODS

Premultiply by A7!,

or

which establishes the result.

11. Each eigenvalue of an idempotent matrix is either zero or unity.

By property 9
Cx x
But when Ais idempotent,
Atx = Ax =x
Thus
A(A — 1)x = 0
and since any eigenvector x is not the null vector,
A=0 or A= 1

12. The rank of an idempotent matrix is equal to its trace.

This follows from properties 6, 8, and 11,

o(A) = p(A) from property 8


= number of nonzero eigenvalues
= tr(A) from property 11
= tr(A) from property 6

4-6 QUADRATIC FORMS AND POSITIVE DEFINITE MATRICES

We have already introduced quadratic forms briefly in Sec. 4-2 and have
seen that
there is no loss of generality in considering only symmetric matrices.
For a 2 x 2
symmetric matrix A and a two-element column vector x, the quadratic form
is
MAN Giri dy i a
For a third-order matrix
WAX = 4),x7 + 2a,x,x, + 2413X\X3
La Xe ay wax
Ste A33X3 2
ELEMENTS OF MATRIX ALGEBRA 151

For the general nth-order case

ON Geeks ae 2dn xe

fey nee

Definitions
If x’Ax > 0 for all x * 0, the quadratic form is said to be positive definite and A is
said to be a positive definite matrix.
If x’Ax > 0 for all x + 0, the form and matrix are positive semidefinite.
Reversing the above inequality signs defines negative definite and negative
semidefinite matrices, respectively. If a form is positive for some x vectors and
negative for others, it is said to be indefinite.
It is important to have tests for positive definite matrices.

1. A necessary and sufficient condition for the real symmetric matrix A to be


positive definite is that all the eigenvalues of A be positive.

To prove the necessary condition assume x’Ax > 0. For any eigenvalue A,
Ax; = AX;
Premultiplying by x’, gives
x’ Ax, = A.x)x;
baa ee
=A;
Since x’Ax > 0 holds for any x + 0, it holds for each eigenvector, and so A, > 0
for all i. To prove sufficiency we assume all A; > 0 and show that x’Ax > 0. Since
a symmetric matrix has a full set of n orthogonal eigenvectors x,,X5,..., X,, any
nonnull vector x may be expressed as a linear combination of the eigenvectors
KUMI Cok ee a COX
Thus Ax = c, Ax, + c,Ax, + -:: +c, Ax,
= CyAGX, td C>A5X5 ct nae. stg Ch

x’Ax (5X4 + CyX > ate ol te ex) (eiAGx, ta €A5X>5 + Pthe sta CoN Xe

c7h, + 2A, 4+--- + 2A,


since
Sabie gies 0 i+j aie
xx,=8,= {1 i=j l, J Lo 2s anes

Since all A, are assumed to be positive, x’Ax > 0.

2. A necessary and sufficient condition for a real symmetric matrix A to be positive


definite is that the determinant of every principal submatrix be positive.
152 ECONOMETRIC METHODS

The principal submatrices of A are a set of n submatrices such as


G3, 4; Giz
aii . eee : a fT a OT. » a,ileal ncrsugeox
Ayn; Any Ag
More conventionally one takes the upper submatrices
4 2 93
4, 42
A, 7 [a,,] A, hs fe ar A,=|421 422 423 shits RADE
43, 432 433
When A is positive definite, x‘Ax > 0 for any nonzero x. Thus we may consider
an x vector whose first r elements are nonzero and whose last n — r elements are
zero, that is,
x =[x’ 0]
Then
,
x’Ax DEATHS,
= [x’, 1 Ae
v|4 eee
ald eee?
<7ATX:
where A has been partitioned by the first r and the last n — r rows and columns
and the asterisks denote the remaining submatrices in A, which get wiped out by
the zero subvector in x. Since
x’Ax > 0
it follows that
Xe AUX 0
Thus by the previous condition all the roots of A, are positive, and so
|A,| > 0
A suitable choice of x vectors then gives the necessary and sufficient condition for °
A to be positive definite as
|A,| > 0, |A,| > 0, |A,| > 0,..., JA] > 0 (4-110)

Finally we state a number of useful theorems on positive definite matrices.

1. If A is symmetric and positive definite, a nonsingular matrix P can be found


such that
A = PP’ (4-111)
We know that the matrix of eigenvectors of A can be used to diagonalize A,
that is, from Eq. (4-100),
X’AX = A
which gives
A = XAX’ (4-112)
When A is positive definite, all its eigenvalues are positive. Thus A may be
ELEMENTS OF MATRIX ALGEBRA 153

factored into
A= ANZA?

Ar
where A? =
ro

ie
Substitution in Eq. (4-112) gives
A =XA'2A12x' = (XA'/2)(XA!/2)/

which gives Eq. (4-111) with


P=XAl”
and P is nonsingular since it is the product of nonsingular matrices.

2. If Ais n X n and positive definite and if P is n X m with p(P) = m, then P’AP


is positive definite.

Clearly, P’AP is an m X m symmetric matrix, and for any m-element vector


y,
y (P’AP)y = x’Ax
where x = Py. Thus x is seen to be a linear combination of the m linearly
independent columns of P, and so x = 0 if and only if y = 0. Thus P’AP is
positive definite.
The final three results are stated without proof.

3. If Ais n X m with rankm < n, then A’A is positive definite and AA’ is positive
semidefinite.
4. If A is n X m with rankk < min(m,n), then A’A and AA’ are each positive
semidefinite.
5. If A and B are positive definite matrices and A — B is also positive definite,
then B-' — A~' is positive definite.+

4-7 MAXIMUM AND MINIMUM VALUES

It is convenient to express the main results on maxima and minima in matrix


terms. Consider a scalar variable y defined as a function of n independent
variables,

y SpfAX iy Kase sura)

+ A proof of this result is given in P. Dhrymes, /ntroductory Econometrics, Springer-Verlag, New


York, 1978, App. A2.13.
154 ECONOMETRIC METHODS

which may also be written


y =f(x)
The first-order or total differential of the function is defined as

dy =f, dx, +f,adx,+--:+f,d&, (4-113)


where

toady ee ee ae elt
Bah

and the dx; indicate arbitrary changes in the x;. For small dx, the first-order
differential gives the approximate value of the resultant change in y. Denoting the
vector of partial derivatives by f and the vector of differentials by dx,

fi dx
ie dx
f= ay. = i dx = ;
ox ; ‘

a dx,

the first-order differential of y is simply

dy = f’ dx (4-114)
If y has a stationary value at a point

collate
then dy = 0 for all points in the neighborhood of x*. For such points dx + 0, and
so from Eq. (4-114) the necessary condition for a stationary value is

f=0
that is, all partial derivatives are zero at the stationary point.
A stationary point may be a maximum, where the value of the function is less
at all points in the neighborhood of x*; a minimum, where the value of the
function is greater at all points in the neighborhood of x*; or a saddle point,
where the value of the function increases in some directions from x* and
diminishes in others. One may distinguish between these possibilities by means of
the second-order differential dy. The second-order differential may be found by
totally differentiating the first-order differential. It is an approximation to the
change in dy as we move away from the point x*. Clearly, for a maximum value
dy will decrease from zero to some negative value, so d*y will be negative, and
conversely for a minimum value d7y will be positive. For a saddle point d*y will
be positive for some dx and negative for other dx.} Totally differentiating Eq.

} It is possible, but extremely rare, to have d*y = 0 for some dx. Such complexities are ignored
here.
ELEMENTS OF MATRIX ALGEBRA 155

(4-113) gives
0
d*y “ays Liat hak, re ef dx. Pax,

0 |

0
++. + a Thax yay cee fax, |an.,

= fi, dx? + 2f\, dx, dx, + 2fi,dx, dx; +--+ + Df andes

Todi tt 2 f,diy axe 2p aa

eens dx?
where
gy
hij = hi = (8x, fori * j, alli,7

and dx? indicates the square of the differential dx,. The second-order differential
is thus seen to be a quadratic form in dx. The matrix of the quadratic form is the
symmetric Hessian matrix of second-order partial derivatives, which we will
denote by

92y fi hie rg = ite

Sere ee eae ue
Tigihel ss ES
and we may write
d*y = dx’F dx (4-115)
Thus d’y is positive or negative as F is positive definite or negative definite. To
summarize, the conditions for a maximum or a minimum at a point x* are as
follows:

First-order Second-order
condition condition

: 0 a7y . : ie
Maximum f=—=0 F = —~ is negative definite
dx dx?
2
Minimum f= ove 0 F= oy is positive definite
Ox ax2

where f and F are evaluated at x*.


156 ECONOMETRIC METHODS

Constrained Extrema
In finding stationary values of y = f(x,, X,,..-, X,) the x’s were assumed to be
independent variables. Thus we could specify n arbitrary differentials
dx,, dx,..., dx, In some problems, however, the x’s may be subject to one or
more constraints, and we have to find a maximum or minimum value of y subject
to the constraints. We will assume for the moment that the function hasa single
maximum or minimum value and state the problem formally as follows:

Find the x* vector which maximizes (minimizes) y subject to the m<n


constraints
g(x) =90 j= 1, 2am nn

Define the column vector

8, (x)
(Ca
8m (X)
and a column vector of m Lagrange multipliers,

A,
A,
Nal ee
r
Using these we define a new objective function as

@ = f(x) — X’e(x) (4-116)


Thus @¢ is a scalar quantity, which is a function of the m + n variables in \ and x.
The first-order condition for a stationary value of ¢ is that all m + n first-order
partial derivatives should vanish, that is,

Op Of, TeOe
ox dx ax (W’a(x))
9 = =0 (4-1 17)
@
aN g(x)

Some care is required in the interpretation of

Oe
3x (N’B(x))
Since

Nig(x) =A7oi(%5%22 Xn) SAO eaicepene..,)


Es BotAD ae (eeaX eee 9)
ELEMENTS OF MATRIX ALGEBRA 157

we have
d(N’g(x))
Ox,
_ , 98,
Pee pases
dg, ocean
Og ROE
ve oe

A(d’g(x)) _ Og, ag, 3 Og wee


Ge ee uae Raysnes
where

98)
Ox;
; 98,
of
ax = Ox;
. a
Pela an

Bn
0g;

Since df/dx is an n X 1 vector, we need to arrange for (0/0x)(’g(x)) to have n


rows. Defining G as the n X m matrix of partial derivatives,

PET oes RE een,


OX, Ox, Ox,
Eee Site gS 20a) 88ra loi |e Goer eee
Oe OX, OX, Ox Mane, Ox,
0g, 98, OS
OX ke Xo Cm
we have
OP oe
Ox Ox oh
and the first-order conditions for a stationary point are

<f_@n=0
i (4-118)
g(x)
=0
The second equation in Eqs. (4-118) ensures that the stationary value satisfies
the constraints. To distinguish between maxima and minima, we must still
examine whether the quadratic form in Eq. (4-115) is negative definite or positive
definite, but now only for dx vectors which do not violate the constraints. Totally
differentiating the jth constraint
o (Xk aes = 0
gives
0g; 98; 98; dx,
Os aes Peg Re emat Ax
n
158 ECONOMETRIC METHODS

There is a similar condition for each constraint. Thus the dx vectors which do not
violate the constraints are given by
G'dx =0 (4-119)
In many cases the F matrix consists only of constants, and so its definiteness can
be established independently of any x values.

PROBLEMS

4-1 Expand (A + B)(A — B) and (A — B)(A + B). Are these expansions the same? If not, why not?
How many terms are in each?
4-2 Given

3 4 l 2
a=(! ; al p= (1 = | c-|-i]
= : l 2 229 4
Calculate (AB)’, B’A’, (AC)’, and C’A’.
4-3 Find all matrices B obeying the equation
Oe il LO O71
f 2 |e4 ki 0 |
4-4 Find all matrices B which commute with

oe Ome
NG [2 |
to give AB = BA.
4-5 Write down a few matrices of order 3 X 3 with numerical elements. Find first their squares and
then their cubes, checking the latter by using the two processes A(A”) and A?(A).
4-6 Prove that diagonal matrices of the same order are commutative in multiplication with each other.
4-7 Let

QO 4
J=/0 1 0
OO)
Write out in full some products JA, where A is a rectangular matrix. Describe in words the effect on A.
Do the same with products of type AJ. Find J?.
4-8 If

OY tk @
YeIlO @ fl
@) @ @
find V* and V*. Examine some products of the type VA, VA, and V’A.
4-9 Given

aoa ae eee)
IWS ||2 8 and selene
T @& 1 Or @ A
Calculate |A|, |E|, and |B|, where B = EA. Verify that |B| = |E||A|.
4-10 Show that

1 l l
a gd 2Cla Ceo Oh)(Gad) CD)
a* b2 Ce
ELEMENTS OF MATRIX ALGEBRA 159

4-11 If (x), y,) and (x3, y)) are points on the x, y plane, show that the equation

XPV al
x y ll=0
X, yn 1

represents a straight line through the two points.


4-12 Prove that the determinant of a skew-symmetric matrix of odd order vanishes identically. (If A is
a skew-symmetric matrix, then A’ = — A.)
4-13 Show that the matrix

© |

Al
al-
Sl-
ale
t+
o
is orthogonal, that is, that Q’ = Q-!.
4-14 If the u; are normal variables with

E(u;) =0 eer

E(u?) =o? Pt Monell

E(ujuj)=0 ~ij7
show that E(u’Au) = o7tr(A).
4-15 Given

et)
peels
Peet —
NO
We

Compute

A = (I, — X(X’X)'X’)
Show that A is idempotent and determine its rank. Find the characteristic roots and the
associated characteristic vectors of A, and hence obtain the orthogonal matrix which diagonalizes A.
4-16 A and B are nonsingular matrices of the same order. Prove that AB and BA possess identical
characteristic roots. Show also that no such matrices can be found to satisfy the equation

AB — BA=I

(Cambridge Economics Tripos, 1967)


4-17 Evaluate the characteristic roots and vectors of

5 -6 -6
A= —] 4 2
m4
4-18 Examine the following quadratic forms for positive definiteness:

(a) 6x? + 49x32 + 51x} — 82x 9x3 + 20x) x3 — 4x,x2


(b) 4x? + 9x3 + 2x3 + 8x2x3 + 6x3x, + 6x\x2
(Cambridge Economics Tripos, 1968)
160 ECONOMETRIC METHODS

4-19 (a) Given that

find A” for n > 1 and B” forn > 1.


(b) If A is defined as

WIN
Wie
Wily wiry
wl
BIN Bl—
wiry
WIN

show that A is orthogonal.


Prove that the product of two orthogonal matrices of the same order is also an orthogonal
matrix.
(UL, 1967)
4-20 X is a square matrix of order n and a is an n X | vector. Find
3 (a’Xa)
ox
(a) When the elements of X are independent.
(b) When X is symmetric.
Note: If X is a matrix whose elements are variables x; , and f(X) is a scalar function, then

_ Of (X)
Bs ax
1S a matrix of the same order as X such that
CHAPTER

FIVE
THE k-VARIABLE LINEAR MODEL

Chapter 2 contains a fairly complete treatment of the two-variable (kK = 2) model.


Some of the algebra of the three-variable (k = 3) model was developed in Sec. 4
of Chap. 3. Section 1 of Chap. 4 indicated the power of matrix algebra to give a
compact representation of the general k-variable model. It is now time to give a
complete statistical treatment of the k-variable model. To facilitate this treatment
we first provide a review of some basic statistical results in matrix form. These
results are extensions of the material on matrix algebra in Chap. 4 and the
statistical material in various sections of App. A.

5-1 PRELIMINARY STATISTICAL RESULTS

Let x denote a vector of random variables X,, X,,..., X,,. Each variable has an
expected value
p= EM) ie pin
Collecting these expected values in a vector p, gives

E(X;) by
E( xX, 2
eH : is : cen
BCI lhe
161
162 ECONOMETRIC METHODS

The application of the operator E to the vector x means that £ is applied to each
element of x. The variance of X;, by definition, is

var( X;)= E((X, ~ 4,)>}


and the covariance between X; and X; is

cov( X;, X;) = E{(X, — m)(X) — 4)))


If we define the vector x — p» and then form

(X,
— 4)
E(x ~ w)(x- w= E ee [2% = my)O% a) OG —
(4, - 1)

E(X, — #1) EG Sy)


CG ho) EC eerie
BOG wa th) ond Basta NigSe age a a
E(X, — ba)(G
— wi) BCX,
— (= fe) Ee)
we see that the elements of this matrix are the variances and covariances of the X
variables, the variances being displayed on the main diagonal and the covariances
in the off-diagonal positions. The matrix is known as a variance-covariance matrix
or, more simply, as a variance matrix or a covariance matrix.t We will generally
refer to it as a variance matrix, denote it by var(x), and

var(x) = E{(x — p)(x —p)} == (5-2)


The variance matrix 2 is clearly symmetric. It is important to determine whether _
= is positive definite or not. Define a scalar variable Y as a linear combination of
the X’s, that is,
Y=(x-yp)ec (5-3)
where ¢is any arbitrary n-element column vector. Squaring Eq. (5-3) and taking
expectations gives

E(Y*) = E{e(x — p)(x — p)’c}


= CE{(x — »)(x — p)'}e
= c’=e
There are two useful points to notice about this development. First of all,
(x — p)’c is a scalar, thus its square may be found by multiplying it by its
transpose. Second, whenever we have to take the expectation of a complicated
matrix expression, the E operator may be moved to the right past any vectors or
matrices consisting only of constants, but it must stop in front of any expression

+ Alternative expressions for the variance-covariance matrix are cov(x) and V(x).
THE k-VARIABLE LINEAR MODEL 163

involving random variables. Since Yis a scalar random variable, E (Y7) > 0. Thus
cae > 0
and 2 is positive semidefinite. But

E(Y?)=0=Y=0
which, from Eq. (5-3), means that the X deviations (X, — p,), (Xo = pS) ne ee
— p,,) are linearly dependent. Thus
2 is positive definite, provided no linear dependence exists among the X’s.

The n random variables will have some multivariate probability density function
(pdf) written
DIS) = Pl BOG Xe)

which is simply some formula or rule giving the likelihood of various combina-
tions of X values. The most important multivariate pdf is the multivariate normal.
The univariate normal distribution is specified once its mean p and its variance o”
are given. The multivariate normal is similarly specified in terms of its mean
vector p and its variance matrix 2. The formula is

p(x) = eatin
1
Qmie
xp] — 5(x— wy E-"(x - w) (5-4)
A compact shorthand statement of Eq. (5-4) is

x ~ N(p, 2)
to be read, “the variables in x are distributed according to the multivariate
normal law with mean vector p and variance matrix 2.” When n = 1, = = o? and
Eq. (5-4) becomes
1
P(X) = v270 exp|— 55 x= 4)
which is the familiar univariate normal density. When n = 2, if we use p to denote
the correlation between X, and X,, the variance matrix becomes

0;2 po,0,
Sa with |2| = 0707(1 — p’)
Dory Bo:
Notice that |=| > 0 unless p* = 1, so that the variance matrix is positive definite
unless there is perfect linear correlation between the two variables, in agreement
with the general result above. Substitution in Eq. (5-4) gives

(X,, %) 1 o| 1 (+ =H)
P , a eer ay. Ntk mame ah)
te 270,051 — p* 2a Ge) Oily

sey yaledNes Ba.)


9; o all
164 ECONOMETRIC METHODS

An especially important case of Eq. (5-4) occurs when all the X’s have the
same variance o” and are all pairwise uncorrelated.+ Then
= =o’l
with

|Z] =o", Zl= eA


0

a Hee (2107)? |- eee 202 »)| (5-5)


Equation (5-5) thus factorizes into
1 ] ‘|
POX Kove %) = 11 a exp|P|— ——(X.
a H;)
— pu;
i=1

= p(X) p(X) OY)


so that the multivariate density is the product of the separate marginal densities,
that is, the X’’s are distributed independently of one another. This is an extremely
important result. Zero correlations between normally distributed variables imply
statistical independence. This result does not necessarily hold for variables which
are not normally distributed. Notice carefully that these results depend on zero
correlations in the population and not on zero sample correlations.
A more general case of this result may be derived from Eq. (5-4). Suppose =
has the form
>) 0
Sc |. | (5-6)
where 2,, is square of order r and 2,, is square of order n — r. The form of Eq.
(5-6) means that each and every variable in the set X,, X,,..., X, is uncorrelated
with each and every variable in the set X,,,, X,, ,..., X,- Applying a similar -
partitioning to x and yp,

(x = pe)"(x — pe) = (ey = wy) (x = hy) + (x, - 2 )'Z39'(X2 — pp)


using Eq. (4-75) for 2~'. Also from Eq. (4-79)

[2] = [21111229]
Making these substitutions in Eq. (5-4) gives

p(x) = | —exp| = F(x ~ Bi) ZnO) = »)]}


rcs
1 1 ee
‘ |(Daj O7se 2 exp|- 5 (X> — Hp )/By'(x2 — +)]}
2

+ The assumption of a common variance is only made for simplicity. All that is required for the
result is that the 2 matrix be diagonal.
THE k-VARIABLE LINEAR MODEL 165

that is,

P(x) = p(x,) p(x)


so that the first r variables are distributed independently of the remaining n — r
variables.

Distributions of Quadratic Forms


Suppose
x ~ N(0,1)
that is, the n variables in x have independent normal distributions, each with zero
mean and unit variance. In other terminology, the X’s are independent standard-
ized normal variables. The sum of squares x’x is a particularly simple example of
a quadratic form with matrix I. From the definition of the x? variable,

x’x ~ x?(n)
for x*(n) is the sum of the squares of n independent standardized normal
variables.
Suppose now that
x ~ N(0, 671) (537)
The variables are still independent and have zero means, but each X has to be
divided by o to yield a variable with unit variance. Thus

XX
St 7 aegis
nae 2)
age)

that is,

eh ~ x7(n) (5-8)

or x’(o71)_
‘x ~ x2(n) (5-9)
Equation (5-9) shows explicitly that the matrix of the quadratic form is the
inverse of the variance matrix.
Suppose now that
x ~ (0,2) (5-10)
where & is a positive definite matrix. The equivalent expression to Eq. (5-9) would
now be
x= 'x ~ x*(n) (5-11)
This result does in fact hold, but the proof is no longer direct since the X
variables are no longer statistically independent. The trick is to transform X’s
into Y’s, which will be independent standardized normal variables. Since 2 is
positive definite, by Eq. (4-111) there exists a nonsingular matrix P such that
= = PP’
166 ECONOMETRIC METHODS

which gives
2) = (Po) Pa and PS (Be) (5-12)
Define an n-element y vector as
y=P 'x
The Y variables are multivariate normal since they are linear combinations of the
XS,
E(y) = P“'E(x) =P '0=0
and var(y) = E{P” 'xx’(P7')’}
=iP => (Bea):
on
from Eq. (5-12). Thus the Y’s are standardized normal variables and

VV AT)
But

yy = x(P))
Ps y= x7 ix
from Eq. (5-12). So
xD 'x ~ x?(n)
which is the result anticipated in Eq. (5-11).
Assume again
x ~ N(0,1)
and now consider the quadratic form x’Ax where A is idempotent with rank
r <n. If we denote the matrix of eigenvectors of A by Q, then

]
]Pst
r terms

Q’AQ=A= 1 (5-13)
0
me n — r terms

0
where A will have r units and n — r zeros on the main diagonal. Define

y = Ox
Thus
x = Qy
since Q is orthogonal. Then

E(y) =0
THE k-VARIABLE LINEAR MODEL 167

and var(y) = E{yy’}


= E{Q’xx’Q}
= QIQ
since Q’Q = I. Thus the Y’s are independent standardized normal variables. The
quadratic form may now be expressed as
x’Ax = y'Q’AQy
=eFY Pe Vy eset
using Eq. (5-13). Thus
x’Ax ~ x7(r)
The general result is

Ifx ~ N(0, 071) and A is idempotent of rank r, F wax ~ x7(r)


o
Independence of Quadratic Forms
Suppose x ~ N(0, o7I) and we have two quadratic forms
x’Ax and x’Bx
where A and B are symmetric idempotent matrices of the same order. We seek the
condition for the two forms to be independently distributed. Because the matrices
are symmetric idempotent,
x’Ax = (Ax)’/(Ax)
and x’Bx = (Bx)’(Bx)
If each of the variables in the vector Ax has zero correlation with each variable in
Bx, they will be distributed independently of one another, and hence any function
of the one set of variables, such as x’Ax, will be distributed independently of any
function of the other set, such as x’Bx. The covariances between the variables in
Ax and those in Bx are given by
E{(Ax)(Bx)’} = E{Axx’B)
= o AB
These covariances (and hence the correlations) are all zero if and only if
AB =0 (5-14)
Since A and B are symmetric, the condition may be equivalently stated as
BA = 0; the one implies the other. Thus two quadratic forms with idempotent
matrices will be distributed independently if the product of the idempotent matrices is
the null matrix.

Independence of a Quadratic Form and a Linear Function


Assume x ~ N(0,o7I). Let x’Ax be a quadratic form with A a symmetric
idempotent matrix of order n and let Lx be an m-element vector, each element
being a linear combination of the X’s. Thus L is of order m X n, and we note that
168 ECONOMETRIC METHODS

it need not be square or symmetric. If the variables in Ax and Lx are to have zero
covariances, we require
E{Axx’L’} = o*AL’ = 0
or equivalently
LA =0 (5-15)

5-2 ASSUMPTIONS OF THE LINEAR MODEL

The first basic assumption of the model is that the vector of sample observations
on Y may be expressed as a linear combination of the sample observations on the
explanatory X variables plus a disturbance vector, that is,

l. y = BX, + £x,+---+ 6x, +u (5-16)

where each vector is a column vector of n elements. The x, vector is a column of


units to allow for an intercept term. Each of the remaining x, vectors (i =
2,3,..., k) denotes the sample observations on a specific explanatory variable.
The B’s are unknown population (model) parameters, but even if we knew their
values, the linear combination (8,x, + --- + B,x,) would not determine the y
vector exactly, for economic relations are stochastic, not exact. Thus u is a
disturbance vector measuring the discrepancies between the linear combination
and any actual sample realization of Y values.
Equation (5-15) may be expressed in matrix form as
y=XBP+u (5-17)
where

Y B, u,
ie | | | B, u,
y= X=/]X, xX, X; 5 = u=
| | |
ne B, Uu,
The central problem is to obtain an estimate of the unknown 6 vector. To make
any progress with this we need to make some further assumptions about how the
observations on Y have been generated.

2. E(u) = 0, that is, E(y) = XB

To illustrate the meaning of this assumption, let us assume that the Y


variables measure family income and various other family characteristics and Y
denotes family expenditure on, say, travel. The first row of the X matrix is some

+ An outline of the various reasons for the introduction of the disturbance term has already been
given in Sec. | of Chap. 2.
THE K-VARIABLE LINEAR MODEL 169

specific set of numbers for family income, size, and composition. Let s, denote a
row vector consisting of these numbers. Then

E(Y,)= s,B
is the average, or expected, level of travel expenditure for this type of family.
However, if we observe the actual travel expenditure of a family with these
characteristics, it may be greater than the expected level, and the expenditure of
another family with the same characteristics may well be less than the expected
value. Or if we observe the travel expenditures of the same family in different
periods of time, these may be expected to fluctuate around the mean value.
However, if the theorist has done a good job in specifying all the significant
explanatory variables to be included in X, it is reasonable to assume that both
positive and negative discrepancies from the expected value will occur and that,
on balance, they will average out at zero, that is,
E(u,) =0
Similar considerations apply to each row of X, and so we have

E(u,) 0
E(u,) 0
E(u)= | =|.
E(u,) 0
3. E(uu’) = o7I

Since E(u) = 0, E(uv’) is a variance matrix. This assumption gives

var(u,) cov(u,;,u,) --- cov(u,,u,) On, BOY er @ BO


COV( U55.,) vat) ee COV(GL, i N= 0 kot oe iO
cov(u,,u,). cov(u,,u,) 9° var(u,,) OmiOnleseamad
This is a double assumption, namely:

e Each wu distribution has the same variance.


e All disturbances are pairwise uncorrelated.

The first property is referred to as homoscedasticity (or homogeneous variances)


and its opposite as heteroscedasticity. If the sample observations related to travel
expenditures of a cross section of households, the assumption of homoscedasticity
would probably not be a reasonable one, since low income families will almost
certainly have low average expenditures on travel and also a low variance of
actual travel expenditure about the average, while high income families will tend
to display both higher mean levels of expenditure and greater variance about the
mean. The second part of this assumption—all disturbances being pairwise
170 ECONOMETRIC METHODS

uncorrelated—is a very strong assumption indeed. Again, in the context of the


travel example it means that the size and sign of the disturbance for any one
family has no influence on the size and sign of the disturbance for any other
family. This is not to deny the possibility of “keeping up with the Joneses” as an
important economic and sociological fact. If such a phenomenon does exist, it
would be more appropriately characterized in the specification of the X variables.
If the sample data related, say, to aggregate travel expenditure over a period of
years, the same assumption means, for example, that unusually heavy expenditure
in one year does not tend to be associated with unusually low (or high)
expenditures in the next year or indeed in any subsequent year.

4. p(X) =k

This assumption states that the explanatory variables do not form a linearly
dependent set. For example, if we had just two explanatory variables, X, and X,,
and this assumption was not fulfilled, there would then exist an exact relationship
X3 =c¢, + ¢,X, (5-18)
which, combined with the hypothesized
Y=8B, + BX,
+ BX, + u (5-19)
gives
Y = (B, + B3c,) + (B, + Bycy) X_ + u (5-20)
The constants c, and c, can be determined exactly, and we can estimate the
intercept and slope of Eq. (5-20), but there is no way to obtain estimates of the
three 8 parameters.

5. X is a nonstochastic matrix.

L
This assumption at first sight seems incongruous. It means that if we take
another sample of n observations, the X matrix of explanatory variables remains
unchanged, the only source of variation then being in the u vector and hence in
the y vector. However, the social sciences are notoriously difficult for being
observational and nonexperimental so that in general the X variables are not
subject to experimental control by the social scientist. There are three main points
to be made about this assumption. First of all, in spite of the remarks above, there
are cases where the X data can be controlled. In a cross-section survey, the sample
design may call for the inclusion of certain numbers of families with specific
characteristics, and sampling is continued until these specifications are met.
Second, even if it is not in fact feasible to control the X data precisely, it is still
useful to be able to make statistical inferences which are conditional on the X
values actually present in the sample. In this light it is very much an assumption
of convenience in that it simplifies dramatically the derivation of several basic
statistical results. Third, once these simple results have been derived, it is possible
to weaken the assumption to allow the X variables to be stochastic, but distributed
THE k-VARIABLE LINEAR MODEL 171

independently of the disturbance term, and then see what modifications of the
earlier results are required.

6. The u vector has a multivariate normal distribution.

Assumptions 2, 3, and 6 may then be combined in the single statement:


u~ N(0, 071) (5-21)

5-3 ORDINARY LEAST-SQUARES (OLS) ESTIMATES

The most frequently used estimating technique for the model outlined in Sec. 5-2
is least squares. The hypothesized model is

y= Ap eu (5-22)
Let b, denote any arbitrary k-element vector. This in turn serves to define a
vector of errors, or residuals,
e, = y — Xb, (5-23)
The least-squares principle for choosing by, is to minimize the sum of the squared
residuals e,e,. From Eq. (5-23)

exe, = (y — Xb,)’(y — Xbx)


= y’y — 2b,.X’y + bi,X’Xb,

Thus See. = —2X’y + 2X’Xb, (5-24)


ok

The necessary condition for a stationary point requires that we set Eq. (5-24)
equal to the 0 vector. Denoting the resultant OLS solution for b, simply by b gives
(X’X)b = X’y (5-25)
These are referred to as the OLS normal equations. Assumption 4 ensures that X’X
is nonsingular. Thus an equivalent expression for b is
b = (X’X) 'X’y (5-26)
The vector of OLS residuals is likewise denoted by e, where
e=y— Xb (5-27)
Using this expression to substitute for y in Eq. (5-25) gives
(X’X)b = (X’X)b + X’e
xje 0
xe 0
Thus Xe=/] |=|].|/=0 (5-28)
172 ECONOMETRIC METHODS

This is a fundamental OLS result. The first element in this equation gives
e=0
that is, the residuals from the OLS regression always have zero mean, provided
that the equation contains a constant term. The remaining elements in Eq. (5-28)
state that the residual has zero sample correlation with each X variable.
To establish that the stationary point does indeed correspond to a minimum
of the sum of squares, differentiate Eq. (5-24) once again with respect to b to
obtain

d*(e,ex)
ab,
—_———. (X’X)
= 2(X’X (5-29)
5-29

From Sec. 4-7 this gives a minimum provided X’X is positive definite. To establish
this, let d be any nonnull k-element vector, and consequently define an n-element
vector ¢ as
c = Xd (5-30)

The assumption that X has full column rank ensures that ¢ is nonnull; otherwise
Eq. (5-30) would express a linear dependence between the columns of X. Thus
c’c = d’X’Xd > 0
and X’X is positive definite.
Returning to Eq. (5-26),

h = (XX) -X’y
and substituting
y=Xfp+u
gives

b =B + (X’X) 'X’u (5-31)


Taking expectations gives
E(b) =B (5-32)

E{(X’X) 'X’u) = (X’X) 'X’E(u) = 0


by assumptions 2 and 5. The OLS estimator is thus a /inear unbiased estimator.
The linearity property refers to linearity in y (or u) as is seen in Eq. (5-26) or Eq.
(5-31), for each element in b is a linear combination of the elements of y (or u),
the weights being functions of the X data which are nonstochastic. The unbiased-
ness is established in Eq. (5-32).
Next we derive the variance-covariance matrix of the OLS estimators. From
Eqs. (5-31) and (5-32),

b — E(b) = b— B = (XX) ‘Xu


THE k-VARIABLE LINEAR MODEL 173

Thus
var(b) = E{(X’X) 'X’uu’X(X’X) '}
(X’X) 'X’o7IX(X’X) ' from assumptions 3 and 4
o?(X’X) !
(5-33)

since I may be suppressed at will and the scalar o* moved in front or behind
matrices. The elements on the main diagonal of Eq. (5-33) give the sampling
variances of the corresponding elements of b, and the off-diagonal terms give the
sampling covariances. The most important result in least-squares theory is that no
other linear unbiased estimator can have smaller sampling variances than those of
the OLS estimator in Eq. (5-33). OLS estimators are thus said to be best linear
unbiased estimators (b.1.u.e.), that is, to have minimum variance within the class of
linear unbiased estimators. This result is known as the Gauss-Markov theorem.
The following proof is somewhat roundabout, but it has the advantage of
establishing a further important result at the same time. Let ¢ denote an arbitrary
k-element column vector of known constants and define a scalar quantity pw as
w= cB (5-34)
If we choosec’=[0 1 O --- OJ, then yu = £,. Thus we can use Eq. (5-34) to
pick out any single element in B. Or if we choose
coma Kop eas el, Mya
then
w= E(Y, 41)
which is the expected value of the dependent variable Y in period n + 1,
conditional on the X values in that period.
We wish to consider the class of linear unbiased estimators of w. Thus define
a scalar m which will serve as a linear estimator of uw, such that
m=ay=aXB+ a'u (5-35)
where a is some n-element column vector. The definition ensures linearity. To
ensure unbiasedness we have
E(m) a’XB + a’E(u)
a’XB
= cB
only if
aX=c’ (5-36)
From Eggs. (5-35) and (5-36),
var(m) = E{a‘uu’a}
=o0°a’/a
which derivation uses the fact that since a’u is a scalar, its square can be written
174 ECONOMETRIC METHODS

as the product of its transpose and itself. The problem is thus to choose a to
minimize a’a subject to the k side conditions a’X = ce’. Define
= a’a — 2 (X’a — c) (5-37)
Here X is a column vector of k Lagrange multipliers, and the side conditions
(5-36) have been transposed to make the multiplications in Eq. (5-37) conform-
able. Differentiating

coe 2a — 2XX=0 (5-38)


da

and —dg
Oy =
2(X’a — c) =0
, — => (5-39)
5-39

Premultiplying Eq. (5-38) by X’ and using Eq. (5-39) gives


c = X’a = X’XX
Thus
d= (XX)
-'e
Substituting back in Eq. (5-38) gives

a = X\ = X(X’X) 'c
and so the desired minimum variance linear unbiased estimator of c’B is
m=aly

= ¢(X’X) “X’¥y
= c’b (5-40)
that is, the unknown B is replaced by the OLS b. It follows directly that +

1. Each OLS coefficient 5; is a best linear unbiased estimator of the correspond- —


ing population coefficient B..
2. The b.l.u.e. of any linear combination of 8’s is that same linear combination
of the b’s.
3. The b.l-u.e. of E(Y,) is
b, F bX, , Sta b; X3 , ta oats Ar DieNoee

The Model in Deviation Form


In Chaps. 2 and 3 it was seen that the two- and three-variable regression
models
could also be treated in deviation form. The essence of the approach
was to first
of all express all data in terms of deviations from sample means and then
estimate
the regression parameters in two stages, the first Stage dealing with
the slope
coefficients and the second stage with the intercept term. The same treatme
nt may
7 As far as the disturbance u is concerned, the derivation of the Gauss-Markov
result has only
required the assumptions of zero mean and zero covariances and has not required the assumptio
n of
normality.
THE k-VARIABLE LINEAR MODEL 175

be applied to the k-variable model by use of the following transformation matrix:


I,
A=I- il (5-41)
where i denotes a column vector of n units. Thus

1 , Lovack 1
A= tell Ped 1
OU Neer coer Raewean
1 real 1

fy =|[¥,; Y, *= “¥s}then

—iy=Y

me Ame
and Ay =y-iY=

Y,-Y
Thus premultiplying any column vector of observations by A produces a vector
showing those observations in deviation form. Two special cases are
Ai = 0 (5-42)
or, more generally, premultiplying any vector of identical elements by A gives the
zero vector. Second,
Ae =e (5-43)
for the residuals have zero mean, and are thus already in deviation form. It is
easily verified that the A matrix is symmetric idempotent.
The OLS estimator b and residual vector e are connected by
y=Xbt+e (5-44)
If we partition the X matrix as
X=[x, X,]
where x,(= i) is the usual column of units and X, the n xX (k — 1) matrix of
observations on the variables X,, X,,..., X,, we can rewrite Eq. (5-44) as
y =x,b,+ X,b, +e (5-45)
where b’ = [b, 4] indicates a conformable partitioning of the b vector into the
intercept b, and the subvector b, of slope coefficients. Premultiplying Eq. (5-45)
by A gives
Ay = AX,b, + e
using Eqs. (5-42) and (5-43). Premultiplying this by X’, yields
X’, Ay = X’,AX,b, (5-46)
176 ECONOMETRIC METHODS

for X5e = 0 from Eq. (5-28). Finally, using the symmetric idempotency of A
means that Eq. (5-46) is equivalent to

(AX, )’(Ay) = (AX, )’(AX, )b, (5-47)


The interpretation of Eq. (5-47) is as follows:

eb, is the subvector of OLS slope coefficients.


e Ay is the y vector expressed in deviation form.
e AX, is the matrix of explanatory variables in deviation form.
« Equation (5-47) is a set of normal equations [compare with Eq. (5-25)] in terms
of deviations, whose solution yields the OLS slope coefficients.
e The remaining coefficient, which is the intercept term, is obtained by premulti-
plying y = Xb + e byi'’/n to yield

by
Yi=3| lax X,]| 2

by

or
b, = Y—b,X,
— b,X, — ++: — b,
X, (5-48)
The sum of squared deviations in the dependent variable, denoted by TSS, is
TSS = y’Ay
This may be decomposed into an explained sum of squares (ESS) and a residual
sum of squares (RSS) in the manner of Chaps. 2 and 3. Return to
Ay = AX,b, +e
Transposing and multiplying,
y’Ay =b,X,AX,b, + e’e (5-49)
(TSS) (ESS) (RSS)

since the cross-product term vanishes in view of X’e = 0. The multiple correlation
coefficient R, »;..., for the k-variable case may then be defined in a number of
alternative ways. The basic definition is
ESS e’e
ick = ass TVS Yay oy
In view of Eq. (5-49) this is equivalent to

Ro biee
X5,AX,b, biX5A
ey (5-51)
y’Ay y’Ay
where the second expression follows from Eq. (5-46). Alternatively, we may start
with the complete OLS regression
y= Xb+e
THE k-VARIABLE LINEAR MODEL 177

Transposing, multiplying, and again using X’e = 0 gives


y'y = b’X’Xb + e’e (5-52)
Using b = (X’X)~'X’y, equivalent expressions for Kq. (5-52) are
y'y = b’X’y + e’e = y’X(X’X) ‘X’y + e’e (5-53)
Comparing Eq. (5-52) with Eq. (5-49), the residual sum of squares, e’e is the same
in each equation, since the OLS regression is unique and it makes no difference
whether we fit the complete regression directly to the original data or transform
the data into deviation form and compute the slope coefficients followed by the
intercept. The left-hand sides differ only in that
yy De
and yay XY, ay y’
= DY? —nY?
= yyays
Subtracting the correction for the mean nY? from both sides of Eq. (5-52) gives

(y’'y — nY”) =(b’X'Xb’ — nY?) + ee


(Tss) (ESS) (RSS)
Thus the previous expressions for R* in terms of sums of squares may all be
computed in terms of the original data, provided only that the correction for the
mean is subtracted from any total or explained sum of squares (but not from the
residual sum of squares).
It is sometimes useful to compute an R’, adjusted for degrees of freedom,
especially when comparing the explanatory power of different numbers of ex-
planatory variables. Adding any extra explanatory variable can never increase the
residual sum of squares and thus can never decrease the R* defined in Eq. (5-50),
since that expression takes no account of the number of explanatory variables
employed. It may be rewritten as

Rin... Brat ag) me


: y Ay/n
The adjusted R? is defined as

Rpp3..-4 ule toh §)


y’Ay/(n — 1)
The rationale behind the adjustment is that k parameters have been used in fitting
the regression plane from which the residual sum of squares is measured, and one
parameter, the sample mean, has been estimated in computing TSS. As will be
seen later, these provide unbiased estimators of the disturbance variance and the
Y variance. Equivalent expressions for the adjusted coefficient are
= n-1
Roe ere os z (1 i Ree)
1-—k n-1
Fy az
n—-k n-k Rio3.0k
178 ECONOMETRIC METHODS

It is thus possible for the adjusted coefficient to decline if an additional variable


produces too small a reduction in 1 — R? to compensate for the increase in
(n — I)/(n — k).

Example 5-1 To help fix some of these concepts, here is a brief numerical
example. The numbers have been kept artificially simple so as not to obscure
the nature of the operations with cumbersome arithmetic. The sample data
are
3 alc ee)
] Pr 4
Vo =a1e8 and Kee lS 26
3 [S24
5 bea 6
where we have already inserted a column of units in the first column of X.
From these data we readily compute

Sacro 20
OX peo) en and X’y =| 76
PS NE eAVAD 109
The normal equations of Eq. (5-25) are then

Selon 25 iD, 20
15.558) tbs) |e= ar
25. Sl “12976; 109

Rather than invert (X’X) we will solve these equations by the elimination
method. In the first step subtract three times the first row from the second
and five times the first row from the third. This gives the revised system
Selene onl leas 20
Oy s10e) koyl Rb arn G
O26 oc tlies 9
Next subtract six-tenths of row 2 from row 3 to get

5s 15) 2954 4) be 20
0 10. 46.8 |\tbal=s lado
Cig 500 4a ee ~0.6
The third equation gives 0.46, = —0.6, that is,

Substituting for b, in the second equation,


10b,
+ 6b;= 16
THE k-VARIABLE LINEAR MODEL 179

gives

b, = 2.5

Finally, the first equation |


5b, + 156, + 256, = 20
gives

The regression equation is thus


Y=4+2.5X, — 1.5X,
Alternatively, transforming the data into deviation form gives
—1 0 0
3 2 eel
Ay = 4 and AX = 2 l
hl Se eee
1 1 1
The normal Eqs. (5-46) now become

fs alle]=[
The observant reader will notice that these are the second and third equations
obtained in the first step of the elimination method above.+
Thus the solutions for b, and b, will coincide with those already ob-
tained. Likewise, b, will be the same as before, for the final equation in the
back substitution above is readily seen to be

Thus the elimination process applied to (X’X)b = X’y is, in fact, equivalent to
transforming the data into deviation form and proceeding in two-step fash-
ion.
To calculate R* we note from the Ay vector that
TSS = y’Ay = 28

+ The reason why may be seen as follows:


n LX, LX;
SOX | GX
SNe Kame Ke
To produce a zero in the (2, 1) position, we must subtract 1 X,/n (= X,) times the first row from the
second. In the (2, 2) position this gives 2X3 — nXz = L(X, — X>)*, and in the (2, 3) position it gives
DX, X;, — nX,X, = X(X, — X,)(X; — X3). A similar argument applies to the transformed third row.
180 ECONOMETRIC METHODS

Calculating the explained sum of squares from bj X‘, Ay gives

ESS = 6. = 15] is = 26.5


Thus
RSS = 28 — 26.5 = 1.5
26.5
and R 2
78
——
0.9464

so that the regression has accounted for almost 95 percent of the variance of
Y. As a check we may calculate the explained sum of squares from b’X’y by
subtracting the correction for the mean,
20
b’X’y = [4 2.5 -13] a = 106.5
109
nY2 = 5(4)° = 80
Thus ESS = b’'X’y — n¥? = 26.5
in agreement with the previous calculation.

Estimation of o7
Finally in this section we derive an estimator of o*, the variance of the dis-
turbance term. As the values of u are not directly observable, it seems plausible to
base an estimate of o” on the residual sum of squares e’e. The only question is
what should the divisor be, and this can be settled by requiring the estimator to
be unbiased. We have
e=y — Xb

=y — X(X’X) 'Xy
= [1- x(x’x)'x’ly
= My (5-54)
where

M = I — X(X’X) 'X’ (5-55)


M is a very important matrix. It is easily verified that it is symmetric idempotent.
It also follows directly by multiplying out that
MX = 0 (5-56)
Returning to Eq. (5-54),
e = M(XB + u)
that is, e = Mu
in view of Eq. (5-56). From the symmetric idempotency of M,
e’e = uMu
THE k-VARIABLE LINEAR MODEL 181

Taking expectations
E(e’e) = E(u’Mu)
= E({tr(u/Mu)} since u’Mu is a scalar
= E{tr(Muu’)} —_ from Eq. (4-15)
=o’ trM by assumption 3
From Eq. (5-55)
tr(M) = tr(1) — tr[X(X’x)~'x’]
tr(1) — tr[(X’X)~'x’x]
=k
Thus if we define

= (5-57)

it follows that
E(s?) =o?
and we have found the desired unbiased estimator. The square root s is often
referred to as the standard error of the estimate, and may be regarded as the
standard deviation of the Y values about the regression plane.

5-4 INFERENCE IN THE OLS MODEL

So far we have not used the assumption that the u’s are multivariate normal, but
this now becomes necessary. We now make the twin assumptions

u ~ N(0, 071)
and X is nonstochastic with rank k

The first is a combination of assumptions 2, 3, and 6, and the second is a


combination of assumptions 4 and 5. We have seen in Eq. (5-31) that

b=B + (X’X) ‘Xu


so b is then multivariate normal, and since we have already established the mean
vector and the covariance matrix, we have the fundamental result

b ~ N(B, 0?(X’X)') (5-58)


From the end of the previous section we also have
e’e = uMu
From the result in Sec. 5-1 on the distribution of quadratic forms with idempo-
182 ECONOMETRIC METHODS

tent matrices it then follows that

(ee) ~x?(n-k) (5252)


oO
The degrees of freedom in Eq. (5-59) come from the fact that
p(M) = tr(M)
since M is idempotent, and we have just shown the trace to be n — k. Finally,
applying condition (5-15) for the independence of a linear and quadratic form to
b and e’e gives
(X’X) 'X’M = 0 (5-60)
since MX = 0. Thus
b is distributed independently of s?.
These results suffice to establish inference procedures for any element of b.
Consider, for example, b;, the estimated coefficient of X, in the OLS regression. It
follows from Eq. (5-58) that :
Di N( i> 0°a;;)
where a,, denotes the ith element on the principal diagonal of (X’X)~'. Thus

BSP (0,1)
Oye
From Eqs. (5-59) and (5-60),

(n—k)s?
= ) ie x?(n a k)

independently of b;. Thus we can proceed directly to form a ¢ variable, that is,

bp al)
oa, sy(n = k)
or

A= b; i B,
~t(n—k) fori=1,2,...,k (5-61)
s\a;;
Result (5-61) may be used to test an hypothesis about B, or set up a confidence
interval for B; in the usual way. However, we will not pursue the details further at
the moment as it is more efficient to develop a general set of inference procedures,
of which tests on asingle coefficient are just one particular application.

Sets of Linear Hypotheses


Consider the set of linear hypotheses about the elements of B, embodied in

RB =r (5-62)
where R is a known matrix of order q X k with q < k, andr is a known q-element
vector. We also assume R to have full row rank, that is, that there are no linear
dependencies between the hypotheses. It is extremely important to understand the
THE k-VARIABLE LINEAR MODEL 183

range of various hypotheses represented by Eq. (5-62). We illustrate them with


some examples.

PRS[Os eae Ql Oi ea] and r=0


Here R contains only a single row (g = 1) with a unit in the ith position and
zeros everywhere else, and ris the scalar zero. On substitution in Eq. (5-62)
we have
8, =0
Thus this specification of R and r sets up the hypothesis that B, is zero.
Choosing a nonzero value for r would set up the hypothesis that 8; is equal to
the specified constant.
Pe Roe08 eh dy ier] and r=0
produce the hypothesis
B, — B; = 0
or B, = B;
Se:Ree Ovi Oy telat lO wes >t iO} and r=1
specify the hypothesis
Papp
4.
OFF I 0 0
R=/|0 0 1 0
Fee wi :

of order (k — 1) X k and
0
0
r=].

0
of order (k — 1) X 1. This is equivalent to the joint hypothesis
B, 0
B; 0

By 0
that is, that the set of explanatory variables X,,X4,..., X, has no influence in
the determination of Y. This is a very important hypothesis. The test of this
hypothesis is often referred to as a test of the overall relation. Notice that the
hypothesis does not include B, = 0, since that involves the additional implica-
tion that the mean level of Y is zero. Our usual concern is whether the
hypothetical explanatory variables help to explain the variation of Y around
its mean value, but the actual level of the mean is of no particular importance.
184 ECONOMETRIC METHODS

5. R=[0 I,] and r=0


Here 0 is a null matrix of order s X (kK — 5) and r is an s-element column
vector. This sets up the hypothesis that the last s elements in f are jointly
zero,
Beste Bee ee
For example, in an equation explaining the rate of inflation the explanatory
variables might be grouped into two subsets—those measuring expectations
of inflation and those measuring pressure of demand. The significance of
either subset might be tested by using this formulation with the numbering of
the variables so arranged that those in the subset to be tested come at the end.

It is thus clear that a procedure for testing the general hypothesis RB = r will
be extremely useful and powerful, since various specifications for R and r will
cover a range of questions.
To develop such a test procedure, we first of all replace the unknown B
vector in Eq. (5-62) by the OLS vector b, obtaining the vector Rb. The more this
vector departs from r, the greater is the doubt cast on the hypothesis. The
problem is to determine the sampling distribution of Rb and devise a practical
test procedure. First of all, we see directly that
E(Rb) = RB (5-63)
and

var(Rb) = E{R(b — B)(b — B)'R’}


o*R(X’X) 'R’ (5-64)
Since b is multivariate normal,

Rb ~ N(RB, o?R(X’X) 'R’)


or R(b — B) ~ N(0, 0?R(X’X) 'R’) (5-65)
If the hypothesis (5-62) is true, we can replace RB in Eq. (5-65) by r, obtaining

(Rb — r) ~ N(0, o?R(X’X) 'R’) (5-66)


We can now apply Eq. (5-11) directly to Eq. (5-66) and write

(Rb — r)'[o?R(X’X) Se 'R’] eal ‘(Rb — r) ~ x2(q) (5-67)


where the degrees of freedom q are given by the number of elements in the Rb
vector.f
The only problem hindering practical applications of Eq. (5-67) is the
presence of the unknown o%, since all other elements are known. However, we

+ To show that the inverse of R(X’X)~'R’ exists, we show that

z7/R(X’X) 'R’z > 0


for z * 0, so that the matrix is positive definite and thus nonsingular. Define y = R’z. Then v is a
THE k-VARIABLE LINEAR MODEL 185

have already shown that

<Feeo ~ x(n k)
independently of b, and hence independently of Rb. Thus we can form an F ratio,
and the unknown o? will cancel out. The basic result is thus, if RB = ris true,

(Rb — r)’[R(X’x)'R’] ‘(Rb — r)/q 4


F(q,n-—k) (5-68)
e’e/(n — k)
The test procedure is then to reject the hypothesis RB = r if the computed F value
exceeds a preselected critical value. Now we must see what this test procedure
amounts to in some of the specific applications indicated above.

Testing a Single Coefficient


Rea (0 eee SO) ie Oe eer 0] and r=0

ith element

sets up the hypothesis


Hy: B,=90
We then have
Rb —r=5,

and R(X’X) 'R’


merely picks out the ith element a,, on the main diagonal of (X’X)~'. Thus the
test statistic becomes
b:2
F=— ~ F(1,n-k) (5-69)
Sa li

If instead of testing the hypothesis


B; = ()

one wishes to test the hypothesis that 8, assumed some specified value,

B; = Bio

linear combination of the rows of R. Since R has full row rank, v + 0. Thus

7/R(X’X)'R’z = v(X'X) ‘v
But (X’X) is positive definite by assumption, and so (X’X)~' is positive definite, since its eigenvalues
are the reciprocals of the eigenvalues of (X’X). Thus
v'(X’X) 'v>0
and so R(X’X) 'R’ is positive definite.
186 ECONOMETRIC METHODS

we simply set r = 8,9, and the test statistic becomes

F = (b, ce Bi)
———

sa,u
This is, of course, the same result as that already derived by a different route in
Eq. (5-61), since t?(n — k) = F(1,n — k).

Testing the Significance of the Complete Regression


Osler 0
R=/0 0 1 0
ORL Olea siete eke 6 mcs

of order (k — 1) X k andr = 0. The hypothesis now is Cr


By = By = + = Be 0
The vector Rb — r now reduces simply to the k — 1 vector of OLS regression
slopes

R(X’X) 'R’ picks out the (k — 1) x (k — 1) submatrix formed by the last k — 1


rows and columns in (X’X)~'. To see what is implied by this matrix, partition the
X matrix as
X=[i X,]
where i is a column of units and X, is the n X (k — 1) matrix of observations on |
all the explanatory variables. Then
ae n iX,
TalEXS IA XEXS
By Eq. (4-66) the right lower k — 1 submatrix in (X’X)~! is
] ve es
(X2X; — X,i-1X, | ='(X5AX,) —
from Eq. (5-41). Thus the F statistic becomes

F
__b4(X4AX,)by/(k — 1) (5-70)
e’e/(n — k)
From the decomposition of the total sum of squares in Eq. (5-49) above this is
seen to be
. ESS/(ke71}
RSS/(n = k) (5-71)
THE k-VARIABLE LINEAR MODEL 187

or, in terms of R?,


ine ee Nh)
(5-72)
BR) (ial)
The joint significance of the complete set of explanatory variables is thus tested by
computing F from any of these three formulas and seeing whether the computed
value exceeds the preselected critical value.

Testing the Significance of a Subset of Coefficients


Specifying
R=[0 I,] and r=0
sets up the hypothesis
Pee sean ec a0
We can always renumber the variables, if necessary, so that the subset of interest
comes at the end. Let us partition X and b conformably so that the complete OLS
regression may be written

b,
y=[X, X,] +e=X,b,+ X,b+e (5-73)
b,

where X, is of order n X (k — s) and denotes the first k — s columns in X, and X,


denotes the last s columns in X. Now
Rb —r=b,
and R(X’X)~'R’ picks out the square submatrix of order s in the bottom
right-hand corner of (X’X)~'. Let us call that submatrix C,,. From the partition-
ing of X above
ee XX, XX,
Os KX OTXAK
and from Eq. (4-66)
all
C,, = (X.X, rs X’,X,(X,X,) 'X,X,)

= (x, [1— x,(xx,)7'x,]x,}


= (X,M,X,) (5-74)
where
M, =I- X,(X,X,) 'X;, (5-75)
Thus the numerator in the test statistic, Eq. (5-68), becomest
by(X,M,X, )b,/s
+Do not confuse this s, which denotes the number of coefficients under test with the square root of
the residual variance defined in Eq. (5-57).
188 ECONOMETRIC METHODS

We will now show that this numerator has a very fundamental and important
interpretation in terms of sums of squares. Suppose y is regressed just on the
subset of variables in X,. Let e, denote the resultant vector of residuals. From Eq.
(5-54) we have
e,= M,y
where M,, is exactly the matrix just defined in Eq. (5-75).
Thus

eve, =residual sum of squares from regression of y on X,


e’e=residual sum of squares from regression of y on[X, X,]
e’e, — e’e= reduction in residual sum of squares due to adding X, to regression
= increase in explained sum of squares due to adding X, to regression

Our purpose is to show that


b/(X,.M,X, )b, = efe, — e’e
Return to Eq. (5-73) and premultiply by M, to get
M,y = M,X,b, + M,X,b, + Moe
= M_X,b. +e
for the definition of M, in Eq. (5-75) implies M,X, = 0 and M.e = esince
X’e = [Xie X‘e]=[0 0]. Transposing and squaring this equation gives
y’'M,y = b(/X,M,X,b, + e’e
but y'M,y = eve,
and the desired result follows. Thus the test statistic, Eq. (5-68), in this case
becomes
pa eee ee)/s ~ F(s,n—k) (5-76)
(e’e)/(n — k)
In words, the test of the joint significance of the subset X, is achieved by the
following steps:

1. Regress y on the variables X, which are not in the subset, and measure the
residual sum of squares e/e,.
2. Carry out the complete regression and measure the residual sum of squares
e’e. The difference efe, — e’e measures the reduction in the residual sum of
squares due to adding X, to the regression.
3. The mean square (e’e, — e’e)/s, associated with the subset, is then contrasted
with the overall mean square e’e/(n — k). If the resultant F value exceeds a
preselected critical value, the hypothesis that the variables in X, have zero
effect on Yis rejected.

The previous test for the joint significance of all the explanatory variables
may also be seen to be of the same form as Eq. (5-76). That test was based on
ESS/(k — 1)
RSH SK)
THE k-VARIABLE LINEAR MODEL 189

From Eqs. (5-49) we have

y Ay =b, X,AX,b, + e’e


(TSS) (ESS) (RSS)

Thus the above Fstatistic could be written

a WAy = ee)/(k = 1)
e’e/(n — k)
and y’Ay, which is the sum of the squared deviations of the Y values, can be
interpreted as a residual sum of squares when Yis regressed only on a vector of
units i, for replacing X, in Eq. (5-75) by i gives

Meee ait
n

This is the A matrix of Eq. (5-41), which transforms a variable into deviation
form. Thus ee, becomes y’Ay in this case.
The test of a single coefficient is merely a special case of the test of a subset.
Thus the ¢ or F test for the significance of a single coefficient may also be
interpreted in a sums of squares context. The test of

Hy: B; =a)
amounts to

. Regress y on all X’s except X;.


— . Regress y on all X’s.
NO

3. Compute the reduction in the residual sum of squares from step | to step 2
and contrast with e’e/(n — k).

Finally, we note another illuminating way of interpreting these various tests.


The regression of y on X, leading to the residual vector e, may be regarded as a
restricted regression. The essence of the restriction is that any coefficients which
are specified by the hypothesis to be zero in the population are actually set at zero
in the sample. Thus in testing 8, = 0, the restricted regression omits X,, so that in
effect b, = 0. Likewise, in testing the overall regression, the restricted regression
leaves out all variables except the unit vector, thus setting b, = b, = --- = b, = 0.
The complete regression may be regarded as an unrestricted regression, since all
the variables are included, and the estimated coefficients come out as the sample
data determine. Thus
ee,
/
residual sum of squares from restricted regression
e’e = residual sum of squares from unrestricted regression
and the test of the significance of a restriction, or the set of q restrictions, is
_ (ee, — e’e)/q
e’e/(n — k) Oe)
190 ECONOMETRIC METHODS

Confidence Intervals
Confidence intervals for a single B coefficient may be readily determined from the
result on the ¢ distribution in Eq. (5-61). Joint confidence regions for two or more
parameters may also be determined. From Eq. (5-65) we have

[R(b — B)]'[o?R(X’X)
'R’] '[R(b —B)]~ x2(q)
and, as usual,
,
ee

oO
= mean)
independently of b. Thus

p_Be
R(b — SOROS
B)]’|R(X’X) = 'R’| -1[R(b —
e’e/(n — k)
Appropriate specifications of R in Eq. (5-78) will yield confidence regions for
various groups of parameters. For example, setting R = J, and equating the
expression in Eq. (5-78) to some critical value F, gives a condition on the
unknown B vector from which a joint confidence region may be determined.

Example 5-2. Example 5-1 was based on


3 [eae Sie
] Yen oh 4
y=/8 and Sa) Te 56
3 Lest ad
5 1 4 6
The estimated regression was
Y=44+2.5X, — 1.5X,
with ESS = 26.5, RSS = 1.5, TSS = 28, and
R? = 0.9464
We also have n = 5 and k = 3. We will now illustrate tests of various hypotheses
with these data. It must be emphasized that the tests are simply meant to
illustrate the use of the formulas of this section. The data have been “cooked” to
give simple numbers, and the sample size is too small to allow any sharp
interpretations.

1. Testing the joint significance of X, and X,

Substitution in Eq. (5-71) gives

ESS
/(ee) e205 31)
a RSS/(w#—)) = 17.67
15/6 = 3)
THE k-VARIABLE LINEAR MODEL 191

From the tables of the F distribution, Fy ;(2, 2) = 19.00, so that the sample F
falls short of the 5 percent critical value. Even though the sample R? is
numerically high, the sample size is so small that it fails to reach significance.
. Testing the significance of X,

Result (5-61) states that


OTB;
t(n —k)
Sya li

where a,, is the ith term on the main diagonal of (X’X)~'. We do not need,
however, to invert the 3 X 3 matrix X’X. In the development of Eq. (5-70) we
showed that the right lower k — 1 submatrix in (X’X)~! is given by
(X‘,AX,)~', which is simply the inverse of the matrix of sums of squares and

xan['9§
products of the variables in deviation form. For this example we have
. 10 6

Thus
, -1_ 1 —1.5
Ce ns he be
giving a, = 2.5. Further, s* = e’e/(n — k) = 1.5/2 = 0.75. Finally, sub-
stituting — 1.5 for b, and 0 for B, gives the test statistic
=
ee Val
¥0.75 ¥2.5
which is insignificant.

Alternatively, we may show that the same numerical value for the test statistic
comes from the stepwise reduction in the residual sum of squares. It is again
simpler to work with the data in deviation form. Regressing Y on X, gives an
estimated regression coefficient of

ae
vs 10

The explained sum of squares due to this regression is then


b,X yx, = 1.6(16) = 25.6
and the residual sum of squares is
Ly? — b,Xyx, = 28.0 — 25.6 = 2.4 = ee,
When the complete regression X, and X, is run, we already know the residual
sum of squares,
RSS = e’e = 1.5
192 ECONOMETRIC METHODS

Thus substitution in Eq. (5-77) gives


w(etmeesg, 2415 is
e’e/(n — k) 1.5/2 ;
which is the square of the above f statistic, as it should be.

3. Coefficients of X, and X, equal in magnitude but opposite in sign

Hy: B, + B;
=0
From the general formulation
RB=r
this gives
R=[0 1 1] and’ r=0
with g = 1. The appropriate test statistic is given by the general result in Eq:
(5-68), namely,

p_ Rb = r)[R(X’x)'R’] ‘(Rb — r)/q


< e’e/(n — k)
We then have

and R(X’X)~'R only involves the elements in the 2 X 2 submatrix in the


lower right-hand corner of (X’X)~'. As we have already seen, this is
(X,AX,)~'. Thus
me eilety IR’et= [1 Heesgy tse
R(X’X) yale 0.5
Thus the test statistic becomes

b, + b,)°
aye) =2.66
0.75(0.5)
which falls well short of any usual critical value for F(1, 2). Thus the data are
not inconsistent with the hypothesis that 8, + 8, = 0.

4. 95 percent confidence interval for B.,


b, = 2.5
2_ ee
= 0.75
pei?
The top diagonal element in (X,AX,)~! = 1. Thus
var(b,) = 0.75
s.e. (b,) = 0.866
to.o2s(2) = 4.303
THE k-VARIABLE LINEAR MODEL 193

Thus the confidence interval is given by


2.5 + 4.303(0.866) = 2.5 + 3.7
that is,
= 206.2;
The fact that the confidence interval includes zero means that b, is not
statistically significant at the 5 percent level.

5. Joint confidence region for B, and 8, Returning to Eq. (5-78) we specify


Eupon tiie G z
ha 1 tineae
Thus

and
no-8)=|7]- [2]=|3-8
[R(x’x) 'R’] | = fe ‘|
Substitution in Eq. (5-78) gives

he
asm tsa’? SIL iS fl 1*5
—1.5 — B,

265-928) 18h 1B Bacl0B, 4B.


ai 1.5
Choosing, say, the 5 percent critical value of F, we have
Pr{F < Fyo;) = 0.95
Then setting
Palos
defines a 95 percent confidence ellipse for the unknown f parameters in F.
For this problem Fo ,;(2, 2) = 19. Setting F = Foy, then gives
108? + 128,
8, + 48? — 328, — 188, + 26.5
ee el
5
that is,

This defines the 95 percent confidence ellipse for 6, and B,, which is sketched
in Fig. 5-1. The ellipse is centered at the estimated point b, = 2.5, b, = —1.5.
There is a strong negative covariance between the two estimates and the
origin lies just inside the ellipse, in agreement with the result of test 1 above.

6. Point and interval forecasts Suppose we wish to forecast the value of Y


associated with X, = 10 and X, = 10. Plugging these values into the regres-
194 ECONOMETRIC METHODS

Figure 5-1 95 percent confidence ellipse for by, b3.

sion equation gives the point forecast


Y, = 4 + 2.5(10) — 1.5(10) = 14
A point forecast is of little use unless supplemented by a
measure of
precision, which enables us to put the forecast in interval form. The
forecast
may be written

b,
¥,=[1 10 10]| 6,| =Rb
b;
THE k-VARIABLE LINEAR MODEL 195

where R=[1 10 10]


The actual Y value in the forecast period will be
Y,= RB + uy
where u, denotes the actual value assumed by the disturbance in the forecast
period. Let us then define the forecast error e, as
e,= Y¥,—- ¥,= —R(b-
B) + u,
It follows immediately that
E(e,) =0
since E(b) = B and E(u,) = 0. Also
var(e;) = E{[—R(b — B)+ u,|[—R(b —B)+ u,|}
= o?R(X’X) 'R’ + 0?
using Eq. (5-64), and also the fact that u, will be independent of the sample
disturbances and hence independent of b. Thus
A

es ~ N(0,1)
o\1 + R(X’X) 'R’
Replacing the unknown o by
s=ee/(n—k)
then gives
ae ~t(n—k)
sy1 + R(X’X) 'R’
and so a 95 percent confidence interval for Y; is

¥, + too258/1 + R(X’X) 'R’ (5-79)


where R is a row vector containing the values of X in the forecast period
prefaced by unity in the first position. In this example
Sucelsi 25
3 [Link] aby ey
25° 81 1129
with
Ma Feds = 80
oxy =| aseeinn i “13|
zB eee ese eS
Thus
26.7 Asn 8.0 1
R(X’X) R’=[1 10 10]] 4.5 WU gp Reamalesya
8.0) teil.5 ZO MLO
1
=[-8.3 —-0.5 200 = 6.7
10
196 ECONOMETRIC METHODS

We also have
5” =1 0:75
and
toons (2) = 4.303

Thus substitution in Eq. (5-79) gives

14 + 4.303V0.75 V7.7 = 14 + 10.34

3.66 to 24.34
or
This is a prediction interval for Y,, the value of Y in the forecast period.
Sometimes an investigator prefers to set up an interval for E (¥;), that is, the
mean or expected value of .Y in the forecast period, the reason being that Y,
contains the disturbance u,, which is essentially unpredictable. We have
Y,= RB + u,

Thus E(Y,) = RB
and the forecast error would now be defined as

er= E(Y;) — Y= —R(b — B)


Following through the steps of the previous analysis then gives a 95 percent
confidence interval for E (¥;) as

¥, + to.o258/R(X'X)'R’ (5-80)
The numerical implementation of Eq. (5-80) gives

14 + 4.30370.75 V6.7

or 4.36 to 23.64

which is a slightly narrower interval than that for Y;.

There is an alternative way of generating either interval forecast, which is


simpler in that it only requires the inversion of a second-order rather than
a
third-order matrix, and which also provides an illuminating way of looking
at
the OLS regression. From Eq. (5-48) we can write the OLS regression
as
Y=Y+b.x.+b,x,+---+b,x,
+e
where, as in Chaps. 2 and 3, x, = X, — X, and so on. This is equiva
lent to
the regression of y on X, where X is now defined as

X=[i AX,]
THE k-VARIABLE LINEAR MODEL 197

i being a column of units and AX, the n X (k — 1) matrix of deviations. The


OLS estimator is then

n 0 i'y

0 X‘,AX, X’, Ay

0 i'y
I
X|-
So (X,AX,)"|| X,Ay
Yj

(X,AX,) 'X,Ay
where we used the result that i/AX, = 0 (the sums of sample deviations being
identically zero). The covariance matrix is
1
a 0
var(b) = o7| ”
0 (X,AX,)
The point forecast may be written
Y,= Y +x,b,
where x;=[x2, **: X,,] is a row vector of the X deviations in the
forecast period and b, is a (k — 1)-element column vector of the OLS
regression slopes. Thus
E(¥,) = £E(¥)+x,B,
and
var ¥,) = var(Y ) a x -E{(b, mtB, )(b, = B,)’}x’,

= o?|+ + x,(X,AX,)
x4
since the matrix var(b) above shows that Y and b, are distributed indepen-
dently. For the problem in hand,
Y=4 X,=3 X,=5
and so
x,=[7 5]
; -1_ 10 —-—1.5
ob %) wee |
and s? = 0.75. Thus the estimated var(Y;) is
0.75(0.2) + 0.75[7 oifea? a Te = 0.75(6.7)
198 ECONOMETRIC METHODS

The point forecast is


25
¥,=4+[7 5}| J=14
mea)
and thus the 95 percent confidence interval for E(Y;) is
14 + 4.303V0.75 V0.67
or 4.36 to 23.64
as before.

Prediction when the X Variables Are Uncertain


The treatment of interval forecasts given above assumes implicitly that the values
of the X variables in the forecast period are known with certainty. In practice,
however, it is more realistic to postulate some uncertainty about the X values. Let

xy= [1 Xap epee eX


indicate the true values of the explanatory variables in the forecast period and

R= [0 Xe tare exes|
the values that the forecaster thinks will be obtained in the forecast period. The
true value Y; is given by

Tee Paes
and the point prediction will now be
Y, = Xb
Thus the forecast error is

cia Yaa,

For simplicity we will drop the f subscript, since there is no ambiguity, and write
the forecast error as

e=u~X(b—B)—(R~
x)
= (UX (DiBaaRee es) (5-81)
If we assume that the forecaster makes unbiased forecasts of the X¥ values, that is,
E(&) =x
and, in addition, that there is zero covariance in the population between forecasts
of x and estimate of B from the sample data, then

E{%'(b — B)} = 0
and so
E(e)=0
Hence the variance of the forecast error is found by squaring Eq. (5-81) and
THE k-VARIABLE LINEAR MODEL 199

taking expectations, to give


of = E{u? + &’(b — B)(b — B)’R + B’(& — x)(& — x)B + cross-product terms)
(5-82)
On taking expectations the cross-product terms vanish because of the indepen-
dence of u, &, and b. For the remaining terms

E{B'(& — x)(& — x)'B} = BE((% — x)(& — x)}B


= B’var(%)B
and
E{&’(b — B)(b — B)’%) = E(tr[%’(b — B)(b — B)’S])
= E{tr[(b — B)(b — B)’&8’]}
= trl E{(b — B)(b — B)') E(&8’}]
Now
var(%) = E{(& — x)(& — x)’}
= E(&%’) — xx’
and so
E(&&’) = var(&) + xx’
Thus
E(8/(b — B)(b — B)’&) = tr[var(b) - (var(X) + xx’)]
and
tr[var(b)xx’] = tr[x’var(b)x] = x’var(b)x
Thus substituting these expressions in Eq. (5-82)
of = 0, + x'var(b)x + B’var(&)B + tr[var(b) - var(&)] (5-83)
If there were no uncertainity about the X values in the forecast period, this
expression reduces to
0, + x’var(b)x
which is the conventional formula for the variance of a forecast. To implement
Eq. (5-83) the various unknowns are replaced by estimated values:

eo,” is replaced by s” = e’e/(n — k) in the usual fashion.


ex is replaced by %.
e var(b) is estimated by the usual OLS program.
- B is replaced by the estimated b.

The main practical difficulty is likely to be in the estimation of var(X), the


variance-covariance matrix of the forecast X values. There may be accumulated
experience in forecasting from which variances and covariances may be estimated.
200 ECONOMETRIC METHODS

Alternatively, the forecaster may have subjective assessments that a forecast value
is very likely to be within, say, 5 percent of the true value, which in turn implies a
figure for the variance.
The remaining practical difficulty about the use of Eq. (5-83) is that we can
no longer determine exact confidence intervals using the ¢ and normal distribu-
tions. The reason is that even if normality is assumed for & as well as u, the
forecast error in Eq. (5-81) is not normally distributed since it involves %’(b — B),
which is the sum of products of normal variables. One may follow the suggestion
of Feldstein to use the Chebyshev inequality to determine an outer-bound
forecast interval.+ The practical procedure is as follows. Letting s; denote the
square root of the estimated value of Eq. (5-83) we can state:

The probability that the observed value of Y in the forecast period will fall
outside the interval Y; + cS, does not exceed 1/ Ca

The researcher can set the value of c to make 1 /c” equal to 0.05 or whatever
is desired. The Chebyshev inequality strictly involves the true o,, but it is a very
conservative statement and unlikely to be seriously affected by the replacement of
o; by sy. If the distribution of y were sufficiently well behaved to be unimodal
and symmetric, the probability of Y;, lying outside the interval i + cs would not
exceed 4/9c?.

PROBLEMS

5-1 Test the hypotheses (N.B. plural) 8, = 1, B) = 1, 8; = —2 in the regression model


Ye Bo a BX, a B,X>, uF BX, + u,
given the following sums of squares and products of deviations from means for 24 observations:

Ly? = 60 Lx? = 10 Dx2 = 30 ie =a


Lyx, =7 Lyx,= -7 Lyx,= —26

Lx,x> = 10 Lx,x3 = 5 Uxx3 = 15


Test also the hypothesis that 8, + 8, + B; = 0. How does this differ from the hypothesis that
[B, 8, B3]=[l 1 —2]? Test the latter hypothesis.
5-2 The following sums were obtained from 10 sets of observations on Y, X,, and X):

LY = 20 XX, = 30 LX, = 40
LY? = 88.2 UXP = "92 DEX a168
LYX, = 59 LYX, = 88 LX, X, = 119
Estimate the regression of Y on X, and X>, and test the hypothesis that the coefficient of X, is zero.
5-3 Let

y be ann X | column vector


X ann
X k matrix

7M. S. Feldstein, “The Error of Forecast in Econometric Models when the Forecast-Period
Exogenous Variables are Stochastic,” Econometrica, 39, 1971, pp. 55-60.
THE k-VARIABLE LINEAR MODEL 201

X{i] be X with the ith column (x;) removed


e,; be the residual vector from the regression of y on X{i]
e; be the residual vector from the regression of x; on X[{i]

Now consider the two regressions:

1. e,; one;
2. yonX

Prove that:
(a) The slope b = e},e;/e;e; from regression 1 and the multiple regression coefficient b; from
regression 2 are identical.
(b) The residuals from the two regressions are identical.
(c) The simple correlation between e,; and e; is the same as the partial correlation between y and
x, 1n regression 2.
5-4 The following regression equation is estimated as a production function for Q:
logQ = 1.37 + 0.632 log K+ 0.452 log L
(0.257) (0.219)
R?=0.98 — cov( bg, b;) = 0.055
and the standard errors aregiven in parentheses.
Test the following nullhypotheses:
(a) The capital K and labor L elasticities of output are identical.
(6) There are constant returns to scale.
(University of Washington, 1980)
Note: The problem does not give the number of sample observations. Does this omission affect
your conclusions?
5-5 Consider a multiple regression model for which all classical assumptions hold, but in which there
1s no constant term. Suppose you wish to test the null hypothesis that there is no relationship between y
and X, that is,

Ho: B= = By =0
against the alternative that at least one of the 8’s is nonzero. Present the appropriate test statistic and
state its distribution (including the appropriate number(s] of degrees of freedom).
(University of Michigan, 1978)
5-6 One aspect of the rational expectations hypothesis involves the claim that expectations are
unbiased, that is, that the average prediction is equal to the observed realization of the variable under
investigation. This claim can be tested by reference to announced predictions and to actual values of
the rate of interest on three-month U.S. Treasury Bills published in The Goldsmith-Nagan Bond and
Money Market Letter. The results of least-squares estimation (based on 30 quarterly observations) of
the regression of the actual on the predicted interest rates were as follows:
r= 0.24 + 0.94 r*+e,, RSS = 28.56
(0.86) (0.14)
where r, is the observed interest rate, and 7;* is the average expectation of 1, held at the end of the
preceding quarter. Figures in parentheses are estimated standard errors. The sample data on r* give

Y*/30=10, (re — 7)? = 52


t t

Carry out the test, assuming that all basic assumptions of the classical regression model are satisfied.
(University of Michigan, 1981)
5-7 Consider the following regression model in deviation form:
Vp = ByX1,+ BoxX2, + u,
202 ECONOMETRIC METHODS

with sample data:

n= 100," Ly? _ 493


3 > Ux? = 30, Exe = 3,
Xx,y=30 Lx.y=20, LUx,x,=0
(a) Compute the OLS estimates of B,, 8), and R?.
(b) Test the hypothesis Hy: $B) = 7 against H,;: B, = 7.
(c) Test the hypothesis Hy: £, = B, = 0 against H,: B, + 0 or B, = 0.
(d) Test the hypothesis Hp: 8, = 7B, against H;: B, + 7B).
(UL, 1981)
5-8 Given the following least-squares estimates:

C, = constant + 0.92Y, + e;,

C, = constant + 0.84C,_, + e5,


|= constant + 0.78Y, + e3,
A

Y, = constant + 0.55C,_, + e4,

calculate the least-squares estimates of B, and B; in

C,= B, + BY,
+ B3C,_, + u,
(University of Michigan, 1981)
5-9 Prove that R? is the square of the simple correlation between y and y, where y = ».(0... Gz
5-10 Prove that if a regression is fitted without a constant term, the residuals will not necessarily sum
to zero, and R?, if calculated as 1 — e’e/(y’'y — nY?), may be negative.
S-11 A researcher wishes to estimate the regression of y on X without an intercept term, that is, X
does not contain a column of Is. Unfortunately, the regression program at hand automatically
computes an intercept term. Douglas M. Hawkins suggests that the program can be “tricked” into
estimating the correct intercept free regression by entering each data point twice—once in its correct
form (y;,x;) and once with the opposite sign (—y,, aXe)
Prove that:
(a) The “trick” regression and the correct regression (with intercept suppressed) yield the
same
coefficients for X.
(b) The residual sum of squares from the “trick” regression is exactly double the value
from the
correct regression.
Compute the ratio of the standard errors of the two regressions.
(American Statistician, 34, Nov. 1980, p. 233)
5-12 (a) Prove that R* increases with the addition of an extra explanatory variable
only if the F
(= 17) statistic for that variable exceeds unity.
(b) Prove that the partial

r=
A F =
t*
FF redheads
where1 is the value of the statistic for testing the significance of the coefficie
nt of the X; to which the
partial r is related, and df is the number of degrees of freedom in the regression.
5-13 Let the regression equation be partitioned as

y= X,B, + XB, +e
Let b, and b, be the usual least-squares estimators. Suppose that E(e)
= X,y, that is, the mean vector
of the disturbances is a linear combination of some of the regressors.
Prove that b, is biased but b, is
unoviased.
(University of Michigan, 1981)
THE k-VARIABLE LINEAR MODEL 203

5-14 Suppose that the m X 1 vector x; denotes m observations on the ith individual (i = 1,..., p)
and x; is the corresponding vector of deviations from the ith sample mean. Let the x, and x; vectors
be “stacked” to give mp X | vectors
x= (x, x, --: x5]
and v= [k R ¥]
Find a matrix D such that Dx = x.
CHAPTER

SIX
FURTHER TOPICS IN THE
k-VARIABLE LINEAR MODEL

6-1 ESTIMATION SUBJECT TO LINEAR RESTRICTIONS

In Chap. 5 we have described the procedure for testing the hypothesis that the
elements of the population vector B obey the set of q (< k) linear restrictions .
embodied in the relations
Hy: RB=r
If Ho is not rejected, one may wish to reestimate the model, incorporating
the
restrictions in the estimation process. One important reason for such
reestimation
is that it will improve the efficiency of the estimates. This produces
an estimator
b, which then satisfies

Rb, =r (6-1)
For example, if the hypothesis of constant returns to scale is not
rejected for a
production function, the reestimation process would yield a produc
tion function
with estimated elasticities which sum to unity.
We must first of all show how to derive an estimator b, which
satisfies Eq.
(6-1). Second, we will use this estimator to cast new light on
some of the test
procedures of Chap. 5, and third, we will look at some important
applications of
the new estimator.
The assumed model, as before, is

y=Xp+u
FURTHER TOPICS IN THE k-VARIABLE LINEAR MODEL 205

We define the scalar function

= (y — Xb,)’(y — Xb,) — 2\’(Rb, — r) (6-2)


where X denotes a column vector of q¢ eee multipliers. Taking the partial
derivatives of » gives
a
ab,? = —2X’y + 2X’Xb, — 2R’X
Op _
and ae 2(Rb, — r)
Setting these partial derivatives to zero gives the equations to be solved for b, and
X, namely,t
X’Xb, — X’y — R’X = 0 (6-3)
Rb, —r=0 (6-4)
Premultiplying Eq. (6-3) by R(X’X)
7! gives

Rb, — R(X’X) 'X’y — R(X’X) ‘RX =0


Using Eq. (6-4) and resurrecting the OLS estimator of Chap. 5, that is,

= (XX) 'Xy
this equation may be solved for \ as

d = [R(X’x)'R’] '(r — Rb)


Substituting back in Eq. (6-3) gives

b, = (X’X) 'X’y + (XX) 'R’[R(X’X) 'R’] '(r — Rb)


that is,
b, = b + (XX) 'R’'[R(X’X) 'R’] '(r — Rb) (6-5)
where b is the unrestricted OLS estimator (X’X)~'X’y. Formula (6-5) defines the
restricted least-squares estimator satisfying the set of q restrictions embodied in
Rb, = r.£ Corresponding to b,, we may define the residual vector
e, = y — Xb,

+ To keep the notation as simple as possible, we have not distinguished between the vectors b, and
X which appear in the objective function, Eq. (6-2), and the specific vectors that emerge as the
solutions to Eqs. (6-3) and (6-4).
+ Provided the restrictions RB =r are true, the variance-covariance matrix of the restricted
least-squares estimator may be shown to be
var(bs) = 02{(X’X) | — (xx) 'R'[R(X’x)'R’]'R(X’X)“'}
See Problem 6-6. We should also note that in some problems it may be simpler to obtain b, by
imposing the restrictions directly on the problem rather than by substituting in Eq. (6-5). For example,
suppose the data are already in deviation form and we wish to estimate
y = Box.
+ B3x3 + u
206 ECONOMETRIC METHODS

which may be written


ex = y — Xb — X(b, — b)
e — X(b, — b)
where e is the OLS residual vector. Transposing and multiplying
e,e, = e’e + (by — b)’X’X(b, — b)
the cross-product term vanishing since X’e = 0. Thus the difference between the
restricted and the unrestricted residual sums of squares may be written
e,e, — e’e = (b, — b)’X’X(b, — b) (6-6)
Substituting for b, — b from Eq. (6-5) and simplifying gives}

ee, — e’e = (r — Rb)’'[R(X’X) 'R’] '(r — Rb) (6-7)


The right-hand side of Eq. (6-7) is exactly the expression in the numerator of the
F statistic for testing Hj): RB =r derived in Eq. (5-68). Thus an alternative
expression of the test statistic for Hj): RB =r is

— (exes— e'e)/q
Het e’e/(n — k) ic)
where ee, denotes the restricted residual sum of squares derived from the vector
b,, which satisfies the q restrictions Rb, =r, and e’e denotes the unrestricted
residual sum of squares from the usual OLS regression. We have already derived
this result for one particular application in Eq. (5-76), but the derivation leading
up to Eq. (6-8) is perfectly general and applies to all cases.
To summarize, the test of the hypothesis that the elements of B obey aset of g
(< k) linear restrictions embodied in

Hy: RB=r
may be carried out by computing the unrestricted OLS vector b and the residual
vector e and then calculating the F statistic, Eq. (5-68),

ahs: Rb)’[R(X’X) 'R’] '(r — Rb)/q


e’e/(n — k)
rejecting Hy if F exceeds a preselected critical value taken from the
F distribution
with g, n — k degrees of freedom. Alternatively one may compute
the restricted
vector b, from Eq. (6-5) or otherwise, and the correspondin
g residual vector

Subject to the restriction 8, + B; = 1. Substituting the restriction


in the equation gives
Y = B)(x3)
x2—+x3+4
so that the regression of y — x; on x2 — x3 yields an estimat
e b,, and b; is then obtained from
b; = 1 — by. For the general version of this approach, see Problem
6-11.
7 As shown in the footnote on page 184, [R(X’X)~ 'R’] | is positive definite. Thus e,e, — e’e > 0,
with equality only when Rb = r. Imposing restrictions cannot
lower the residual sum of squares.
FURTHER TOPICS IN THE k-VARIABLE LINEAR MODEL 207

e, = y — Xb,. The test statistic is then

b, — B
re (bet b)’X’X(b, — b /a
yeR(bsTih) (6-9)
e'e/(n — k)
or equivalently
F= (exe, — e’e)/q
e’e(n — k)
One of the most useful applications of these formulas is in tests of structural
change.

6-2 TESTS OF STRUCTURAL CHANGE

Example 6-1 Suppose we have data on two variables,


Y = consumption expenditure
X = disposable income
The data cover two distinct subperiods, n, observations relating to wartime
years and n, observations relating to peacetime years. Suppose we wish to
investigate whether there is any change, or shift, in the consumption function
between the wartime and peacetime periods. Such a change is referred to as a
structural change or structural break. Let us denote the consumption func-
tions by
Y=a,+8,X+u — wartime function (6-10a)
Y=a,+8,X+u peacetime function (6-10b)
This is the unrestricted form of the model, allowing intercepts and slopes to
be different in the two periods. This model would be set up in matrix form as
follows:
Y, ai
Y, LLIeX = 10 0 Uy
Lacy X,0 MO 0
S10 @ 5p) 6) 6. @ 8 wie 8 eofe 6 6 Qa,

Y,, Lane 8 0 B u,
=---|=|-------<-- Ve ea (E11)
Te OP 10s 1 PEA Aa atae2 Un +1
Yr? OO Ae By Uy, +2
Oi ie)coy, eee
ny +ny Unitny

where the wartime observations have been listed first and the peacetime
208 ECONOMETRIC METHODS

observations last. More compactly Eq. (6-11) becomes


a
yi Xie OU eri ay, :
y= [hf=[7 4 a, Caenan (6-12)
2
where the data matrix X is block-diagonal.t Notice that each of the sub-
matrices X, and X, has a column of units in the first position followed by a
column of observations on income, and fB indicates a column vector of the
four structural parameters. Applying OLS to Eq. (6-12) gives

ay

b= |i]a, = (xx) xy
b J

b,

i (GX) 0 aca
0 (XEX3) ssa
(XX) 1,
if: (6-13)
(X’,X,) X5Y>
These estimates are seen to be identical with those obtained by applying OLS
separately to Eqs. (6-10a) and (6-106). One merely sets the data up in the
form of Eq. (6-12) and a single regression will produce all four regression
parameters. Using Eq. (6-13) one can then obtain the vector e of ny +n,
residuals, and e’e gives the unrestricted residual sum of squares.
Now set up the null hypothesis of no structural change. This may be |
formulated as
a) a
qs Hi 7% a ae
or, putting it in the RB = r framework,

a)
li, 60. 3h EO 0
Ho: K [nO | a, -(¢]
B,
so that
R = [I - I] and ‘r=0 (6-15)
+ When using computers the student must take care to understand the propertie
s of the program
being used. If the program automatically estimates an intercept, feeding
in the block-diagonal X
matrix would produce a linear dependence between the column of units supplied
by the computer and
the first and third columns of X. Thus one must either feed in X as
it stands and suppress the
automatic intercept, or else allow the automatic intercept and modify the
X matrix in a way to be
discussed later in this section.
FURTHER TOPICS IN THE k-VARIABLE LINEAR MODEL 209

where I is a unit matrix of order 2. The restricted model may thus be


formulated as
a
+u (6-16)
|:ae | | X,3 ||
|8
The contrast with the unrestricted model in Eq. (6-12) is that the X, matrices
are now stacked vertically, so that only two parameters are required to
describe the relation.
We now have two alternative procedures for testing H,. Using the
unrestricted b computed from Eq. (6-13) and the R matrix and r vector
defined in Eq. (6-15), we can calculate the F statistic defined in Eq. (5-68).
Alternatively, we may compute the e, vector that comes from fitting the
restricted model (6-16) by OLS and substitute either in Eq. (6-8) or in Eq.
(6-9). The second procedure is the simpler, but we will illustrate both with the
following data. Again these are artificially simple numbers, designed to keep
the arithmetic to the minimum and highlight the methods.
Wartime data
1 1 az
2 1 4
Vs ne X, =| 6
4 110
6 1pad3

Peacetime data
1 ] 2
3 1 4
3 ] 6
5 1 8
6 Wess lO)
NG wee 1
if bala
9 Lat
9 eels
pl Be

The sample sizes are


n,=5 n, = 10 n= 15
and we have
Pet de OD eh wet lO i
eee: | (Son) ae 1540
5 Pef<325 " = 35 Fetes eel 154 a
(x, X,) -z,|
, ental ora
3] C972) fs 380i be Ly deste = —

erates eae 60
XY, ela 292 lee
yiy, = 61 Y2¥2 = 448
210 ECONOMETRIC METHODS

Substitution in Eq. (6-13) gives the unrestricted OLS estimator

ae — 0.062500
ye Ei m (X/X,) XY, ae 0.437500
b, (Xexey 1X25; 0.400000
0.509091
Thus the estimated regressions are
Y = —0.0625 + 0.4375X wartime
and Y = 0.4000 + 0.5091X peacetime
These point estimates give the wartime function a smaller intercept and lower
slope than the peacetime function. The residual sum of squares from the
wartime regression is
ere; = yiy, — bi Xiy,
= 61 — [-0.0625 0.4375]| oF = 61 — 60.3125
= 0.6875
Similarly for the peacetime regression

e@5€) = Y,¥, — BL X4y,

= 448 — [0.4 0.509091}] ,60


60
|= 448 — 445.5273
= 2.4727
Thus the unrestricted residual sum of squares is
e’e = ele, + ee,
= 3.1602
In fitting the restricted model, Eq. (6-16), the data matrix is now

giving

(X4X,) = (X{X, + X,X,)


X4y = Xiy, + X4y,
From the data,

(sents 15
(X4Xx)
».¢ xX =
145 te ane eee
f =>
sel
The restricted coefficient vector is

[¢]- 1 1865 ee ae ese


b 6950 | — 145 15 || 968 0.524460
Thus the estimated common regression is

Y = —0.0698 + 0.5245.
FURTHER TOPICS IN THE k-VARIABLE LINEAR MODEL 211

and the restricted residual sum of squares is

ee, = (61 + 448) — [—0.069784 0.524460]| a = 509 — 502.4435


= 6.5565
Thus substitution in Eq. (6-8) gives
_ (6.5565 — 3.1602)/2
= 5.91
3.1602/(15 — 4)
Notice that the number of degrees of freedom in the numerator is 2, since
there are two restrictions embodied in Hp, and in the denominator n — k =
15 — 4 since four regression coefficients are estimated in the unrestricted
regression. From the tables of the F distribution,
Fyo5(2,11)=3.98 and Foo9(2,
11)=7.21
Thus the hypothesis of no structural change would be rejected at the 5
percent level of significance, but not at the 1 percent level.
The alternative approach is to calculate only the unrestricted b vector
and substitute directly in Eq. (5-68). This requires the evaluation of
R(X’X)~'R’. From Eas. (6-12) and (6-15),
: =I
-]] (X{X,) 0 I
R(X’X) 'R=[I ; ae
=|

(XX ,)0 ot (X4X5)5


1.27916667 —0.12083333
— 0.12083333 0.01553030
with
ry) hp] 7! 2.949641 22 .965567
[R(X’x) R| eee oe

— 0.062500
= ut 0.437500 as — 0.462500
Reg Lege! ~ 0.400000 eee
0.509091
and
r = 0. Thus

(r — Rb)'[R(X’X)'R’] '(r — Rb) = 3.3969


and from the previous calculations,
ee, — e’e = 6.5565 — 3.1602 = 3.3963
so that the two numbers agree, subject to rounding errors in the calculation.
The Fstatistic from Eq. (5-68) is

SS
3.3969/2
3,1602/11 5.91
A as before
f
212 ECONOMETRIC METHODS

Example 6-2: Tests of change in the regression slope Example 6-1 showed
how to test the hypothesis
i a, a,
PE tienes
The restricted and unrestricted models are pictured in Fig. 6-1.
Sometimes the investigator is more interested in testing for the homo-
geneity of the regression slope, the values of the intercept term being of no
particular importance. The null hypothesis is now specified as
Hy: By = 8 (6-17)
The @ parameter is free to take on different values in the two subperiods. For
instance, in simple Keynesian theory the size of the national income multi-
plier depends only on the marginal propensity to consume £ and not at all on
the intercept a. Thus the H, in Eq. (6-17) is equivalent to asking whether the
income multiplier is the same in each subperiod. The restricted and unre-
stricted models may then be set up as follows:
Restricted Unrestricted
a)
. a 3
vie 0 x, A aw eto a 0 9071) 8, coe
Y2 0 i, x, B Y2 Oe OF exe
B,
(6-18)
where i, denotes a column vector of n, units, i, a column vector of nN, units,
x, a column vector of the n, observations on wartime income, and X, a
column vector of the n, observations on peacetime income. OLS may then be |
applied directly to each model in Egs. (6-18) and H, tested by comparing the

aS (a 2,8)

(a1,8))

(a,8)

Ue Xi ov ~ X
(a) (d)
Figure 6-1 (a) Restricted model; (b) unrestricted model.
FURTHER TOPICS IN THE K-VARIABLE LINEAR MODEL 213

residual sums of squares from the restricted and unrestricted models in the
usual way. The two models are shown in Fig. 6-2.
The unrestricted model in Eqs. (6-18) is exactly the same as that in Eq.
(6-12), so we already have
e’e = 3.1602.
For the restricted model,
sea la ee gs
eG Ls hex
so

ny; 0 ix, iy;


XX, =| 0 Ny ix, X4y = 1,Y>
Xi, XQI, XX) + XOX, XiY) + X4Y5
5 0 35 15
=| 0 10 110 =! 60
357 =REO, & 1865 968
Thus
, , , ah ,
exe, = yy — y’X(X4X,4) X4ay
5 0 aot e15
= (61 + 448) —[15 60 968]] 0 10 110 60
359.110), 1865 968
1 6550 3850 = 3503 //215
= 509 — 15 60 a 3850 8100 —550]} 60
20,500
= 350 a0) 50 || 968
= 509 — 505.5098 = 3.4902

>
=e

x O = OG

(a) (bd)
Figure 6-2 (a) Restricted model; (6) unrestricted model.
214 ECONOMETRIC METHODS

The test statistic for the null hypothesis that 8, = B, is then


3.4902 — 3.1602

and Fo ;(1, 11) = 4.84. Thus there is no evidence of a significant difference in


the regression slopes in the two periods. Since the F statistic in Example 6-1
was on the borderline of significance, this suggests that any change between
the two periods lies in the intercepts rather than the slopes.
Before leaving this example, let us note a much simpler way of calculat-
ing e,e,, based on deviations from the subperiod means. We have earlier
introduced in Eq. (5-41) the A matrix which transforms a vector of n
observations into deviation form. Let us define A, and A, to be such matrices
for use with vectors of n, and n, observations, respectively. Premultiplying
the restricted model by the block-diagonal matrix
A, 0
0 A;
then gives
a
ae “|, 0 Ax, of a
A>Y> 0 0 A;x, B

since A,i,; = 0 and A,i, = 0. The 8 parameter is thus estimated by a simple


regression of
|
Aiy, a A\x,
A>Y2 A}X>
Each vector consists of two subvectors, the first being the deviations of the
wartime Y (or X) from the wartime sample mean and the second the
deviations of the peacetime Y (or X) from the corresponding peacetime —
sample mean. Denoting the vectors of deviations by ¥ and X, respectively,
~~

p- X’x
and iat — KKK)
exe, = VF Weep eyes
KF
Computing the deviations for the two subperiods and evaluating these
expressions gives
203
b= Ale 0.4951
and
2
ee, = 104 -— (203)"
410
3.4902 as before
Example 6-3: Testing for structural change in the intercept The null
hy-
pothesis is now
Hy: a, =a, (6-19)
We must be very careful in the specification of the restricted and unrestr
icted
models. By analogy with Example 6-2 it might seem reasonable to specify
the
FURTHER TOPICS IN THE k-VARIABLE LINEAR MODEL 215

restricted model as

y my beky ix 1 “elles
| : hi 0 Fe 4 uae Gey)
with the unrestricted model as before. Model (6-20) imposes a common
intercept but specifies different slopes. If the functions have different slopes,
they must intersect at some X value. There may be cases where it is relevant
and important to test that the intersection occurs at X = 0, as is implied by
specifying Eq. (6-20) as the restricted model. However, this is not usually the
case, and the most common practice is to test Hy, subject to the assumption of
a common regression slope. Thus the restricted and unrestricted models
become

pI ples Bile tal]


Restricted Unrestricted

i, x ise oe ey et
y2 In Xo yp 0 i, x, B

Notice that the unrestricted model in this example is the restricted model of
Example 6-2 [see Eqs. (6-18)], and the restricted model here is the same as the
restricted model in Example 6-1. Thus from our previous calculations the
relevant sums of squares are
ee, = 6.5565
and e’e = 3.4902
Thus the test statistic for Hj: a, = a, conditional on a common 8, is
6.5565 — 3.4902
F=~3.4902/12 = 10.54

and Fj 99(1, 12) = 9.33 so that the difference in the intercepts is significant at
the 1 percent level. The models are shown in Fig. 6-3.

4 Yi

x O Fe
(a) ()

Figure 6-3 (a) Restricted model; (b) unrestricted model.


216 ECONOMETRIC METHODS

Y Ys
' f

Y,
Y, d
Ya an

Y,

“1
O > xX Oo Xe
(a) | (b)
Figure 6-4

This test is a simple example of the analysis of covariance, which has


widespread applications. Suppose, for instance, that Y denotes the yield of wheat
per acre and X indicates hours of sunshine. One set of observations relates to
strain 1 of wheat and the other set to strain 2. The crucial question is whether one
strain shows a significantly different yield than the other, but suppose that the
experimental plots sown with the two strains have not received equal amounts of
sunshine. The difference between the sample means would then not only reflect
any possible difference between the strains, but also the difference due to
the
variable hours of sunshine received. In the analysis of covariance, hours
of
sunshine would be termed an intervening variable. The problem is depicted
graphically in Fig. 6-4.
In Fig. 6-4a the single line denotes the assumed positive relationship betwee
n
yield and hours of sunshine, and the ellipses denote the samples from each
strain.
The difference between the sample means is here due solely to a differen
t set of X
values for the two varieties. In Fig. 6-4b strain 2 is assumed to have
a greater yield
than strain 1, a difference indicated by the difference between the
intercepts,
deny
The observed difference between the sample means is then d plus
or minus any
differential effect due to sunshine. In testing the hypothesis
Hy: a, =a,
proper allowance must be made for any possible interference
from the intervening
variable, but this is precisely what is achieved by the test of the
models in Eqs.
(6-21), namely, the test of Hy, conditional on the assumption of
a common 8B.
Summary
These three examples are based on a hierarchy of models:
I }%] [hs x} Qa common regression
Y2 ba B ae for both periods
FURTHER TOPICS IN THE k-VARIABLE LINEAR MODEL 217

é a
I spe a hs Xj ns re differential intercepts,
Yy> OI xX, B common slope

O47 |e
Yoo Ay OF OR, differential intercepts,
Ill = : +
Y> 0) Oi, xy differential slopes
B,
Fitting each model by OLS produces a residual sum of squares. There are
three basic tests on the differences between the various residual sums of squares.
These are the following:
e Test of differential intercepts—model I contrasted with model II
e Test of differential slope coefficients—model II contrasted with model III
e Test of differential regressions—model I contrasted with model III (slopes and
intercepts)

The tests outlined above have implicitly assumed that the disturbance vari-
ance o” is the same in each period. Schmidt and Sickles have investigated the
effect of departures from this assumption on the significance level of the test.f For
equal-sized samples there are modest increases in the true significance level over
the nominal level, even for very large departures from the assumption of equal
variances. For instance, with n, = n, = 25 the true significance level only rises to
0.059, compared with a nominal value of 0.05, when one variance is 100 times the
other. If the X variable is a linear trend, the true significance level rises to 0.063
for a tenfold increase in the variance and to 0.084 for a one-hundredfold increase.
When the sample sizes are unequal, the true significance level shows a greater
departure from the nominal level, and it may now be less or greater than the
nominal level. Full details are given in the reference.

Example 6-4: Tests of structural change (k variables) The previous three


examples have only been concerned with a two-variable model for two
subperiods. The tests need to be generalized in two directions, namely,
extending to k variables and also making comparisons between more than
two subperiods. In this example we make the first extension.
The unrestricted model is now
Y x, 0 B; +u (6-22)
Y2 : 0 X,}/B,
where X, is of order n, X k, X, is of order n, X k, and B, and B, each
denotes vectors of k coefficients. This model has exactly the same matrix form
as Eq. (6-12), the only difference being in the number of variables in the
X,,X, matrices. Let us partition X, and X, by the first column of units and
the remaining k — 1 columns of observations on the explanatory variables as

+P. Schmidt and R. Sickles, “Some Further Evidence on the Use of the Chow Test under
Heteroscedasticity,” Econometrica, 45, 1977, pp. 1293-1298.
218 ECONOMETRIC METHODS

follows:
X, 7 [i, XT]

and Xo = [i> x3]


We may then construct the same hierarchy of models as in the two-variable
case. This now gives

Vite tiie sn common regression


: Vale mci exs oem for both periods
Ox differential intercepts,
II HEF ; 4 + common vector of
Y2 Cae oan p* regression slopes

a
Il Yu ie ee pee Ota a differential intercepts,
f ¥ | O15 01 XS TR . differential slopes
By
where we have partitioned the k-element B vector as

a
B, 3
STB a
By,
Application of OLS to each model will yield a residual sum of squares (RSS) .
with an associated number of degrees of freedom as indicated by
Model I RSS, n—k
Model IT RSS, Ne ike ab
Model III RSS, R= 2k
where n = n, + n, indicates the total number of observations in the com-
bined samples. The test statistics for various hypotheses are then as follows:

Hy: a, = a): Test of differential intercepts


RSS,
— RSS,
RSS (Gee) ~ F(1,n—k-—1)
iF oes
eee (6-23)
Hy: BY = Bx: Test of differential slope vectors

< (RSS, — RSS,)/(k — 1)


=2E)ee
RS/S, ee lata 2k O24)
FURTHER TOPICS IN THE k-VARIABLE LINEAR MODEL 219

Hy: B, = B,: Test of differential regressions (intercepts and slopes)


(RSS, — RSS, )/k
F= ~ F(k,n — 2k) (6-25)
RSS,/(n — 2k)
The degrees of freedom in the numerator are simply obtained as the difference
in the degrees of freedom of the residual sums of squares in the numerator.
This is equal to the number of restrictions involved in going from the
unrestricted to the restricted model. For example, only one restriction is
imposed in going from model II to model I, and k — 1 restrictions (equality
of regression slopes) are imposed in going from model III to model II.
However, a further test is possible in the k > 1 case, which did not arise
in the two-variable model. We may now test whether a subset of coefficients is
stable over the two periods. For example, most wage equations take the form
Wage change = f (market pressure, expectations of inflation)
One might wish to test whether the reaction to the market pressure variables
has changed between two periods, or alternatively, one might hypothesize
that the reaction to inflationary expectations is different in “high” inflation
and “low” inflation periods. The principle of the test is the same as in all the
previous examples, as shown by the following rule.
Fit the restricted model with the subset of coefficients whose stability is
being tested, taking the same value in each subperiod, and compute the
residual sum of squares e,e,. All other coefficients are left to vary between
the two subperiods. Then fit the completely unrestricted model, where all
coefficients are free to vary, with the residual sum of squares e’e. The test of
the stability of the subset is then based on

oe (exes — e'e)/q
e’e/(n — k)
where g indicates the number of coefficients in the subset. Formally the
restricted model is set up as

Bi,
Yi Xi 0 ue B
= +u 6-26
Ki |0 XX, Xx» B, ( )
D

where X,,=7, X (k — q) matrix of observations in period | on the variables


not being tested
X,, =n, X (k — q) matrix of observations in period 2 on the variables
not being tested
Xj) =", X q matrix of observations in period 1 on the variables in the
test
X55) =n, X q matrix of observations in period 2 on the variables in the
test
B,,, Bo; =coefficient vectors of X,, and X,,, respectively
B, =common coefficient vector for X,, and X,,
220 ECONOMETRIC METHODS

Example 6-5: Tests of structural change (n, < k) A special problem arises if
one of the subperiods has fewer observations than the number of parameters
to be estimated in the model. Let us assume that we have n, (> k)
observations in one subperiod and n, (< k) observations in the other. There
is no difficulty about the restricted model in which one set of k parameters is
estimated for the n (= n, + n,) sample observations, namely,

3 be X,
b,
+ ey
Y2 X,

and e,e, has n — k degrees of freedom. If n, = k, the unrestricted model can

[3
be fitted and will have a residual vector

where e@; = y, — Xb,


denotes the residual vector from the first regression, and 0 is a k-element
residual vector from the second regression, in which the regression plane fits
the k observations exactly. The residual sum of squares e’e has

degrees of freedom. If n, < k, all k parameters cannot be determined for the


second period, but the residual vector is still 0 since an infinite number of
hyperplanes of dimension k can be passed through a set of less than k
observations. Thus the unrestricted residual sum of squares is still
e’e = e'e, with n, — k degrees of freedom
Analogy with the previous tests suggests that the appropriate test of the null
hypothesis that the n, additional observations belong to the same structure as
the first n, observations is based on}
“3 (€,€y0 ee, )/n,
(6-27)
eje,/(n, — k)
where n, is given by (n — k) — (n, — k), the difference between the number
of degrees of freedom of the sums of squares in the numerator. Thus the
practical procedure is as follows:

*Fit the regression to all n, + n, observations, giving the residual sum of


Squares e,e,.
¢ Fit the regression to the n, observations, giving the residual sum of squares
€7e1:
¢Compute the F statistic defined in Eq. (6-27) and reject the hypothesis
of a common structure if F exceeds a preselected critical value from
TG. Isa ha).

7 This is only a heuristic proof. For an exact derivation of Eq. (6-27), see F.
M. Fisher, “Tests on
Equality between Sets of [Link] Two Linear Regressions: An Expository Note,”
Econometrica,
28, 1970, pp. 361-366. An alternative proof is given in Sec. 10-1.
FURTHER TOPICS IN THE k-VARIABLE LINEAR MODEL 221

Example 6-6: Tests of structural change (x variables, p periods) The final


extension is going from two periods to more than two periods. For example,
we might wish to test whether a Phillips curve has the same structure prior to
World War I, between the two world wars, and post-World War II. But the
test need not be across periods. We might examine the stability of a relation
across countries, industries, social groups, or whatever.
The usual hierarchy of three models may be set up:

y
i, Xf
y2 i Qa
I eee X3 fa +u
o * we
Y, Pp P

Y) oa
: i, 0 OF xe
it ‘i = 7) 0 Xi] - | t+u
0 0 1, XS a>
Yp B*

Oy
Eo)

; i, 0 Oe Xt 0 Dal eee
Tt erty, os eal NOs Se a Onl laa
: 0 i ee 0 AS Bx
.
P

By
differential intercepts, differential slope vectors

Here i, is the column vector of n; units (i = 1,2,...,p) and X} is the


n, X (k — 1) matrix of observations on the explanatory variables in class i
G= TV,2s p):
The residual sums of squares from the three models have the number of
degrees of freedom

Boks Bp shirt ele and epi

where
LW2. ar
222 ECONOMETRIC METHODS

Table 6-1

Class
l 2 3 4
Observation Y X x XG VG X va XE

l 22 29 30 15 12 16 23 5
Pp. 22 20 32 9 8 31 25 25
3 20 14 26 l 13 26 28 16
4 24 21 26 6 25 35) 26 10
5 12 6 37 19 7 12 23 24 Ys X

Sums 100 90 150 50 65 120 125 80 440 340


Means 20 18 30 10 13 24 25 16 DD 1a

denotes the total number of sample observations. The various hypotheses


may then be tested by contrasting residual sums of squares in the usual way.t+

Example 6-7 This illustration is based on p = 4classes, but for simplicity of


calculation we have kept k = 2 (Table 6-1).
Denoting the data matrix in model I by X,, the residual sum of squares
for model I is

RSS, = y'y — y’X,(X,X,) 'X4y


n Do athe Says
= YY? (SY x]
Ae ees wx ¥:
where the summations are over all 20 observations,

10,876 — [440 7288] ~ me] AO


=

a) 340 7462 7288


1174.1
Using the same approach, the residual sum of squares for model II is
given by
RSS V5 Y Spee a |
n 1 0 0 0 a EX —] Da Y

Ny 0 0 dix a

x| 0 Oat ne Ob) aX Seay


0 0 0 Nye dX ay
DEA Ay aXe ee ee ar
where 2; indicates summation over observations in class i and Y indicates
7 It should be emphasized that the tests for structural change discussed in this section assume
that
the researcher has strong a priori views about when or where the potential change(s) occurred.
Additional complexities arise when the switch point(s) may be unknown. There is also the possibil
of a transition phase between
ity
regimes. There is some discussion of these topics in Sec. 10-4.
FURTHER TOPICS IN THE k-VARIABLE LINEAR MODEL 223

summation over all classes,


RSS, = 10,876 — [100 150 65 125 7288]
Sar). 05-0 90 }-1] 100
Giies OP veh) 50 150
On £0 50 0 25120 65
0," 0 On rao 180 125
O0mrI0%, 1200" 805 7462 7288
=9251.0
As shown in Example 6-2, this sum of squares may be more easily calculated
by first of all expressing the data in the form of deviations from class means,
pooling all the deviations and calculating RSS, as the residual sum of squares
from the regression of the Y deviations on the X deviations. Table 6-2 shows
the data in deviation form. The residual sum of squares from the regression
of the Y deviations on the X deviations is then
= =)\]2
Low lene zy - Eula =Wy = XO
y ee ak)
(134 + 117 + 177+ 0)
ap AEE 2008) page mnaenagIE B02
= 251.0 as before
Finally, we need to obtain the residual sum of squares from the com-
pletely unrestricted model, model III. This is the sum of the residual sums of
squares obtained by fitting a linear regression to each class separately. These
are most simply obtained from Table 6-2. They are

(134)°
eye, = 88 — 55, — 26.925

i (117)
e794 = 26.897

Ok =se 206 — Coane


ese, ~455— = 123.987

eels
/
ang
0 ——
8
giving
RSS, = 195.8
The various tests may be set up in an analysis of variance framework as
shown in Table 6-3.
The test for a common regression slope is

7. BSS
= RSS)/3 _ 184 |
7 RSS,/12 163 eke
224 ECONOMETRIC METHODS

og eal
li-6 0 9- g ZOE
ee
here
p 0
Z- 0 € I Z- 81

ee EO
8— L Z II Z1- Z8E

a
€ LLI
On) ee
I- S— 0 ZI9- 902
a
SseID
ee
re
¢ I- 6- y- 6 p07
a
Z LU
ee a
Se a) 0 a p- $- L r6

oy
eS
= Il Z p- € Z1- 67
hy)
I rel
ay
re

Z Z 0 P g— 88
Ap ere
a
Ce
LE
hawx
a uoNeAIISqO
ee ie
7-9
=
FIQUL By)
1)
i I é € p ¢ ia li ii
FURTHER TOPICS IN THE k-VARIABLE LINEAR MODEL 225

Table 6-3

Mean
Model Residual sum of squares Degrees of freedom square

I RSS, = 1174.1 “n-—k=18


II RSS,= 251.0 Jo Ue ae ans) 16.7
III RSS; = 195.8 n— pk = 12 16.3
RSS, — RSS, = 923.1 3 307.7
RSS, — RSS; = 55.2 B 18.4
RSS, — RSS; = 978.3 6 163.0

which is insignificant. Thus we do not reject the assumption of a common X


effect in all classes. The test for common intercepts (conditional on a common
slope) is
a (RSS7— RSS; )/3 9307.7
RSS, /15 Sig
and Fy 9(3, 15) = 5.42. Thus we conclude that a “class effect” is established,
that is, that the levels of the regressions do appear to differ between classes.
The test for overall homogeneity of regressions across classes is
_ (RSS, — RSS,)/6 _ 163.0 _
iB
RSS,/12 16.3 =

and Fo 99(6, 12) = 4.82, so that this too is a highly significant result, but it
would appear that the significance is due to variation in the intercepts and
not in the slopes.

6-3 DUMMY VARIABLES

Dummy variables have already made their appearance in the previous section, but
we have not explicitly labeled them as such. For example, the unrestricted model
in Eqs. (6-21) specifies a consumption function which has different intercepts, but
a common slope, in wartime and peacetime periods. The specification is repeated
here
: a)
= if , au %}+u (6-28)
yz 0 i, X, B

where the subscript 1 refers to wartime and the subscript 2 to peacetime. This
model may be written as
Y,= 0,D,,+ a)D,,+ BX,+u, t= 1,2,...,0 (6-29)
D,, and D,, are dummy variables whose sample values are given in the first two
226 ECONOMETRIC METHODS

columns of the data matrix in Eq. (6-28). That is,


yee 1 if t indicates a wartime observation
Bb a\ if ¢ indicates a peacetime observation
and
ee 0 if ¢ indicates a wartime observation
ee || if ¢ indicates a peacetime observation
Notice that the model in Eq. (6-29) has no general intercept term. If one runs a
regression of Y on D,, D,, and X with a computer program that automatically
produces an intercept term, the estimation procedure will break down (or possibly
give nonsense coefficients, which are merely ratios of rounding errors) since D,
and D, sum to the unit vector. The practical alternatives are

1. Run Eq. (6-29) with the general intercept suppressed.


or
2. Reformulate Eq. (6-29) as
Y,=y¥, + Do, + BX, + u, (6-30)
and run with the standard OLS program.

Comparing intercepts in Eqs. (6-29) and (6-30) gives

Equation Equation
(6-29) (6-30)
Wartime intercept a, vA
Peacetime intercept Q> Yi + Y2

Thus the relation between the a’s and the y’s is

Vitus and UD ee Oey


The model of Eq. (6-30) may then be put in matrix form as
a a
tl : By ey (6-31)
Y2 by hy Xe B

The choice between the two estimation procedures is of no great importance, but
it is very important to be clear about precisely what is being tested in either
model. For instance, testing the significance of D, in Eq. (6-30) is, in effect, testing
the hypothesis
Hy: a,—a,=0
which is testing whether the peacetime and the wartime intercepts are significantly
different, whereas testing the significance of D, in Eq. (6-29) is asking whether the
peacetime intercept is significantly different from zero.
FURTHER TOPICS IN THE k-VARIABLE LINEAR MODEL 227

The dummy variables may also be allowed to interact with the X variable.
Consider
Y =a,D, + aD, + B,(D,
+ B,(D,X)
X)+4 (6-32)
where the subscript ¢ has been omitted for simplicity. Equation (6-32) implies two
separate relations, namely,
Y=a,+B,X+u wartime function
Y=a,+fB,X+u peacetime function
Thus performing a single regression of Y on D,, D,, D,X, and D,X with the
general intercept suppressed is equivalent to fitting separate regressions to the two
subperiods. An alternative formulation of Eq. (6-32) is

Te a) + (a, 4,)D, + BX + (8, 6) Dx Pu (6-33)


This corresponds to a regression of Y on D,, X, and D,X with a constant term.
One advantage of Eq. (6-33) is that testing the significance of the D,X variable is
a test of the hypothesis

Thus we see that, in the two-variable model, the tests for homogeneity of
intercepts and homogeneity of slopes are equivalent to tests of the significance of
single coefficients in an appropriately specified regression equation using dummy
variables.
Dummy variables may also be usefully applied in more complex models. For
the data of the last numerical illustration we may specify

Y =a, + (a, —4,)D, + (a; — a)D, + (ay= a,)D, + B,X

+ (By — B,)D,X+ (B; — B,)D,X+ (B, — B,)D,X+ u (6-34)


where
Diz {1 for an observation in class i i= 2,3,4
é 0 otherwise
Equation (6-34) allows intercepts and regression slopes to vary across all four
classes. The choice of the class not to be represented by a dummy variable is
arbitrary, but once it is made, the coefficients of all the other variables are
differences from the coefficients of that class.
Estimating Eq. (6-34) from the data of Table 6-1 gives
Y =11.7959 + 12.4688D, — 9.9163 D, + 13.2041D, + 0.4558X
(256) 2 (2.19) (- 1.41) (2.13) (1.19)
+ 0.1177D,X + 0.0076D,X — 0.4558 D,X
(0.32) (0.02) (— 1.38)
with R* = 0.7468 and 12 degrees of freedom.j The figures in parentheses are t
ratios, and we see that none of the coefficients of the D,X variables is significantly
+ I am indebted to G. Gujarati for discussions of this point and also for the calculations.
228 ECONOMETRIC METHODS

different from zero. This, of course, confirms the homogeneity of regression slopes
established earlier by the F test. Imposing the assumption of a common regression
slope gives the revised regression
Y =13.4882 + 12.8969D, — 9.1726 D, + 5.742D, + 0.3621X
(4.79) (4.68) (—3.42) (2.20) (3.04)
with R? = 0.7341 and 15 degrees of freedom. All three dummies are significantly
different from zero at the 5 percent level, thus establishing that the intercepts in
the second, third, and fourth classes are different from the intercept in the first
class, again in agreement with the earlier F test on intercepts. One advantage of
this type of dummy variable setup is that in cases where the tests examine the
Joint significance of a subset of variables the dummy variables can indicate which
variables may have made the most important contribution to the overall signifi-
cance of the group.
The dummy variables specified above play an important role in describing
temporal effects (where the classes refer to different time periods), spatial effects
(where the classes refer to different regions or countries), industrial effects (where
the classes refer to industries), and so forth. Suppose we have qualitative variables
such as

e Education (none, grammar, some high school, high school diploma, some
college, college degree, advanced degree, foreign education)
¢ Marital status (unmarried, married 1 year, 2 years, 3 years, 4 years, 5—9 years,
10—20 years, over 20 years)
¢ Sex (male, female)
e Race (white, black, other)

Only the last two are truly qualitative variables. Education might be treated as a .
cardinal variable, measured by years of formal education, and likewise, duration
of marriage is a cardinal variable. In both cases, however, we may use groupings
of a cardinal variable to define a qualitative variable. If a qualitative variable is
thought to influence some dependent variable, we may use the categories of that
variable to classify the sample observations into various classes, and the preceding
method of analysis applies. There are, however, some slight complications if we
wish to use two or more qualitative variables in a single equation.

Two or More Sets of Dummy Variables


Suppose we wish to incorporate two qualitative variables in a regression equation,
each such variable being represented by a set of dummy variables. To be specific,
suppose the variables are educational level (3 classes) and sex (2 classes). We then
define

jae ] if observation relates to education level i, p= be


, 0 otherwise
FURTHER TOPICS IN THE kK-VARIABLE LINEAR MODEL 229

and
ee 1 if observation relates tosexj, j= 1,2
0 otherwise
Suppose we then wish to examine the relationship between hours spent in reading
nonfiction Y and these two qualitative variables. It is instructive to examine first
of all what happens if we have only one set of dummy variables in the model. A
linear model for the influence of E on Y would be written

Y=a,E£, + a,£, + a,£,+ 4 (6-35)


The data matrix is
i, 0 O
E=1.0 i, 0
0 0 i,
so that
Nye. Ore 20
FE=|.0 an, 0
Op .0j6on;
and the estimated OLS vector is

a, %
a,)= yy
a3 Ne

so that the OLS regression coefficients are simply the mean values of Y in each of
the educational classes. If we used the alternative formulation

Y=a,+ajE,+ af E,+u (6-36)

consistency with Eq. (6-35) requires


ay = a, — a,
*
and a} =a,
— a,

The data matrix is now


i. 0:40
Bt =shige inal,
PeeOyt,

and it is simple to show that the OLS coefficient vector ist

a Y,
eee
BV i
i = =
a3 Y, 1
+ See Problem 6-7.
230 ECONOMETRIC METHODS

Table 6-4 E(Y|E,, S))


Educational level

Educational level
E,

If we now specify Y as a function of both education level and sex, we might


be tempted to write
Y= aE, + a,Ey 4+-0,E, + BS) + B.S. 4 u (6-37)
Thus there is an expected value of Y for each combination of E and S. These
conditional expected values are shown in Table 6-4. An immediate difficulty with
Eq. (6-37) is that the OLS program (even with the intercept suppressed) will break
down, for the column vectors corresponding to E,, E>, E;, S,, and S, form a
linearly dependent set. Thus we cannot find unique estimates of the five parame-
ters in Eq. (6-37). As we will see, however, this does not imply that we cannot find
unique estimates of the sums of those parameters appearing in Table 6-4.
One way out of the difficulty is to reformulate Eq. (6-37) as
Y=pt+a,EF,+0,£, + BS, + u (6-38)
where the dummy variables are the same as before, but the a, 8 parameters will
not have the same meaning (or values) as in Eq. (6-37). The set of conditional
means for Eq. (6-38) is shown in Table 6-5.
The parameter p represents the expected number of hours spent in reading
nonfiction for people in the category (E,, S,). The a5, 3; parameters measure
differential effects for E, and E,, respectively, compared with E ,- The differential
effect for E, compared with E, is a, — a. Similarly, 8, represents a differential
effect for S, compared with S|. Notice that the differential sex effect is the same
at all educational levels, and similarly the differential educational effects are
invariant to sex. The four parameters of Eq. (6-38) may be estimated uniquely by

7 We might choose any one of the six cells to be represented by p. The differenti
al effects would
then be measured from that cell, but the numerical estimates of the conditional means will be
invariant to the starting position. See Problem 6-8.
FURTHER TOPICS IN THE k-VARIABLE LINEAR MODEL 231

Table 6-6 Hours spent in reading nonfiction Y

12, 14
10 20

OLS, and from these we obtain unique estimates of the expected values in Table
6-5.

Example 6-8 Hypothetical data on hours spent in reading nonfiction is


shown in Table 6-6 for ten subjects classified by sex and educational level.
Taking those data row by row, model (6-38), set out in full, gives
Y Ee Ets

R
a, |+u (6-39)
N

Tw)oO

Pm
pe
ep
et
ee SO
[Link]
or [OS
Ore
OE
IS OS)
eS =)
SS
SIS

Letting X represent the data matrix (often referred to as a design matrix in


experimental contexts) in Eq. (6-39),
10 ese ed 117
ST seeing 2 40
Ae 4 4]
the estimated equation is
Y = 11.14 + 3.31E, + 12.54E, — 7.35S,
The corresponding estimates of the expected number of hours are shown in
Table 6-7.

Interaction
The main drawback of Eq. (6-38) and the estimates to which it gives rise is the
built-in assumption that the differential effect of each factor is constant across the
levels of the other factor. Thus Table 6-7 shows that hours for S, are 7.35 lower
than for S,, irrespective of the level of education. Conversely, E, shows 3.31 more
232 ECONOMETRIC METHODS

Table 6-7 Estimated mean hours

11.14 14.45 23.68


S19 7.10 16.33

hours than E,, and E, shows 12.54 more than £,, irrespective of whether we are
in the S, row or the S, row. This implies the absence of any interaction between
the two factors. If, however, it is to be expected that the differential sex effect
varies with the level of education, then an interaction effect exists, and we need to
see how to incorporate it into the model and estimate it.
Returning to Eq. (6-38), we would now expand the relation to read

Y=p+a,E, + a,E, + BS, + y,(£,S,) + y;(£,8,) +4 (6-40)

There are only two possible interaction variables in this case, and they are found
by multiplying each E level by each S level. The conditional expected values are
now shown in Table 6-8.
The first row is the same as in Table 6-5, but the second row incorporates the
interaction effects. Thus the sex differential is

B, for E,
By + ¥p for E,
By +7; for E,

Likewise, the E,/E, differential is


a> for S;
a, t+ y, for S,

and the £,/E, differential is


a, for S,
a, + 3 for S,

Table 6-8 E(Y|E;, S;)

B+ Q, b+ a
p+ a, + Bo + yp w+az+ B+ ¥5
FURTHER TOPICS IN THE k-VARIABLE LINEAR MODEL 233

Table 6-9 Estimated mean hours

1333 13.00
0.50 10.00

Referring back to the data matrix in Eq. (6-39), the data matrix for this problem
would now be
E, E; So ES) E3S>

1
1
|
1
tea
ee 1 1
1 1
1 |
is ise
1 ea |
Thus
10Ne O24 GL 117
aoe EGhe (uetdee 0 36
Oe nee2 ae Oe 40
ee eed UAT Bee a Hl See Unt oet
Tete colts Ie O 10
es Os eel Otel 20
The OLS equation is now
Y = 13.33 — 0.33E, + 6.67E, — 12.835, + 9.83( E,S,) + 12.83(£;S,)
and substitution in Table 6-8 gives the estimated number of mean hours shown in
Table 6-9.
Compared with the previous regression, where no interaction effect was
incorporated, we now have a large negative sex effect (— 12.83) at E,, which is
reduced to —3.00 at E, and eliminated completely at E,. This last result is an
automatic consequence of our data, where in the interests of simplicity we had
only one observation in each of the £; cells and also in the E,, S, cell. The
regression values, with interaction, then coincide with these observations. This has
also distorted the estimate of the E, differential effect to give a small negative
number (— 0.33) for S,, but the calculations do illustrate the principles involved.+

+ This section has only dealt with dummy variables on the right-hand side of the equation. For a
discussion of the application of dummy variables to the left-hand-side variable see Sec. 10-5,
Qualitative Dependent Variables.
234 ECONOMETRIC METHODS

6-4 SEASONAL ADJUSTMENT

Dummy variables also play an important role in problems of seasonal adjustment.


These problems are of two kinds. First there is the conventional and long-stand-
ing problem of deseasonalizing a given quarterly or monthly time series, and
second there is the problem of estimating an econometric relationship between
variables that are available in both unadjusted and deseasonalized forms.
Suppose we have 4n quarterly observations on a variable Y, such as unem-
ployment, imports, or food prices. Such variables are likely to display a pro-
nounced seasonal movement, and for purposes of economic intelligence and
policy it is important to produce a “deseasonalized” series, from which one can
better assess whether unemployment, say, is really increasing or decreasing. There
are several methods of deseasonalizing series in practice, but here we are only
concerned with applications of dummy variables.
Let us define a 4n X 4 matrix D,

VO OVO
OV Sn Ona)
OFS.O eae 0
D=s(0O
08 02
1 0250-550
Ome? OVO
OF 052.0: el
This is the sample matrix for four dummy variables defined by

rks 1 if ¢ occurs in quarteri i= 1,2,3,4


1. 0 otherwise
If we regress y on D, we obtain

YaDb ay. (6-41)


where b is the vector of the OLS coefficients and y* the vector of residuals. From
the analysis of Chap. 5,

yo = My (6-42)
where

M =I-D(DD) 'D’ (6-43)


and M is symmetric idempotent with the property

MD = 0 (6-44)
The series y* cannot serve directly as a deseasonalized series for two reasons.
First of all, it sums to zero, and it would seem plausible to require a deseasonal-
ized series to have the same sum as the original, unadjusted series. Second, as
FURTHER TOPICS IN THE kK-VARIABLE LINEAR MODEL 235

shown earlier for model (6-35),

Yy,
Y,
b=|_
Y,
Y,
where Y, (i = 1,..., 4) is the mean of all ith-quarter Y values. Thus y* merely
consists of deviations of the Y values from the quarterly means. But if the series
contains trend and/or cyclical components, the elements of b will be an amalgam
of trend, cyclical, and seasonal effects. Thus subtracting b year by year from the
actual Y values will not yield satisfactory estimates of a deseasonalized series. The
remedy is to introduce into the regression a polynomial in time of sufficiently high
order to represent the trend and cyclical components, so that the coefficients of D
will be a more satisfactory estimate of the seasonal component. Thus one
computes the regression
y =Pa+Db+e (6-45)
where

1 12 er 12

2 22 2?
Paine B32 3?
4 4? 4P

4n (4n)° (4n)?
The deseasonalized series would now be defined as
y* =y — Db (6-46)
Jorgenson has argued that if the P and D matrices are properly specified, then a
and b will be best linear unbiased estimates of the systematic and seasonal
components, since Eq. (6-45) is then a straightforward example of ordinary least
squares.} The estimates of a and b are given by

]-[>» po] [os| a


, iD 1-1f p”

Applying the results for the inverse of a partitioned matrix,


b = (D/ND) 'D/Ny (6-48)
where
N=I-P(PP) 'P’ (6-49)
+D. W. Jorgenson, “Minimum Variance, Linear, Unbiased Seasonal Adjustment in Economic
Time Series,” Journal of the American Statistical Association, 59, 1964, pp. 681-725.
236 ECONOMETRIC METHODS

Table 6-10 Quarterly seasonal component of the U.K. Index


of Industrial Production, 1948-1957

Seasonal component
Method by by b; bg

Moving average (additive) 3.28 0.77 eal 3.08


Regression on D 1.85 0.35 SINS 4.95
Regression on [P D] (p = 4) 4.87 0.36 = IN Z95
Regression on [P D] (p = 6) 3.35 0.95 =o 3.25

Substituting in Eq. (6-46) gives


Yaeay
where

T =I- D(D/ND) 'D’/N (6-50)


Thus the deseasonalized series can still be expressed as a linear transformation of
y. However, in contrast with the M matrix defined in Eq. (6-43), the T matrix is
not symmetric, though it is idempotent and does satisfy the condition TD = 0.
As a numerical illustration of these methods we made several estimates of the
quarterly seasonal component in the U.K. Index of Industrial Production for the
period 1948-1957. The results are shown in Table 6-10. The centered four-quarter
moving average is a flexible method for removing trend and cycle, and we will
take the estimates of the seasonal component in the first row of Table 6-10 as a
standard by which to judge the various regressions. It is seen that the simple
regression on seasonal dummies alone gives misleading estimates of the seasonal
component, apart from the pronounced dip in the third quarter which is well |
picked up by all methods, and it is only when we use a sixth-degree polynomial
that the results agree closely with those obtained from the moving average
method.

Estimation of Econometric Relationships


Faced with the choice between using raw data or seasonally adjusted data, one
should think carefully about the basic decision process underlying any behavioral
relation being estimated. For example, in the study of production decisions it is
often found that firms attempt to base production rates on “smoothed” sales
figures, so that the appropriate regression might be actual production on desea-
sonalized sales. A salaried worker may have an income with no seasonal compo-
nent, but consumption expenditures with a strong seasonal component due to
vacation and Christmas spending. The appropriate model would then regress
actual consumption on actual income plus a set of seasonal dummies. Income
itself may have one seasonal pattern and consumption a different seasonal pattern
with deseasonalized consumption a function of deseasonalized income. If one
FURTHER TOPICS IN THE k-VARIABLE LINEAR MODEL 237

then wished to explain actual consumption, the appropriate regression is


Actual consumption = deseasonalized consumption + seasonal component
= f (deseasonalized income) + seasonal component
= f(deseasonalized income, dummy variables)
In many cases theory may give no clear guide to the appropriate regression,
and as the data are often available in both unadjusted and deseasonalized form, it
is sometimes difficult to decide in which form to incorporate variables in the
regression. In practice, however, the problem of specification turns out to be less
important than might have been expected, because of an important set of results
due to Lovell.} To illustrate one of Lovell’s basic results, consider the least-squares
regression
y = Xe, + Db, + e, (6-51)
This may be interpreted as a regression of unadjusted Y values on unadjusted X
values and a set of seasonal dummies. However, it is more instructive to consider
a more general specification first of all, and simply regard X and D as a
partitioning of the set of explanatory variables in the regression model,

y= (x p][,']+e
The OLS coefficients are then given by
c, = XX XD|~'[X’y
|
b, | |DX DD] [Dy eee)6-52
Applying Eq. (4-68), the first element in this inverse matrix is

(xX — X’D(D’D)'D’x)| = (X’Mx)"!


where
M =I-D(D'D) 'D’
This is the M matrix already defined in Eq. (6-43), which we know to be
symmetric and idempotent and to have the property MD = 0. The remaining
element in the first row of the inverse matrix is then
— (X/MX) 'X’D(D’D) |
Equation (6-52) may then be solved for c, as
c, = (X’MX) 'X’y — (X’MX)_'X’D(D’D) 'D’y
that is,
c, = (X’MX)_ 'X’My (6-53)

+™M. C. Lovell, “Seasonal Adjustment of Economic Time Series,” Journal of the American
Statistical Association, 58, 1963, pp. 93-1010. The basic result goes back to R. Frisch and F. V.
Waugh, “Partial Time Regressions as Compared with Individual Trends,” Econometrica, 1, 1933, pp.
387-401.
238 ECONOMETRIC METHODS

Now consider the transformed variables


y* = My and X* = MX (6-54)
From Eq. (6-54) it follows that y* is the vector of residuals after y has been
regressed on D. Similarly, each column in X* is the vector of residuals after the
corresponding X variable has been regressed on D. If y* is regressed on X“, the
estimated coefficient vector is
(X’M’MX)_ 'X’M’My = (X’MX)_ 'X’My

in view of the symmetry and idempotency of M. Thus we have the important


result that if we partition the explanatory variables in a regression into two blocks
denoted by
[xX D]
the estimated coefficients of the: X variables are exactly the same, whether we run
the full OLS regression of y on X and D or first “correct” y and X for the effect of
D and regress the Y residuals on the X residuals. More formally, if we calculate
the two regressions
y = Xe, + Db, + e,
y* = X%, + e,
the result is

Cy © (6-55)
This result is, of course, symmetrical with respect to X and D, and the D matrix
need not consist of dummy variables; it is merely any subset of explanatory
variables. However, Lovell is concerned with seasonal adjustment, and D is then
appropriately an n X 4 matrix of quarterly seasonal dummies.
Two further basic results from Lovell are that the regressions
y= X°c, cs e3

and y = X°c, + Db, + e,


also yield identical vectors of coefficients for the X variables, that is
C, =, =¢,=¢, (6-56)
The proofs are simple. Regressing y on X* gives

c,; = (X’MX) 'X’My


= ¢)
and regressing y on [X* D] gives
| . ae Tata
b,| |D’MX DD D’y
_[XMx
0
0 ]-'[xmy
‘DD Dy using Eq. (6-44)
FURTHER TOPICS IN THE k-VARIABLE LINEAR MODEL 239

Thus
c, = (X’MX)_'X’My

These results raise some further questions. We have already seen that if D is
merely a matrix of seasonal dummies, then y* and X°, defined in Eqs. (6-54), are
not properly deseasonalized series. On the other hand, if properly deseasonalized
series are obtained by using the transformation matrix T defined in Eq. (6-50),
this matrix, though idempotent and orthogonal to D, does not have the symmetry
property used in the above proofs. Furthermore, many official series are not
deseasonalized by least-squares methods at all, but by moving average or other
methods. Thus the Lovell results cannot be expected to hold exactly when y* and
X“ indicate properly deseasonalized series. Nonetheless some experimental calcu-
lations with various equations from the Oxford econometric model of the United
Kingdom indicate agreement to several decimal places between estimated coeffi-
cients, whether the regression has been run with raw data and dummy variables or
with deseasonalized variables produced by moving average methods or by least-
squares regressions on D or on[P_ D].} The years covered by the model showed
fairly steady growth and negligible cyclical oscillations. One would not expect
such close agreement if the cyclical effects were very strong, and in practical work
one should not allow this theorem to be a substitute for careful thought about the
proper specification of the relationship.

6-5 MULTICOLLINEARITY

We have seen in Chap. 5 that the OLS estimator is


b = (XX) 'X’y
and that its variance matrix is
var(b) = 02(X’X)|
Thus the sampling variances depend not only on the disturbance variance 0”, but
also on the sample values of the explanatory variables. Consider the following
hypothetical matrices.
x’X (X’X)7! \X’X|

Le altel ae
ee Ombvcetyi Wika eae oe ls 30?
# as ca ae | Pes
+ A. Georgopoulou and J. Johnston, “Seasonal Adjustment of Economic Time Series,” University
of Manchester, discussion paper.
240 ECONOMETRIC METHODS

In case 1 the two explanatory variables are orthogonal and the coefficients of the
X’s in the multiple regression equation would be the same as those given by the
simple regressions of Y on each X in turn. Orthogonal variables may be set up in
experimental designs, but they are the exception, not the rule, in economic data.
Cases 2 and 3 display increasing correlation between the two explanatory
variables, as evidenced by the increasing numerical value for the off-diagonal
(covariation) term. This is also reflected in the dramatic fall in the value of the
determinant. This is described as a situation of collinearity (or multicollinearity)
between the explanatory variables. Three important effects are illustrated in the
sequence of matrices:

1. The sampling variances of the estimated OLS coefficients increase sharply


with increasing collinearity between the explanatory variables. Taking case 1
as the base, they are more than five times as great in case 2 and 50 times as
great in case 3. Thus in any specific application individual coefficients are
likely to differ substantially from their true values.
2. Greater covariances between the explanatory variables produce greater sam-
pling covariances for the OLS coefficients. Comparing the off-diagonal terms
in X’X and (X’X)~! shows that a positive covariance for the X’s gives a
negative covariance for the b’s, and vice versa. Again, in a specific applica-
tion, if b, is below B,, b; is most likely to exceed 6, and vice versa (provided
the X’s are positively correlated).
3. Small variations in the data (for instance, dropping or adding a few observa-
tions) may produce substantial variations in the OLS coefficients. Suppose the
normal equations for case 2 are

by 0.952 A aly
0.9b,+b,=2.9 ? 3
—_ = a

Now suppose the X; variable is somewhat more highly correlated with X, and
we have normal equations for case 3 as

b, + 0.99b, = 2.8 ee
O99 b= ae aa
The only numerical change between the two sets of equations is a 10 percent
(or less) increase in two coefficients, yet the solution values change dramati-

+ Assuming the variables to be in deviation form


P.
Xoy

p= [2 Oy |oe xoyel L exexe


0 X3X3 x3y X3Y
/
X 3X3
where each element in the right-hand-side vector is the slope coefficient in a simple two-variable
regression.
FURTHER TOPICS IN THE k-VARIABLE LINEAR MODEL 241

cally.t Notice, however, that b, + b, = 3 in the first case and b, + b, = 2.964


in the second. Thus the sum of the coefficients appears to be estimated fairly
precisely, even though the individual coefficients are subject to large errors of
estimation. Even this happy result is dependent on the covariance between
the b’s being negative (i.e., covariance between the X’s being positive) for
var(b, + b;) = var(b,) + var(b,) + 2cov(b,, b;)
Thus increasing collinearity increases both var(b,) and var(b,). However, it
also increases the numerical value of cov(b,, b,) and, provided this covariance
is negative, var(b, + b,) may not increase at all. For example, case 2 gives
var(b, + b,) = 2(5.26 — 4.74)
07 = 1.040?
and case 3 gives
var(b, + b,;) = 2(50 — 49.5)o? = 1.0007

For simplicity these three important points have been illustrated for the case
of two explanatory variables. It is important to establish that similar results hold
for the k-variable case and to discuss how multicollinearity may be detected and
what may be done about it. However, before doing that, we will discuss the
limiting case of exact, or complete, multicollinearity.

Exact Multicollinearity and Estimable Functions


In the case of two explanatory variables exact collinearity is represented by

where we are working with the variables in deviation form. Then

Kok Ei “|
eae
with |X’X| = 0 and p(X’X) = 1. This is simply a breakdown of the assumption
that X has full column rank, and so we cannot obtain the unique OLS vector
defined by

b= |= (X’X) 'X’y
The normal equations
(X’X)b = X’y (6-58)
however, will admit an infinity of solutions for

X’y = Ex2y| 1]
Qa

+ This is only a hypothetical example, but the literature of applied econometrics is full of examples
of small changes in the data base producing substantial changes in estimated coefficients. For one
example, see J. Johnston, “An Econometric Model of the United Kingdom,” Review of Economic
Studies, 29, 1961, pp. 29-39.
242 ECONOMETRIC METHODS

so that the rows of X’y exhibit the same linear dependence as the rows of X’X.
The set of equations in Eq. (6-58) is thus consistent, and there is an infinity of
solution vectors. Taking the first equation in Eq. (6-58), we have
Ex7(b, + ab;) = Expy
and the second equation is
abx3(b, + ab,) = abx,y
Both equations reduce to

b, + ab, =
x7y
2
(6-59)
2

Thus no matter which arbitrary solution to Eqs. (6-58) we take, the linear
combination b, + ab, will always have the same numerical value. We then define
B, + aB; as an estimable function, where we notice that the a in the estimable
function is the parameter defining the linear dependence between x, and x3.
The same result may be derived by writing the model in deviation form as
y = Bx. + Bx, + (u— @)
and substituting Eq. (6-57) to get
y = Bx,+ (u-@) (6-60)
where
B = B, + af, (6-61)
The B parameter may be estimated by applying OLS to Eq. (6-60) to give

poy (6-62)
Ds
which is the same expression as that already obtained in Eq. (6-59). The expected |
value of y for a given x, (and x;) is
E(y)= ByX_ + B3x3
= (B, + aB;)x,
= Bx,
Thus E(y) can be estimated uniquely since 8B can be estimated uniquely by Eq.
(6-62).

Example 6-9 Suppose x, = 2x, and the sample data are


10 20 5
a Sal 20 40 rfc XY= (Er110
The normal equations (6-58) are
106, + 20b, = 5
20b,
+ 40b, =10
FURTHER TOPICS IN THE K-VARIABLE LINEAR MODEL 243

Table 6-11

Estimate of
b; by b> 5 2b; E(y|x2 = 20)

0 0.5 0.5 10
l —1.5 0.5 10
—] DS 0.5 10
110) 20.5 0.5 10

with solution
7 >= 20D;
B 10
Taking some arbitrary values for b, gives Table 6-11.
The linear combination b, + 2b, is invariant to the solution chosen for
the normal equations, and it is readily seen to be equal to
z DX y Ce
b 0.5
re, al?
Likewise, the regression value for any given x, is invariant to the normal
solution vector.
To summarize, even though £, and £, cannot be estimated, a certain
linear combination of B, and f, can be estimated, and E(y) can also be
estimated for any given x, value.

The nature of the problem may also be illustrated geometrically. In Fig. 6-5a
the standard OLS case is shown. The x,,x, vectors are not perfectly collinear,
and they span a two-dimensional subspace in &”. Dropping a perpendicular from

(a) (d)

Figure 6-5
244 ECONOMETRIC METHODS

y to that subspace splits y into


Yoo yeet <
where
§ = Xb = bx, + 3x3
The regression vector § is a unique linear combination of the column vectors
X5,X,. By contrast in Fig. 6-5b the x,,x, vectors only span a one-dimensional
subspace (line) in ®”. The § vector is still unambiguously determined by dropping
a perpendicular from y to the line, but § cannot be expressed uniquely in terms of
x, and x3.
In the general case perfect multicollinearity exists if p(X) < k. Suppose
o(X) =r. There is then at least one set of r linearly independent columns in X.
Let one such set be assembled in the first r columns, so that we partition X as
X=[X, X,] (6-63)
where X, =n X r matrix of rank r
X, =n X s matrix of the s = k — r remaining columns in X
Each column vector in X, may then be expressed as a linear combination of
the columns of X,. Thus we may write
X, = XW (6-64)
where W is an r X s matrix, each column of which gives the coefficients of the
linear combination for the corresponding vector in X,. The numerical values of
the elements of W can, in principle, be determined. Combining Eqs. (6-63) and
(6-64) gives
X = X,[I,. W]
= X,Z (6-65)
where
zZ=(l, W] (6-66)
The linear model may then be written
y=Xfp+u
= X,ZB + u
=XB+u (6-67)
where
B, = ZB (6-68)
Note carefully that B, indicates a vector of r linear combinations of the original B’s,
the coefficients of those linear combinations being given by the rows of Z, as
defined in Eq. (6-66). The elements of B. may be estimated by applying OLS to
Eq. (6-67), since the X, matrix has full column rank. The estimator is thus

B, = (X,X,) 'X/y (6-69)


Likewise, E(y) can be estimated by the regression vector

9=XB (6-70)
FURTHER TOPICS IN THE k-VARIABLE LINEAR MODEL 245

and the estimates in Eqs. (6-69) and (6-70) will have the usual OLS properties.
The operational procedure would be to identify the largest submatrix in X with
full column rank, denote it by X,, and substitute in Eqs. (6-69) and (6-70). If
there is more than one such submatrix, } will be invariant to which is chosen.
As in the case of two explanatory variables, an alternative procedure is to
derive any solution by to the normal equations _

(X’X)by = X’y
and compute
§ = Xby (6-71)
The numerical values for § in Eqs. (6-70) and (6-71) will be identical. One needs
to determine the linear dependencies in the X data (as in the W and Z matrices)
in order to determine the precise linear combinations of £ coefficients that are
being estimated in B,, but such combinations are not usually of any economic
significance. Furthermore, the use of Eq. (6-70) or Eq. (6-71) for forecasting
outside the sample observations rests on the same linear dependencies holding
among the X’s in the forecast period.

Near Multicollinearity
The prevalent case in so much econometric work, especially with time series data,
is one of high but not exact multicollinearity. This raises three questions:

1. What effects to expect from muiticollinearity


2. How to detect the degree of multicollinearity
3. What remedial action to take

Effects
Provided the X matrix has full column rank, the OLS estimates exist and will still
be the best linear unbiased estimates. This property, however, is now cold comfort
since the sampling variances of the estimates increase alarmingly with rising
collinearity. To prove this in the general case, partition the X matrix as
», [x, X;]

where x ,=column vector of observations on the ith explanatory variable


X= submatrix of observations on the k — 1 other explanatory variables

Then
* Xi XA,
ee Nine XK
Applying Eq. (4-68), the leading term in (XX) is

[(xx,) — x/X,(X%X,) (Xx, = (Mx,7!


246 ECONOMETRIC METHODS

where
M, = 1 - X,(X;X,)°X;
Thus the sampling variance of the OLS estimate of ; is

ue (6-72)
2
0
var(b,) =

But from Eq. (5-54) it is seen that


; residual sum of squares from the regression of the ith explanatory
x, Mx,=__. :
it~ variable on the other k — | explanatory variables
The residual sum of squares decreases with increasing collinearity between the ith
explanatory variable and the remaining explanatory variables, and thus the
sampling variance of b, increases. It is clear from Eq. (6-72) that not all
coefficients will be affected similarly by collinearity. The denominators in the k
sampling variances are the residual sums of squares from the multiple regressions
of each explanatory variable in turn on all the other explanatory variables, and
these can vary considerably from one to another, as is illustrated in the following
numerical examples.
Suppose we have three explanatory variables X,, X,, and X,, all measured in
deviation form. We show four illustrative X’X matrices, and in each case the
determinant, the inverse, and the values of the squared multiple correlation
coefficients obtained when each explanatory variable in turn is regressed on the
remaining explanatory variables.

X’X (X’X) re R334 R3 04 R45

LOR O02a0 OO 0
1. Ove) 0 OP 0) 0 0 0
ORO! 0 0 1
det = 50

LO oeSD O25 i022 0


2: 2 whe Oe 1:2: ~=2.0)|" 0.5000. °0;8333 .0:8000
oe vel 0 =.) 5.0
det ='5

LO” O43 1.0 0 — 3.0)


3. Gee 0 10 —2.0] 0.9000 0.8000 0.9286
Bele | 7.0 — 20 14.0
aceu—eal

10 TAS 5337 y= 73" 407


4. hi See —7.3 10.3 0.7} 0.9812 0.9802 0.2500
1S lel —0.7 0.7 3
det = 0.75
FURTHER TOPICS IN THE kK-VARIABLE LINEAR MODEL 247

es 1 shows perfectly orthogonal variables. The sail variances are given by


o* times the elements in the principal diagonal of (X’X)~'. In the orthogonal case
these variances are inversely proportional to the amount of variation in the
corresponding explanatory variable. For example, X, displays 10 times the
variation of X,, and the sampling variance of its eostticedtt is one-tenth of that
for the coefficient of X,. To standardize the comparisons, the elements in the
principal diagonal of X’X have been kept constant throughout, but the cases
display increasing collinearity, as reflected in the declining value of the determi-
nant. In case 2 the sampling variances are all larger than in case 1, but they still
retain the same order in that

var(b,) < var(b,) < var(b,)


However, they have been increased by varying factors. We notice that var(b,) is
now 25 times var(b,), contrasted with 10 times in the orthogonal case. Case 3
shows a large increase in var(b,), which is now as large as var(b,). Case 4 is the
most interesting of all in that var(b,) is now the smallest of the three sampling
variances, and indeed it is not much larger than in the orthogonal case, whereas
var(b,) and var(b;) are each more than 50 times as large as in the orthogonal
case.
Careful study of the R?’s will show that there is an association between the
size of R* and the extent to which the corresponding sampling variance is
increased over the orthogonal case. The relationship is, in fact, a precise one and
may be set out as follows: Let
TSS, total sum of squared deviations for X;,
RSS, I = residual sum of squares when X; is regressed on the other k — 1
explanatory variables
R? = square of multiple correlation coefficient from the same regression

1 -
RSS,

TSS,
Then Eq. (6-72) may be rewritten as
o2
var(b,) =
RSS, ~ TSS,(1 — R?)
Letting b,, denote the estimate cf B, in the orthogonal case,

var(b,,)= TSS,

Thus, if TSS, is held constant, the magnification of the sampling variance with
increasing collinearity is given by
var(b)) nil
(6-73)
var(®,,) 1 — R?
The orthogonal case is not meant to be a feasible target, but is used as a
248 ECONOMETRIC METHODS

Table 6-12 Magnification of sampling variances

R?l 0.5 0.8 09 095° 0.96 0.97 0.98 0.99 :0.999

b
RCC i bs eh ahs
var(b;,)

benchmark from which to measure the relative magnification of the sampling


variance of different coefficients. Some illustrative calculations from Eq. (6-73) are
shown in Table 6-12, and the graph of the function is drawn in Fig. 6-6.
As the table and the figure show, the relationship is highly nonlinear, and the
magnification factor increases dramatically as R? exceeds 0.9. The formula also
reveals why different coefficients fare differently in a regression. For example, in
case 4 the magnification effect was very serious for b, and b, and almost negligible
for b,, which is exactly in line with the pattern of R?’s. In that data X, and X, are
highly correlated with one another, but X, is not closely correlated with either X,
or X,, or with any linear combination of them.
The three main effects already listed for the case of two explanatory variables
will thus carry over to the general case, namely

e Very large sampling variances


e Greater covariances
e Great sensitivity of estimated coefficients to small data changes

A common result is to find regressions possibly with a very high overall R*, but

var (b;)/var (dio)


A

100 |-

80

60

40F

0.5 0.6 0.7 0.8 0.9 1.0 E


Figure 6-6
FURTHER TOPICS IN THE K-VARIABLE LINEAR MODEL 249

with some (or many) individual coefficients apparently insignificant. The high R*
arises when the y vector is close to the hyperplane generated by the x, vectors and
the apparently insignificant coefficients arise because the x ; vectors are nearly
linearly dependent. It is also possible to find a high R? and highly significant ¢
values on individual coefficients, even though multicollinearity is serious. This can
arise if individual coefficients happen to be numerically well in excess of the true
value, so that the effect still shows up in spite of the inflated standard error
and/or because the true value itself is so large that even an estimate on the
downside still shows up as significant. The multicollinearity would likely show up
in varying parameter estimates as some sample observations are dropped or
added. For any regression, however, comparison of the R?’s shows which
coefficients are likely to be most seriously affected by collinearity.

Detection
Computer programs often print out |X’X|. As our numerical examples illustrate,
the determinant declines in value with increasing collinearity, tending to zero as
collinearity becomes exact. While a useful warning signal, we have no calibration
scale for assessing what is serious and what is very serious, and again it gives no
guide to the relative effects on individual coefficients. Similar remarks apply to the
computation of the eigenvalues of X’X. Since
[XX] =A\Ap--- Ay
a small determinant means that some (or many) of the eigenvalues will be small.
But again knowledge of the eigenvalues is of little direct help in assessing effects
on individual coefficients.
The most useful single diagnostic guide is the R?’s, as shown above. In a
sense TSS, determines the minimum sampling variance that might be achieved for
b,; in that in the orthogonal case
ae

Any collinearity in the sample data will raise all sampling variances, but the
relative magnifications for different coefficients will be indicated by a comparison
of the R?’s.
Belsley, Kuh, and Welsch suggest the combined use of two diagnostic tools to
detect which coefficients are most likely to be affected by the collinearity.t The
first statistic is the condition number of the X matrix, defined by
r max
K(X) =
r min

+ The precise relationship between var(b;) and the A’s is derived below.
+ D. A. Belsley, E. Kuh, and R. E. Welsch, Regression Diagnostics, Identifying Influential Data and
Sources of Collinearity, Wiley, New York, 1980, chap. 3.
250 ECONOMETRIC METHODS

where A,,,, and A,,, denote the maximum and minimum eigenvalues of X’X,
respectively. If the X matrix has been standardized so that each column has unit
length, then «(X) is unity when the columns of X are orthogonal and rises above
unity with collinearity between the columns. Various applications with experi-
mental and actual data sets suggest that condition numbers in the range of 20 to
30 are probably indicative of serious collinearity problems, and a fortiori for
numbers in excess of that range. A condition index may be computed for each
eigenvalue, starting at unity for A, = A,,;, and rising to «(X) for A; = A,,.x- Thus
a given data matrix may yield one or more condition indexes in excess of a
“danger” level. The second and related diagnostic tool is the regression coefficient
variance decomposition. If X isn X k and V is the k X k matrix that diagonalized
X’X, then
(X’X)V = VA
where A is the diagonal matrix of the eigenvalues of X’X. Thus
var(b) = o?(X’X)| = o2VA~'V’
and
One2 ©:2 v;2
var(b,)
( ) = 07) =!
r, +2
v5 4... 4 8
rx i= aeek

where 0,;, U;2,---, U;, are the elements of the ith row of V. From this one may
compute the proportions of var(b;) associated with each A. The two-step procedure
recommended by Belsley, Kuh, and Welsch is

1. Compute the A,’s and identify any A; which gives a condition index in excess
of the “danger” level (say, 20 to 30).
2. For each of those selected A,’s inspect the proportions of the sampling
variance of each 5, associated with that eigenvalue. Coefficients with propor- .
tions in excess of, say, 0.50 are likely to have been adversely affected by the
collinearity in the X matrix. Reference should be made to Belsley, Kuh, and
Welsch for detailed examples of the technique.

Remedies

More data is no help in multicollinearity if it is simply “more of the same.” What


matters is the structure of the X’X matrix, and this will only be improved by
adding data which are less collinear than before. However, there is often no easy
way for an econometrician to get better data. The data are produced by the
functioning of the economic system, and the collinearities reflect the nature of
that system. One hopeful approach in some areas is the joint use of both
time-series and cross-section data, which we will take up in Chap. 10. A related
approach is to feed in estimates of some parameters which may be available from
an independent, relevant study. The classic example is the analysis of demand
functions, where an estimate of the income elasticity obtained from cross-section
studies is fed into the estimation of the price elasticities from a time-series sample.
FURTHER TOPICS IN THE k-VARIABLE LINEAR MODEL 251

This is a sequential approach using one set of cross-section data and another set of
time-series data, rather than a joint (or simultaneous) set. The latter requires that
the observations relate to a common set of decision units.
The general framework for the incorporation of prior estimates of some
parameters may be set out as follows. Partition X, B, and b as

X=[X, X,] e-|F| »-(?|

where X, is the n X r submatrix consisting of the first r columns of X, X, is the


submatrix consisting of the remaining s = k — r columns, and B and b are
partitioned conformably. Suppose that a previous study provides the estimated b,
vector with an estimated variance matrix V,. We will assume b, to be unbiased,
that is,
E(b,) = B,
and we will take V, to be approximately the true variance matrix

E{(b, — B,)(b, — B,)’}


The problem now is to estimate the remaining unknown parameters in B,. The
procedure is to “correct” y for the X, data by forming

¥s— ¥.7 Xb, (6-74)


and then perform an OLS regression of y, on X,. The result is

b, = (X,X,) 'Xys (6-75)


Writing
y=X,B,+ X,B,+u
and substituting in Eqs. (6-74) and (6-75) gives

b, = B, + (X,X,) ‘Xiu — (X,X,) 'X'X,(b, — B,)


Taking expectations
E(b,) = B,
since E(u) = 0 and E(b,) = B,.+ Then

var(b,) = E{(b, — B,)(b, — B,)’}


='0?(X/X,) |+ (X¢X,) 'XXV,X,X,(X-X,)"' (6-76)
on the assumption that the two sets of data are independent. The first term in Eq.
(6-76) is the conventional variance matrix for an OLS regression involving X,,
and the second term shows the elements by which this must be adjusted because
of the sampling variation in the b, coefficients, used in calculating y,. The only

+ Notice that this operation involves taking expectations over two different sets of data. E(u) refers
to expectations over the current sample data and E(b,) to expectations over the data underlying the
prior estimate b,.
252 ECONOMETRIC METHODS

remaining practical problem is the estimation of 07. Defining


e = y, — X,b, = y — X,b, — X,b,

the estimate is e’e/(n — k), where we divide by n — k rather than by n — rsince


e depends on k estimated parameters.
Some authors suggest dealing with multicollinearity in a rather mechanical
and purely numerical fashion. For example, a currently fashionable technique is
that of ridge regression.} The ridge estimate of B is defined as

bp = (X’X + cl) 'X’y (6-77)


where c > 0 is an arbitrary constant. The rationale for the estimator may easily be
seen by referring back to the X’X matrices given in the numerical illustrations.
Increasing the diagonal elements and leaving the off-diagonal elements unchanged
may be expected to reverse the sequence of effects shown in those examples where
the off-diagonal elements have been increased relative to the diagonal elements. It
follows directly from Eq. (6-77) that

E(b,) = (X’X + cl) 'X’XB (6-78)


and

var(bp) = 07(X’X + cl) 'X’X(X’K + cl)| (6-79)


The ridge estimator is thus biased, but it may be shown that the variances of the
elements of bp are less than those of the OLS estimator.t This raises the
possibility that a ridge estimator may have a smaller mean-square error (MSE)
than the OLS estimator. The main difficulty centers on the selection of a
numerical value for the arbitrary scalar c. In their original article Hoérl and
Kennard suggested trying various values of c in an attempt to see if the bz vector
stabilized. Schmidt, in the source cited, establishes conditions for c to minimize
E(bp — B)’(be — B), the sum of the MSEs of the ridge estimators. These condi-
tions, however, depend on unknown parameters. Using sample estimates of these
parameters to determine c would yield an estimator with complicated and as yet
unknown sampling properties, so that inferences about B could not be made. The
ridge technique essentially consists of an arbitrary numerical adjustment to the
sample data, and one does not really know how to interpret the resultant
estimators.]

7 See A. E. Hoerl and R. W. Kennard, “Ridge Regression: Biased Estimation for Non-Orthogonal
Problems,” Technometrics, 1970, pp. 55-68, for an exposition of the theory; and A. E. Hoerl and
R. W. Kennard, “Ridge Regression: Applications to Nonorthogonal Problems,” Technometrics, 1970,
pp. 69-82, for two illustrations.
+ See P. Schmidt, Econometrics, Marcel Dekker, New York, 1976, pp. 48-55, for this result and a
very useful discussion of the theory of ridge regression.
§ For a definition of MSE see Chap. 2, pages 27-28.
{| For further discussion see the series of papers in Journal of the American Statistical Association,
75, 1980, pp. 74-103.
FURTHER TOPICS IN THE k-VARIABLE LINEAR MODEL 253

Another approach to improving the MSE involves the suggestion that one or
more explanatory variables be dropped in order to improve the MSE of the
remaining coefficients. To illustrate the approach consider just a three-variable
model,
y = Bx. + Bix, +-u (6-80)
where the variables have been expressed in deviation form.} Let us denote the
coefficients resulting from the application of OLS to Eq. (6-80) as
b,.; = OLS estimate of B,
b\32 = OLS estimate of B,
From the properties of the OLS model we know these estimators are unbiased
and their sampling variances are$
o2
var(b
( 123) SS rx3(1 fn r3) (6-81)

ee
var(b,39) = (6-82)
eal _ 7)
Clearly, as 73 gets close to unity, both sampling variances increase dramatically.
Now consider the simple regression of y on x, and denote the slope coefficient by
yx
by = 5
LXS
Substituting for y from Eq. (6-80) gives
Yx,u
bi, = By + bs)
B; + a (6-83)
Dx
where
b =
x53
eee

a x5

+ Strictly speaking, when the relation is written in deviation form, the disturbance term is u — i,
but this slight complication has no effect on any of the derivations in which we are interested and so
may be ignored.
+ These are derived from the general formula

vaD> Ea(xk) = =, eS nail


bi3.2 Lxe Lx? - (2xx3) —Ux2x3 x3

which gives
a7 DxF i o
var(b\23) =
Lx3Dx} — (Lx_x3)° tsEx3(1 = 3)
where r>, is the simple correlation coefficient between x, and x3. A similar derivation yields the result
for var( 5,32).
254 ECONOMETRIC METHODS

From Eq. (6-83) it follows that


E(b,.) = By + bb; (6-84)
and}
mo
var(b,,) = Ee (6-85)

Thus b,, is a biased estimator of 8,, unless x, and x, are orthogonal so that
b,, = 0. However, comparison of Eqs. (6-81) and (6-85) shows that b,, has a
smaller sampling variance than b,,;. The possibility then exists of a tradeoff
between bias and variance. The crucial question is under what conditions b,, may
have a smaller MSE than 5), 3.
As shown in Eq. (2-23),
MSE = sampling variance + square of bias
Thus
2
MSE(b,,) = sear
0
b2, B2
2

and
o2
MSE(6,,
3) =
Exel ‘a ra)
A little algebra then showst
MSE(b,,)
MSE(b,,) |* ee)
——~—*~ = 14+ 73(7? - 1) 6-86

where

2 = a a
Bs = —3_
Bs (6-87)
0° /Ex3(1—rZ) — var(b,39)
This 7° statistic is the ratio of the square of the true (but unknown) B, to the true
(not the estimated) variance of b,,,. From Eq. (6-86), if 7? < 1,

MSE(b,,) < MSE(b,,


3)
Thus if one were mainly interested in obtaining as accurate an estimate as
possible of 8,, and if one felt confident that 7? was less than unity, it might seem
sensible to drop x, from the regression and carry out a simple regression of y on

7 Notice that, in this case, var(b,7) is the variance about a biased expectation. From first principles

var(b,.) = E{[bis = E(by)'|) = E{[ by = p= bsB3]°)

OF (=) os
xe xs

+ See Problem 6-9.


FURTHER TOPICS IN THE k-VARIABLE LINEAR MODEL 255

X. The snag of course is that r” is unknown, and the consequences of mistakes


about its value could be fairly serious. If, for example, B, is three times its true
standard error and r;, is around 0.8, then MSE(b,,) will be over seven times as
large as MSE(5,, 5).
In view of Eq. (6-86) it might seem plausibleto drop x, from the regression if
the estimated t value is numerically less than 1, that is, if
b2

1Sp eS 6-88
57x (il ~ 1) ( )

Thus one may define a conditional omitted variable (COV) estimator of B, as


bi if
F< 1
ecav— bas ifF>1 (6-89)
Other COV estimators may be defined using critical F values other than unity.
Feldstein has investigated the MSE of boy relative to MSE(b,,,;) for various
values of r,,, various values of 7, and also for several critical F values, including
unity and the conventional F),;., When |t| > 1, sampling fluctuations can still
give F <1, and consequently the COV estimators are inferior to OLS.
Feldstein’s main conclusion is

OLS is preferable to any of the COV estimators unless the researcher has a
strong prior belief that tT < 1.4

Feldstein also investigates the properties of a weighted (WTD) estimator


which is simply a linear combination of b,, and b,,;. We define

bwrp = Abi23 + (1 — A)bi2 (6-90)


It may be shown that the value of A which minimizes MSE (bw 7p) is§
2
X= (6-91)
1+ 7?
This is the same unknown 7” statistic already defined in Eq. (6-87). The WTD
estimator could be made operational by computing the ¢? statistic defined in Eq.
(6-88), hence computing
t2

1+ 7?
and substituting this value of A in Eq. (6-90). Feldstein’s simulation experiments
show the WTD estimator to be generally superior to the various COV estimators
in his study, but to be inferior to OLS when |7| > 1.5. Thus exhaustive study of

+. S. Feldstein, “Multicollinearity and the Mean Square Error of Alternative Estimators,”


Econometrica, 41, 1973, pp. 337-346. See especially Tables I, I, and II.
+ M.S. Feldstein, op. cit., p. 344.
§ See M. S. Feldstein, op. cit., or work it out directly in Problem 6-10.
256 ECONOMETRIC METHODS

the three-variable case suggests that, even in the presence of high correlation
between x, and x3, the best procedure is probably the straightforward OLS
regression of y on x, and x3. Only if the investigator has really strong prior beliefs
that B, is less than /var(b,,,), should x, be dropped from the regression. Even
this nonstartling advice to drop a variable when you are fairly sure its coefficient
is “small” is only helpful if the investigator is mainly interested in the other
coefficient, B,.
Even though these results on the three-variable case are not very helpful,
considerable work has been done on extensions of the approach to the k-variable
case. As we have already seen in Chap. 5, setting a coefficient or group of
coefficients at zero is a special case of imposing a set of linear restrictions on the
coefficients. Thus the question arises whether the imposition of a set of restric-
tions will result in estimators which are better in some MSE sense than the
unrestricted OLS estimators, even though the restrictions may not, in fact, be
true.
The first problem is the generalization of the MSE criterion to a number of
estimators. Consider the usual linear model
y=Xfp+u
with the set of g (< k) restrictions embodied in
RB =r
As seen in Eq. (6-5), the estimator embodying these restrictions is

b, = b + (X’X) 'R’[R(X’X)
'R’] ‘(r — Rb)
where
b = (X’X) ‘'X’y
is the unrestricted OLS estimator. We may define the MSE matrix for b, as

MSE(b,) = E{(b« — B)(bs — B)’} (6-92)


This is a symmetric k X k matrix with the MSEs of the individual coefficients
displayed on the principal diagonal. The typical off-diagonal term is

E(x) 8)\( yj atB)) bf


which is essentially a covariance defined in terms of the true £,, 8; values rather
than in terms of the expected values of the estimators. One might then say that b,
is better in MSE than b if
c/MSE(b,)¢ < ¢’MSE(b)c (6-93)
for any nonnull k-element vector c.f This is a very strong criterion, requiring that
+ Notice that Eq. (6-93) is equivalent to the condition MSE(c’b,) < MSE(c’b) for

MSE(c’b,) = E{c’(bs — B)’} = E(c’(by — B) (bs — B)’c) = e/MSE(b,)e


Jv en 2 wy,

and similarly for MSE(c’b). Thus Eq. (6-93) requires that the MSE of any linear combination of the
elements of b, be no greater than the MSE of the same linear combination of the elements of b.
FURTHER TOPICS IN THE K-VARIABLE LINEAR MODEL 257

any quadratic form in MSE(b,) be less than or equal to the corresponding


quadratic form in MSE(b). A much weaker criterion would be

tr MSE(b,) < tr MSE(b) (6-94)


that is, that the sum of the MSEs of the restricted estimators be less than or equal
to the sum of the MSEs of the unrestricted estimators.
The problem of determining the conditions under which Eq. (6-93) or Eq.
(6-94) might hold has been investigated in a number of papers by Wallace,
Toro-Vizcarrondo, and Goodnight.} It can be shown that the restricted estimators
(even when the restrictions are incorrect) will have smaller variances than the
unrestricted OLS estimators. However, taking expectations of Eq. (6-5) shows
that

E(b,) = B + (X’X) 'R'[R(X’x)


'R’] '(r — RB)
so that b, will be a biased estimator if the restrictions are not correct. This is a
generalization of the tradeoff between bias and variance in the previous simple
example. Wallace and Toro-Vizcarrondo show that the strong MSE criterion
(6-93) will be satisfied if

A 20° aa:
(6-95)
As in the previous simple case, this condition involves the true but unknown B
vector and the unknown o?”. If these are replaced by their OLS estimators and the
resultant value of the statistic in Eq. (6-95), denoted by A, it is easy to see that

F=

2%,
+ C. Toro-Vizcarrondo and T. D. Wallace, “A Test of the Mean Square Error Criterion for
Restrictions in Linear Regression,” Journal of the American Statistical Association, 1968, pp. 558-572;
T. D. Wallace and C. E. Toro-Vizcarrondo, “Tables for the Mean Square Error Test for Exact Linear
Restrictions in Regression,” Journal of the American Statistical Association, 1969, pp. 1649-1663;
T. D. Wallace, “Weaker Criteria and Tests for Linear Restrictions in Regression,” Econometrica, 40,
1972, pp. 689-698; J. Goodnight and T. D. Wallace, “Operational Techniques and Tables for Making
Weak MSE Tests for Restrictions in Regressions,” Econometrica, 40, 1972, pp. 699-709.
£ Notice that dropping x, from the model
y = Box.
+ B3x3+u
is equivalent to imposing the restriction

Ont
B, =
|B;
and with these specifications of r and R condition (6-95) becomes

ae
2var(bi32) 2
or tT? < 1 as derived in Eq. (6-87).
258 ECONOMETRIC METHODS

where
F ae (e,ex 8 e’e)/q

e’e/(n — k)
is the sample statistic, defined in Eq. (6-8), for testing the null hypothesis
H,: RB=r
When H, is true, A = 0, and F has the central F distribution with g, n — k degrees
of freedom. The test of H, is made, as we have seen, by comparing the sample F’
with a preselected critical value from the central F distribution. The basic result
of Toro-Vizcarrondo and Wallace (1968) is that when Hy is not true, the F
statistic, defined above, follows the noncentral F distribution with degrees of
freedom g, n — k and noncentrality parameter A, defined in Eq. (6-95). Thus the
test for the improvement in MSE is to compare the sample F statistic with a
critical value from the noncentral F distribution with X = 0.5. Critical points of this
distribution are tabulated in Wallace and Toro-Vizcarrondo (1969). The practical
procedure is as follows:

1. Compute the usual F statistic, based on the difference in the residual sums of
squares from the restricted and unrestricted regressions.
2. If F > F(q,n — k)oos, say, in the table by Wallace and Toro-Vizcarrondo,
reject the hypothesis that the restricted estimators are better in MSE. If the
sample F is less than the critical value, use the restricted estimators.

The above procedure is for the strong MSE criterion, embodied in Eq. (6-93).
Wallace (1972) has shown that the weaker MSE criterion (6-94) will be satisfied if

AS5Z

where g is the number of restrictions. The appropriate critical values of F are


tabulated in Goodnight and Wallace (1972). As an indication of how these
procedures would work, consider the following critical F values, taken from the
appropriate tables:

Noncentrality parameter
A=0 A = 0.5 A=q/2

F(3, 20)9.95 3.10 4.06 5.73

If one were testing Hj: RB =r at the 5 percent level with g = 3 and


n — k = 20, Hy would be rejected if the sample F exceeded 3.10, and one would
conclude that the restrictions were not true. However, a sample F as high as 4.06
in the case of the strong MSE criterion, and as high as 5.73 in the case of the
weak MSE criterion, would still lead to the imposition of the restrictions and the
use of the restricted estimator on the Wallace criterion.
FURTHER TOPICS IN THE k-VARIABLE LINEAR MODEL 259

These procedures, as in the COV estimator of Eq. (6-89), rest on a prior


significance test. Their actual performance in repeated applications would need to
be evaluated as Feldstein did for the COV estimator in the three-variable case. It
is also doubtful whether, in practice, econometricians would wish to impose
restrictions, which seem unlikely to be true, in order to improve the estimators in
the MSE sense. For example, suppose the estimation of a production function
leads to the rejection of the hypothesis of constant returns to scale, but the sample
F value does not exceed the critical value for the weak MSE condition. The
restricted estimates of capital and labor elasticities may have lower MSEs than
the unrestricted estimates, but they will sum to unity and incorrectly indicate
constant returns to scale. The investigator must make a value judgment as to
whether this kind of tradeoff is desirable.
The upshot of this discussion of multicollinearity is not very comforting.
Some data sets contain very little information and do not enable one to disentan-
gle the relative effects of variables with much precision. The multiple correlation
coefficients among the explanatory variables will indicate those coefficients which
are likely to be most adversely affected by the collinearity, and one should not
readily drop these variables from a regression because of low 1f statistics. Re-
stricted estimators may have greater precision, but at the cost of a bias. Only the
accumulation of more and better data sets will yield more precise estimates of
complex interrelationships.

6-6 SPECIFICATION ERROR

Strictly speaking the term specification error covers any mistake in the set of
assumptions underpinning a model and the associated inference procedures, but it
has come to be used particularly for errors in specifying the data matrix X.f
There are two problems involved in specifying X. The first is knowing which
variables (such as income, relative prices, etc.) to include, and the second is in
what mathematical form each variable is to be included. So far we have blithely
assumed such knowledge to be readily available. In practice it is not. Economic
theory can normally indicate the set of explanatory variables corresponding to
any assumed model (utility maximization, cost minimization, etc.), but theory
cannot usually indicate the precise form of the relationship. In less favorable
situations where there is no clearly articulated theory there may be no clear guide
to relevant explanatory variables. On top of all this one may not be able to obtain
measurements on appropriate variables and, hence, have to use proxy variables in
their place.
To establish the effects of misspecification of X, let us suppose that the true
model is
y=XBp+u (6-96)
+ See H. Theil, “Specification Errors and the Estimation of Economic Relationships,” Review of
the International Statistical Institute, 25, 1957, pp. 41-51.
260 ECONOMETRIC METHODS

with
E(u)=0 and = E(w’) =o7I
The model specified by the investigator is
y=X,B+u (6-97)
where, of course, some variables may be common to both X and X,. The
investigator thus computes the estimated coefficient vector

b, = (XiX,) Xsy
Substituting for y from Eq. (6-96) gives

by = (X4Xx) 'X4XB + (X4X4) Xu


Thus

E(by) = (X4Xa) -X4XB (6-98)


and the expectations of the estimated coefficients are seen to be not the true
population parameters but rather linear combinations of those parameters. We
may distinguish a number of different possibilities.

Case 6-1: Exclusion of relevant variables Suppose that the X, and X matrices are
Rei (X) Xre ex Ay
x= [x, Xp. 0 XX yp x;,| =[X, X,]
The investigator has correctly included the first r explanatory variables but
mistakenly omitted the remaining k — r variables. It follows directly that

(X,X4) XX = (X,X,)'[X{X,_ XX,]


A (1, (XX) aN x]
Thus
E(05;)
= 8,+.0;74) 8 eh pee
aie jie eee aE
where a, ,,|,---, @;,, are the elements in the ith row of (X{X,)~'X{X,. The
columns of this last matrix are seen to be the OLS coefficients obtained when each
excluded variable in turn is regressed on the set of included variables. Thus even
though the investigator has managed to include a number of the true explanatory
variables, their coefficients will be biased, and the bias is seen to be some linear
combination of the true coefficients of the excluded variables. This of course
destroys the conventional b.l.u.e. property of OLS estimators. The conventional
inference procedures are also undermined, not only because of Eq. (6-98), but
also because the disturbance variance cannot be correctly estimated. When yis
regressed on X, = X, =[x, X, -:+: x,], the residual vector is M,y, where

M, =1— XX XX,
is a symmetric idempotent matrix of rank and trace equal to n — r. The residual
FURTHER TOPICS IN THE k-VARIABLE LINEAR MODEL 261

sum of squares is
RSS = y’My
Writing Eq. (6-96) in partitioned form as
y = X,B, + X,B, +u
and substituting in RSS gives

RSS = (X,B, + u)’M,(X,, + u)


since M,X, = 0, and so
RSS = uM + B;XM,X,B, + 2B;X,Mu
Thus
E(RSS) = E(u/Muu) + B; X’,M_,X,B,

= 07(n—r) + BiX5M,X,B,
and so
RSS 1 ros
E| ma | = o2 apw B2X2Mi X28 (6-99)

The matrix of the quadratic form in Eq. (6-99) is the matrix containing the sums
of squares and the cross products of the residual vectors obtained when each
excluded variable in X, is regressed on the set of included variables X,. Apart
from a constant divisor it is a variance-covariance matrix and thus positive
semidefinite, so Eq. (6-99) establishes that the residual variance estimated from
the specified regression of y on X, will, on average, overestimate the true
disturbance variance. As in Eq. (6-98), the bias involves the true but unknown
coefficients of the excluded variables. The bias in the regression coefficients would
disappear if the included and excluded variables were orthogonal, X/, X, = 0, but
the estimated disturbance variance would have expectation
E| RSS ]
]meh=o°+
esta. ye
oerrarBy XX 2B, >o 2
iat,
so that faulty inferences would still be made.

Case 6-2: Inclusion of irrelevant variables The X, and X matrices could now be
specified as

X, = [X, X,]

Xi [X,]
where X, isn X k and X, (the matrix of irrelevant variables) is n < s. When each
true variable in X, is regressed on [X, X,], the least-squares fit will force the
coefficient of that same variable on the right-hand side to unity and all other
coefficients to zero. Thus
ot I
(XEXe) XX = |‘|
262 ECONOMETRIC METHODS

where 0 is a null matrix of order s X k. Thus the coefficients of the variables in X,


will be unbiased estimates of the true parameters, and the coefficients of the
variables in X, will have zero expectations. The residual variance will also be an
unbiased estimate of o”. The residual sum of squares from the regression of y on
X, is
RSS = y’My
where
M = I~ X4(X4X4)
X4
Since the true model is, by assumption,
y= X,B, +u
RSS = (X,B, + u)’M(X,B, + u)
u’Mu + 28; X)Mu + 6B;X{MX,B,
= uMu
since MX, = 0 as MX, = M[X, X,]= 0. Thus
E(RSS) = o?trM
=(n-—k-s)o?
and so
E(s”) = 0?
where
5
ie RSS ee
n—-k-s
Notice that although the true model only contains k variables, the correct divisor
in s* isn — k — s, where k + s is the number of variables actually included in the
misspecified model.

It would seem from the discussion of these two cases that it is more serious to
omit relevant variables than to include irrelevant variables since in the former
case the coefficients will be biased, the disturbance variance overestimated, and
conventional inference procedures rendered invalid, while in the latter case the
coefficients will be unbiased, the disturbance variance properly estimated, and the
inference procedures will be valid. This constitutes a fairly strong case for
including rather than excluding variables from a regression equation. There is,
however, a qualification to this view. Adding extra variables, be they relevant or
irrelevant, will lower the precision of estimation of the relevant coefficients. This
point has already been illustrated for a simple model in the previous section on
multicollinearity. Suppose the true model is
y= Bx, +u (6-100)
and the assumed model is
y = Bx. + Byxz4+ u (6-101)
FURTHER TOPICS IN THE K-VARIABLE LINEAR MODEL 263

The sampling variance of the estimate of B, obtained by applying OLS to Eq.


(6-101) is, as shown in Eq. (6-81),

var(.b,) = Eee Ak
yx3(1 = 13)
whereas the correct sampling variance, under Eq. (6-100), is o*/Lx3. More
generally if X, indicates the set of true explanatory variables, X, the set of
irrelevant variables, and if y were regressed just on X,, the variance matrix for the
estimated coefficients would be

o*(X,X,) | (6-102)
When y is regressed on [X, X,], the variance matrix for the coefficients of the
variables in X, is

oN Ki KER K Kae KEK} (6-103)


The diagonal elements in Eq. (6-102) will be smaller than those in Eq. (6-103). In
the event of substantial collinearity this drop in precision may be serious, but
subject to this qualification, including irrelevant variables would seem a less
serious problem than the exclusion of possibly relevant variables. Riddell and
Buse derive all the main results for Cases 6-1 and 6-2 in a unified fashion by
treating them as special cases of restricted least squares.

Case 6-3: The general case The general case relates to the mistaken use of the X,
matrix instead of the X matrix, as specified in Eqs. (6-96) and (6-97). The residual
sum of squares from the regression of y on X, is
e’e = y’Myy
where

M, = 1— X,(X4X4)
X%
Substituting for y from Eq. (6-96) gives
e’e = (XB + u)’M, (XB + u)
= uM,u + B’X’M,XB + 2B’X’M,u
and

PNIS Oeak ame b Miase) (6-104)


e’e =
2D, Fe
1 ry
-| 4

If no specification error were made, the M, matrix would become


M, =I — X(X’X) 'X’
and the quadratic form on the right-hand side of Eq. (6-104) would vanish. For
any specification error at all the matrix of the quadratic form, being essentially
+ See W. C. Riddell and A. Buse, “An Alternative Approach to Specification Errors,” Australian
Economic Papers, 19, 1980, pp. 211-214.
264 ECONOMETRIC METHODS

the variance-covariance matrix computed from the residual vectors obtained when
each variable in X is regressed on X,, is positive semidefinite. Thus the expected
value of the residual variance computed from the regression of y on X, will
exceed o7 and would only fall to «? when X, = X. This provides a rationalization
for the common practice of searching among regressions to find the minimum
residual sum of squares (or maximum R?’), though, of course, in any specific
application sampling fluctuations might yield a lower residual sum of squares for
X, than for X.
The result obtained in Eg. (6-98) that specification error leads to biased
estimates of the population parameters must be interpreted with care. Suppose,
for example, that y indicates observations on the rate of inflation, X the set of
explanatory variables in a “fiscalist’” theory of inflation, and X, the set of
explanatory variables in a “monetarist” theory. A fiscalist will estimate Eq. (6-96)
and a monetarist Eq. (6-97). Monetarists will have little interest in the “news”
that their monetary coefficients are biased estimates of the coefficients of fiscalist
variables, nor would fiscalists be interested in the reverse information. Even if one
model really is the “true” model, the substantial correlation existing among
economic data may well help the “wrong” theory to put up a reasonably good
statistical showing. We are touching on the very difficult problem of the choice
between alternative models, which we will discuss in some more detail in Chap.
12?

PROBLEMS

6-1 A data matrix of full column rank is partitioned as

X=[X, X,]
where X, is n X k, and X, is n Xk. Show that the upper left-hand block in (X’X)~' may be ©
expressed as
7 —1
(X,M,X,)

where

M, =I XK, (X5X5)Xs
Give a least-squares interpretation of M,X, and hence of X{M,X,.
6-2 The following estimated equation was obtained by OLS regression using quarterly data for 1958
to 1976 inclusive:

Y= 2.20 + 0.104 x,; — 3.48x,. + 0.34x,3


(3.4) (0.005) (2.2) ~=(0.15)
Standard errors are in parentheses, the explained sum of squares was 109.6, and the residual sum of
squares 18.48.
(a) Test the significance of each of the slope coefficients.
(b) Calculate the coefficient of determination R2.
(c) When three seasonal dummy variables were added and the equation was reestimated, the
explained sum of squares rose to 114.8. Test for the presence of seasonality.
(d) Two further regressions, based on the original specification, were computed for the subperi-
ods 1958, quarter 1, to 1968, quarter 4; and 1969, quarter 1, to 1976, quarter 4, yielding residual sums
FURTHER TOPICS IN THE K-VARIABLE LINEAR MODEL 265

of squares of 9.32 and 7.46, respectively. Test the following hypotheses:


(i) The error variances are identical in the two subperiods.
(ii) The coefficients are identical in the two superperiods.
(UL, 1981)
6-3 The following regression was estimated from 16 quarterly observations (t ratios in parentheses):

Y= OT 090X743 Siete 6.559 yt 2183935 R* = 0.68


(3.7) (0.27) (3.37) (3.40) (3.37)
where S;, = | in the ith quarter and 0 otherwise. Explain the implied pattern of seasonal variation and
interpret the result.
(UL, 1980)
6-4 A production function model is specified as
Y; = By + BX; + B3X3; + 4;
where Y; = log output, X,,; = log labor input, and X;,; = log capital input. The data refer to a sample
of 23 firms, and observations are measured as deviations from the sample means

Ex}; = 12 Lx;X3; = 8
Ex3; = 12 LyiX2; = 10
Yy7=10 Ly, x3, = 8
(a) Estimate B,, 83, their standard errors, and R?.
(6) Test the hypothesis that B, + B; = 1.
(c) Suppose now that you wish to impose the a priori restriction that 8, + B; = 1. What is the
least-squares estimate of 8, and its standard error? What is the value of R? in this case? Compare
these results with those obtained in (a) and comment.
(UL, 1979)
6-5 A set of cross-section data on family income y and expenditure c is partitioned into subsets of
observations, relating to families headed by:

1. Manual workers
2. Salaried workers
3. Self-employed

A regression of log c on log y is computed for each subsample and for the full sample, yielding:

A
B : s? 10

Manual workers 1.02 0.24 102


(0.06)
Salaried workers 0.91 0.46 104
(0.1)
Self-employed 0.76 0.30 26
(0.08)
All families 0.86 0.39 232
(0.05)

Here B is the slope coefficient (standard errors in parentheses), s* is the residual variance, and T is the
sample size.
Test the hypotheses that:
(a) The elasticity of c with respect to y is the same for all occupational classes.
(b) Its value is unity.
Interpret your results and give some possible explanations for the observed differences.
(UL, 1979)
266 ECONOMETRIC METHODS

6-6 On the assumption that the elements of B obey the restrictions


RB =r
show that the variance-covariance matrix of the restricted estimator b,, defined in Eq. (6-5), is

var(by) = 0°{(X’K)'
—(XX) 'R'[R(X’X)'R’] 'R(X’X)'}
ox = = | a

6-7 The model


Y=a, + a} E,+ afEB, + u
is estimated by OLS, where E, and E, are dummy variables indicating membership of the second and
third educational classes, respectively. Show that the OLS estimates are

aye | Y,
as|}=|%- Y,
a} Yerdi
where Y, denotes the mean value of Y in the ith educational class.
6-8 Rework the estimation problem based on the data in Table 6-6, using any other cell as the starting
position and confirm that one obtains the same numerical estimates of the expected number of hours
as those given in Table 6-7.
6-9 Prove the result on MSEs stated in Eq. (6-86).
6-10 Derive the result given in Eq. (6-91).
6-11 The set of restrictions RB = r, with appropriate partitions of R and B, may be reformulated as
R\B, + R2B, =r
where R, is g X q and nonsingular and R, is q X (k — q). Show that the restricted estimator b,,
defined in Eq. (6-5), may be obtained in two stages, namely:
(a) Regress the vector (y — X,Rj 'r) on the matrix (X, — X,R; 'R,) to obtain an estimate b, of
Bo.
(b) Substitute this estimate in

B, = RK;'(r — RB.)
to obtain an estimate of B}.
CHAPTER

SEVEN
MAXIMUM LIKELIHOOD ESTIMATORS AND
ASYMPTOTIC DISTRIBUTIONS

7-1 REVIEW AND PREVIEW

Chaps. 5 and 6 have set out the main features of the k-variable linear model. It is
very important to emphasize that the results obtained so far depend upon the
particular set of assumptions made in specifying the model. It will be helpful to
review those assumptions and results very briefly as this sets the stage for the
remainder of the book, which is concerned with the many problems that arise in
econometrics when various assumptions underpinning the simple model that we
have considered so far have to be revised and extended.
The k-variable linear model with n sample observations was specified as
y=X 8B
+ U4
yen ey ay d
(n X 1)(n X k)(k X 1) (nx1)

with two crucial sets of assumptions, namely, assumptions about the X matrix
and assumptions about the disturbance vector u, that is,

1. X is of full column rank and nonstochastic


2. u has the properties E(u) = 0 and var(u) = oI

or

3. u~ N(0, 071)
267
268 ECONOMETRIC METHODS

The combination of assumptions 1 and 2 yields the result that the OLS
estimator b = (X’X)~!X’y with var(b) = 07(X’X)' is a best linear unbiased
estimator of B. The development of inference procedures required an assumption
about the form of the distribution of the disturbance term, and the combination
of assumptions 1 and 3 resulted in a comprehensive set of exact, finite sample
inference procedures—tests of coefficients, confidence intervals, analysis of vari-
ance procedures, tests of structural change, and so forth.
The above assumptions are very restrictive, and parts of Chap. 6 examined
some issues relating to the X matrix. Sections 6-3 and 6-4 indicated various
applications resulting from the incorporation of dummy variables among the X’’s.
Section 6-5 examined the problems that arise when the X variables are highly
correlated, and Sec. 6-6 discussed the problems involved in specifying the X
matrix, that is, in knowing which variables in what functional form should
comprise the columns of X. None of these issues violates the basic assumption
that X was nonstochastic. It is, however, very important to relax this assumption.
Also important is the relaxation of assumptions about the disturbance term. We
will see in Chap. 8 that many real-world situations would preclude var(u) from
having the extremely simple form set out in assumption 2, and it is important to
develop appropriate estimators for these more complicated situations. We also
need to ask what are the effects of removing the normality assumption for the
disturbance term. Finally we note that when X is nonstochastic, there is no
question of any statistical dependence between the X’s and the u’s, but when the
nonstochastic assumption is removed, this now becomes a possibility to be
investigated, and, in fact, this particular problem has generated some of the major
developments in econometric theory.
In tackling this broader range of complex problems, the least-squares princi-
ple alone cannot always yield an appropriate estimator. We need, therefore, to
introduce the powerful maximum likelihood principle. Furthermore, in many of
the new problems it proves excessively laborious and often impossible to derive
exact finite sample results, but it is possible to derive results which hold in the
limit, or asymptotically, as the sample size becomes infinitely large. Thus we need
a simple introduction to asymptotic theory, and this is attempted in Sec. 7-2,
followed by an introduction to maximum likelihood estimators in Sec. 7-3.

7-2 SOME REMARKS ON ASYMPTOTIC THEORY

Asymptotic theory is concerned with the behavior of random variables as the


sample size tends to infinity. To illustrate, let x,, denote the mean of a random
sample of n observations drawn from some population of x values. Or let b,
indicate the estimated slope of an OLS regression based on n pairs of sample
observations. Both x, and b, are random variables with probability density
functions (pdf) denoted, say, by f(X,|",0*) and f(b,|a, B,o2). The first pdf
assumes that the x distribution involves just two parameters, the mean p and the
variance o*. The second pdf involves the three parameters of the two-variable
linear model of Chap. 2.
MAXIMUM LIKELIHOOD ESTIMATORS AND ASYMPTOTIC DISTRIBUTIONS 269

The crucial question in asymptotic theory is how random variables such as x


or b, and their pdf's behave as n — oo. For our purposes there are two main
aspects of this behavior, the first relating to convergence in probability and the
second to convergence in distribution.

Convergence in Probability
A basic result in elementary statistics states that, if the x’s have been drawn at
random from some distribution with mean p and variance o?,
o2
Ee =o Toands var) ae
Thus x,, is an unbiased estimator of » for any sample size, and the variance tends
to zero as n increases indefinitely. It is then intuitively clear that the distribution
of X,, whatever its precise form, becomes more and more concentrated in the
neighborhood of pu as n increases. Formally, if one defines a neighborhood around
pf. as p + «, the expression
Pri — e < X, < po + 2) = Pri|x, — p| < 2}
indicates the probability that x, lies in that interval. The interval may be made
arbitrarily small by suitable choice of e. Since var(x,,) declines monotonically with
increasing n, there exists a number n* and a 6 (|8| < 1) such that for all n > n*,
Pri{|x, — p| <e}> 1-8 (7-1)
The random variable x, is then said to converge in probability to the constant wp.
As n increases, the probability of x, lying in a specified interval becomes larger,
that is, 6 becomes smaller. Thus an equivalent statement is
lim Pr{|x, — p| < e}= 1 (7-2)

In words, the probability of x, lying in an arbitrarily small interval about can


be made as close to unity as we desire by letting n become sufficiently large. A
shorthand way of writing Eq. (7-2) is

plim x, =p (Ges)

where plim is an abbreviation of probability limit. The sample mean is then said
to be a consistent estimator of the population mean p. By a similar argument the
reader may easily show that, in the two-variable regression, b, is a consistent
estimator of B, since it was shown in Chap. 2 that

(2,) == B (0,) =85


o2
E(b,) and ~_svar(b,) = —

These two examples are very simple in that the estimators are unbiased for all
sample sizes. Suppose we have another estimator, m,,, of . such that
Cc
270 ECONOMETRIC METHODS

where c is some constant. This estimator is biased in finite samples, but

lim E(m,) =
n— oo

and m, is said to be asymptotically unbiased.f Provided var(m,,) goes to zero as n


increases, it may be shown, by use of Chebysheff's theorem, that m,, is a
consistent estimator of p.
Chebysheff’s theorem states that for a random variable X with finite mean
and variance, and o”, and for given A > 0,
]
Pr{|x — p| > Ao} < 2
Applying the theorem to this example gives

nl m,— (1+ alle vvar(m,) |< e

Setting e = Ajvar(m,,) , this becomes


var(m
Pr ne (u+<)}|>el peat)2
E

Taking the limits of each side as n goes to infinity gives


lim Pr(jm, — p| > e) = 0 (7-4)
Thus m,, is a consistent estimator of pm, since Eq. (7-4) is equivalent to the
definition of a consistent estimator in Eq. (7-2). So a sufficient condition for
consistency is that an estimator should be asymptotically unbiased and have a
variance which converges to zero.
One of the great advantages of probability limits is their simplicity of
operation, as illustrated by the following examples:

plim(x?) = (plim x)’

plim(x~!) = (plim
x)7!

ol) fate
whether or not x and y are independently distributed. Probability limits may also
be extended to vectors and matrices. It simply means taking the probability limit
of each element of the vector or matrix, provided of course that such probability
limits exist. Operation with these probability limits is again extremely simple. For

7 An alternative definition of asymptotic unbiasedness will be given shortly in the discussion of


convergence in distribution.
$ Consistency, however, does not necessarily imply asymptotic unbiasedness. The standard coun-
terexample is given in W. P. Sewell, “Least-Squares, Conditional Predictions and Estimator Proper-
ties,” Econometrica, 37, 1969, pp. 39-43.
MAXIMUM LIKELIHOOD ESTIMATORS AND ASYMPTOTIC DISTRIBUTIONS 271

example,
plim(AB) = plimA - plimB
plim(A~!) = (plim A)!
As an illustration, recall the OLS coefficient vector from the basic model in Chap.
5, namely,
b = (X’X) 'X’y
= B + (X’X) Xu
=e (2xx) (2x)

ae
The matrix

consists of the mean squares and mean cross products of the explanatory
variables. If the X matrix is constant in repeated samples, thent

im (Lxex) = (ex]
n>o \Nn n

If the explanatory variables are stochastic, it can be shown that the sample
moments will converge in probability to the population moments. Thus we write

plim(—x’x} => (7-5)


where 2 is a given symmetric, finite, positive definite matrix. It remains to
evaluate
eel
plim( x,
1
plim{5X u,|
plim{— X’u}= eee
n

whit)
plim(—5X.,u,

The element inside the first parentheses is 7. Since E(#) = 0 and var(z) = o*/n,
it follows that plim(#) = 0. For the ith element

E(—EX,u,] = 0
n
which holds both for the case where X is fixed and also for stochastic X on the
assumption of zero covariance with u. Also

var( EX,,u, = —o* : poe


n n
+ For the proof see H. Theil, Principles of Econometrics, Wiley, New York, 1971, pp. 364-365.
272 ECONOMETRIC METHODS

In view of Eq. (7-5), the probability limit of £X7/n is a constant. Thus the
probability limit of the variance is zero. Repeating the argument for the other
terms,
et ielee t \es
plim(—X'u} = 0

and so
: NEES Niwan We e e
plimb = B + plim(—xx| plim(—-X’u

=B+2°'-0
=8
which proves the consistency of the OLS estimator.

Convergence in Distribution
Return again to the sample mean x, If the population from which the x’s are
drawn at random may be characterized by

x ~ N(p, 07)
then x,,, being a linear combination of normal variables, has a normal pdf. Thus

2 07
x,~N [H,al

and f(xX,,) is normal for every n. The limiting distribution is found by examining
what happens to f(X,,) as n goes to infinity. Since var(X,,) goes to zero, the whole
mass is concentrated on the point p in the limit and the distribution is said to be —
degenerate. A simple transformation of x,, however, can lead to a limiting
distribution which does not collapse on a single point. Consider

Zn Vk)
Clearly, E(z,,) = 0 and var(z,) = 0”. Thus

f(z,)
is N(0,67) — foranyn
that is, the limiting distribution and all finite sample distributions are identical
since the parameters of the distribution do not involve n.
The real application of these ideas comes in situations where finite sample
pdf's either cannot be derived at all or are very difficult to derive and manipulate,
but a tractable limiting distribution can be obtained. The limiting distribution may
then be taken as an approximation for the unknown or intractable finite sample
distribution. As an illustration, suppose the random variable X has mean p. and
variance o7, as before, but the distribution of X is no longer normal. A
MAXIMUM LIKELIHOOD ESTIMATORS AND ASYMPTOTIC DISTRIBUTIONS 273

fundamental result, the central limit theorem, states that}

the limiting distribution of z,, = vn (X,, — p) is N(O, o”).

Thus irrespective of the form of f(x), the limiting distribution of z, 1s still normal,
though the quality of the approximation to any finite pdf will be influenced by the
extent to which f(x) departs from normality. Alternative ways of expressing this
result are

vn (X,, — ) converges in distribution to N(O, 07)

Vn (%, — n) > N(0,


62) (7-6)
This result is often expressed loosely as “x,, is asymptotically normally distributed
with mean p and variance o*/n” and o7/n is then referred to as the asymptotic
variance of x,,. A shorthand version of this statement is
o2

ae AN (1 =| (7-7)
with
o2
asy var(X,,) = an

where AN indicates asymptotically normal and asy var, asymptotic variance. As


already emphasized, the limiting (asymptotic) distribution of x, is degenerate.
The practical import of Eq. (7-7) is that in cases where f(x,,) is intractable we are
taking it, for sufficiently large n, to be approximately normal with mean p and
variance o*/n. These procedures extend directly to multivariate situations, and
we will give an example in Sec. 7-4 with a treatment of the k-variable linear model
when the disturbances are nonnormal.
If we consider the class of consistent and asymptotically normal estimators,
the one with the minimum asymptotic variance is said to be asymptotically
efficient. The mean of the asymptotic distribution provides an alternative measure
of the asymptotic expectation of an estimator, and hence of asymptotic bias.
Previously we implicitly defined the asymptotic expectation (AE) of an estimator
@ as
AE(6) = lim E(6)
no

The alternative definition is

AE(6) = mean of the asymptotic distribution of 6


Similar definitions apply to second- and higher-order moments. In many cases the
two definitions are equivalent, but there are instances where it is important to
+ For a proof see H. Theil, op. cit., pp. 367-369.
274 ECONOMETRIC METHODS

distinguish between limits of sequences of moments and the corresponding mo-


ments of a limiting distribution. Sometimes the moments of a finite sampling
distribution may not exist or cannot be established, although a limiting distribu-
tion with well-defined moments does exist.

7-3 MAXIMUM LIKELIHOOD ESTIMATORS

We will illustrate the principle of maximum likelihood (ML) estimation in the


context of the linear regression model.£ Let us retain the assumption of a fixed
nonstochastic X matrix. The model
y=XBP+u (7-8)
then defines a transformation from u to y. The assumption of a multivariate
density function for u implies a multivariate density function for y, which may be
written
ou
p(y) = p(u) dy
where |du/dy| indicates the absolute value of the determinant formed from the
matrix of partial derivatives§
du, du, du,

du, Ou, du,

ae ox an Dia ae Sa sua

dy, Ay, Oy,


In the case of Eq. (7-8) this matrix is seen to be the identity matrix whose
determinant is unity. Thus
P(y) = p(u)
If we further assume, as before, that u is multivariate normal with mean vector 0
and variance matrix oI, then formula (5-5) gives

P (2102)"/* P hes

and so

p(y) = enn] -sy - xy xB)]


(2102)"/*
(79)
+ For details see H. Theil, op. cit., pp. 375-378.
For a general account of estimators the reader might consult P. G. Hoel, Introduction to
Mathematical Statistics, 4th ed., Wiley, New York, 1971, pp. 196-200; L. D. Taylor, Probability and
Mathematical Statistics, Harper and Row, New York, 1974, pp. 197-230; or M. G. Kendall and A.
Stuart, The Advanced Theory of Statistics, vol. 2, Griffin, London, 1961, pp. 1-67.
§ See App. A-9, Change of Variables in Density Functions.
MAXIMUM LIKELIHOOD ESTIMATORS AND ASYMPTOTIC DISTRIBUTIONS 275

Equation (7-9) involves both the observations on y and the unknown parameters
B and o*. Writing p(y) in the form L(y; B, 0”) soe neees that it is the
probability cap for the y’s, given the parameters B and o? . Alternatively,
writing it as L(B, 0°; y) stresses that for given y it can be fecarticd as a function
of the parameters. It is termed the likelihood function and is conventionally
denoted by the symbol L.
The ML principle is to choose as estimators of B and o? the values which
maximize the likelihood function, given the sample data y. Letting 0’ = [B’ 0°]
denote the vector of unknown parameters and 6 the ML estimator, 6 is obtained
as the solution of the equation
OL
00
=0 (7-10)
In practice the derivation of the ML estimators is often simplified bymaximizing
the log of the likelihood function, that is, by finding 6 as the solution to

a(InL ys
(7-11)
305% i
Since
A(inL) 1 OL
00 Taeieo
the same vector 6 is obtained as the solution to Eqs. (7-10) and (7-11) for any
BO;
Taking the natural logarithm of the likelihood in Eq. (7-9) gives

In L = ~Fin(2m) — 5In(o) ~ —(y ~ XB)'(y — XB)


Differentiating partially with respect to B and o” and evaluating these derivatives
at the ML estimators gives the specific form of Eq. (7-11) as

apes
(In L)
Se 2X’y
/
+ 2X’xB) = 32l (XY, — XXB)
ywA\)
=

7-12
AO)= 5 + al ~ XB)
d(In L)
0
(y-xB) =1 an x8) = 1)

The simultaneous solution of these k + 1 equations gives


= (XX) ky (7-13)
and
2 _ eeA (7-14)

The ML B is seen, in this case, to be idence with the OLS b. The estimate of 07,
however, differs from the unbiased s* of Eq. (5-57) by the factor (n — k)/n,
which illustrates the fact that ML estimators are nos necessarily unbiased. In une
application B is an unbiased estimator of B, but 6? is a biased estimator of o?
276 ECONOMETRIC METHODS

Properties of ML Estimators
ML estimators have a number of desirable properties, some of which hold for
finite samples and some of which only hold asymptotically. Of the finite (small
sample) results one of the most important is the following:

If a minimum variance bound (MVB) estimator exists, it is given by the ML


method.

The minimum variance bound (MVB), developed in the remarkable Cramer-


Rao theorem, establishes a minimum for the variance of an unbiased estimator. It
is important to note that the theorem relates to the class of unbiased estimators
and not just to the subset of /inear unbiased estimators. Furthermore the theorem
establishes a lower bound for the variance, but there may, of course, be situations
where the bound cannot be attained, that is, where one can derive a minimum
variance unbiased (MVU) estimator, but its variance will exceed the MVB. The
bound is derived from the likelihood function.
Consider first of all the case of a single unknown parameter @, a density
function f(y|@), and a random sample of n observations from this density. The
likelihood function is then

L(6ly) = T(x)
Let 6 denote an unbiased estimator . Then the Cramer-Rao theorem states

var(@) > wetilleA ona (7-15)


z
E (dln L ) E 2 L
0*In
00 062
where either of the expressions on the right-hand side indicates the MVB. For the
multiparameter case, let 8 denote an unbiased estimator of the vector 6 of, say, k
unknown parameters. Now we have a variance-covariance matrix for the elements
of 6, denoted by var(@). The multidimensional equivalent of
e|0 2 aS
06?
in Eq. (7-15) is now the symmetric matrix
d7In L
Ls -E| al
Gl Ls Wes ns ee aan Og Lae
902 0, 00, 90, 90,
aes G7 lnSaad vs ee eee luigi
38,00, aa2 90, 00, (7-16)
sp ER
MAXIMUM LIKELIHOOD ESTIMATORS AND ASYMPTOTIC DISTRIBUTIONS 277

R(6) is often referred to as the information matrix. The multidimensional version


of the Cramer-Rao theorem now states

var(6) — R~'(0) isa positive semidefinite matrix

Thus the MVB, for any 6,, is given by the ith element on the principal diagonal of
R-'(8).+
As an illustration of this result, let us return to the k-variable linear model.
The first-order derivatives of the likelihood function were given in Eq. (7-12).
Differentiating these again we obtaint
OF (in Tey, Mal,
ORORT g2
a(InL)_ n__ (y— XB)‘(y
—XB)
(02) 204 o°
iD

a aie = (Xy — X’XB)


Taking expectations of these second-order derivatives and reversing signs gives

-1{ 2002)
dB op’
_ Ly
0
2 2
ee |eME) eZ ea
d(02)” 2o0 o° 204

since

E(y — XB)’(y — XB) = E(u'u) = no?


d?(In |
and Se
|3B 002

since

E(X’y — X’XB) = E{X’(y — XB)} = E(X’'u) = 0


Substituting in Eq. (7-16) and inverting the resulting matrix then gives

[8 OAC XX) 0
R-!
\ 4= 204 (7-17)
oO 0 ia
n
We see immediately that the ML (OLS) estimator of B attains the Cramer-Rao
MVB, since var(B) = var(b) = 07(X’X) !,which is identical with the top left-hand
+ Derivations of the Cramer-Rao MVB may be found in P. G. Hoel, op. cit., pp. 362—365;. L. D.
Taylor, op. cit., pp. 209-213, and M. G. Kendall and A. Stuart, op. cit., Chap. 17.
+ Note that Eq. (7-12) contains B and a? because the first-order derivatives had been equated to
zero. We now ignore the equalities, replace B by B, o* by o”, and differentiate again.
278 ECONOMETRIC METHODS

submatrix in R~!. The same result does not hold for either the OLS or the ML
estimator of 0”. The OLS estimator is

and it was established in Eq. (5-59) that

3l (ee)
,
FpNe 2 kaka)
Thus
2 0° 2
SS 7a ae (aah)
Recalling that the variance of a x* variable is equal to twice its number of degrees
of freedom,

O 20°
var(s?) = Rien, Es

which, for any finite n, is somewhat greater than the variance term given in R~!.
There is, in fact, no unbiased estimator of o” which can attain the MVB. The
derivation of the variance of the ML estimator, 6? = e’e/n, is left as an exercise
for the reader, but in any case it is a biased estimator of or
A second important feature of ML estimators is their invariance property,
which holds for any sample size and may be stated as follows:

The ML estimate of a function g(®) is (6), where 6 is the ML estimate of 9.

We have seen that the ML estimate of o* in the regression model is


6° = e’e/n. Thus the ML estimate of o is simply /e’e/n.
The most important result about ML estimators relates to their /arge sample
or asymptotic properties.

Under certain regularity conditions, ML estimators are consistent, asymptoti-


cally normally distributed, and asymptotically efficient.

Specifically, if 6 denotes the ML estimate of @,

plim6 = 6 (7-18)
§ ~ AN(0,R-') (7-19)
where R has already been defined in Eq. (7-16) as

d7InL
a =2 30 00”
+ Reference may be made, for example, to Kendall and Stuart, op. cit., for a comprehensive
statement of the underpinning assumptions and the derivation of this and other results.
MAXIMUM LIKELIHOOD ESTIMATORS AND ASYMPTOTIC DISTRIBUTIONS 279

Thus the ML estimators, besides being consistent and asymptotically normal, are
efficient in that the asymptotic variance matrix reaches the Cramer-Rao lower
bound.

7-4 SOME ASYMPTOTIC RESULTS FOR THE k-VARIABLE


LINEAR MODEL

In this section we will relax two of the assumptions underpinning the k-variable
linear model, namely, the normality of the disturbance term and the nonstochas-
tic nature of the X matrix.

Nonnormal Disturbances
Let us retain assumptions | and 2 of Sec. 7-1, that is,

1. X is nonstochastic of full column rank k.


2. E(u) = 0 and var(u) = 07I.

but dispense with the assumption of a normal distribution for the u’s. Under
assumptions | and 2 the OLS b is still a best linear unbiased estimator of B with
variance matrix o*(X’X)~'. Moreover, as already shown, b is a consistent estima-
tor of B. Thus even when the disturbances are nonnormal, OLS is still a very
acceptable technique for deriving point estimates. The difficulty is that the various
exact inference procedures outlined in Chaps. 5 and 6 are no longer strictly valid
since their derivation depended on the assumption of normality. However, one
may conjecture that the procedures are reasonably robust for moderate depar-
tures from normality.t More importantly the tests can be given a large sample
justification. This requires the use of two theoretical results.
First, if X is nonstochastic of full column rank k, E(u) = 0, var(u) = o7I, the
elements of X are uniformly bounded, and lim,_,,(1/n)X’X = &, a finite,
symmetric, positive definite matrix, thent
1
- (X’u) ~ AN(0, 72) (7-20)
n

+“... it has been shown that these tests are not very sensitive to departure from normality. If the
errors are not normally distributed but have a variance, it is generally true that only trivial errors are
made in the powers or the levels of significance if we retain the formulae which are strictly applicable
in the case where the errors are normal.” E. Malinvaud, Statistical Methods of Econometrics, 2nd ed.,
North-Holland, Amsterdam, 1970, p. 99. See also Malinvaud’s discussion on pp. 296-302. Additional
references on this topic are P. Schmidt, Econometrics, Marcel Dekker, 1976, pp. 55—64; A. C. Harvey,
The Econometric Analysis of Time Series, Wiley, New York, 1981, pp. 112-117; and G. G. Judge,
W. E. Griffiths, R. C. Hill, and T. C. Lee, The Theory and Practice of Econometrics, Wiley, New York,
1980, Chap. 7.
+ For a proof see P. Schmidt, op. cit., pp. 56-60.
280 ECONOMETRIC METHODS

Second, let (X,,, Y,,} denote a sequence of pairs of random variables, where X,,
has a probability limit and Y, a limiting distribution, that is,
plim X, =c
and

ee
then

(X,¥,) > £(c¥)


or, in words, the product X,Y, has the limiting distribution f(cY).7 If, in
particular, Y, has a normal limiting distribution

¥, > N(u, 0?)


then
D
X,Y, > N(cp, c?o7)

The multidimensional version of the same result is as follows: Let y, denote a


k X 1 vector with limiting distribution
D
y, > N(p,
&)
and z,, an r X 1 vector defined by
ant teNy (7-21)
where H, is an rr X k matrix with probability limit
plimH, =H
then

z, > N(Hp, HOH’) (7-22)


Returning now to the main argument

»—p=(4xx) (4xu} n
maak

n
Thus

vn (b — B) = [Exx) [xu]
vn
This is seen to be in the form of Eq. (7-21), and a direct application of Eq.
(7-22) gives the result that Vn (b — B) has a limiting normal distribution with zero

+ See C. R. Rao, Linear Statistical Inference and Its Applications, Wiley, New York, 1965, pp. 101
ff., for this and other important limit theorems.
MAXIMUM LIKELIHOOD ESTIMATORS AND ASYMPTOTIC DISTRIBUTIONS 281

mean vector and variance matrix given by

Coy Sy =o S|
Thus we may write

in(b - B) 2 N(0, 0257") (7-23)


or

b= AN|B, o> | (7-24)


In a practical application = is unknown and is replaced by the sample estimate
((1/n)X’X) and o? is estimated by s* = e’e/(n — k), where e = y — Xb. Thus the
estimated asymptotic variance-covariance matrix for b is

asy var(b) = s?(X’X)|


which is the finite sample estimator of Chaps. 5 and 6. It can be shown that Eq.
(7-24) essentially ensures that all the conventional ¢ and F tests are valid
asymptotically.f It is a moot point whether in the test of a single restriction
critical values should be taken from N(0, 1) rather than t(n — k), and in the test
of a set of g restrictions, whether one should consult the x*/q distribution or
F(q, n — k). However, it is not a matter of great importance since

Hue) > WO)


ae
and Fags ak)
q
and, moreover, under the assumptions made in this section, the procedures are
only valid asymptotically so that one should not place undue emphasis on precise
levels of significance.

Stochastic X Matrix
The explanatory variables in an econometric relation are not usually subject to
control by the economic researcher, the secretary of the U.S. Treasury, or anyone
else. Rather they are mostly the outcome of the functioning of some
economic/social system. Let us, therefore, characterize the X;, (i = l,..., k;
t = 1,..., n) as possessing some multivariate density function g(X). We will make
two crucial assumptions about this density function, namely:

1. The parameters of g(X) do not involve the parameters B and o? of the


regression model.
2. X and u are independently distributed. This is the strong assumption of full
independence. It implies, for example, that the disturbance u, is independent
of X,, for all i= 1,..., & and all ¢ = 1,..., or, in other words, that u, is

+ For details see P. Schmidt, op. cit., pp. 60-64.


282 ECONOMETRIC METHODS

independent of all past values, the current values, and all future values of all
explanatory variables. This strong assumption would be violated if a lagged
value of Y, such as Y,_,, appeared among the explanatory variables, for u,_,
influences Y,_,, so Y,_, is not independent of u,_,. Furthermore u,_,
influences Y,_,, which in turn influences Y,_,. Thus Y,_, is dependent on
U,_1,U,_7,---, but is independent of u,, u,,,, and all later u’s.f

Formally we may express assumption 2 as


p(u|X) = p(u)
E(u|X) = E(u)
E(y|X) = E(XB + u|X) = XB + E(u)
E(uu’|X) = E(uu’)
The likelihood of the sample observations on both Y and X’s may be written
p(y, X) = p(y|X) - g(X)
= p(u|X) - g(X)
= p(u) - g(X)
We also retain the assumption:

3. u~ NO, 071
Thus the log likelihood becomes

InL=-— 5in(27) = 5In(o?) — el ——(y — Xb)'(y — Xb) + In g(X)


207
This differs from the likelihood for the fixed X model only in the additional term.
in In g(X). Since this term does not involve B and o”, the ML estimators of these
parameters are the same as those already given for the fixed X model in Eqs.
(7-13) and (7-14). Thus the ML estimator of B, which still equals the OLS
estimator, will at least have desirable asymptotic properties when the X’s are
stochastic. Furthermore, reworking the development leading up to Eq. (7-17) now
gives

8 o*E(X’X)' 0
Rela = 0 20+ (7-25)
ae
so that the asymptotic variance matrix for B(= b) is o7E(X’X)~', which is the
MVB.
Turning to the small sample properties it is easy to show first of all that b is
an unconditionally unbiased estimator. From Chap. 5,

b=8 + (X’X) 'X’u


+ This case is treated in Sec. 9-2.
MAXIMUM LIKELIHOOD ESTIMATORS AND ASYMPTOTIC DISTRIBUTIONS 283

Thus
E(b) = B + E{(X’X) 'X’u)
= B+ E{(X’X)” 'x’\ - E(u) since X and u are independent
=B since E(u) = 0
For the variance-covariance matrix
var(b) = E{(b — B)(b — B)’}
= E{(X’X) 'X’uu’X(X’X) |}
= E,{ Eyx(X’X)
u|x "X’uu’X(X"X) "|
where E,,,, indicates the expected value in the conditional distribution of u given
X, and £, indicates the expected value in the marginal distribution of X.+ Thus
var(b) = E,{0?(X’X)")
= 0°E(X’'X)' (7-26)
This shows that the finite sample variance attains the Cramer-Rao lower bound
and differs only from the corresponding formula in the stochastic case in that
(X’X)~! is replaced by E(X’X)~!.
Formula (7-26) may be established in an alternative and instructive fashion.
It was shown in Eq. (5-33) that the variance matrix for b, given some X matrix,
was o°(X’X) |. Emphasizing the conditional nature of this variance we can write
var(b|X) = 07(X’X)|
or letting S = X’X,
var(b|S) = o*S~!
Now suppose that the random X’s can, in principle, generate a finite number of S
matrices S,, S,,..., S,, with probabilities p,, p,,..., p,,- Then the unconditional
variance matrix for b, determined from the marginal distribution, is

var(b) = i=1¥ var(bIS,) - p,


= 0° OE S; Pi
i=l
=o°-E (S32)

= 6°E(X’X)' asin Eq. (7-26)


It is clear that the same argument goes through when the X’s are continuous.
Unfortunately both elements in Eq. (7-26) are unknown. Intuition suggests
replacing
ee
ice iter ae as

+ See App. A-8, Expectations in Bivariate Distributions.


284 ECONOMETRIC METHODS

and E(X'X)o) bye Ose


It can be shown that the result is an unbiased estimator of Eq. (7-26) for
E{s?(X’X) ') = Ex Bias? (XR) iy

o°E(X’X) |
To summarize the position so far, when the X’s are stochastic but indepen-
dent of the u’s, the ML (= OLS) estimators for the fixed X case are still ML for
the stochastic case, and thus all the conventional test procedures are still justified
asymptotically. For finite samples the OLS (ML) b(B) is unconditionally unbi-
ased, and var(b) attains the Cramer-Rao MVB. Moreover the conventional
estimator s*(X’X)~' is an unbiased estimator of that MVB.
The only remaining question concerns the finite sample validity of the
conventional inference procedures. The basic result is that all confidence interval
statements and significance levels derived from the usual formulas are still correct,
but the probabilities of type IJ errors and the widths of confidence intervals will
be different. Confidence intervals and hypothesis tests are derived by calculating
probabilities from the sampling distribution of some appropriate test statistic
under H,. The test statistic is in general some function of the sample observations
and may be denoted by f(y, X). If, for example, the test statistic is found to follow
the ¢ distribution, then we may make the probability statement
Pr{ —t, 2 < tly,
X) < ta} = 1-4 (7-27)
This statement may be used to derive a confidence interval or, equivalently, to
determine acritical region for a test of Hp at the a level of significance.
The statement in Eq. (7-27), however, has been derived under the assumption
of a fixed X matrix, and so it is a conditional probability statement. We need to—
find what unconditional probability statement can be made about f(-) when the
stochastic nature of X is allowed for. Let A denote the event
A: = bestia anete 7,

and let us suppose that the pdf for X gives a finite number of X matrices,
Xsse e
with associated probabilities

Pix Pyro Pre atl) prea

A restatement of Eq. (7-27) then gives


Pr{A|X;}=1—a i=1,...,m
and the unconditional probability of A is

Pr( A) = [Link](A|X,) = ( Ee yet ae l-a (7-28)


i=1 i=]
MAXIMUM LIKELIHOOD ESTIMATORS AND ASYMPTOTIC DISTRIBUTIONS 285

Thus a confidence interval computed in the usual way will have the same
confidence coefficient for random X as for fixed X, and a hypothesis test will have
the same significance level in each case. The assumption of a discrete distribution
for X is only a simplification, and it is clear that the argument carries through for
continuous distributions.+ Notice that it has not been necessary to derive the
sampling distributions of the estimators to establish the above result. These
distributions will be more complicated than those already obtained in Chap. 5. As
an illustration consider the ML estimator ® of the slope coefficient B in a
two-variable model. Letting x denote the vector of sample observations on the
explanatory variable, we know that the conditional distribution of B, given x, iS

f(Blx) = {a2
See
When xis stochastic, the marginal distribution of is

f(B) = f(BIx) - g(x)


where g(x) is the marginal distribution of x. Clearly, the marginal distribution
depends on the distribution of x, but the crux of the matter is that, provided x
and u are fully independent and the parameters of g(x) do not involve B or 7, the
precise form of g(x) does not affect the probability statements underlying
confidence coefficients and significance levels.

PROBLEMS

7-1 Derive the mean and the variance of the ML estimator 6” = e’e/n of the disturbance variance for
the regression model y = XB + u with u~ N(O, 071). (Hint: If w ~ x?(r), then E(w) =r and
var(w) = 2r.)
7-2 Prove that the OLS estimator s* = e’e/(n — k) and the ML estimator of o7 in the regression
model of Problem 7-1 are both consistent.
7-3 Suppose that we have n independent observations y,, y2,.--, ¥,, Say, incomes, drawn by simple
random sampling from a Pareto distribution which has the following pdf:

10,000°
P(yla)=————_—y = 10,000; « > 0
y

What is the ML estimate for a?


(University of Washington, 1979)
7-4 Consider a regression equation
y, =at Bx; t+e; Dal alt

where x, is nonstochastic and ¢, €,...,€, are independently and identically distributed. The
distribution of e, is
f(e,) = Ne OT? eos eee

+ A complete derivation of this and other results is given in F. A. Graybill, Theory and Application
of the Linear Model, Duxburg Press, Mass., 1976, chap. 10.
286 ECONOMETRIC METHODS

Suppose A is not known. Set up the likelihood function for y,, y3,..-, ¥, and describe a way to obtain
ML estimates of a, B, and X.
(University of Michigan, 1977)
7-5 Consider the uniform density
f(X) = lfa 0<X<a
What is the ML estimator for a? Does it attain the Cramer-Rao lower bound? Compare the
asymptotic efficiency of the ML estimator for a with the alternative estimator derived from using the
sample mean, and prove the consistency of that estimator for a.
(University of Chicago, 1975)
CHAPTER
EIGHT
GENERALIZED LEAST SQUARES

A basic assumption underpinning the methods outlined in Chaps. 5 and 6 is that


E(uv’) = 071 (8-1)
This is described as the assumption of spherical disturbances. It involves the
double assumption that the disturbance variance is constant at each observation
‘point and that the disturbance covariances at all possible pairs of observation
points are zero. We now seek to do four things, namely:

1. To indicate some of the more important cases in which assumption (8-1) may
not be fulfilled
2. To determine the properties of OLS estimators if they are (perhaps inad-
vertently) applied, even though the underpinning assumption about the
disturbances is not valid
Ww . To develop tests of whether assumption (8-1) has broken down

4. To develop appropriate estimation procedures for cases where assumption


(8-1) does not apply

8-1 SOURCES OF NONSPHERICAL DISTURBANCES

If the sample observations relate to households or firms in a cross-section study,


the assumption of a common disturbance variance at all observation points is
often implausible. For example, if Y refers to family expenditure and X to family
income, the variance about the Engel curve is likely to increase with the size of X.
287
288 ECONOMETRIC METHODS

Similarly if Y denotes profits and X is some measure of firm size, the same
property is to be expected. The specification of the disturbance variance matrix
would then be
a; 0 0
E(u’) =) 0) tops PO (8-2)
Qu 0 0.
which is the standard case of heteroscedasticity. Formulation (8-2) still assumes
that the disturbances are pairwise uncorrelated.
Suppose, to take a different example, that an investigator is studying the
relationship between wage change and the level of unemployment and that he
measures wage movements in terms of four-quarter overlapping changes. That is,
the annual rate of wage change in quarter ¢ is specified as
WwW, — W-4
W,—4

where w, is the level of the wage index in quarter t. The observed change in the
index, w, — w,_4, is the result of some groups securing a wage change in quarter
t — 3, some others in quarter ¢t — 2, and so forth. If one assumes one fourth of the
labor force to secure a wage change in each quarter, the dependent variable is an
average of these separate changes, and the disturbance term in the macrowage
equation is similarly an average of the separate quarterly disturbances and so
might be specified as
Dea (eee They ete eat) (8-3)
where the e’s indicate the disturbances in the wage change equations for the
separate groups. Let us assume
E(e)=0 and E(ee’)=o/1 (8-4)
It then follows from Eq. (8-3) that |
E(u?) = 402
which does not depend on f, so that the {u,} series is homoscedastic. Further

E(u,u,_\) = 0,
E(u,u,_>) = %60,
E(u,u,_3) = 766,
and
E (uu) =0. = forse
=4
Thus the variance matrix following from Eqs. (8-3) and (8-4) is

1
E(uu’) = raed

(8-5)
GENERALIZED LEAST SQUARES 289

This again is a departure from Eq. (8-1), but in contrast with Eq. (8-2) there is just
one unknown in Eq. (8-5), namely, 02. This is an example of autocorrelated
disturbances. The autocorrelation arose from temporal aggregation over individ-
ual disturbances, which were themselves uncorrelated over time.
To continue the wage change model, it was customary in many early studies
of the Phillips curve for researchers to employ four-quarter overlapping changes.
These, however, have some unfortunate side effects. In consequence it is now
more customary to specify a model of the form

eee hx Bu, (8-6)


W,— |

where the dependent variable is now expressed as a one-quarter change in the


wage index, and x’, denotes a vector of explanatory variables, such as expected
price inflation, the reciprocal of the unemployment rate, and so forth.+ With the
specification (8-6), however, it becomes less appropriate to specify zero correla-
tions for various disturbances. In particular, neighboring disturbances might be
positively correlated. Suppose, for instance, that an “unusually large” settlement
was secured by the workers settling in period ¢ — 1, where unusually large means
“larger than would normally be associated with the vector x ,_ ,.” If this leads to a
greater than usual “push” by the workers negotiating in period ¢, then we might
specify

u, = pu,_, + &, (8-7)


where p is some parameter, and the {e,} series has the usual simple properties
specified in Eq. (8-4). Equation (8-7) specifies a first-order autoregressive (Markov)
scheme— autoregressive since u is related to lagged values of itself, and first-order
because the maximum lag in the autoregression is 1.
Let us introduce the lag operator L such that, when applied to any variable
Xis
Ex, =X,
Lx, = L(Lx,) = Lx; "= %,_2
and, in general,
Ka Sean!
Thus Eq. (8-7) may be rewritten as
(1 — pL)u, = «,
ort
]
u,= ren

(1+ pL+p'L?+---)e,
+ Note that we retain our convention of using u to indicate the disturbance term in Eq. (8-6). It
does not indicate the unemployment rate.
¢ The lag operator may be treated as a scalar for purposes of algebraic manipulation. For any
nonzero constant a we have (1 — a) '=1+a+ a* +--+. Replacing a by pL gives the result
stated above.
290 ECONOMETRIC METHODS

that is,
u, =e, + pe, + pe,» + °°" (8-8)
Squaring both sides of Eq. (8-8) and taking expectations,
02
var(u,) = E(u?) = ; age lp| <1 (8-9)

The right-hand side does not involve ¢, thus the {u,} series has a constant
variance, 0” = o7/(1 — p’).
Using the definition of u, in Eq. (8-8) and the properties of e, assumed in Eq.
(8-4), it is simple to establish that

E(u,u,_1) = po?
E(u,u,_>) a p°o*

and, in general,

E(u,u,_,) a p'o° (8-10)


Thus the variance matrix for a disturbance following a first-order autoregressive
scheme is
1] p p- pie

E(uu’) = 6°] p l p eeu tae (8-11)


pn pos pr? 1

If p were known, this expression, like Eq. (8-5), would involve only one unknown.
There are many other ways, as we shall see later, in which nonspherical
disturbances may arise, but these three examples illustrate some important
patterns. The general nonspherical disturbance matrix may be specified as
E(uv’) = 0° (8-12a) |
or E(uu’) = V (8-125)
The choice of specification depends on whether or not we wish to single out an
unknown scalar, which multiplies all the elements in the matrix as in Eqs. (8-5)
and (8-11). In either case, since we are dealing with a disturbance matrix, 2 and V
are assumed to be positive definite matrices.

8-2 PROPERTIES OF OLS ESTIMATORS UNDER


NONSPHERICAL DISTURBANCES

Our assumed model is now


y= XBp+u
where X is taken as a nonstochastic matrix with full column rank,

E(u)=0 and E(u’)=V_ (oro?)


GENERALIZED LEAST SQUARES 291

The OLS estimator of B may be expressed as usual as

b =8 + (X’X) 'X'u
Thus E(b) =8B
so that OLS is still unbiased. The variance matrix is given by
var(b) = E{(b — B)(b — B)’}
E{(X’X) 'X’uu’X(X’X) '}
O2( XX) XOX). (8-13)
Thus the conventional formula o7(X’X)~' no longer measures the sampling
variances of the OLS estimators, and any application of it is potentially mislead-
ing. More importantly, even if one could use Eq. (8-13) to estimate the sampling
variances, the substitution of these numbers in the conventional ¢ formulas and
confidence interval formulas is strictly invalid since the assumptions used in
deriving those inference procedures no longer apply. For the same reason the
optimal minimum variance property of OLS no longer holds. We will illustrate
these points for various specific departures from spherical disturbances in later
sections and also discuss various specific tests for departures from spherical
disturbances. Now it is more important to turn to the development of a more
appropriate estimator.

8-3 THE GENERALIZED LEAST-SQUARES ESTIMATOR

Suppose we premultiply the assumed model


y= XBp+u
by some n X n nonsingular transformation matrix T to obtain
Ty = (TX)B + Tu (8-14)
Each element in the vector Ty is then some linear combination of the elements in
y, and so forth. The variance matrix for the disturbance in Eq. (8-14) is
E(Tuu’T’) = 0° TOT’ (8-15)
since E(Tu) = 0. If it were possible to specify T such that
TOT’ =I (8-16)
then we could apply OLS to the transformed variables Ty and TX in Eq. (8-14),
and the resultant estimates would have all the optimal properties of OLS and
could be validly subjected to the usual inference procedures.
It is, in fact, possible to find a matrix T which will satisfy Eq. (8-16), for it
was shown in Eq. (4-111) that if & is a symmetric positive definite matrix, a
nonsingular matrix P can be found such that
Q = PP’
Since P is nonsingular,
POPS = (8-17)
292 ECONOMETRIC METHODS

Comparison of Eqs. (8-16) and (8-17) shows that the appropriate T is given by
haw
and it easily follows that
Om [= P’- Ip- ]

a (P~ ea 1

= TT
Applying OLS to Eq. (8-14) then gives

b, = (X’'T’TX) 'X’TTy
= (x‘2~-'x) 'xQ-'y (8-18)
with the variance-covariance matrix given by
var(by) = 07(X’Q>'X)| (8-19)
The estimator b, is defined to be the generalized least-squares (GLS) or Aitken
estimator. Since Eqs. (8-15) and (8-16) imply that Eq. (8-14) satisfies the assump-
tions required for the application of OLS, it follows that by, is a best linear
unbiased estimator of B in the model y = XB + u with E(uw’) = 07Q.
Alternatively Eq. (8-15) may be written
E(Tuu’T’) = TVT’
and setting TVT’ = I gives T’T = V_' so that the GLS estimator may also be
written as

b, = (XV
'X) Ux’v7y (8-20)
with

var(by) = (X’V~'x)7! (8-21


Comparing Eqs. (8-18) and (8-20) shows that it makes no difference to b, whether
var(u) is formulated as o7Q or as V, but care has to be taken to select the correct
expression for var(b,), as a comparison of Eqs. (8-19) and (8-21) shows.t
If the further assumption of normality for the u’s is added, it may be shown
that the GLS estimator is also an ML estimator. We specify
u ~ N(0, 072)
Thus the likelihood function is

L = p(y|X) = ae
DR
(27) o"|Q|'7
xp — <5(y— XB)@""(y — xB)

and the log likelihood is


n n l I
In L = — FIn( 27) ~ FIn( o?) — >In|Q | — mar t — XB)'Q° '(y — XB)

+ Note that the transformation matrix T, defined in Eq. (8-16), differs by a scale factor from that
defined by TVT’ = I.
GENERALIZED LEAST SQUARES 293

Maximizing In L with respect to B implies minimizing the weighted sum of


squares
(y — XB)'Q"'(y — XB) = y'Q>'y — 2B’X'Q~!y + B’X’'R-'XB
Differentiating with respect to B and equating to zero gives

b, = (X-'X) 'x/Q-'y
as in Eq. (8-18).
An unbiased estimator of o* may be derived from the application of OLS to
Eq. (8-14). It is
ee (Ty — TXb,)’(Ty — Txb,)
n—k

_ (y
Xb,)’T’T(y
— —Xb,)
n—k

ee (y — Xb,)’Q° '(y — Xb,)


n—k

_Lae
yQ°'ty — bx’ Q"'y
ae (8-22)

On the assumption of normality for the disturbance term all the inference
procedures of Chaps. 5 and 6 carry through for this model. Thus the test of
Hy: RB=r
is based on

ete Rb,)’[R(X’Q-'X)'R’]'(r — Rb,)/q


g2

having the F(q, — k) distribution under the null hypothesis, where b, is the
GLS estimator defined in Eq. (8-18) and s” the variance estimator defined in Eq.
(8-22).
The above formulas are only operational if the elements of Q are known. In
some exceptional cases this may be so, but in most practical cases it is not. We
must therefore proceed to the development of operational procedures for such
cases, but there is, in fact, no single procedure which is generally applicable. One
must look for the procedure which is best suited to the features of each specific
problem in turn, and that is done in the remaining sections of this chapter.

8-4 HETEROSCEDASTICITY

We have already mentioned in Sec. 8-1 the possibility of heteroscedastic dis-


turbances in cross-section studies. Heteroscedasticity may also arise in dealing
with grouped data. Suppose the model is
Y,=a+ BX, + u, pa
294 ECONOMETRIC METHODS

where the u, are homoscedastic with zero covariances. However, suppose we only
have access to data which have been averaged within m groups, where n, indicates
the number of observations in the ith group. The form of the model appropriate
to the data is now

and clearly

var(u;) = — i=1,...,m

Thus

a 0 0
ta
a7 0 = 07 eae Sn as 0 (8-23)
nS

Os0 =n
where Q is known and the GLS estimator can easily be computed.

Example 8-1 We have taken the same X, Y data as in Example 2-1, only now
it is assumed that they relate to group means. The n, column indicates the
number of observations in each group. The overall means are easily computed
from

Seen)
X= oa mes0 he 4.04

ea 00
Y= a ras) ae 8.00

which are almost identical with the simple means of 4 and 8 in Table 2-1. We
assume that Eq. (8-23) is the appropriate assumption about var(w), that is,

var(u) = 0?Q = 0” ie
GENERALIZED LEAST SQUARES 295

Thus
ny 12
Nn» 6
Qu! = = 1]
10
£5 11
It may then be seen that

ee acy 1 Ny me
EK cae dl ds > :
ee X, - cries
Ns 1 X,

un, &n,X,
un,X, &n,X?

i
and X’2>'y = es
un, X.Y,
Formula (8-18) for the GLS estimator now simplifies to

Ln; HEN, |Dey


b, —

ae es rn;X,Y,

which is a form of weighted least squares. Applying the data from Table 8-1
gives
506,4 + 2026,, = 400
202b, + 1254b,4= 2388
with solution b,, = 0.88 and b,, = 1.76. To obtain the sampling variance of
these estimates, substitute for Q~' from Eq. (8-23) in Eq. (8-22) to obtain for

Table 8-1

X; i nj; n,X; ni, n,X? n, X,Y, ni¥/

2 4 2 24 48 48 96 192
3 7 6 18 42 54 126 294
1 3 11 11 33 1] 33 99
5 9 10 50 90 250 450 810
9 Iba 11 99 187 891 1683 3179

Sums 50 202 400 1254 2388 4574


296 ECONOMETRIC METHODS

this example
bn iY,
(n os k)s? aa ny? as ar by «|
un, X.Y,

= 4574 — [0.8791 1.7626]|


400
2388
= 13.2712
Mae) ye)P
Thus S = 4.4237
es
Notice that the n which occurs in the denominator of the variance formula,
Eq. (8-22), is the number of sample points. It is not the total number of
observations underlying the sample points. In this example, the latter number
is Ln, = 50, but n = 5. Finally, substitution in Eq. (8-19) gives
var(by) = s2(X’Q~!X) 7!
, 50a, 202) ae
= 4.4237]50) a
ie 0.057271 —0.009225
< 4.4231| — 0.009225 Perea
=| 0.2533 eae
—0.0408 0.0101
Thus var(b,*) = 0.2533
var(b,
«) = 0.0101
This example might have been treated equivalently by finding the T
matrix satisfying T’T = Q7'. Given Q~', the T matrix is simply
a
oe

ns,
Thus the data of Table 8-1 could have been recorded as

Xp 2¥12 7 3y68 ely esy0 oy te


¥ 4Vi2 S16" Sythe 29) On gy
and OLS applied to these five pairs of numbers.

A different variant of a cross-section study is one with replication of the Y


variable for given values of X. Suppose, for instance, that agronomists are
investigating the variation of crop yield in response to varying applications of
fertilizer. Let X,,..., X;,..., X,, denote the different fertilizer dosages chosen for
GENERALIZED LEAST SQUARES 297

the experiment. For dosage X,,n, plots are chosen, and Le iawn)
denotes the resultant set of n, yields: A model for the linear effect of fertilizer on
yield would then be specified as
ys Oe BX u,, Lesa of ole oy ert i (8-24)
Denoting the vector of disturbances in the ith application by u;, we make the
conventional assumptions
E(uj)=0 and E(uwi)=0671, i=1,...,m (8-25)
Thus Eg. (8-25) allows the disturbance variance to be different in different
applications, but assumes homoscedasticity and zero covariances within applica-
tions. However, an additional assumption is now required to cover the relations
between disturbances in different applications. We assume these to be uncorre-
lated, that is,
E(u.) = 0 Dare
ee loser: (8-26)
The complete model may now be written

Yi x, uy
¥
=| X
ial+
Ihe mM)
(8-27)
where
ee
X,=
nex,. ieee mn

ie
A more compact form of Eq. (8-27) is
y=XBp+u
where y’ = [y; y; °°: ¥,,], and so on. Assumptions (8-25) and (8-26) produce
a block-diagonal form for var(u), namely,

a I,,, 0 oss 0
var(u) = 0 vg ves 0 (8-28)
‘ neta: : a

Notice that each X, submatrix has only unit rank, since the same dosage is
applied to all plots within the group, but the X matrix has full column rank.
Model (8-27) is a special case of a more general model, which may be written

yi XxX, u, |
Se eS Efe (8-29)
Yin xXm um
298 ECONOMETRIC METHODS

where each X, is of order n, X k, the rows of X; are not required to be identical,


and the assumption is still made that var(u) has the block-diagonal structure given
in Eq. (8-28). For example, ¥,, might represent investment expenditure by firm /
in year j and the X’s the variables thought to influence investment expenditures.
The study would thus cover m different firms in various years, that is, a pooling of
time-series and cross-section data. Model (8-29) assumes a common set of
reaction coefficients B for all firms, but assumption (8-28) would allow the
disturbance variances to differ across firms.

Testing for Heteroscedasticity

Test for the equality of variances. In the case of replicated data, model (8-27), and
u, ~ NCO, GLa) a standard test for the equality of variances is available. The
hypothesis of homoscedasticity is
Ay: o¢=o0;=--- =O"
The test is conducted as follows.

1. For each class or group, compute the within-group sample variance

one ake i ile


i

where v; = n; — 1, and
n;
Y. = ie ii;
U Nn;

2. Then compute the pooled variance

Phe eePe pes


eT
5
Vv

where

va 2 (apa)
i=1
and the quantity

Q’ = vins* — = vIn s?
i=l
Under the null hypothesis Q’ will be approximately distributed as x*(m — 1).
However, the approximation will be improved by dividing Q’ by the scaling

+ The zero covariances incorporated in Eq. (8-28) may be an oversimplification for this model. See
Sec. 8-6, Sets of Equations, and Sec. 10-3, Pooling of Time-Series and Cross-Section Data.
+See M. G. Kendall and A. Stuart, The Advanced Theory of Statistics, vol. 2, Griffin, London,
1961, pp. 234-236.
GENERALIZED LEAST SQUARES 299

constant
] elem!
C=
ane y; |
to give

eee
2="C
. If Q > xho5(m — 1), say, then the null hypothesis would be rejected at the 5
percent level of significance.

Example 8-2 The data of Table 6-1 do not fit this test exactly since the X
variable is not constant within each class. However, we will ignore this
discrepancy and use the Y data of Tables 6-1 and 6-2 to illustrate this test.
We have pv; = 4 for each i and

p= y= 16 m =4
=|

Thus

C=14+ 75(1- 7g) = 1.108


3(3) 16
=
S52 alge os Sy5 == 2a)
23.5 535 = 51.5: Sy1 = 4.5:
Ins? = 3.0910 Ins? =3.1570 Ins?=3.9416 Ins? = 1.5041
Lyn s7 = 46.4868
2

eae ook = 25.3750


yln s* = 51.7402
Thus
2S T402) 2555790
g 1.1042 Sec
which exceeds

Kees (3) =a 781

and the hypothesis of homoscedasticity would be rejected on these numbers.


Let us repeat, however, that these calculations were for illustrative purposes
only since the within-group variability of X violates a basic assumption
underlying the test. A more appropriate procedure would be to replace the
within-group Y variances by the residual variances around the within-group
regressions of Y on X, that is, compute

gh (y, — X;b;)'(y; — X:b;) i=1,...,m


J Ni ik
300 ECONOMETRIC METHODS

where

b, = (X/X,) . 'X/y,

The Breusch-Pagan test.} Situations with replicated observations are rare in


practice. The more common situation is one of single observation points. The
Breusch-Pagan test for heteroscedasticity is a very general test in that it covers a
wide range of heteroscedastic situations, and it is also a very simple test in that it
is based on the OLS residuals. The model is
y=XBp+u
where the disturbances u, are assumed to be normally and independently distrib-
uted with variance
07 = h(z’a) (8-30)
where h(-) denotes some unspecified functional form, a is a p X 1 vector of
coefficients unrelated to B, and z, is a p X 1 vector of variables thought to
influence the heteroscedasticity. The first element in z; is taken to be unity. Thus
the null hypothesis
Hyp: @=a,=::: =a,=0
specifies homoscedasticity since then 67 = h(a,), which is constant over all i. The
other Z variables may consist partially, or even exclusively, of X variables, that is,
the heteroscedasticity may be governed by the explanatory variables in the
structural relationship. The test is a large sample or asymptotic test and is
conducted as follows.

1. Fit the OLS regression of y on X and obtain the vector e of OLS residuals.
2. From e compute
n 2
2 ee ete:

n
and the series

ee
Saris faile sen
6
3. Specify the variables in the vector z,. Notice that the functional form h(-) in
Eq. (8-30) does not have to be specified, merely the variables in the linear
combination z,a. Then fit the regression of g, on z’, and compute the
explained sum of squares (ESS) from the regression.
4. The quantity Q@ = ESS/2 is, under the null hypothesis, asymptotically distrib-
uted as x*(p — 1). Thus if Q > x{5(p — 1) one would reject the hypothesis
of homoscedasticity at the 5 percent level.

The Goldfeld-Quandt test. An especially simple and finite sample test, which is
applicable if it is thought that one of the X variables is the basic explanation of
7 T. S. Breusch and A. R. Pagan, “A Simple Test for Heteroscedasticity and Random Coefficient
Variation,’ Econometrica, vol. 47, 1979, pp. 1287-1294.
GENERALIZED LEAST SQUARES 301

heteroscedasticity, is the Goldfeld-Quandt test.} Suppose it is suspected that o, is


positively related to one of the X variables, say X,. The test procedure is then as
follows.

pd. Reorder the observations by the values of X,.


2. Omit c central observations.
3. Fit separate regressions by OLS to the first and last (n — c)/2 observations,
provided, of course, that (n — c)/2 exceeds the number of parameters to be
estimated.
4. Let RSS, and RSS, denote the residual sum of squares from the two
regressions, with the subscript 1 denoting that from the smaller X, values and
2 that from the larger X, values. Then
RSS,
DemRSS,
will, on the assumption of homoscedasticity, have the F distribution with
((n — c — 2k)/2,(n — c — 2k)/2) degrees of freedom. Under the alternative
hypothesis R will tend to be large. Thus if F > Fo,;, one would reject the
assumption of homoscedasticity at the 5 percent level.

The power of the test will depend, among other things, on the number of
central observations excluded, and will clearly be small if c is too large (so that
RSS, and RSS, have very few degrees of freedom) or too small (so any possible
contrast between RSS, and RSS, is reduced). A rough guide is to set c at
approximately n/3.£

The Glesjer test. None of the previous tests yields any specific estimate of the
form of heteroscedasticity which could then be inserted in var(u) to help derive
the GLS estimator. A test which helps in this direction is that due to Glesjer.§ It
is suggested, however, only for the case where a single variable Z is presumed to
determine the heteroscedasticity. The Z variable may, of course, be one of the
explanatory X variables in the structural relation. The test proceeds as follows.

1. Fit the OLS regression of y on X and derive the residual vector.e.


2. Regress the absolute value of the OLS residuals on Z", that is,
le,| = 6) + 6,Z7 + error (8-31)

+S. M. Goldfeld and R. E. Quandt, “Some Tests for Homoscedasticity,” Journal of the American
Statistical Association, vol. 60, 1965, pp. 539-547; or S. M. Goldfeld and R. E. Quandt, Nonlinear
Methods in Econometrics, North-Holland, Amsterdam, 1972, Chap. 3, for a more general discussion.
+See A. C. Harvey and G. D. A. Phillips, “A Comparison of the Power of Some Tests for
Heteroscedasticity in the General Linear Model,” Journal of Econometrics, vol. 2, 1973, p. 312.
§ H. Glesjer, “A New Test for Heteroscedasticity,” Journal of the American Statistical Association,
vol. 64, 1969, pp. 316-323.
q Alternatively one might use e? as the dependent variable.
302 ECONOMETRIC METHODS

As it stands, this relation is nonlinear in 6), 6,, and A. Glesjer suggests trying
regressions for a few specific values of h, such as 1, —1, +. The estimated slope
coefficient 5, is then used to test the hypothesis that 4, is zero, although the
conditions required for the validity of the usual significance test will not, in fact,
be satisfied by this regression.t Acceptance of Hy: 5, = 0 implies homoscedas-
ticity and its rejection, heteroscedasticity.

Estimation under Heteroscedasticity

1. Example 8-1 has illustrated a simple case of GLS estimation for grouped
data, where the Q matrix was known.
2. Another simple case occurs where one of the explanatory variables de-
termines the heteroscedasticity. This may have been determined by a Glesjer
type regression or postulated on a priori grounds. Suppose the heteroscedas-
ticity is modeled by
0, = 0° X;, Ca en al (8-32)

where X; is the explanatory variable thought to be the source of the hetero-


scedasticity. The variance matrix of the disturbance term then takes the form

Xj ) oar 0

VAT(
41) =O 1s Ogee N pas meonn
Osman O Xe

1
xee 0 0

T = 0 val
ies 0.

wage state cap ae 7c


0 0
X,
er

Thus the original relation


Y= By By Xe + BX Goal Pala Peo

would be transformed for estimation purposes to


y |~ Pilg
1 X, u
a} lx.) +a(3 Dea aaa oes
x} + Bo{|—t|+---+B4---4 ipa ee

(8-33)

7 See the discussion in S. M. Goldfeld and R. E. Quandt, Nonlinear Methods in Econometrics,


North-Holland, Amsterdam, 1972, pp. 92-94.
GENERALIZED LEAST SQUARES 303

and the inference procedures of Chap. 5 could then be validly applied to the
transformed variables in Eq. (8-33). Notice, however, that B; is estimated by
the intercept in the transformed relation, and the original intercept By is
estimated by the coefficient of 1/ X;. Equivalently, the diagonal matrix

Ql = diag ek ss
Lies eke
may be inserted along with the original y, X data in Eqs. (8-18) and (8-19).
3. In cases 1 and 2 we have assumed the elements in the Q matrix to be known
exactly. In many realistic cases these elements have to be estimated and the
estimates then substituted in the GLS formulas. This is sometimes referred to
as a two-stage Aitken estimator (2SAE), or as a feasible GLS (Aitken)
estimator. For example, in the case of replicated data the within-group,
sample Y variances could be estimated and substituted in Q. Or if a
Glesjer-type assumption postulated
0, = 6) + 6X,
and these parameters were estimated from
2 = 8) + 6,X, + error
the estimated disturbance variance matrix would be
var(u) = diag(S, + §,X,,6, + 8,X,,..., 6) + 8,X,) n

An unfortunate consequence of replacing unknown disturbance parameters


by estimated values is that we no longer have exact finite sample results for the
resultant estimators. In general the small sample properties of feasible GLS
estimators are unknown. There is, however, some evidence from Monte Carlo
studies on the relative performance of feasible GLS estimators compared with
OLS.+ Furthermore, under fairly general conditions the feasible GLS estimators
will have the same asymptotic distribution as the GLS estimators, so that the
conventional tests may be given an asymptotic justification.£
The importance of adjusting for heteroscedasticity depends on the extent of
the departure from homoscedasticity. No general pronouncements can be made,
but as a very simple illustration consider a two-variable model
Y,=a+t+ BX,+ u,

where X, takes on the values 1, 2, 3, 4,5. Let b denote the OLS estimator of 6 and
b, the GLS estimator and let us assume further that the nature of the hetero-
scedasticity is
0,Dig = 0°X; eee:

+ For a summary and detailed references see G. G. Judge, W. E. Griffiths, R. C. Hill, and T. G
Lee, The Theory and Practice of Econometrics, Wiley, New York, 1980, Chap. 4.
+ For the condition under which there is asymptotic equivalence see P. Schmidt, Boman.
Marcel Dekker, New York, 1976, Chap. 2, or H. Theil, Principles of Econometrics, Wiley, New York,
1971, Chap. 8. Unfortunately these conditions need to be checked out for each specific application.
304 ECONOMETRIC METHODS

A straightforward application of Eqs. (8-13) and (8-19) then gives


var(b,) _ 0.69
Tua Toa 0.56 (8-34a)

and if the form of the heteroscedasticity is


OS Ors
the corresponding result 1s

mene)
var(b)
=aie (8-34b)
Thus in this illustration, the efficiency of the OLS estimator ranges from 56 to 83
percent of the GLS estimator, depending on the postulated range for the
heteroscedasticity. Finally, we may note that in the heteroscedastic case, and in
other cases where GLS estimation is appropriate, there is no unique measure of
goodness of fit. A measure may be based on weighted sums of squares, using On!
(or V_') as a weighting matrix, or on sums of squares of the transformed vector
Ty, although the latter is inappropriate if the transformed relation does not have
an intercept term. For details of these and other measures the reader should
consult the article by Buse.

8-5 AUTOCORRELATION

Definitions
The autocorrelation, which is the focus of this section, is that of the {u,} series.
There may or may not be autocorrelation in the explanatory variables, but for the
moment we are only concerned with possible autocorrelation in the disturbance.
term. When present, it results in some or all of the off-diagonal terms in the var(u)
matrix being nonzero. This in turn destroys the optimal properties of OLS and
gives rise to another application of GLS.
We assume, as usual, zero mean for the series, that is,
E(u,) =0 for all ¢
The autocovariance at lag s is defined by
y,= E(u,u,,) s=0,+1, +2,... (8-35)
At zero lag we have simply the constant variance of the series

Yo = Eu? = of
The autocorrelation coefficient at lag s is defined by

oes seen De ee (8-36)

+ A. Buse, “Goodness of Fit in Generalized Least Squares Estimation,” The American Statistician,
vol. 27, 1973, pp. 106-108.
GENERALIZED LEAST SQUARES 305

We note that the y’s and p’s are symmetrical in 5 and have been assumed to be
independent of the ¢ subscript, that is, these coefficients are constant over time
and depend only on the length of lag s. The variance matrix for the disturbance
term may then be written as

Yo a Y2 haem ele
var(u) =|] ¥ Yo 7% Wes
Yn=1 Yn-2 nes Yo

l P| P, Pas
a a P) | P| se pt Pn—2 (8-37)

Pn—1 Pn —2 Pn—3 |

The above exposition has implicitly assumed a temporal, or time-series,


framework, but the same phenomenon may arise with cross-section data, where it
is often referred to as spatial autocorrelation. Suppose a sample of six observa-
tional units is represented by Fig. 8-1. The units might actually be contiguous as
in the case of adjoining states, or “nearness” might be defined in terms of some
other variable. For instance, if the sample units were households, the first
household might have an income close to the incomes of the second and third
households but not close to the incomes of the remaining households. If one
hypothesized that the disturbance for the ith unit was related to the disturbances
of contiguous units, then

u, = f(u, U3)
Uy = f(u,, U5, U4, Us)
and so on, and we would have some nonzero terms in the off-diagonal positions in
var(u).
Estimation of var(u) as in Eq. (8-37) from any finite sample is impossible
since the number of unknowns exceeds the number of observations, nor is any
relief afforded by increasing the number of observations, as it brings a concom-
itant increase in the number of unknowns. The practical procedure is to secure a
reduction in the number of unknown parameters by postulating some structure

a Figure 8-1 A sample of “contiguous” units.


306 ECONOMETRIC METHODS

Qs
A
1.0

UES) =

a I L 1 : Figure 8-2 Correlogram for AR(1) with


1 2 3 4 ¢ = 0.5.

for the disturbances. In time-series applications simplified structures are typically


autoregressive (AR) processes, moving average (MA) processes, or joint autore-
gressive, moving average (ARMA) processes. We have already had an example of
an autoregressive process in Eq. (8-7), namely,

U, a pu,_| ty e; || = ]

where the coefficient of the lagged term is denoted by ¢ as we now wish to use p
to denote an autocorrelation coefficient. This is a first-order AR(1) process, and
the result already established in Eq. (8-10) gives us the autocorrelation function
(ACF) of the process as
p,=¢ s=0,1,2,... (8-38)
Thus the autocorrelations decay exponentially and will oscillate in sign if ¢ is
negative. The graph of the autocorrelation function is called the correlogram, and
a typical correlogram for the AR(1) process, with positive @, is shown in Fig. 8-2.
The AR(2) process is defined as

u, = o\U,_, + $2U,_2 + &, (8-39)


In these and all other applications the {e,} series is always assumed to be
“well-behaved,” that is, E(e)= 0 and E(ee’) = 671. The condition |¢| < 1
ensured that the AR(1) process had a finite variance. The resultant {u,} series was
an example of a stationary process. In the second-order case the conditions for
stationarity aret

Ip.| < 1 ¢, +o, <1 ¢, — >, < 1


To establish the autocorrelation coefficients of the AR(2) process, multiply Eq.
(8-39) by u,_, and take expectations, giving

Xa Pivean © PoVee2 s>0


since e, has zero covariance with all previous u’s. Dividing through by the

+ A stationary process has a constant and finite mean and variance and a set of covariances which
are independent of time and are functions only of the lag length.
+G. E. P. Box and G. M. Jenkins, Time Series Analysis Forecasting and Control, revised edition,
Holden-Day, San Francisco, 1976, p. 58.
GENERALIZED LEAST SQUARES 307

variance y, of the series gives


Ps = PP, + $2 0,2 s>0 (8-40)
Setting s = 1 and using p) = 1 and p_, = p,, it follows that

Se
ca 8-4]
P| 1 = 5 ( )

Setting s = 2 and using Eq. (8-41) then gives

=o, > +
$1 >
Pr = ie (8-42 )

These first two autocorrelations, in conjunction with the recurrence relation


(8-40), will yield the higher order autocorrelations. The stationarity conditions
ensure that the autocorrelation function decays as the lag length increases. To
obtain the variance of the wu series, square both sides of Eq. (8-39) and take
expectations, giving

0,[1 oa 21920; | Say


Substituting for p, from Eq. (8-41) and simplifying gives the result

2 ie (1 cs >) ‘ o,.
i (8-43)
u
-@3]
-¢)°1
(1 + @)[( B

The general AR process of order p is defined as


Ble Oe Oot 5 ee Gy atte, (8-44)
where the {e,} series is well behaved, and suitable conditions are imposed on the
¢’s to ensure stationarity.
The general MA process of order gq is defined as
uae + Oe) 7 + 0,65 Fe: o> + Os (8-45)

An example of a fourth-order process was given in the context of a wage change


model in Eq. (8-3), and the corresponding variance matrix in Eq. (8-5) showed
that the autocorrelation coefficients are zero for all lags greater than the order of
the MA process. Consider the MA(1) process
u, =e, + Oe,_, (8-46)
It follows directly that
% = 6, = E(u?) = (1 + 67)?
and ile Eu) = 60,
esd Ying 0
so that Ce aAad?

and all further autocorrelation coefficients are zero.


A finite-order MA process may be converted into an infinite-order AR
process and vice versa. For instance, using the lag operation L, Eq. (8-46) may be
308 ECONOMETRIC METHODS

rewritten as
u,=(1+ OL )e,
giving (1+ OL) ‘u, =e,
or (1 — 6L + 67L? —---)u,
=e,
or = Ou), — Oreo Ou ee tre.
which is an infinite AR process with the restriction that the coefficients are given
by the successive powers of 6 with alternating signs. Similarly the AR(1) process
may be written
(1 — oL)u, = &,

or u,=(1-$L)'e,=(1+@L
+ @L? +---)e,
or Ue Er Dey + ge, 5+ se

which is an infinite MA process with the coefficients given by the successive


powers of ¢. The general ARMA( p, q) process is defined by
Uy ply te eee ae ey tea eee tee (8-47)
This may be written more compactly, using polynomials in the lag operator, as

o(L)u, am O(L)e, (8-48)

where o(L)=1 =o ob) =o EF


and O(L)=1+0,L+6,07+--+ 6,13
The mixed ARMA formulation enables quite complicated processes to be repre-
sented by a suitable choice of /ow-order polynomials.
Returning to the spatial example in Fig. 8-1, a common hypothesis about the:
structure of the autocorrelation ist
u; = pw; ju, ie: (8-49)
J
where

Va = a :
HE tess
and
es ] if units i, j are contiguous
? 0 otherwise
In this formulation p is a scalar indicating the overall strength of the autocorre-
lation, and the weights w,, are essentially dummy variables which allow any
disturbance to be affected by contiguous disturbances. The matrix formulation of

7 See, for example, R. L. Martin, “On Spatial Dependence, Bias and the Use of First Spatial
Differences in Regression Analysis,” Area, vol. 6, 1974, pp. 185-194.
GENERALIZED LEAST SQUARES 309

Eq. (8-49) is
= eWut+e (8-50)
and for Fig. 8-1 the W matrix would be

10g SF 0:0 0
pb 20 go
Pe ee Oba Ome
DSTA ie, Osada
Oe Omi ac)
Os Oe ea tw) m0
From Eq. (8-50) the variance matrix of the disturbance term is
var(u) = o2(I — pW) ‘(I — pW)’! (8-51)
In Eq. (8-51) W is generally a known matrix, but p and o? are unknown scalars.
However, we will not consider the resultant estimation problems here.+

Reasons for Autocorrelated Disturbances


Some possible reasons have already been mentioned in Sec. 8-1. A general source
of autocorrelated disturbances is the fact that the disturbance represents the net
influence of omitted explanatory variables. Economic theory cannot prescribe an
exhaustive list of explanatory variables to be included in a relation and, in any
case, data limitations often curtail the number of variables that can be included.
Exclusion of variables would not of itself impart autocorrelation to the dis-
turbance term unless the excluded variables were autocorrelated. Even then
autocorrelation in one explanatory variable might offset that in another. However,
economic variables tend to be nonrandom over time and also to move roughly in
phase so that excluded variables may impart autocorrelation to the disturbance
term. A second source of autocorrelation may be a misspecification of the form of
the relationship. Suppose the true relationship is represented by line A in Fig. 8-3,
and the linear function B is fitted to the data. The sample points will be scattered
around A and so the residuals from B will tend to be positive for X < X,
negative for X, < X < X,, and positive again for X > X,. Such a case might be
spotted simply from an inspection of the scatter diagram, and a transformation to
a quadratic or other nonlinear function would yield random disturbances. Even in
the case shown in Fig. 8-3 the X variable would need to be reordered
in increasing size for the correlation between adjacent residuals to show up. In
multiple regression situations misspecifications of the form of the equation cannot
usually be detected by graphical means, and some experimentation with different
functional forms may be required to judge whether any autocorrelation that
shows up may bea reflection of specification error.
+ For a discussion of estimation procedures see A. J. Cliff and J. K. Ord, Spatial Processes, Models
and Applications, Chapman and Hall, London, 1981.
310 ECONOMETRIC METHODS

|
|
|
|
|
| |
| L >~X
x; Xx) Figure 8-3

A third possible source of autocorrelated disturbances may be measurement


error in the dependent variable. Economic statisticians typically have various
formalized routines and procedures for estimating (or some say guesstimating)
economic magnitudes. The sequential publication of revised estimates is eloquent
testimony to the fact that the creators of the series believed their products to
contain some error, and indeed a series becomes definitive simply when the
statisticians stop revising it, which is not to say that it is then free of error. It is
unlikely that the estimating procedures produce errors which are random from
period to period and so, letting the y vector denote the observed Y values and y,
the true Y values generated by the mechanism XB + u, we have
yY=ys
+v= XB + (u+v)
where v is a vector of measurement errors. In the observed relationship the
disturbance term is u + v, which may exhibit autocorrelation through u or v or
both.

Consequences of Autocorrelation for OLS


The consequences of applying OLS toa relationship with autocorrelated dis-
turbances are qualitatively similar to those already derived for the heteroscedastic
case, namely, unbiased but inefficient estimation and invalid inference proce-
dures. It is possible to illustrate some of these factors quantitatively for certain
simple cases. Consider the model
iia Bx, i Uu, (8-52)
with uy =) pus, ee)
where
E(e)=0 and E(ee’) =o0/1
If OLS is applied to Eq. (8-52), we know from Eggs. (8-11) and (8-13) that
var(OLS b) = 02(X’X) 'X’QX(X’x) ' (8-53)
where
] p pe Ou 1

Q='\) p
Si
I Dian
: cer sie ah ae oe: : (8-54)
GENERALIZED LEAST SQUARES 311

and, as may be verified by multiplying out,

l =i 0 Me 0 0 0
—p 1+ 9° a) see 0 0 0

ees) te i et ne DE ete
0 0 0 —p l+p”? —-p
0 0 0 0 = ]
Substituting Eq. (8-54) in Eq. (8-53) for the model of Eq. (8-52) gives

var(OLS b)

== f2 eal n—-1
Dek oae =o
Sees Wake
ep tee =e
t=1%7 ai ae yy

(8-56)
If p were known and GLS was applied to Eq. (8-52), then substitution of Eq.
(8-55) in the general formula

var(b,) = 62(X’Q-'X)"'
gives

var(GLS b,) = O, ete: (8-57)


BP ee Pte
hy i/ ieee
Comparison of Eq. (8-57) with Eq. (8-56) shows that the efficiency of OLS is
measured by the ratio of the two terms in parentheses and thus depends not just
on p but also on the nature of the x variable. Let us suppose that x follows a
stable AR(1) scheme with parameter A. As the sample size n gets very large, the
terms involving x are approximated by the autocorrelation coefficients of x, which
are simply the successive powers of A. Thus the asymptotic efficiency of OLS is
given by
1 — 9p?
asy eff(OLS b) =
(1 + p? — 2pA)(1 + 2pdA + 2p?d? + ---)

ee) Cee) (8-58)


(1 + p? — 2pA)(1 + pA)
Table 8-2 shows illustrative values of this asymptotic efficiency for selected values
of p and A. Looking first at the right-hand side of the table, where p and X are
both positive, it is clear that p is the dominant parameter. Efficiency declines from
90 percent to about 10 percent as p rises from 0.2 to 0.9, with variations in A
having a relatively minor effect. The diagonal entries are equal to those in the first
row since if A = 0 or if p =A, the efficiency measure simplifies to (1 — p”)/
(1 + p*). Looking at the left-hand side of the table, the efficiencies are symmetri-
cal across the first row where the {x,} series is random. The remaining rows show
that A now exerts a much stronger effect and that the combination of a positive A
312 ECONOMETRIC METHODS

Table 8-2 Asymptotic efficiency of OLS 5 (percent)

r =a) 9) mal sG == 35) cle 0.2 0.5 0.8 0.9

0.0 10.5 22.0 60.0 925 92.3 60.0 22.0 10.5


0.2 12.6 25.4 63.2 92:9 903 58.4 19.8 9.1
0.5 18.5 34.4 71.4 94.6 93.5 60.0 18.4 7.9
0.8 8509) 56.2 85.4 97.5 96.6 71.4 22.0 8.4
0.9 52.8 71.8 92.0 98.7 98.1 81.3 29.3 10.5

and negative p can moderate the dramatic declines in efficiency shown in the
right-hand side of the table. These calculations are, of course, only illustrative, but
they indicate the possibility of a serious loss in efficiency if OLS is applied in the
context of autocorrelated disturbances.
A second problem with the application of OLS is that the conventional
formula on the computer for var(b) will, in this example, estimate o7/L7x?,
whereas Eq. (8-56) shows that this is no longer the true variance. As the sample
size gets very large, the ratio of the conventional formula to the true variance is
given by
Leak
eae Ok

Thus the proportionate bias that the conventional program will impart to the
estimation of the true sampling variance of the OLS 6 is, in the limit,

—20N
Asymptotic proportionate bias = rear (8-59)

Table 8-3 shows values of this statistic for selected values of p and A. Again it is
instructive to consider the table in two halves. A positively autocorrelated
disturbance in conjunction with a positively autocorrelated {x} series implies
underestimation of the sampling variance by the conventional OLS formula. If
p = A = 0.9, the estimated variance will only be about one-tenth of the correct
number, which would cause a serious overestimation of t statistics and significance
levels in conventional inference procedures. On the other hand, different signs for
p and A will cause an overestimation of the sampling variance. Casual empiricism

Table 8-3 Percentage bias in estimating var(b)

p
rN — 0.9 S05) a2) 0.2 0.5 0.9

0 0 0 0 0 0 0
0.2 43.9 22D 8.3 ale) SUS SNS
0.5 163.6 40.0 DoD, Sal Se — 40.0 SOD)
0.9 852.6 163.6 43.9 30:5 = 6271 SOO)
GENERALIZED LEAST SQUARES 313

indicates that the predominant situation in applied studies is a conjunction of


positive autocorrelation in both disturbance and explanatory variable so that
underestimation of var(b) is the more likely situation.
The comparison involved in Table 8-3 has implicitly taken o7 as known. In
practice it must be estimated from the sample data, and here again there is a
possibility of bias if the disturbance term is autocorrelated. We saw in Chap. 5
that
e’e = wu — u’X(X’X) 'X’u
Thus E(e’e) = E(wu) — E{u’X(X’X) 'X’u}
Now E(u'u) = no?
and

E{u’X(X’X)— 'X’u) =hEN{tr wX(X’X)'X’uJ) since the expression in


brackets is a scalar
E{tr[X(X’X) 'x’uu’]}
= 02tr{X(X’X) ‘X’Q)
= o2tr{(X’X) 'X’2X) (8-60)
For the simple model being analyzed here

Est= :x; X41 n—-2

E(e’e) = oh~ p+ 2p Sate e2 ty 2p Ly Xi 49 n

n n 2 n 2:
es 1x; ak, ee }

If p = 0, then E(e’e) = (n — 1)o2, which confirms the unbiased estimator s* =


e’e/(n — 1) of Chap. 5. If we assume the {x,} series to be a stable AR(1) process
with parameter A, then for large n

E(e’e)
,
= o;2, [xvee
Fete
aN

If p gue \ have the same sign, then s* will have a downward bias as an estimator
of [Link], for instance, p = 0.9 = A and n = 101,
E(s*) = 0.9150;
Thus when p and J have the same sign, this bias accentuates the bias analyzed in
Table 8-3. It is clear that autocorrelated disturbances are a potentially serious
problem, and it is very important to be able to test for their existence.

Tests for Autocorrelation


Suppose that in the model y = XB + u one suspects that the disturbance term
follows an AR(1) scheme, that is,
u, = ou,_; tives
314 ECONOMETRIC METHODS

The null hypothesis of zero autocorrelation would then be set up as


Hy): ¢=0
and the alternative hypothesis as
H,: ¢=#0
The null hypothesis is about the u’s, which are unobservable. One therefore looks
for a test of the null hypothesis using the vector of OLS residuals, e = y — Xb.
This raises several difficulties. We know from Chap. 5 that
e = Mu
where
M = I — X(X’X) 'X’
is symmetric, idempotent of rank n — k. Thus the variance-covariance matrix of
the e’s is \
E(ee’) = 62M

Thus even if the null hypothesis is true, so that E(uu’) = 071, the OLS residuals
will display some autocorrelation, for the off-diagonal terms in M do not vanish.
More importantly M is a function of the sample values of the explanatory
variables, so that it is impossible to derive an exact finite sample test on the e’s
which will be valid for any X matrix that might ever turn up.

Durbin-Watson test. These problems were treated in a pair of classic and path-
breaking articles.t The Durbin-Watson test statistic is computed from the vector
of OLS residuals e = y — Xf. It is denoted in the literature variously as d or DW
and is defined as

eee =. ext
a=
8-61
Biaaes (
Figure 8-4 indicates why d might be expected to measure the extent of first-order
autocorrelation. The mean residual is zero, so the residuals will be scattered
around the horizontal axis. If the e’s are positively autocorrelated, successive
values will tend to be close to each other, runs above and below the horizontal
axis will occur, and the first differences will tend to be numerically smaller than
the residuals themselves. Alternatively if the e’s have a first-order negative
correlation, there is a tendency for successive observations to be on opposite sides
of the horizontal axis, so that first differences tend to be numerically larger than
the residuals. Thus d will tend to be “small” for positively autocorrelated e’s and
“large” for negatively autocorrelated e’s. If the e’s are random, we have an
in-between situation with no tendency for runs above and below the axis or for
alternate swings across it, and d will take on an intermediate value.

7 J. Durbin and G. S. Watson, “Testing for Serial Correlation in Least Squares Regression,”
Biometrika, vol. 37, 1950, pp. 409-428; vol. 38, 1951, pp. 159-178.
GENERALIZED LEAST SQUARES 315

(a)

Figure 8-4 (a) Positive autocorrelation; (b) negative autocorrelation.

The Durbin-Watson statistic is closely related to the sample first-order


autocorrelation coefficient of the e’s. Expanding Eq. (8-61),
n 2 2
et wae; a5 De, a 20 Ceri
d
Leer
For large n the different ranges of summation in numerator and denominator
have a diminishing effect and
d= 2(1-r) (8-62)
where r = Ye,e, ,/Le; is the coefficient in the OLS regression of EROMse ee
Formula (8-62) shows heuristically that the range of d is from 0 to 4:

d < 2 for positive autocorrelation of the e’s


d > 2 for negative autocorrelation of the e’s
d = 2 for zero autocorrelation of the e’s
The hypothesis under test is, of course, about the properties of the unobserv-
able u’s, which will not be reproduced exactly by the OLS residuals, but the
above indicators are nonetheless valid in that d will tend to be less (greater) than
2 for positive (negative) autocorrelation of the u’s. For a random useries the
expected value of d is
2(k — 1)
E(d)=2+ (8-63)
ook
where k is the number of variables in the regression.
Because of the dependence of any computed d value on the associated X
matrix, exact critical values of d cannot be tabulated for all possible cases. Durbin
and Watson established upper (d,,) and lower (d, ) bounds for the critical values.
The tabulated bounds are to test the hypothesis of zero autocorrelation against
the alternative of positive first-order autocorrelation. The testing procedure is as
follows.
1. If d<d,, reject the hypothesis of nonautocorrelated u in favor of the
hypothesis of positive first-order autocorrelation.
316 ECONOMETRIC METHODS

2. If d > dy, do not reject the null hypothesis.


3. If d; < d < dy, the test is inconclusive.

If the sample value of d exceeds 2, we wish to test the null hypothesis against
the alternative hypothesis of negative first-order autocorrelation. The appropriate
procedure is to compute 4 — d and compare this statistic with the tabulated
values of d, and d, as if one were testing for positive autocorrelation. The
original DW tables covered sample sizes from 15 to 100, with 5 as the maximum
number of regressors. Savin and White have published extended tables for
6 <n < 200 and up to 10 regressors.} The 5 percent and 1 percent Savin-White
tables are reproduced in App. B-S.
There are two important qualifications to the use of the Durbin-Watson test.
First it is necessary to have included a constant term in the regression. Second, it
is strictly valid only for a nonstochastic X. Thus it is not applicable when a lagged
dependent variable appears among the regressors, and indeed it can be shown
that the combination of a lagged Y variable and a positively autocorrelated
disturbance term will bias the Durbin-Watson statistic upward and thus give
misleading indications.t Even when the conditions for the validity of the Durbin-
Watson test are satisfied, the inconclusive range is an awkward problem, espe-
cially as it becomes fairly large at low degrees of freedom. A conservative
practical procedure is to use d,, as if it were a conventional critical value and
simply reject the null hypothesis if d < d,. The consequences of accepting Hy
when autocorrelation is present are almost certainly more serious than the
consequences of incorrectly assuming it to be absent, which is one reason for the
procedure.§ Second, it has been shown that when the regressors are slowly
changing series, as many economic series are, the true critical value will be close
to the Durbin- Watson upper bound.
When the regression does not contain an intercept term, d is bounded by
dy Sas,
where d, is the upper bound of the conventional Durbin-Watson tables.
Farebrother has provided extensive tabulations of both lower and upper 1 percent
and 5 percent significance points for d,,.|

+N. E. Savin and K. J. White, “The Durbin-Watson Test for Serial Correlation with Extreme
Sample Sizes or Many Regressors,” Econometrica, vol. 45, 1977, pp. 1989-1996.
¢ M. Nerlove and K. F. Wallis, “Use of the Durbin-Watson Statistic in Inappropriate Situations,”
Econometrica, vol. 34, 1966, pp. 235-238.
§ A comprehensive Monte Carlo study relevant to this question is J. K. Peck, “The Estimation of a
Dynamic Equation Following a Preliminary Test for Autocorrelation,” Cowles Foundation Discussion
Paper, no. 404, September 9, 1975. After studying the properties of regression estimators following
different significance levels for d, the author recommends using a significance level much more
likely
(than the conventional levels) to reject Hy when it is true. This is in the same spirit as
using d,, as the
critical value.
| H. Theil and A. L. Nagar, “Testing the Independence of Regression Disturbances,”
Journal of
the American Statistical Association, vol. 56, 1961, pp. 793-806; and E. J. Hannan
and R. D. Terrell,
“Testing for Serial Correlation after Least Squares Regression,” Econometrica,
vol. 36, 1968, pp.
133-150.
|| R. W. Farebrother, “The Durbin-Watson Test for Serial Correlation when There
Is No Intercept
in the Regression,” Econometrica, vol. 48, 1980, pp. 1553-1563.
GENERALIZED LEAST SQUARES 317

The inconclusive range of the Durbin-Watson statistic can be narrowed if


explicit account can be taken of the form of any regressors in addition to the
constant term, since that reduces uncertainty about the X matrix. King has
presented tabulations of d, and d,, for three classes of linear regression models,
namely:

1. Regressions with a full set of quarterly seasonal dummy variables


2. Regressions with an intercept and a linear trend variable
3. Regressions with a full set of quarterly seasonal dummies and a linear trend
variablet
The Wallis test for fourth-order autocorrelation. Wallis has pointed out that
many applied studies employ quarterly data, and in such cases one might expect
to find fourth-order autocorrelation in the disturbance term. The appropriate
specification is then
U, = day + &, (8-64)
and the null hypothesis would be
Hy: %=0
To test the null hypothesis, Wallis proposes a modified Durbin-Watson statistic
2
ct eee, ey C4)
d, n 2
(8-65)
Dane,

where the e’s are the usual OLS residuals. Wallis derives upper and lower bounds
for d, under the assumption of a nonstochastic X matrix. The 5 percent
significance points are tabulated in App. B-6. The first table is for use with
regressions with an intercept, but without quarterly dummy variables. The second
table is for use with regressions incorporating quarterly dummies. As shown in
Chap. 6, one may employ a constant term and three quarterly dummies or use
four quarterly dummies without a constant term.
Further significance points at 0.5, 1.0, and 2.5 percent levels are provided by
Giles and King.§ The same authors also point out that if one is testing Hy) against
the alternative hypothesis H,: 9, < 0, the test statistic 4 — d, may be correctly
referred to the critical values 4 — d, ,, and 4 — d, ,, where d, ,, and d, , are the
5 percent values tabulated by Wallis, only in the case where seasonal dummies
have been included among the regressors. For the case where an intercept but no
seasonal dummies have been employed, these critical values are inappropriate and
the authors provide a revised set.

+™M. L. King, “The Durbin-Watson Test for Serial Correlation: Bounds for Regressions with
Trend and/or Seasonal Dummy Variables,” Econometrica, vol. 49, 1981, pp. 1571-1581.
+K. F. Wallis, “Testing for Fourth Order Autocorrelation in Quarterly Regression Equations,”
Econometrica, vol. 40, 1972, pp. 617-636.
§ D. E. A. Giles and M. L. King, “Fourth-Order Autocorrelation: Further Significance Points for
the Wallis Test,” Journal of Econometrics, vol. 8, 1978, pp. 255-259.
q M. L. King and D. E. A. Giles, “A Note on Wallis’ Bounds Test and Negative Autocorrelation,”
Econometrica, vol. 45, 1977, pp. 1023-1026.
318 ECONOMETRIC METHODS

Durbin tests for a regression containing lagged values of the dependent variable.
As has been pointed out, the Durbin-Watson test procedure was derived under
the assumption of a nonstochastic X matrix, which is violated by the presence of
lagged values of the dependent variable appearing among the explanatory vari-
ables. Durbin has derived a large sample (asymptotic) test for the more general
case.} Consider the relation

Yi aa BiY,_| phew Bi ves a B41 X eae ey Bax ip U, (8-66)


with
u,=ou,_,+e, and e~ N(0,o71)
The basic result is that under the null hypothesis, Hy: = 0, the statistic

h= Nica = AN(0, 1) (8-67)

where n= sample size


var(b,)= estimated sampling variance of the coefficient of Y,, in the OLS
regression of Eq. (8-66)
r=L72€,€;_|/L7a2e;_1, the estimate of from the OLS regression of €,
on e,_,, the e’s in turn being the residuals from the OLS regression of
Eq. (8-66)

Thus the test procedure is as follows.

1. Fit the OLS regression denoted by Eq. (8-66) and note var(b,).
2. From the residuals compute ror, alternatively, if the Durbin-Watson statistic
has been computed, we may use the approximation r =~ 1 — d/2.
3. Substitute in the formula for A, and if h > 1.645, reject the null hypothesis at
the 5 percent level of significance in favor of the hypothesis of a positive
first-order autocorrelation.
4. A similar one-sided test for negative autocorrelation can be carried out for
negative h.

The test breaks down if it should happen that n - var(b,) > 1. Durbin showed
that an asymptotically equivalent procedure is the following.

1. Estimate the OLS regression of Eq. (8-66) and obtain the residual e’s.
2. Estimate the OLS regression of
et one r= A yee ee, st

3. If the coefficient of e,_, in this regression is significantly different from


zero
by the usual OLS test, reject the null hypothesis Hy: ¢=0.

7 J. Durbin, “Testing for Serial Correlation in Least Squares Regression when


Some of the
Regressors are Lagged Dependent Variables,” Econometrica, vol. 38, 1970, pp. 410-421.
GENERALIZED LEAST SQUARES 319

Breusch-Godfrey test. The procedures considered so far test the significance of a


single autocorrelation coefficient. One might expect these tests to have reasonable
power in the presence of more general forms of autocorrelation. For instance, if

io Png t Dola gat ce $4; p + &,


the d test might well show ¢, to be significantly different from zero. But ?,
represents just a part of the autocorrelation now present, and one might find an
insignificant value for the first-order statistic. This, however, sheds no light on the
significance of $),..., ,. Thus a more general test is clearly desirable. Such a test
has been developed, apparently independently, by Breusch and by Godfrey.+
They postulate the usual model y = XB + u, where the X matrix may include
lagged values of the dependent variable. The null hypothesis is
H): u~ N(0, 021)
Two alternative hypotheses are considered. One is that the {u,} are generated by
an AR( p) process,
Cod Away SP tO Ue E, (8-68)
The other hypothesis is that the {u,} are generated by an MA( p) process,
UCP a Oe ey arco dy Oe (8-69)
where in each case {e,} is well-behaved. The test is based on the OLS residual
vector e, and is essentially a test of the joint significance of the first p autocorrela-
tions of these residuals. A remarkable feature of the test is that the same test
statistic applies for either alternative hypothesis. The components of the test
Statistic are
e = y — Xb, the usual n X 1 vector of OLS residuals
6? = e’e/n, the ML estimate of o?
0 0 0
e; 0 0
e> e| 0

E,= |e, e CAG Was? sa wae ye


ej
€2
eee : ee: Lea : oe

The test statistic is


= Bal
Pen, [E,E, — E,X(Xx’Xx) ‘XE, | Eve/67 (8-70)

+ L. G. Godfrey, “Testing Against General Autoregressive and Moving Average Error Models
when the Regressors Include Lagged Dependent Variables,’ Econometrica, vol. 46, 1978, pp.
1293-1302; and T. S. Breusch, “Testing for Autocorrelation in Dynamic Linear Models,” Australian
Economic Papers, vol. 17, 1978, pp. 334-355.
320 ECONOMETRIC METHODS

which, under the null hypothesis, is asymptotically distributed as y?(p). The


asymptotic properties of the test would not be affected if ¢? were replaced by the
usual unbiased estimator s* = e’e/(n — k). Significantly large values of / would
lead to the rejection of the null hypothesis, but would not indicate which of the
alternative hypotheses, Eq. (8-68) or Eq. (8-69), should be regarded as the more
appropriate. Godfrey shows that when the X matrix contains only exogenous
variables, the test statistic is asymptotically equivalent to
cas
te 2 tis 2 +--+ +7?)
2
where

Lae ~ Limit 1€ © 7-3

a 1&;
bee De ie

is the 7th autocorrelation coefficient of the OLS residuals.


The test statistic defined in Eq. (8-70) may seem to imply a burdensome
amount of computation, but it can be expressed in a simpler form. Suppose that
the OLS regression of y on X has been computed. The X matrix may contain
lagged values of the dependent variable. Denoting the residual vector from this
regression by e = y — Xb, it follows that
{e,} has zero mean
and eX = 0
Suppose now that e is regressed on the matrix [E, X] where E, is as defined
above. The R? from this regression is given by
Roe ESS
e’e
since no correction is required for the mean of the dependent variable. Moreover,
from Eq. (5-53),
E,E, E,X = ,

ESS=e'[E, xX] sc
E,

GD
Using e’X = 0, this simplifies to

ESS = e'E,|E,E, — E,X(X’X)'X'E,|


, , / y —ly, i
l
‘Ere
,

Thus l= Eas
62

and since 67 = e’e/n, we have


l=nR? | (8-71)
The test procedure is thus as follows.

1. Fit the OLS regression of y on X to obtain e.


2. Regresse on [E, X] and obtain R?.
3. Refer nR? to the x°(p) distribution and reject the null hypothesis if a
significantly large value is found.
GENERALIZED LEAST SQUARES 321

In practice the second step in the test procedure might be described as

2. Regress €, on €,_,,..., € t—p? and x, (that is, the ¢th row of X).

Since there are only n values of e available, this regression might be carried
out using only the last n — p observations. The Breusch-Godfrey procedure sets
€o, @_},--- at zero. Asymptotically it does not matter which route is taken, and it
is a moot point whether it matters in finite samples.

Estimation with Autocorrelated Disturbances


The Durbin test for a regression and the Breusch-Godfrey test for the presence of
autocorrelated errors are applicable when lagged values of the dependent variable
appear in the X matrix. In discussing estimation procedures we will restrict
consideration in this section to nonstochastic X matrices. Additional remarks on
estimation in the presence of lagged dependent variables will be made in Sec. 9-2.
Consider again the model
y=Xp+u
with E(u)=0 and E(u’) =07Q
From the discussion in Sec. 8-3 it is clear that GLS estimation may be achieved if
it is possible to find a transformation matrix T of known parameters such that
T’T = Q"' and then apply OLS to the transformed variables Ty and TX. As an
illustration, suppose we have just a two-variable regression where the disturbance
follows an AR(1) scheme, that is,
Y=a+bX,+u, t= 1,...,4 (8-72)
and U.= pu,_, + €,
with |p| < 1 and well-behaved e’s. The form of & for this model has already been
given in Eq. (8-11) and its inverse in Eq. (8-55) as
l —p 0 0 0 0

Sle etree aboca Oe


Dole 6 0 0 =O) lata op
0 0 0 0 =p 1
Consider first an (n — 1) X n transformation matrix T, defined by
=e. alana: 2-72 0 O
T, =| 0 =o ih
hOlLe 400 =the
Multiplication then shows that T,T, gives an n X n matrix which, apart from a
proportionality constant, is identical with Q~' except for the first element in the
leading diagonal, which is p” rather than unity. Now consider the n X n matrix T
obtained from T,, by adding a new first row with 1 — p* in the first position and
322 ECONOMETRIC METHODS

zeros elsewhere, that is,

1— OF) 50 Or 0)
T= —p ene 0 an0
0 —p 1 Ot, 50
ND Ace eae ae

Multiplication shows that T’T = (1 — p?)Q™~'. The difference between T, and T


lies only in the treatment of the first sample observation. Applying T, to Eq.
(8-72) gives the transformed model

Ty 0X, i x x a2
Y; — pY, eee a(1 — p) eA :
: =i fal X; — px. B tiles (8-73)

Me oi Pa
iL :
En

so that only n — 1 transformed observations are used in the OLS estimation. The
variables in Eq. (8-73) are sometimes referred to as quasi first differences, and
the intercept term being estimated is now a(1 — p). The variance matrix of the
disturbance term in Eq. (8-73) is an (n — 1) X (n — 1) matrix,

var(e) = (1 — p?)o7I
Application of T to Eq. (8-72) gives the transformed model

67 21 Vere
x2 —p
| 1 ji-pe jl-p-X, a é
ere
= en Nee la] + ‘i (8-74)
Sonat ere
ep jee : ie En

The /1 — p° factor is required to make the transformed disturbances homo-


scedastic. From Eq. (8-9),

G5. C,laeece)
for an AR(1) scheme. Thus var /1 —p- u,| =o.
If p were known, GLS estimation could be achieved by applying OLS to Eq.
(8-74), or the process could be approximated by using Eq. (8-73). The difference
between the two procedures can be important when the sample size is small. The
extensions to include additional explanatory variables and higher-order AR
processes are simple. Additional X’s are treated in exactly the same way as the
single explanatory variable in the example. If the disturbance term followed an
AR(2) scheme,

Uu, a U,_ | oe $yU,_> hge,


GENERALIZED LEAST SQUARES 323

the transformed variable would take the form

ty ae ee t=3,...,0
Special transformations would also be required for the first two observations.+
The assumption, however, of a known value for p is unrealistic. It is a
parameter to be estimated along with a, 8, and 07. Lagging Eq. (8-72) one period
and subtracting from Eq. (8-72) gives
Y,=a(l — p) + BX, — BpX,_, + pY,_, + «, f= Dome Ween (S275))
The disturbance in Eq. (8-75) satisfies the assumptions required for OLS. How-
ever,
Le? = f(a, B, p) (8-76)

which is a function of just three unknown parameters, while a straightforward


application of OLS to Eq. (8-75) will yield four estimated coefficients, namely,
ee
Deal =)
b,=8
Ds abe
b,
=p
where the circumflex denotes an estimate of the corresponding parameter. These
four equations will, in general, be inconsistent in that they do not yield a unique
set of estimates of a, B, and p. Thus a nonlinear constraint, b, = —b,b,, would
have to be imposed on the estimation process. To put the same point another
way, minimizing Le? with respect to a, B, and p gives equations which are
nonlinear in the parameters and thus cannot be solved analytically.
The basis of an iterative estimation process can be found by rewriting Eq.
(8-75) in two equivalent fashions, namely:

ela Pl,
ay OCIS p) + BCX, — pX,_1+)
&;
and
lat 4X) ipYep Oa Ag) te,
Starting with any value for p, the quasi first differences in the equation of step 1
could be computed, and OLS applied to it would then yield estimates of a and P.
These estimates in turn can be used to compute the Y, — a — BX, series. Regress-
ing this series on itself lagged one period in the equation of step 2 yields a revised
estimate of p, which can then be fed back into the equation of step 1, and the
process continues.
This is known as the Cochrane-Orcutt iterative process, and versions of it are
incorporated in almost all social science computer packages.t There is a variety of

+ For details see F. B. Lempers and T. Kloek, “On a Simple Transformation for Second-Order
Autocorrelated Disturbances in Regression Analysis,” Statistica Neerlandica, vol. 27, 1973, pp. 69-75.
+D. Cochrane and G. H. Orcutt, “Application of Least Squares Regressions to Relationships
Containing Autocorrelated Error Terms,” Journal of the American Statistical Association, vol. 44, 1949,
pp. 32-61.
324 ECONOMETRIC METHODS

starting positions and rules for termination. If the initial value of p is set at zero,
step | is then simply the OLS regression of Y, on X,, which yields the OLS
residuals e, = Y, — a — bX,. In step 2 e, is regressed on e,_,, without an intercept
term, to obtain an estimate r of the first-order autocorrelation coefficient. Alterna-
tively r may be computed from the Durbin-Watson statistic, which is a routine
output in an OLS package, as
r= 1—td
The estimated r is then used to compute the series {Y, — rY,_ ,) and {iy 7 Agen
which are used in a repeat of step 1. The process may be stopped any time the
Durbin-Watson statistic in step 1 indicates random residuals. This frequently
occurs after one complete iteration. Alternatively one can stop the process after
successive estimates of the parameters differ by less than some prescribed amount.
It is clear that step 1 in the Cochrane-Orcutt process is the use of model
(8-73) and the associated transformation matrix T,. A modification of the process
is to use model (8-74), where: the first term gets explicit treatment. This is often
referred to as the Prais-Winsten method.} The Prais-Winsten modification may be
expected to improve the efficiency of the estimation, especially in small sample
sizes. Yet another modification is to use a method suggested by Durbin for
obtaining the initial estimate of p.t This is to fit Eq. (8-75) by OLS without
worrying about the nonlinear restriction and take r as the coefficient of Sere
Monte Carlo study by Griliches and Rao suggests that a two-step estimator
consisting of the Durbin estimate of p followed by the Prais-Winsten treatment of
the transformed variables performs somewhat better than any of the other
variants over a fairly wide range of parameter values.§
The two-step Durbin estimator extends easily to more than one explanatory
variable and to higher order autoregressive schemes. For example, suppose the
model is
Yea By + Xa He Bex ee
with u, = >\U,_. + ou,_, + &,
Combining the two equations gives
¥e= $,¥,-, + &Y0 FBX, +: - + + BEX $B, Xo, 7
OUP Xie $2 By Xy4_9 Se $B, X,, 1-2 +1 — 6; - )B, + e;
Let $, and ¢, denote the coefficients of Y,_, and Y,_, when this regression
is '
fitted by OLS. The transformed variables are then computed as
(Ye OY 1oon eal (Xe ey ee
Lie 2s a es bo
and OLS is applied to these to obtain estimates of the B’s.

7S. J. Prais and C. B. Winsten, “Trend Estimators and Serial Correlati


on,” Cowles Commission
Discussion Paper, no. 383, Chicago, 1954.
+ J. Durbin, “Estimation of Parameters in Time Series Regression Models,”
Journal of the Royal
Statistical Society, ser. B, vol. 22, 1960, pp. 139-153.
§ Z. Griliches and P. Rao, “Small Sample Properties of Several Two Stage
Regression Methods in
the Context of Autocorrelated Errors,” Journal of the American Statistica
l Association, vol. 64, 1969,
pp. 253-272.
GENERALIZED LEAST SQUARES 325

The iterative procedures described above will, in general, converge to a


solution vector since at each stage one is minimizing a quadratic function in the
unknowns. There remains the question of whether one has reached a local or a
global minimum of the sum of squares function. This may be investigated by
performing a grid search over the permissible range of p values. For an AR(1)
scheme stability requires |p| < 1. Thus a grid of p values might be specified
ranging from —1 to +1 by increments of 0.1. For each value step 1 of the
Cochrane-Orcutt process is applied and the residual sum of squares computed.
The p value and the associated a and £ values with the minimum residual sums of
Squares are then chosen as the estimates. Or a finer grid may be imposed around
this p value and a further grid search carried out to achieve greater precision.
When some parameters of the variance matrix have to be estimated, we have
a further example of feasible GLS estimation. Thus our conventional test proce-
dures no longer have an exact finite sample justification but are only justified
asymptotically. The tests should be based on the final least-squares regression
computed in either a two-step or an iterative procedure, as the usual standard
error formulas will, in general, yield consistent estimates of the asymptotic errors.
The justification for this remark is provided at the end of the next section on ML
estimation.

Maximum Likelihood Estimation


A full ML procedure for a regression equation with an AR(1) disturbance has
recently been proposed, and the algorithm has been incorporated in some
computer packages.t The model considered is
Y=xXBp+u
with
u,=pu,,te, E(e)=0 E(ee’)=o071
The likelihood function is then
1 1
L(e) = pea SS “)
(2702) oe 20,
Using the transformation matrix T defined above, we have
Tu=e
where u and e are both n X 1 vectors. Changing variables in the likelihood
function gives

L(u) = L(e)| <°


where |de/du| indicates the absolute value of the determinant formed from the
matrix of partial derivatives of the e’s with respect to the u’s. In this case

7 |detT| = 1 — p*

+C. M. Beach and J. G. MacKinnon, “A Maximum Likelihood Procedure for Regression with
Autocorrelated Errors,” Econometrica, vol. 46, 1978, pp. 51-58.
+ See App. A-9, Change of Variables in Density Functions.
$26 ECONOMETRIC METHODS

and so

In L(u) x — “in o, + at —p*)-meee (8-77)


2 2 Doe
where & means is proportionate to.
Thus a procedure which minimizes e’e, even one that incorporates the
Prais-Winsten treatment of the first observation, is not full ML since it has not
taken account of the term in 1 — p’ in the likelihood function. Beach and
MacKinnon devised an iterative procedure for maximizing Eq. (8-77), which has
now been incorporated in White’s SHAZAM program and also in new versions of
the time-series processor (TSP).} They also conducted some sampling experiments
which suggest that their procedure may yield better estimates than conventional
procedures, such as the Cochrane-Orcutt process, and may also be computa-
tionally less expensive. Some further experiments conducted by Harvey and
McAvinchey compare the full ML procedure not just with the two-step
Cochrane-Orcutt process, which [Link] comparison in the Beach-MacKinnon
study, but also with the iterative Cochrane-Orcutt and with the two-step Prais-
Winsten procedures.t Their study uses the root-mean-square error (RMSE) of
estimators as the principle of comparison and confirms the results of the Beach-
MacKinnon experiments, but it also brings out a number of important points.

1. The two-step Prais-Winsten method is as efficient as full ML estimation for


the parameter values underlying their experiments.
2. The iterative Cochrane-Orcutt process is sometimes inferior to two-step
Cochrane-Orcutt, especially when the explanatory variable is basically a time
trend.
3. The two-step Cochrane-Orcutt process is in turn inferior to the two-step
Prais-Winsten method when the explanatory variable is trending.
4. OLS has RMSEs only about 3 to 4 percent in excess of full ML estimation °
with trending data, but its relative performance deteriorates when X is a
stationary random series.

A more recent study by Park and Mitchell confirms the main findings of
Harvey and McAvinchey and adds some additional findings.§

1. Their range of estimators includes an iterative version of Prais-Winsten, with


p estimated from the least-squares residuals, and they find this to be the best
of the feasible estimators.

7+K. J. White, “A General Computer Program for Econometric Methods—SHAZAM,”


Econometrica, vol. 46, 1978, pp. 239-240.
¢ A. C. Harvey and I. D. McAvinchey, “The Small Sample Efficiency of Two-Step
Estimates in
Regression Models with Autoregressive Disturbances,” Discussion Paper no. 78-10,
University of
British Columbia, April 1978. See also, A. C. Harvey, The Econometric Analysis
of Time Series, Wiley,
New York, 1981, pp. 196-199.
§R. E. Park and B. M. Mitchell, “Estimating the Autocorrelated Error Model with
Trended
Data,” Journal of Econometrics, vol. 13, 1980, pp. 185-201.
GENERALIZED LEAST SQUARES 327

2. They also investigate how well the various estimators perform in hypothesis
testing by looking at the number of type I errors in 1000 trials at the 0.05
significance level. The results are only reported for positively autocorrelated
disturbances, but the message is very clear. All estimators seriously under-
estimate standard errors, making estimated coefficients appear to be much
more significant than they actually are. This is, of course, to be expected for
OLS, but it is also fairly substantial for two-stage Prais-Winsten (2SPW),
iterative Prais-Winsten (ITERPW), and Beach-MacKinnon ML (BM). For a
sample size of 20, p = 0.8, and GNP as the trending explanatory variable, the
number of type I errors reported are OLS (449), 2SPW (251), ITERPW (246),
and BM (258). These numbers should be contrasted with an expected range
of 37 to 63. Thus it would be advisable to apply more stringent significance
levels than usual in testing coefficients in models with autocorrelated dis-
turbances.

Beach and MacKinnon have extended their ML approach to accommodate


an AR(2) process.t The treatment of relationships where the disturbance term
follows an MA or ARMA process is less well developed than the AR case.t
As an illustration of the derivation of the asymptotic errors for ML estima-
tion of a relationship with an AR(1) disturbance, consider the simple model

Y= at BX, + u, pata
ee

with up =) pu," + e; lp| <1

and e ~ N(0, 621)


We need to evaluate the information matrix. We are only concerned with
asymptotic results, and so the treatment of the first sample observation does not
matter since its effect becomes negligible as the sample size increases. Thus we
may write the model as

OP ee Ope oo (Yo BX Neto ey td Xn) 1 eee

The log likelihood function is then

AT = n—1
; In(277) —
n—-1
Ino? — ales be
7 20, t

The first-order partial derivatives with respect to the unknown parameters a, £, p,

+ C. M. Beach and J. G. MacKinnon, “Full Maximum Likelihood Estimation of Second-Order


Autoregressive Error Models,” Journal of Econometrics, vol. 7, 1978, pp. 187-198.
+ For an account of recent developments see A. C. Harvey, The Econometric Analysis of Time
Series, Wiley, New York, 1981, Chap. 6. See also A. C. Harvey and I. D. McAvinchey, “On the
Relative Efficiency of Various Estimators of Regression Models with Moving Average Disturbances,”
in E. G. Charatsis, Ed., Proceedings of the Econometric Society European Meeting, Athens, 1979,
North-Holland, Amsterdam, 1981, pp. 105-118.
328 ECONOMETRIC METHODS

and o, are
dln ES le
Le,
0a a,

diInL _ ]
Goel ee oX,_1)&,
Op €

dln L a ee
Lu; = 18;
dp a,
dln L Lk n—-1 1
do2€ 20, DGe
where all summations are over ¢ = 2,..., n. Setting these derivatives to zero and
solving for the parameters gives conditional ML estimators (conditional, that is,
on X, which is taken as fixed). We may note in passing that the first three
equations give \
L(Y, 6Y,,) = (n= la + BEX, 16X,5,)
(x "¥ 6Y,_,)(X, ae 6X,_,) a aX X, A 6X,_,) si pLex aq pXRa)s

p=
By Oe Be Oss BX iy)

Ye Sills Bx
which are the equations of the iterative Cochrane-Orcutt process, the first two
being the least-squares equations on the quasi first differences and the third the
first-order autoregressive coefficient of the estimated residuals.
Turning to the second-order partial derivatives

PnL_ _(n—1)(1-p)
da? 0
07 In L l 2
ad pre Pee)
071 DYE ESM ae
dp” Ce
d*In L Sele!
Se AT yy?
d(02) 20, 0,

The cross partial derivatives are

07 In L l=)
da Op Etoy ne Exp p Xe)

d7In L l
da dp care G2 ee ai (1 a p)Xu,_,)

Oana
Nieie ] ele De
0a do, 02
GENERALIZED LEAST SQUARES 329

d7In L 1
OB dp ga eet — Ue, — eX.)
eine: 1
dp aoe ae gael %,5 oxX,_;)

07 ln L 1
os
ede
Taking the negative of the expectations gives the information matrix

= oxy)
(n= 1)@= eyo © = p)ELY 0 0
ear ote erie x Li)" 0 0
0 (n= 1)o? 0
ee o2 0
0 0 fia
0
202€
The crucial feature of this information matrix is its block-diagonal nature.
Asymptotically the estimates of the regression parameters a and £ are distributed
independently of the estimate of the autocorrelation parameter p and of the
estimate of 0”. Referring back to Eq. (8-73), the data matrix for this regression is
given by the (m — 1) X 2 matrix X,, where
at l—p 1h". oes 1—p
mee PAS py GSS pay re OG a Xe
with unknown parameters a and B. The 2 X 2 submatrix in R is easily seen to be
X4,X,. Since the asymptotic variance matrix is given by R7' and since this has the
same diagonal form as R, we have
A
Qa
asyvar |=o (X-Xe).
B
which is consistently estimated by the usual least-squares procedures, justifying
the remark at the end of the previous section. We also see from R~! that
1—p°
asy var(6) =
no
remembering that o7 = o7/(1 — p’).

Prediction in the Presence of Autocorrelated Disturbances

If the model
Y,=a+ BX, +4, uU,= pu,.4 + &,
has been estimated from n sample observations, the best prediction of Y in period
n + 1 is no longer
x, staal =a+t bX, n+1
where a and bare estimates obtained by any of the above methods, since this
prediction sets the disturbance term at zero, and the AR(1) process implies
330 ECONOMETRIC METHODS

E(u,,,,) = eu, Both elements in pu, are unknown, but might be estimated by
ri, =r(Y, —a— bX,)
The suggested predictor is then
Y¥,,,=a+bX,,,+ 1,
n

This, in fact, would be a best linear unbiased predictor if » were known and rset
equal to p since it is the predictor that would emerge from the relation
Ye pYeie SP) HRC ee eee
which may be rewritten as

Y= BX pe re = BAe ae
giving the predictor.}
yen+1 = CDA eee Ue

8-6 SETS OF EQUATIONS


Sets of equations occur in various branches of economic theory. In the theory of
consumer behavior the decision maker faces a given money income M andaset of
prices P,,..., P,. The assumption of utility maximization leads to a set of demand
equations
O= fi(Pin Poe) Caer
where Q, indicates the optimal rate of consumption of the ith commodity. Theory
imposes various conditions on these demand equations. The assumption of a
specific form of utility function will impose yet further conditions. For example, if
one postulates an indirect addilog utility function

the ith demand equation ist


w= Da( 7]
* a M\"
eee

a,b, M*'P- i lee:


Sor BT oIED
Or Pee ee 7 (8-78)
Sea) ed j
where we have inserted a multiplicative disturbance term e*' to prepare the way
for empirical estimation.§ Expenditure Z, = P,Q, on the ith commodity is then
given by
a;b,M*'P-
et
= ee i= Teer (8-79)
v4;b,M”) Pi

7 For a detailed treatment of this topic see A. S. Goldberger, “Best Linear


Unbiased Prediction in
the Generalized Linear Regression Model,” Journal of the American
Statistical Association, vol. 57,
1962, pp. 369-375.
+ For this result and indeed for an elegant and lucid presentation
of the theory and measurement
of demand systems see L. Phlips, Applied Consumption Analysis, North-H
olland, Amsterdam, 1974.
§ The e in the disturbance term indicates the mathematical
constant e = 2.71828, and should not
be confused with the use of the same symbol for OLS residuals.
GENERALIZED LEAST SQUARES 331

The expenditure equation (8-79) is nonlinear in the a’s and b’s. However, if the
logarithm of the ratio Z,/Z 18 taken,
M M
ingZ in Z; = A; 4 bn| | - bl Fe + Uj, (8-80)

where A;, = In ite


J a,b,

and = tence,
This is clearly an estimable equation. Given r commodities, there are r(r — 1)/2
such equations, but most are redundant. As an illustration, for commodities i and
k we have
In Z, — In Z, = A,, + btn| | = bain|5] Tate (8-81)
P: P,
Subtracting Eq. (8-80) from Eq. (8-81) gives

In Z,~IZ, = Ay +n 5M -bain|5M J+
a Pi,
_ a,b;
where A, = or

and Os, Ey Ey
Thus of the three possible equations for commodities i, 7, and k only two are
independent. Given any pair of equations, the third follows by subtraction. For r
commodities there are just r — 1 independent equations, and for estimation
purposes one may select any set of r — 1 independent equations. Thus one might
write the system
M, \ M,
In-Z,,— In Z,, = Ay, + 0, In| —— | — dyin Pitti.
; Pi, 12

M, M,
; Pi, P3,

Ree en gl St ie Bites Maret ueN eager URINE A VaR ee (8-82)

lWZ. —InIn Z,, = A,, + 6)1 n Be


i bin P,
“ + Urry

The sample observations on the first equation in Eqs. (8-82) may then be
written as
y, = X,B, + u,
where

| | | ae €11 — a
©]
; M M 12 22
X,=|i in(
5 | In | B,=] 4, u, =

| f
332 ECONOMETRIC METHODS

Likewise, defining Y,, = In Z,, — In Z;,, the sample observations on the second
equation in Eqs. (8-82) may be written as
y. = X,B, + u,
where

oil muses
! z |M ae Eyn 889
X,=|i in|
> | In|5) Bp,=| 5, u,=
P Ps
b;

| | | Ein — ©3n

The complete model implied by Eqs. (8-82) is then


y, = XB, + u,
y, = XB, + u, (8-83)
@)\@) ef80, ce) Fonte ma: ia:\ie> elle

where m = r — 1. This set of equations may be written equivalently

y; X, B, u;
y xX B u
a oar ec \aolee (8-84)
OF aS

y=XB+u (8-85)
Because of the block-diagonal form of X the application of OLS to Eq. (8-85),
treated as a simple regression, would be exactly equivalent to the application of
OLS to each of the m equations in Eqs. (8-83) separately. However, the applica-
tion of OLS to Eq. (8-85) would not be optimal for two reasons. First of all the u
vector is not homoscedastic. From the structure of the u’s
U6) ie. (me eee
Thus var(u;) = var(e,) + var(e,,,) — 2cov(e;, €,,,)
Even if the original e’s are contemporaneously uncorrelated,
var(u;) = var(e,) + var(e,,)
,
and var(u,) = var(e,) + var(e;,)
Thus the u’s would only be homoscedastic if the e’s were homoscedastic,
but
there is no a priori reason for the disturbance variances in the various expendi
ture
equations to be equal.
A second reason for the nonoptimality of OLS is that the off-diagonal
terms
in var(u) will not be zero.
E(u,u;) x E(e, rh €41)(& Be a)

naE(¢7) aE eres) E(ee;,;) a E(e;41€41)


Even if the covariances of the e’s vanish, E (u,u;) does not vanish.
Thus var(u) is
not spherical, and GLS is the appropriate estimation procedure.
GENERALIZED LEAST SQUARES 333

These conditions on the disturbances may be embodied in the set of assump-


tions
Bua j= el i= 1,5,7
and Bluey) Sol ey, f= Use.
The first condition allows the disturbance variance to be different in the various
equations, but within each equation the assumptions of homoscedasticity and
zero covariances are still imposed. The second condition allows for nonzero
covariances between the disturbances in different equations and assumes that, for
any pair, the covariance is the same at each sample point; all lagged covariances,
however, are assumed to be zero. Collecting these variances and covariances in
the symmetric positive definite matrix

Suter ye ee iy
Z= |), 97m

Onl Gn2 Gnm

the variance matrix for the u vector in Eq. (8-85) may be written}
var(u) = V=Ze@I (8-86)
Thus a set of demand equations should almost certainly be considered as a group
and estimated by GLS because of the nature of the variance matrix of the
disturbance term. In addition, theoretical considerations will impose restrictions
across equations. For example, in the addilog demand system above the second
parameter, 5, in each B, vector is constrained to be equal across all m equations.
This constraint has not been imposed in the specification (8-84). Implementation
of that system would allow a different estimate of the coefficient b, to be made for
each commodity. One may wish to test for constancy of b, across commodities
and to reestimate the system with constancy imposed.+
A second illustration of sets of equations with cross-equation restrictions and
connections between the various disturbances is found in sets of “share” equa-
tions approximated by transcendental logarithmic functions, which has recently
become the dominant methodology in the estimation of various substitution
elasticities, especially in the field of energy economics.§ Consider a production
function
Q=f(X, X%,.-., X,)
where Q denotes the rate of output and X; (i = 1,..., 7) the rate of input of the
ith productive factor. If one assumes the firm to face a given set of factor prices

+ See Eqs. (4-76) and (4-77) for the definition of a Kronecker product and its inverse.
+ For an illustration of the estimation and testing of three different demand systems see R. W.
Parks, “Systems of Demand Equations: An Empirical Comparison of Alternative Functional Forms,”
Econometrica, vol. 37, 1969, pp. 629-650.
§ See, for instance, E. A. Hudson and D. W. Jorgenson, “U.S. Energy Policy and Economic
Growth,” Bell Journal of Economics, vol. 5, 1974, pp. 461-514; E. R. Berndt and D. O. Wood,
“Technology, Prices and the Derived Demand for Energy,” Review of Economics and Statistics, vol.
57, 1975, pp. 259-268; and J. M. Griffin and P. R. Gregory, “An Intercountry Translog Model of
Energy Substitution Responses,” American Economic Review, vol. 66, 1976, pp. 845-857.
334 ECONOMETRIC METHODS

P,,..., P,, one formulation of the firm’s decision problem is to choose the input
mix to minimize the cost of producing a given output Q. This gives rise to a set of
factor demand functions

X= fF(Pie sees
0), > sie
Denoting the optimal inputs by X*, the optimal (minimal) cost level is
c* = PAX =f (Pie On)

Differentiating C* with respect to the factor prices gives}

oC =r
OP.l

} This result is an application of Shephard’s lemma. (R. W. Shephard, Theory of Cost and
Production Functions, Princeton University Press, Princeton, NJ, 1970, p. 170.) The lemma may be
illustrated for a two-factor production function Q = f(X,, X2). Suppose the firm is required to
produce some stated output Q at minimum cost, given factor prices P, and P,. If we define

$= (PX, + PyX,) -A[ F(X, X) - Q]


where A is a Lagrange multiplier, we then seek the minimum of ¢. The first-order conditions
are
dg
SMS LTO

Beh
dg
=0 (1)
dp
pr ~f% X%) = 2'= 0
The solution of these equations gives the cost-minimizing factor demands
X* and X*, expressed as.
functions of P,, P;, and Q. The minimum achievable cost is then given by

Ge PX eh Xe (2)
Differentiating Eq. (2) partially with respect to P, gives

SOE ae OI ee
GPa TOP,
Thus Shephard’s lemma requires that

P, 0X; fe OXF ni
OP, OP,
and, similarly, that

OX* OXF ot
Prop, + 2 OP,

Differentiate the system (1) totally, setting dP, = dQ = 0. This


gives
Afi, dX, + Afi, dX + f,ddrd = dP,
A fo, AX, + fon AX, + fy dX =0
(3)
i dX, tio XG) =0
GENERALIZED LEAST SQUARES 335

Further
ac* P, aa P, X7*

OP Cre rcs
ES dlnC* = PX Lae
Cit? Gt oe
where S; denotes the cost share of the ith factor, that is, the proportion of total
cost absorbed by the ith factor. Since C* depends on the factor prices and output,
the cost shares will be functions of the same variables, that is,

Sete (Pitch Oe geht er


By estimating the parameters of the share equations one may be able to estimate
the parameters of the cost function. All depends on the functional form pos-
tulated for the cost function.
A currently favored specification is the transcendental logarithmic (or trans-
log) function. This is a very flexible form, capable of approximating a wide
variety of functional forms. As an illustration, the production function for the
industrial sector of an economy is often specified as

Q=f(K,L,
E, M)
where the inputs distinguished are capital K, labor L, energy E, and materials M.
Assuming constant returns to scale plus exogenous factor prices P,, P,, P;, and
Py, and imposing symmetry on the second-order partial derivatives, gives the
translog cost function
nC =a, + n@e a,inP, +a, in/P, + a,in Peay,in Py,

+ $Bxx(In Px) + Buz (In Px )(In P,) + Bxg(In Px )(In Py)


+ Bxag(In Py )(InPy) + $B, (In P,)” + By -(In P,)(In Pr)
+ Byy(In P,)(In Py) + 3Bee(In Pe) + Bey (In Pe)(In Py)
+ 3Byy (in grate

with the solution

dX at= aie aP,

]
dX, = Aoi dP,

where A is the determinant of the 3 x 3 matrix of coefficients on the left-hand side of Eq. (3). Thus
ox} OX a
P OP, + P, aP, = _ (Ahh aaP,f?)

= 0 from the first two equations in Eqs. (1).

+ See, for example, L. R. Christensen, D. W. Jorgenson, and L. J. Lau, “Transcendental


Logarithmic Production Frontiers,” Review of Economics and Statistics, vol. 55, 1973, pp. 28-45.
336 ECONOMETRIC METHODS

Differentiating InC with respect to the logs of the prices gives the cost share
equations
Sx = ay + By,in 2, + p,,in P, + Bp, Pe tbe eae
S, = a, + Bein Pe Bein P, + Bplt Pee eee
S; =o; + By,ln PP, +6, ,In P, + Bp, Prt Bay bay
Sy Oy. + Bey
ln Pe By 5,0, 5B py, eee ans ee
Since the shares must sum to unity,
Ay ta, ta; + ay = 1
and the B’s sum to zero in each column (and row). Imposing the rowwise B
constraints on the first three share equations gives the system
P P P
Se ae Bexln(5 . Brctn|5 |a Breln|5 |

P P P
S, =a, + Brctn|5 |+ Balm 5"|a Bretn|5 | (8-87)
Py Py Py

P P P.
S,; = a, + Beetn|5 + Bretn|5 |+ Beeln|5 |
Because of the symmetry in the 8’s there are just nine independent parameters in
this system. Estimation of these, in conjunction with the summation conditions on
the a’s and £’s, will yield estimates of all the coefficients of the cost function
except a.
For the translog cost function the Allen partial elasticities of substitution are
given by
B.,+ SS.
6, = i+]
J S,S;

ee ee
and se fa eles
S2

and the factor price elasticities by

115 = 8/5;
Since the four shares sum identically to unity, one must expect nonzero contem-
poraneous covariances between disturbances in different equations, and there is
also no a priori reason to expect the same disturbance variance in different share
equations. However, this system differs in one major aspect from the addilog
demand functions in Eqs. (8-82). In Eqs. (8-87) the same set of explanatory
variables appears in each share equation, but that is not true in Eqs. (8-82), and
we will return to the significance of this point below. At the next level
of
disaggregation a production function could be specified for the energy sector with
various specific fuels as inputs and the parameters estimated from a set of energy
cost share equations. There have been many applications of this cost share
approach in recent years. However, a word of caution is required. As the
GENERALIZED LEAST SQUARES 337

derivation made clear, a basic assumption underlying the derivation of the share
equations is that in each observation period in the sample there has been afull
and complete adjustment of the input mix to the factor prices ruling in that
period so that the minimum cost level C* is achieved. This is an implausible
assumption for many production processes, and actual cost shares probably
represent various Jagged adjustments to changing factor prices. The assumption of
instantaneous adjustment is likely to produce seriously biased estimates of the
various elasticities.

Feasible GLS Estimation


Returning now to the general set of equations set out in Eqs. (8-83) to (8-86), the
GLS estimator of B is

by = (X’V~'X) 'x’v-ly
From Eqs. (8-86)

where o'/ denotes the i, jth element in 2~'. Substituting for V~! in the formula
for b, gives

eg Xiy,
Nyx Ry’yX 3 Cee X ee el ein
BS alenee us: Rete WEE) oh een (8-88)
eX X ork X oF en
oxy:

and the associated variance matrix is

var(b,) = (X’V~!x)' (8-89)


The obvious operational difficulty with Eq. (8-88) is that = is unknown.
Zellner has proposed the construction of a feasible estimator as follows.+

1. Apply OLS separately to each equation in Eqs. (8-83), obtaining the vectors
of sample residuals e,,e,,..., €,, where

é= (T= x(x) Xylem


2. The diagonal elements o,; of 2 are estimated by
/

e7e;
i

+A. Zellner, “An Efficient Method of Estimating Seemingly Unrelated Regressions and Tests for
Aggregation Bias,” Journal of the American Statistical Association, vol. 57, 1962, pp. 348-368.
338 ECONOMETRIC METHODS

and the off-diagonal elements s,, by


> ee,
/

Caen) Cae ene


Se —_ a ——_

where k, denotes the number of columns in X;.f The denominator in these


estimates may alternatively be taken simply as n since the usual test proce-
dures will now only be valid asymptotically. Thus an estimated © matrix is
computed and substituted in Eq. (8-88) to give a feasible estimator. The usual
significance tests based on an estimated version of var(b,) now have an
asymptotic justification rather than small sample validity.

This estimator is often referred to as SURE (seemingly unrelated regression


equations) estimator after the title of Zellner’s original paper. This title is
something of a misnomer, since the most natural application of the technique is to
sets of equations which are indeed theoretically related, as in the two examples.
The gain in efficiency yielded by the Zellner estimator over OLS increases directly
with the correlation between disturbances from the different equations and
decreases as the correlation between the different sets of explanatory vari-
ables increases. Indeed the GLS estimator reduces to OLS if either (1) the o, j are
all zero or (2) the X; are identical.{ Even if the true correlation between equation
disturbances is zero, the sample OLS residuals may yield nonnegligible covari-
ances, and one might mistakenly compute GLS estimates. The result will be
estimates with somewhat greater standard errors than those of the OLS coefficients.
This will even be true for very small disturbance correlations, but as these
correlations increase, the efficiency of the GLS over the OLS estimates rises
substantially.§

Tests of Linear Restrictions


In order to see how to test a set of linear restrictions in the SURE model we must
extend the test developed in Chap. 5 under the OLS assumptions to fit the new
GLS assumptions. As we have seen in Sec. 8-3, the GLS estimator may be
obtained by applying OLS to the transformed equation

Ty = (TX)B + Tu
where
TT=V'!
Making the appropriate substitution of Ty for y and TX for X in the OLS test

7 In the two illustrative examples the X; matrices had an equal number of columns, but there is no
need to impose such a condition generally. The exposition also assumes an equal sample size in each
regression, but this is merely a simplification and need not be imposed generally.
+ See Problem 8-2.
§ J. Kmenta and R. F. Gilbert, “Small Sample Properties of Alternative Estimators of Seemingly
Unrelated Regressions,” Journal of the American Statistical Association, vol. 63, 1968, pp. 1180-1200.
GENERALIZED LEAST SQUARES 339

Statistic for Hj): RB = r, now gives the statistic

(r — Rb,)’[R(X’V-'x)
'R’] ‘(x — Rb,)/q
eV 'e/(n — k)
(8-90)
where b, is the GLS estimator, g is the number of restrictions embodied in the
null hypothesis, and e = y — Xb,. Under the null hypothesis this statistic follows
the F(q, n — k) distribution.
The SURE model specified in Eqs. (8-84) to (8-86) gives a special case of this
Statistic. There are m separate equations with n observations on each, giving
N = mn observations in all. There are k; variables in X,, and the estimation of the
unrestricted model, Eq. (8-84), will thus yield estimates of K = Lk, parameters.
Finally the V matrix has the special form shown in Eq. (8-86). Thus the test
Statistic becomes

(r — Rb,)’{R[X’(=~! @ DX] 'R) (r — Rb,)/g


e(=' @De/(N- K)
(8-91)
Finally the unknown & in Eq. (8-91) has to be replaced by 3%, containing the s;,,
defined above, and the test now has only an asymptotic justification.
As an illustration of the construction of the R matrix consider an addilog
system, Eqs. (8-82), with just four commodity groups and hence three estimated
equations. Application of the SURE technique to Eq. (8-84) will give the GLS
vector b,, containing nine estimated parameters, namely,

b, = [Ap 5 by Ai bP) b, Aig 5?) b,|’


Thus each of the equations gives an estimate of the b, parameter, as indicated by
the superscript. The null hypothesis is that the true value of this parameter is the
same in all three equations. This translates into a two-element constraint, namely,

bi? — 52 = 0
bi) — 62 =0
Thus R=
S _ o So aS So | = So

a
and

If the null hypothesis is not rejected and one wishes to reestimate the system
with the constraint imposed, one may take the formula for the restricted OLS
estimator, given in Eq. (6-5), and replace X by TX to obtain

Dex = by + (X’V~'X)'R[R(X’V-'X)7'R’] '(r— Rby) (8-92)


Equivalently one may arrange the columns of the data matrix in such a way that a
direct application of GLS gives an estimator obeying the constraints. For a
340 ECONOMETRIC METHODS

three-equation version of Eqs. (8-82) the arrangement would be

Aj)
Aj;
y) i 0-10 x) x, 0 0 Ais u,
Yo) = (0 7h Omex, 0 aX, 0 b, | +]u,] (8-93)
y3 0 Ou x, 0 0 Xl Oe u;
by
b,

where i denotes a column vector of n units and x, the n observations on


In(M/P,), and so on. Repeating the x, vector in a single column of the data
matrix ensures just a single estimate of its coefficient. The efficient estimation
procedure for Eq. (8-93) is still the Zellner SURE technique since the disturbance
variance matrix has the form of Eq. (8-86). There is a moot point whether the
estimates of the elements of 2% in the first stage of the technique should be
obtained from the application of OLS to each equation separately, as previously
described, or from the application of OLS to Eq. (8-93), which incorporates the
restrictions implied by the null hypothesis. The two procedures are equivalent
asymptotically.
Looking now at the cost share equations (8-87), unrestricted estimation of the
equations would be achieved by OLS, even with a nonspherical disturbance
matrix, since the matrix of explanatory variables is identical in each equation.
However, the test of the symmetry restrictions still requires the computation of
the test statistic (8-91): the GLS b, in that formula is now replaced by the OLS b,
but the elements of the 2 matrix must be estimated from the OLS residuals as
before. The specification of the appropriate R matrix and r vector for the test of
the symmetry conditions is left as an exercise for the reader.j Equations (8-87)
already embody summation restrictions on both a’s and B’s. One may wish to
test these restrictions before looking at Eqs. (8-87), or one may wish to test the
complete set of summation and symmetry conditions. Finally if one wishes to
estimate Eqs. (8-87) with the symmetric restrictions imposed, the Zellner SURE
technique is again required, as in the addilog example. The details are left as an
exercise.
The choice of which of the four share equations to drop in obtaining the set
of three equations (8-87) is an arbitrary one. The SURE estimates are not
invariant to the choice of equation to drop. However, iteration of the SURE
technique will produce parameter estimates that converge to the ML parameter
estimates, which are unique and independent of the equation omitted.§

7 See Problem 8-3.


+ See Problem 8-4.
§ See the very useful discussion and references to other relevant papers in E. R. Berndt and L. R.
Christensen, “The Translog Function and the Substitution of Equipment, Structures and Labor in
U.S. Manufacturing, 1929-68,” Journal of Econometrics, vol. 1, 1973, pp. 81-113.
GENERALIZED LEAST SQUARES 341

The iterative process goes as follows.

1. Compute the s,; from the OLS residuals, as described above, and hence
obtain 2.
2. Compute the elements of =~! and substitute in Eq. (8-88) to compute b,.
Wo. Using by compute a new set of residuals e, = y — Xb,.
4. Partition e, into the subvectors corresponding to each equation and use these
subvectors to compute new s,> thus starting the process over again.

The SURE process may be further complicated to allow for autocorrelation


in the disturbance terms, but we will not pursue that topic here.

PROBLEMS

8-1 Derive the results on the efficiency of the OLS estimator under the two forms of heteroscedastic-
ity, given in Eqs. (8-34a) and (8-34d).
8-2 Prove that the SURE estimator in Eqs. (8-88) reduces to the application of OLS to each equation
separately if
(a) 9,;; = 0 for all i + j
or
(b) Xj = X,=-:-- =X,,

8-3 Specify the R matrix and r vector for testing the symmetry conditions in the set of equations
(8-87).
8-4 Consider the four cost share equations prior to Eqs. (8-87) and explain how to test the full set of
summation restrictions (on a’s and B’s) and symmetry conditions. Which, if any, of these restrictions
might be satisfied exactly by the estimated coefficients?
8-5 Consider a heteroscedastic model (for which all other classical assumptions hold)
Y,,=a+ BX, + uj; b= Nose (m
> 1)

a= cen t (1; > 2)


Suppose var(u;) = 07. A sample estimator of o? is

Geereile),
; Namal
where
ny
ad ey;
n.!

Determine E(s7).
(University of Michigan, 1981)

+ See R. W. Parks, “Efficient Estimation of a System of Regression Equations when Disturbances


Are Both Serially and Contemporaneously Correlated,” Journal of the American Statistical Association,
vol. 62, 1967, pp. 500-509, for a treatment of first-order serial correlation; and see G. G. Judge, W. E.
Griffiths, R. C. Hill, and T. C. Lee, The Theory and Practice of Econometrics, Wiley, New York, 1980,
Chap. 6, for more general cases.
342 ECONOMETRIC METHODS

8-6 In the model


Vip = AX,, + Uy,

Yar = Bx2, + Ur,


the x;, are nonrandom exogenous variables and the u,, are serially independent random disturbances
that are normally distributed with zero means and second moments

E(uZ,)=1 E(u3,)=2 — E(uyua,) = 1


for all values of ¢t. The sample second moment matrix below was calculated from 20 sample
observations:

ie) 2 eee

yi 10 -1 1-1

(a) Find the best linear unbiased estimates of the parameters a and B.
(6) Test the null hypothesis
Hy: a=B8
against the alternative H;: a = B.
(University of Michigan, 1980)
CHAPTER

NINE
LAGGED VARIABLES

We will use the term “lagged variables” to cover the inclusion on the right-hand
side of the regression equation of lagged-values of the explanatory variables, the
X’s, and/or lagged values of the dependent variable Y.

9-1 SOURCES OF LAGGED VARIABLES

Realistic formulations of economic relations often require the insertion of lagged


values of the explanatory variables. For instance, a rise in “permanent” income is
likely to have an effect on consumption, which is distributed over a number of
time periods, or a change in investment allowances may be expected to result in
changed investment allocations, which, in turn, will have an effect on actual
investment spending spread over a number of time periods because of production
and other lags.
In general let us suppose that a causal variable X, exerts a distributed lag
effect on Y as follows:
Period t ital bree £3

X,

J lege i

Effect on Y 5) X, 5,X, 6,X, 0;.X,


343
344 ECONOMETRIC METHODS

Assuming the lag pattern to persist through time, any Y, is seen to be built up as
the sum of effects from current and previous values of X. Thus the lagged effect
assumed above would generate the relation
Y, =p Ts 8yX,at: 5, X,_, + 6, X,_> i 6; X,_3 atUu,

where we have also allowed for an intercept and a disturbance term. In practice
one does not usually have any strong a priori information about the maximum
length of lag, and one formulates the general relation
Y= e+ D(L)X, +4, (9-1)
where D(L) is a polynomial of some degree s in the lag operator, that is,
D(L)=6,+6,L+---+6L' (9-2)
If X has remained constant at some level X for s periods, then, apart from
disturbances, Y will have reached an equilibrium value
Yop
+ D)X
where D(1) indicates the value of the polynomial when L is replaced by unity, and
is simply the sum of the 6, coefficients, namely,

D(1) = y6;
i=0
If X changes in period ¢ by an amount AX, and is then held constant at the new
level, Y will gradually adjust from Y to a new equilibrium. The changes are
Period t Bote Lay eh
Change in Y by AX, 5,AX, 6,AX,
The coefficient 6) (= AY,/AX,) thus represents the impact multiplier for X.
Partial sums of the 6’s indicate intermediate multipliers. The 6’s may also be
standardized by dividing by their sum D(1). Partial sums of the standardized 8’s
then indicate the proportion of the total effect achieved by a certain period. For
example, knowledge of the 6’s enables one to estimate how many periods must
elapse before, say, 90 percent of the total effect is achieved. An important concept
is that of the median lag, which is the number of periods required for 50 percent
of the total effect to be achieved. When all the 6’s are positive, another useful
Statistic is the mean lag defined as
ya 0t0y Oy ees a eas
so) 89 $0, FO, +> + 8
From Eq. (9-2) it is seen that differentiating D(L) with respect to L gives
DL )= 64 20, b--2-= $56,124
Thus Mean lag = DAY)
D(1)
As an illustration suppose an estimated version of Eq. (9-1) yields
D(L) = 0.10 + 0.25L + 0.35L? + 0.15L3 + 0.0514
LAGGED VARIABLES 345

Then D(1) = 0.90


so that the final or total effect of a unit change in X is a change of 0.9 in Y. Also
D’(1) = 0.25 + 0.70 + 0.45 + 0.20 = 1.60
and the mean lag is computed as 1.6/0.9 = 1.78 periods. The standardized
coefficients and their cumulated values are as follows:

Period 0 l 2 3 4

Standardized coefficients 0.11 0.28 0.39 0.17 0.05


Cumulated values 0.11 0.39 0.78 0.95 1.00

The median lag would be computed by interpolation as


0:50 =-0.39 :
1+ 0.78 — 0.39 =i 1228 periods

In practice the maximum lag s may have to be fairly large to provide an


adequate representation of the relationship between Y and X. It is frequently
possible to achieve a more parsimonious representation (that is, using a smaller
number of parameters) by postulating a distributed lag on both Y, and X, as in
ALL )(Y, Sp) = BCL) Xo a, (9-3)
where ACE) il =a re Ll (9-4a)
BCE) Base pled ae Be (9-45)
and it is expected that p + q will be less than s. The stability of Eq. (9-3) imposes
conditions on the a’s, which may be expressed in the form that the roots of A(L)
lie outside the unit circle.f Relation (9-3) may be rewritten as
B(L)
Mp ofh X, + u, (9-5)
A(L)
where we are assuming the disturbances to be related by v, = A(L)u,. Compari-
son of Eqs. (9-1) and (9-5) gives
B(L) _7 PY)
FE) (9-6)
,
As an example, suppose that
A(L)=1-—a@L (9-7a)
and B(L)=8,+ BL (9-7b)
Thent

Be) Ee dC) +'(


A(L) sae By + B,
0 10 l a,(0,B
+)L l 10o + By)
1 L* + 0?(0,8
1 10) + BL?
l +-::

+See C. E. P. Box and G. M. Jenkins, Time Series Analysis: Forecasting and Control, tevised
edition, Holden-Day, San Francisco, 1976, pp. 53-54.
+Note that(1 — a@,L)~'=1+a,L+a7Ll?4+---.
346 ECONOMETRIC METHODS

and the correspondence between the a’s, 8’s, and 6’s is


5p ad Bo
5, = a By + B, (9-8)
6, = a6,
6, = a6,
and so on. Two first-order polynomials A(L) and B(L) thus generate an infinite
polynomial for D(L), but there are, of course, implied restrictions on the 6’s. As
shown by Eq. (9-8), the first two 5’s are “free” and subsequent 6’s decline
exponentially. They decline since the stability requirement that the root of
A(L)=1-a,L=0
lie outside the unit circle ensures that a, (= 1/L) has modulus /ess than unity.
The same condition ensures that the infinite sum of the 6’s converges. From Eq.
(9-8) this sum is seen to be

D(1) = fy + “ufoA
_ Bot Bi
|e a,

B(1)

Clearly, extending the power of the B(L) polynomial would extend the number
of “free” 5 coefficients before the exponential decline sets in. The mean lag may
also be derived from the A(L), B(L) polynomials. Since D(L) = B(L)/A(L),
DULY BCE) ACE)
DCB) CBRE) Sea)
and so

Mean lag = BAD ad) (9-9)


BQ) A(1)
This expression is often easier to compute than the equivalent expression in terms
of the 6’s. For the above example
B, ihe Sees a, Bo
+ B,
Mean lag =
Bo +B, 1—a (1—a)(By + B)

The Koyck Scheme


We have seen that, starting with an equation like Eq. (9-1), which involves only
lagged X values on the right-hand side, one may be led by considerations of
parsimonious parameterization to reformulate it as in Eq. (9-3), which introduces
lagged values of the dependent variable Y among the regressors. When the A(L)
polynomial is just of the first degree, as in Eq. (9-7a), we have an example of a
LAGGED VARIABLES 347

Koyck scheme of declining exponential weights.f The simple Koyck scheme has
the coefficients on the X’s declining exponentially from the start, that is,
6, = a,6,_, i i cee
This corresponds to the specification
A(L)=1-a@L and B(L)=82,
and the relationship may be formulated as
¥, = p+ OX, + a6) X,_, + a76)X,_. +--+ + u, (9-10)
or, equivalently, as

(1 — a, L)(¥, — #) = BX, + ©,
which may be written
Vo Gl Oy) chee
19 ataby Aaa (9-11)
Equivalence between Eqs. (9-10) and (9-11) requires

By = 85
and v, = (1 — a,L)u, =u, - a,u,_, (9-12)
Thus if the original disturbances {u,} in Eq. (9-10) are serially independent, the
transformed disturbances {v,} in Eq. (9-11) are serially dependent, which has
implications for the estimation procedures to be considered in Sec. 9-2. Since
A(1)=1—a,, A(1) = —a,, B(1) = By, and B’(1) = 0, the mean lag for the
simple Koyck process is a,/(1 — a@,). As has already been indicated, raising
the degree of the B(L) polynomial, while retaining A(L) = 1 — a,L, increases
the number of “free” coefficients before the Koyck exponential decline comes
into play.
So far we have considered the distributed lag effect of just a single explana-
tory variable. Suppose there are two explanatory variables, each with a Koyck lag.
There may be no a priori reason to expect an identical decay parameter in each
lag. Thus the relation might be formulated as
Ye dt BX, Pa BX eo pkey ct oe yz,
+ 05yZ,_, + a3yZ,_, +++ +4, (9-13)

or Ln peereet ee
which gives
YAS pet (a, ta, )Y, we, ¥_5 + BX, 56%, + ¥Z,— o,yZ2,.5
(9-14)
where u* = w(1 — a )(1 — @,)
and OR Uy ioe (a, aH Oy) U,_| + A ,U,_4

+ L. M. Koyck, Distributed Lags and Investment Analysis, North-Holland, Amsterdam, 1954,


348 ECONOMETRIC METHODS

so that, compared with the single-variable Koyck scheme in Eq. (9-11), we have
two lagged values of Y and lagged values of each explanatory variable. For
estimation purposes the essential point to notice about Koyck schemes is that
they may be formulated either with only lagged values of explanatory variables on
the right-hand side, as in Eqs. (9-10) and (9-13), or with lagged Ys appearing on
the right-hand side, as in Eqs. (9-11) and (9-14). The former have nonlinear
restrictions on the parameters combined with presumably “well-behaved” dis-
turbance terms, while the latter have a dramatic reduction in the number of
right-hand side variables, but “complicated” disturbance terms and sometimes
restrictions on the coefficients [as in Eq. (9-14) but not in Eq. (9-11)].

Adaptive Expectations
Lagged dependent variables may also appear among the regressors in various
expectational models. A firm may base its production rate Y, not on the current
sales rate X,, but on the expected, permanent, or trend sales rate X*. Thus one
may specify
Y,=a+t BX* + u, (9-15)
where a disturbance u, has been included to allow accidental over- or under-
achievement of the production target. Equation (9-15) is not usually statistically
operational since there is a dearth of published information on expected or
forecast sales rates and similar variables. It is therefore customary to add an
auxiliary hypothesis about the formation of expectations, and one of the most
widely used (if not, indeed, abused) schemes is that of adaptive expectations,
which is that expectations get updated each period on the basis of the latest
information about the actual value of the variable. The formal specification is
Xe Xe = (1 A) CX, NE ee 0 er el (9-16)
In this formulation X* indicates the expectation formed at the end of period ¢,
when the information about the current level X, has become available. If expec-
tations were formed at the beginning of the period, X, in Eq. (9-16) should be
replaced by X,_,. If A = 0 in Eq. (9-16), the expected value adjusts period by
period to the current observation and all previous history is irrelevant. If A = 1,
an expectation, once formed, continues unchanged, irrespective of current or
earlier observations. The intermediate and more realistic case of A being a positive
fraction means that expectations get adjusted each period by some proportion of
the discrepancy between the latest observation and the expectation for that
period. Low values of A imply substantial adjustments in expectations, and large
values imply slowly changing expectations.
Equation (9-16) may be reformulated as
(TL WA "(les Ai)X
aN
or P,bs a To AL (9-17)

This in turn may be written


X* = (1—A)X,+ A(1-A)X_, + MU = A) Fe:
LAGGED VARIABLES 349

so that the adaptive expectations hypothesis gives the current expectation as a


Koyck-weighted combination of the current and all previously observed values of
the variable in question. Substitution of Eq. (9-17) in Eq. (9-15) then gives
eM Et] en
ath
Or

Y, 7% a(] ia A) RY wy BC a A) X, ay (u, a Au,_;) (9-18)


This equation is formally identical to the simple Koyck scheme in Eq. (9-11) in
terms of the variables included, the MA(1) disturbance process, and the fact that
the parameter of the MA(1) process is also the coefficient of the lagged dependent
variable.

Partial Adjustment
Another process which can generate lagged dependent variables among the
regressors is that of partial adjustment. Consider the adjustment of gasoline
consumption to a substantial price rise such as that engineered by OPEC in
1973/1974. Initially the scope for economies in consumption, even in the face of
very substantial price rises, was limited by such factors as

1. The existing geographical distribution of residences and work places


2. The existing stock of vehicles
3. The existing supply of alternative transport systems

In the short run, economies could be made in shopping and vacation trips,
car pooling on work trips, and so forth. In the longer run, one expects adjust-
ments in the more fundamental factors, such as the fuel efficiency of the vehicle
fleet. Such adjustment has its own costs and, in any case, must take time to be
achieved. Thus one may postulate Y*, the optimal consumption rate appropriate
to a gasoline price of X,, with income and other factors being held constant, as
*-a + BX,
t (9-19)
For reasons such as those suggested one would not expect actual consumption Y,
to adjust completely to X, in period ¢. Instead, a partial adjustment process is
frequently specified as

WoO, Ga Are = Foye uy! 065A 21 (9-20)


Notice that no disturbance term has been inserted in the calculation of the
optimal Y*, but it would seem essential to include one in the specification of the
actual Y, in Eq. (9-20). An alternative form of Eq. (9-20) is
(1 —AL)Y, = (1 —A)¥* + u, (9-21)
which, in turn, gives

We ear N (LN) Ve aN Uae A) Yds (uy Aue,)


350 ECONOMETRIC METHODS

so that the current consumption rate is a Koyck-weighted combination of current


and all previously desired rates. Substitution of Eq. (9-21) into Eq. (9-19) gives
(1 -—AL)Y,=a(1 —A) + BUI — A)X, + u,
or Y,=a(l =A) +AXY,_, + BOL HA)A, $4, (9-22)
Again notice the formal equivalence of Eq. (9-22) to the adaptive expectations
equation (9-18) and the Koyck-weighted lag scheme in Eq. (9-11). The only
difference is that in Eq. (9-22) the disturbance term may have simpler properties
than in the other two cases.
Demand functions are frequently specified in constant elasticity form. Thus
Eq. (9-19) could be respecified as
Yy* = Ax? (9-23)

ii)-(Fe)"~
The partial adjustment process would then have to be specified conformably as
Y y* =I

a ts

Combining the two relations produces


y, a Ae Ya kee en

which, using lowercase letters to denote natural logarithms, gives


y= a(l =A) +Ay,_,
+ BU Ade, toy, (9-25)
Since Eq. (9-25) is double logarithmic, the coefficients represent elasticities. Thus
the short-run (or impact) elasticity of Y with respect to X is B(1 — A), while the
long-run (or full adjustment) elasticity is seen from Eq. (9-23) to be £B. If the
estimated form of Eq. (9-25) is denoted by
Y= lg + CY + CyX,
then
Estimated short-run elasticity = c,
Estimated adjustment parameter = c,

Estimated long-run elasticity = i —Cy 6,

The partial adjustment process specified in Eq. (9-20) has been widely used in
applied work because of the simplicity of the resultant estimating equation, such
as Eq. (9-25). Nonetheless it implies a pattern of adjustment that may sometimes
be implausible. Suppose X had been constant at X sufficiently long for Y to have
settled at the desired level, Y= a + BX. In period t we assume X to become
X +AX and then to remain at the new level indefinitely. The new desired Y is
given by Y = a + B(X +AX), and the adjustment to that level implied by Eq.
(9-20) for a A value of, say, 0.5 and a negative B is shown in Fig. 9-1.
In the first period one-half of the total desired adjustment is achieved; in the
second period one-half of the remaining adjustment is accomplished, and so
LAGGED VARIABLES 351

Figure 9-1

forth. Thus the maximum adjustment is achieved in the first period, and each
successive adjustment is a fraction A of the previous adjustment. This might be a
plausible reaction pattern for, say, the consumption of broiler chickens in
response to a significant price change, but it is less plausible for the consumption
of gasoline since that consumption is mediated through durable equipment.
A further difficulty with the simple partial adjustment process arises when Y*
is a function of more than one explanatory variable. Suppose, for example, the
optimal level of energy consumption depends on both the relative price of energy
and the level of output in the economy. Applying the partial adjustment process
to actual energy demand imposes the same adjustment parameter on each
explanatory variable. Even if the form of the adjustment process is similar for
each variable, the speed of the process may well be different. Thus at given prices,
one might expect energy consumption to move more or less in step with output,
but to react much more slowly to price changes.

Combination of Adaptive Expectations and Partial Adjustment


Suppose X* represents “permanent” or long-run income, and Y* the correspond-
ing level of “permanent” or long-run consumption.f One might then write

Y= aap x (9-26)

This is not an operational equation since there are no direct observations on the
variables. However, the adaptive expectations hypothesis may be used to explain
X* and partial adjustment to explain the adjustment of Y to Y*. Thus combining
Eqs. (9-17) and (9-21) with Eq. (9-26) and allowing the A parameter to be different

+ See M. Friedman, A Theory of the Consumption Function, Princeton University Press, Princeton,
NJ, 1957.
352 ECONOMETRIC METHODS

in the two processes gives


(VA Ee) (aang) certo
=a(1—A,)+A(1 —A,) X* + yu,

a(1 — A,) B(1 —A,)U = Az) A, tu;


ae LS
or

Y, aa a(l ‘ai A,)( ra A) da (A, ats Ar)¥-1 a NADY,»

+B(1 vr A,)C a A) X, sta (u, me A5u,-1)

The parameters A, and A, appear symmetrically in the systematic part of Eq.


(9-27). Thus if one ignores the structure of the disturbance term and runs a
regression of Y, on Y,_,, ¥,_,, and X,, the resultant coefficients would not yield
estimates of the separate lag parameters A, and A,. The sum A, + A, and the
product A,A, can be estimated directly, and hence the term (1 — A,)(1 — A.) is
estimable and so are a and £. However, taking account of the structure of the
disturbance term can lead to estimates of the A’s, as will be shown in Sec. 9-2.

9-2 ESTIMATION METHODS

Let us begin with the estimation of the distributed lag function (9-1), that is,
Yi pot 0g ch OX ict On a aes (9-28)
where, for simplicity, we restrict consideration to the lagged values of a single
explanatory variable. We usually cannot expect theory to indicate the maximum
length of lag, but one would ordinarily expect significance tests on the 5’s to give
some indication both of the maximum lag length and of any delay in the initial
transmission of an effect from X to Y. The validity of such significance tests
depends on the properties of the disturbance process {u,} and the associated
estimation methods. If E(u) = 0 and var(u) = o7I, then, in principle, OLS would
be an appropriate estimation technique. In practice, however, its application is
likely to be plagued by collinearity between the regressors, leading to great
imprecision in the estimates of the 6’s.

Almon Lags
A general strategy for dealing with this collinearity and the associated imprecision
is to reduce the number of parameters to be estimated by the assumption of some
pattern for the 6’s. The Koyck scheme of Sec. 9-1 is perhaps an extreme example
of such a pattern. The Almon lag scheme provides a more flexible method for
reduced parameterization.t Under the Almon scheme one rules out the direct

7S. Almon, “The Distributed Lag between Capital Appropriations and Expenditures,”
Econometrica, vol. 30, 1962, pp. 407-423.
LAGGED VARIABLES 353

6) 6)

+— — Le vy
—=6
Si

(a) (2)

Figure 9-2

approach of attempting to estimate all (s + 1) 6’s and assumes instead that the
6’s can be approximated by some function 6, = f(i), as in Fig. 9-2b. The basis of
the approximation is Weierstrass’s theorem, which states that a function continu-
ous in a closed interval may be approximated over the whole interval by a
polynomial of suitable degree, which differs from the function by less than any
given positive quantity at every point of the interval.}
As an illustration suppose we postulate a third-degree polynomial, that is,
CD) =a,+ ait Osh + Ont
Then approximately
8) = f(0) = a
6, =f(l) =a, t+a,t+a,+
a,
6, = f(2) 3 ao aie 2a, ote 4a, ate 8a, (9-29)

5, = f(3) = ay + 3a, + 9a, + 27a,

6, = f(s) =a) + sa, + s*a, + 57a,


Substituting Eq. (9-29) in Eq. (9-28) and rearranging gives
ey a NG te AC et Ngo A ea here ee)
aWeoie Ow Ola RY,GF eee 12D ae ay)
+a5(X,4 + 4X,_. + 9X,_, +.-*> + s°X,_,)
+a,(X,_, + 8X,_, + 27X,_, +--+ + 5°X,_,) +u, (9-30)
Thus four new regressors are formed as linear combinations of the lagged X’’s.
The regression of Y on these variables yields estimates of the a’s, which in turn

+R. Courant, Differential and Integral Calculus, vol. 1, 2d edition, Blackie & Son, Glasgow, United
Kingdom, 1937, p. 423.
354 ECONOMETRIC METHODS

yield estimates of the 6’s from Eq. (9-29). The sampling variances and covariances
of the §’s can be computed from those of the &’s and significance tests carried out
on the 5’s. Defining W, as the matrix of coefficients in Eq. (9-29),

Orne 0
anor | ] 1
Woes | eer.
Byes 27
hy4 vs)
Bs3
where the subscript 3 indicates the use of a third-degree approximating poly-
nomial. Equation (9-29) then becomes
5 = W,a (9-31)
and, given 4,
5 = W,4a (9-32)
The matrix form of the original equation (9-28) is
y=ip+ X8+u
Using Eq. (9-31),
y =in + XW,a + u
An OLS regression of y on [i XW,], where XW, is the matrix of observations on
the “new” regressors shown explicitly in Eq. (9-30), gives the estimated coefficients
| plant ixXw, |-'| ivy
&| | WsX’'i WiX’XW, WiX’y
with
as4 .
var(&) = 02 |W,xX’XW, — TW5X'HXW, | (9-33)

From Eqs. (9-31) and (9-32)


E(8) = W,E(&) =8
and
var(5) = W; - var(&) - W; (9-34)
Substitution of Eq. (9-33) in Eq. (9-34) then gives the matrix of sampling
variances and covariances for the 6’s.
The above would be very useful if, in fact, one knew the appropriate degree
for the approximating polynomial. In practice the determination of that degree is
an important problem, even given an assumption about the maximum lag length.
The problem may be approached in two ways. From Eq. (9-30) it is seen that the
coefficient of the last “new” regressor a, is the coefficient of the highest power in
the approximating polynomial. Testing the significance of a, is, in effect, asking
whether we need a third-degree polynomial. However, finding a3 insignificant
does not necessarily imply that higher-order a’s would also be found insignificant.
LAGGED VARIABLES 355

The recommended procedure would be to start with a fairly high degree of


polynomial, say, fourth or fifth, test the last coefficient for significance, and keep
reducing the degree of the polynomial until the last coefficient is found significant.
The disadvantage of this procedure is that, in order to carry Out the tests, various
“new” regressors have to be computed which may not in fact be required in the
final regression.
The second approach avoids this computational difficulty. Tests of the degree
of the approximating polynomial can be based on the unrestricted OLS estimates
of Eq. (9-28) for some assumed value of s, and once the degree has been
determined, the Almon estimators can be found by an application of restricted
OLS estimation. Consider again a third-degree approximation given by
0: == a + o,f +517 2 +-a,1 3
Taking the first difference of this function gives a polynomial of the second
degree, and so on, for each successive difference until+

MS, = 0
But Aé; = 6, — 8,_,
A’6, a (4, Oe (On 8-2)
= 0, 20; jt 0,25
A°5, = 6, — 36,_, + 36,_, — 8,_,
A*6, = 6, — 46,_, + 66,_, — 46,_,+6,,
Thus the assumption of a third-degree polynomial places a set of linear restric-
tions on the 6’s. The full set of restrictions is
6, — 46, + 66, — 46, + 6) =0
6, — 46, + 66, — 46, + 6, =0
oP Bese eraehvelel ce. (ef a°h¢. ise) 0/0, ie)is) orled euiem eo)Seuuey (s)fe Ta 6 (9-35)
64S
Ss RY
65.) 4b + 8 =O

(Sy eees WO
=o :
te Ott Ol Py) sar O41 3

=a — a (i
—1) — a)(i — 1)’ — «(i — 1)’
= (a, — a, + a;) + (2a, — 3a3)i + 30,17
The second difference of 5; is found by repeating the first difference operation. Thus
: 2
A*8, = (2a) — 3a3)i + 3a3i* — (2a) — 3a3)(i — 1) — 3a,(i - 1)
= (2a, — 643) + 603i
The degree of the polynomial in i decreases by | with each differencing. The third and fourth
differences are then
M36, = 6a;

and Mos 0
356 ECONOMETRIC METHODS

No restriction involves the intercept term w. Thus the restrictions (9-35) may be
expressed as
R35 |=0 (9-36)
where R, is the.(s — 3) X (s + 2) matrix
0 | = 4 Gea 4 1 0 Sas 0
0 0 bo 4 6 a4 1 vee 0

0
A second-degree approximating polynomial would imply the set of s — 2 linear
restrictions given by
[
R.[5|-°
where R, is the (s — 2) X (s + 2) matrix
Orman Sear ] 0 ee 0
0 Os Ss ] ves 0
R,= : : (9-37)
0
Notice that the nonzero elements in the rows of the R matrices are given by the
appropriate set of binomial coefficients with alternating signs.¢ If r denotes
the degree of the approximating polynomial, the nonzero elements in R, are the
coefficients of L in the polynomial (1 — L)’*', but in reverse order. However,
since the restriction sets linear combinations of the 6’s equal to zero, we can
multiply the rows of R, by —1 and get the coefficients in natural order.
For a given maximum lag s the sequential procedure for finding a suitable
degree of the approximating polynomial would be as follows.

1. Start with a polynomial of fairly high degree, say, the fourth or fifth.
2. Set out the corresponding R matrix and test the null hypothesis
LL
Hy: R|5|=0

by substituting in Eq. (5-68) the results of the unrestricted OLS estimation of


¥,=p+6)X,+8,X,_,4+-°-+6,X,,
+4,

+ The binomial coefficients may be simply obtained from Pascal’s triangle


l Almon-polynomial

l 3 3 ] second degree
] A ate ibs 4 1 third degree
where an internal element in any row is the sum of the pair of elements immediately above.
LAGGED VARIABLES 357

Under the null hypothesis the resultant test statistic has the F(s — r,n — 5 —
2) distribution, where r is the degree of the approximating polynomial.
3. If the null hypothesis is rejected, the initial polynomial has not been of
sufficiently high degree.
4. If the null hypothesis is accepted, proceed to the next lower degree and test
the new set of linear restrictions, proceeding in this way until the null
hypothesis is rejected.

If the null hypothesis is accepted, say, for R, but rejected for R,, the
appropriate procedure is to find a third-degree approximating polynomial. This
may be done by computing the four “new” regressors specified in Eq. (9-30),
estimating the a’s by OLS and then using Eq. (9-32) to estimate the 6’s.
Alternatively one may use the formula for the restricted estimator given in Eq.
(6-5), and inferences may be made by using the variance matrix given in the
footnote to Eq. (6-5).
The above procedure is conditional on some assumed value for the maximum
lag s. It may be repeated for various values of s and a judgment made by looking
at the overall fit and the significance of the higher-order 5’s.
An implication of the Almon procedure, which does not seem to have
attracted much attention, is that it is likely to yield biased and, indeed, incon-
sistent estimates. Write the original model, Eq. (9-28), for simplicity as
y=Xd+u (9-38)
If the 6’s do not lie exactly on the approximating polynomial, then a formula
such as Eq. (9-31) has to be amended to
8=Wa+v (9-39)
where v is an r X 1 vector of errors involved in the use of an rth-degree
approximating polynomial. Notice that v is independent of time and is a vector of
unknown constants, which does not vanish with increasing sample size. Substitut-
ing Eq. (9-39) in Eq. (9-38) gives
y = XWoa + (Xv + u) (9-40)
In Eq. (9-40) there is obviously some correlation between the explanatory
variables XW and the expanded disturbance term Xv + u, which would lead one
to expect inconsistency in the estimation of a and hence of 6. Looking directly at
the estimator of 6,
§ Wa
w(W a W’X’y
w]= w{(5xx}8
1 1
i Xu} n n

Assuming

ee a5
358 ECONOMETRIC METHODS

and

plim|
hi (x
— X’u =0

we have
plim§ = W[W’2,,W] 'W’S,,8 (9-41)
Substitution of Eq. (9-39) in Eq. (9-41) gives
plim§ = 8 + W[W’S,,W] 'W’S_v (9-42)
so that the Almon estimator is inconsistent unless the unknown 6’s lie exactly on
the chosen polynomial, in which case v = 0. The finite sample bias of the Almon
estimator can be serious if one fits a polynomial of too low degree. This bias,
combined with the smaller sampling variation (as compared with unrestricted
OLS), can sometimes give sampling distributions for the Almon estimators which
fail to contain the true 6 parameter altogether or else have it located near an
extremity of the distribution.
Computer packages with Almon lag estimators usually offer the facility of
including end-point restrictions such as 6_, = 0 and/or 6,, , = 0. Since 6_, is the
notional coefficient of X,,, and that variable has no effect on Y,, it might seem
sensible to incorporate that end-point constraint. As Dhrymes and Schmidt and
Waud have pointed out, that is a fallacious argument.} Setting 6_, = 0 implies a
restriction on the a’s and hence on the 8’s, which in turn is a restriction on how
X,, X,_},---, X;_, affect Y,. For a second-order polynomial the implied restriction
is
My — a, tay=0
Such arestriction could, of course, be tested by estimating the a’s and using the
variance matrix in Eq. (9-33). The purpose of the Almon polynomial is to give a.
good approximation to the unknown 6’s over the interval 0 to s. Its behavior if
extrapolated outside that interval is irrelevant. The second end-point restriction,
8, ,, = 0, may not produce much distortion in the approximation if the coefficients
are decaying with increasing lags but, again, it implies a restriction on the a’s and
6’s, and there seems little valid reason for imposing it.

Direct Estimation of a Koyck Lag


If one assumes Eq. (9-28) to obey a simple Koyck lag, the relation becomes
Y¥,= w+ 6X, + o8X) + @bX
5 Fe os la} <1 (9-43)
with four parameters to be estimated, namely p, 6, a, and oe. The lag is now
infinite, but the coefficients decay exponentially. The relation (9-43) may be

7 P. J. Dhrymes, Distributed Lags: Problems of Estimation and Formulation, Holden-Day, San


Francisco, 1971, pp. 232-234; P. Schmidt and R. N. Waud, “Almon Lag Technique and the Monetary
versus Fiscal Policy Debate,” Journal of the American Statistical Association, vol. 68, 1973, pp. 11-19.
LAGGED VARIABLES 359

rewritten as
Y=
+ 8(X, 4 0X. + +: + at 1X) + a'8(X
+aX,
, to) +a,
or Y,=pe+ OxX* + a'y + u, (9-44)
where
AP SX aXe aX,
0

and y=68) a'X_, = E(%—-


1p)
i=0
The y parameter may be regarded as the expected difference between Y and p in
the period preceding the first sample observation. If u ~ N(0, 671), the applica-
tion of OLS to Eq. (9-44) would yield ML estimates. The matrix formulation of
Eq. (9-44) would be

Le Aga he u
y=/]1 X¥ a? |/s]+u (9-45)
i. y% oe a" Y

where the X*’s may be built up recursively as


X* = X
XZ = X, + aX, = X, + aXF
X¥ = X, + aX, +. 0°X, = X, + aX¥
Since two columns of the data matrix depend on the unknown a, one can proceed
with a grid search over the interval 0 < a < 1. For each specified value of a the
data matrix in Eq. (9-45) is computed and OLS applied, the final choice of
regression being based on the minimum residual sum of squares. The standard
errors for f and 6 from the OLS program would only be correct if a were known
exactly, which is not the case. The asymptotic standard errors can be obtained
from the information matrix, which ist
OX t*
Ae ES *
Sa t
z(0 rere ft
7]
0Xx*
ane Laxt z(0 f + tally) XP
1 0a
x 2 ax? (9-46)
RxXOD
Go, Yat! zoe + tal"Ja!
0a
ox* 2
zs + ta’-'y}
0a

This matrix is symmetric and so only the upper triangular portion has been
shown. The unknown parameters in Eq. (9-46) would be replaced by their

+ See Problem 9-4.


360 ECONOMETRIC METHODS

estimated values, and the inverse would give the estimated variance matrix for the
parameters.

Estimation with a Lagged Dependent Variable


Instead of estimating the simple Koyck scheme directly, as above, one might use
the derived relation
1 u(l a a,) Fe Yo Boke (u, ae au,_1)
already established in Eq. (9-11). The adaptive expectations scheme is formally
identical, as shown in Eq. (9-18). The partial adjustment model, derived in Eq.
(9-22), gives
Y,=a(1 —A) +AY,_; + BU —A)X, + 4,
Both relations incorporate a lagged Y among the regressors and differ only with
respect to the properties of the disturbance term. We must now examine the
estimation problems occasioned by the lagged Y value, and we shall do so under
various assumptions about the disturbance term.

Lagged Dependent Variable and Well-Behaved Disturbances


Consider the relation
T= By Bo Xap, eee, (9-47)
where we assume the w’s to be independently and identically distributed with zero
mean and variance o2. ut The relation (9-47) may be rewritten to show the
dependence of Y, on the stream of current and previous values of X and u, that is,

(1 — B,L)Y, = B, + BX, + u,
giving
Y= a+ B(X, + BX. + BFX,at <1) +O, (9-48)
where a=
B,
[ees

and o= (1 = BiE\gea,
If X were held constant at some level X and Ydenotes the corresponding level of
Y, then

E(Y )= By B, x
ete ees
provided |,| < 1. If |6;| = 1, E(Y) would explode. In practice a {Y,} series may
have explosive tendencies, which are held in check by various “floors” and/or
“ceilings.” A model of such a process would be highly nonlinear, and the
statistical treatment of such models is still in its infancy. We therefore impose the
constraint
|B3| < 1 (9-49)
LAGGED VARIABLES 361

It is also clear from Eq. (9-48) that expressions such as (LY,7/n) and GAY)
will involve linear combinations of quantities such as

] ]
eke yekmir

] ]
and year pele Ce ere

The additional assumption is then made that the X, are bounded and that the
above quantities have finite limits as n tends to infinity.
The model of Eq. (9-47) may be written in matrix form as

y=ZBp+u (9-50)
where

te eX Y
Leis Y,

To make Eq. (9-50) operational Y) has to be known.; The Z matrix is stochastic


since the Y’s are stochastic, even though the X’s may be assumed to be exogenous
and nonstochastic. However, the case is not an exact parallel of the stochastic
data matrix considered in Sec. 7-3. There the strong assumption of full indepen-
dence between the disturbance and the explanatory variables was valid. In this
model it is clear from Eq. (9-47) that, while wu, is independent of X, for all ¢ and all
s and also independent of Y,_, for positive s, it is not independent of Y,, and since
Y, in turn influences Y,,,, u, is not independent of Y,,,, Y,,5,.... This ap-
parently small difference has an important effect on the estimates of Eq. (9-50).
The underpinning assumptions for Eq. (9-50) may now be stated:

1. E(u) = 0 and E(w’) = 021


2. E(X,u,) = E(Y,_,u,) = 0 for all¢
3.
]
plim{—-2°2Z) =2 ZZ

a symmetric positive definite matrix

Assumption 3 follows from the stability assumption on £, and the assump-


tion about limiting values for the second-order moments of X.£ The Mann-Wald

+If it is not, the effective sample size is n — 1, and the statistical inference procedures are
conditional on Y, with n— 1 observations, rather than conditional on Y) with n observations.
Asymptotically, of course, it makes no difference.
+ For a complete derivation see E. Malinvaud, Statistical Methods of Econometrics, 2d edition,
North Holland, Amsterdam, 1970, pp. 540 ff.
362 ECONOMETRIC METHODS

theorem can then be applied to give the results.+

plim(— Zu) =0 (9-51)


and

[ez] 5 N(0, 023,.) (9-52)


vn
The OLS estimator of B in Eq. (9-50) is

B = (ZZ) 'Zy
=B+ (ZZ) ‘Zu
Thus

(b-B)-(52z)
-
vn(B-B) 1
= [|,272} ——el Zu (9-53)
9.53
Using Eqs. (9-51), (9-52), and (7-24) gives

vn (B — B) ~ AN(0, 073;,')
or B ~ AN(B, 02722] ZZ
(9-54)
Thus even without the assumption of normality for the u’s the OLS estimators
will be consistent and asymptotically normally distributed. The unknown variance
matrix in Eq. (9-54) can be consistently estimated by the usual formula s?(Z’Z) ~!.
If, in addition, the u’s are normally distributed, the estimators are also ML and
efficient. These results extend simply to the general case of various lagged Y
values and several X’s. Thus there is substantial justification for the continued
use of OLS in relationships containing lagged dependent variables, provided the
disturbance term is serially independent. The estimators will, however, be subject
to finite sample bias, and one should also recall the problems of testing for
autocorrelated disturbances in this case.

Lagged Dependent Variable and Autocorrelated Disturbances


Suppose now that we repeat the relation (9-47)
Y,=1B itp kc bs) ee = Nee toin
as before, but the u’s are now assumed to follow an AR(1) scheme§

u, = pu,_, + e, |p| eal (9-55)

where E(e)=0 and E(ee’) = 071

+H. B. Mann and A. Wald, “On the Statistical Treatment of Linear Stochastic Difference
Equations,” Econometrica, vol. 11, 1943, pp. 173-220, especially pp. 185-190.
+ See Sec. 8-5.
§ Note that this is different from the error structure in Eq. (9-11) associated with the Koyck lag;
the latter [an MA(1) error] is considered below.
LAGGED VARIABLES 363

This new assumption has an important effect. From Eq. (9-55) it is seen that u ra
influences u,, but from Eq. (9-47), in period ¢ — 1, u,_, influences Y,_,. This sets
up a dependence between u, and Y,_, in Eq. (9-47), that is,
E(Y,_\u,) = 0
From Eqs. (9-47) and (9-55) it follows that}

po, 2
plim|“ZY, 1}
“T= Be (9-56)
The consequence is that the application of OLS to Eq. (9-47) will yield incon-
sistent estimates of all parameters. This is so because

plim(B) = 6 + =; - plim(—-Z/u]
and
Ae (es
plim yee 0
aL a ecrc i 0
plim(= u| = plim(-EX,u,} = poz

plim(—Y,_.«,] s
] pea

It only takes one nonzero element in plim((1/n)Z’u) in general to render all


elements in B inconsistent. There are two main methods of obtaining consistent
estimators in this model, namely, instrumental variables and ML.

Instrumental Variables
Consider
y=Zpt+u
Premultiply by Z’ to give
Z'y = Z'ZB + Zu (9-57)
The OLS estimator b of Chap. 5 may be obtained from this equation simply
by setting Z’u = 0, giving
Z'y = Z'Zb (9-58)
On the assumption that

plim(22) =2,, and plim(—2'u) =0


we can divide Eqs. (9-57) and (9-58) by n, take probability limits, and equate the
right-hand sides to find
2,,plim(b) = 2,8
so that
plim(b) = B
which is the standard result on the consistency of the OLS estimator.
+ See Problem 9-5.
364 ECONOMETRIC METHODS

In the present model the assumption that plim((1/n)Z’u) is the zero vector
cannot be sustained. Suppose, however, that one can find an n X k matrix W
containing variables which are thought to be contemporaneously uncorrelated
with the disturbance term. That is, we assume

E(W,u,) =09 hf = 1k tee (9-59)


Premultiplying the model by W’ and setting W’u to the zero vector, by analogy
with the OLS procedure, gives the instrumental variable (IV) estimator byy,

W’y = (W’Z)byy

which, on the assumption that W’Z is nonsingular, may be written

by = (WZ) ‘Wy (9-60)


On the further assumptions that

plim| kd = ,, | a nonsingular matrix (9-61)

and plim(—W'u] =0 (9-62)

it is easy to see that

RF tial asl cial


plim(b,,) = B + plim(—-w’Z] plim(—W'u)

=B+2,)-0

=B
so that the IV estimator would be consistent.
The variables in W are referred to as instruments. Some of them may simply
be variables from the original Z matrix. In the present model there is no need to
replace X, since it is already assumed to be independent of the disturbance term.
In addition to being uncorrelated with the disturbance term, the instruments
should not be totally uncorrelated with the explanatory variables since W’Z
would then be a null matrix and the estimating technique would break down. If,
in fact, W’Z is “nearly” null, the IV technique will give very poor results.
In Eq. (9-47) we need just one instrument, and it is customary to select X,_ 1
as the instrument for Y,_,. The appropriate matrices are then

1 XX Les
We | Pi lle Xo tvs
Lag ae Dey oer
LAGGED VARIABLES 365

and the IV estimator is

n DEX; e¥eg —1 Nahe

Dinesh Are Be Nas gE GYS, rX,Y,


LX,_| LX,X,_| LX, 1¥,-4 LX,_1Y,

where all summations run from t = 1 tot = n.t


IV estimators are, in general, biased in finite samples and their variances
difficult to establish.{ It is, however, possible to derive a fairly simple and
important asymptotic result. The result requires three assumptions. The first is
Eq. (9-59),
E(W,,u,)=0 — foralli,t
that is, that the instruments are contemporaneously uncorrelated with the dis-
turbances. The second is that the instruments possess finite probability limits for
all second-order moments, that is,
fal
plim( ww] =2, Ww

a symmetric positive definite matrix. The third is that


E(u)=0 and = E(w’)=o71
Because of Eq. (9-55) this last assumption is not true for model (9-47). However,
we will ignore this complication for the moment. Under the above three assump-
tions the Mann-Wald theorem applies so that
oe (RL ga\ TS
plim{—W'u) = 0

and [wl asN(0, 62%,,,,)

The IV estimator of Eq. (9-60) is

by =B + (WZ) 'Wu
Thus

bp )= [Gwz) [zw
n

+ Should the values Xj and Yo not be available, the first row is dropped from W and X, the
summations run from t = 2 to t = n, and n is replaced by n — | in the formula for byy.
+ Contrast the assertion by P. J. Darymes, Econometrics—Statistical Foundations and Applications,
Harper and Row, New York, 1970, p. 297: “All IV estimators, no matter what the choice of
instruments, are unbiased and consistent.”’ This statement comes after a passage in which the only
explicit assumptions relate to probability limits. The IV estimators are consistent. A possible
explanation of the incorrect assertion about unbiasedness is given in App. A-8, Expectations in
Bivariate Distributions, where the matter is discussed in detail. The same type of error can also affect
the derivation of results about finite sample variance matrices, as in formula (6-4-12) of Dhrymes.
366 ECONOMETRIC METHODS

Recalling assumption (9-61) that

plim(—W'] = 2,,,

a nonsingular matrix, an application of Eq. (7-24) gives

Vn (By — B) ~ AN(0, 6,
2,22
yyBuz)
or |

by ~ AN(B, 022Ea/En2c2'] (9-63)


Under the full assumptions the IV estimator would be a consistent and asymptot-
ically normal estimator of B. The variance matrix would be estimated by the
formula
est var(b,y) = s2(W’Z) '(W’W)(Z'W)' (9-64)
where
eeae ie
aNXbyy)(y SNE
— Xb,y)
iaak
The statement in Eq. (9-63) is not strictly valid for the model (9-47) since the
u’s are not independently distributed in consequence of Eq. (9-55). It also follows
that Eq. (9-64) would not be the appropriate formula for estimating the sampling
variances of the IV estimators, though it is often applied for want of anything
better. Results (9-63) and (9-64) hold for the IV estimator under the full set of
three assumptions outlined above, and we shall have need of them subsequently.
The main use of the IV estimator in this model is to provide a consistent
estimator as a starting point in an iterative ML technique.

Maximum-Likelihood Estimator
Combining Eq. (9-47) with the AR(1) disturbance process in Eq. (9-55) gives

Y= 8, ~ Bip + Bax, — Bop Xeiet (Bx + 0) Ye Ol ene


If one assumes
e ~ N(0, 621)
then ML estimators of the B’s and p would be given by the values minimizing
Le?. However, the first-order conditions would not yield linear equations in the
estimators, since there are five variables in Eq. (9-65) but only four parameters to
be estimated. An iterative Cochrane-Orcutt procedure may be used, based on two
alternative ways of rewriting Eq. (9-65), namely,

(Ye— PY,2) 8 Che) ita BoGkaa eXen) 48s Vay ee eames


(9-66a)
(Y= By BX, — BYE) = Oe Bi B22 Bae ie,
(9-665)
LAGGED VARIABLES 367

Given a starting value for p, the transformed variables in Eq. (9-66a) could be
computed and OLS applied to yield estimates of the B’s. These estimates in turn
could be used to compute the transformed variables in Eq. (9-666) and OLS
applied to produce a revised estimate of p with the iterations continuing till
convergence.
Setting up the log likelihood for Eq. (9-65) and differentiating gives the
information matrix}
2
mda Clem OE XS palGlicyee Ri) 0 0
a DX? E(XX*Y* ,) 0
1
2
Bp. ry*2, oa 0
R B; => 30

p °c no, 0
o2 = 2)
. p
ae
20,
(9-67)
where

Ms PA) and Pe si ele


This matrix is symmetric, and we have just shown the upper triangular portion.
As usual it involves the unknown parameters, but it would be calculated using the
estimated parameters.
It has recently been shown that the iterative Cochrane-Orcutt process may
lead to inconsistent estimates.¢ The basic point is that, while for a finite sample
the Cochrane-Orcutt estimators will always converge to some fixed point, that
point may correspond to a local minimum of the sum of squares rather than the
global minimum, and the probability limit of the fixed point will not be the true
parameter vector. To ensure consistency of the iterative process, one must start
with a consistent estimator. Thus starting the process by setting p to zero in Eq.
(9-664) would be inappropriate since that corresponds to estimating the B’s by
applying OLS directly to Eq. (9-47), which is an inconsistent procedure. On the
other hand, the process could be started consistently by computing, say, the IV
estimators of the B’s as described earlier.
An alternative approach to minimizing Le? in Eq. (9-65) is to use a grid
search over the permissible range of p values. Thus a set of p values is specified in
the interval (— 1, 1). Each value is used to compute the quasi first differences in
Eq. (9-66a), and OLS is then applied to minimize Ue?. A fine enough grid should
distinguish the global minimum from any local minima. If necessary a finer grid

+ See Problem 9-6.


+R. Betancourt and H. Kelejan, “Lagged Endogenous Variables and the Cochrane-Orcutt
Procedure,” Econometrica, vol. 49, 1981, pp. 1073-1078.
368 ECONOMETRIC METHODS

may be imposed around the p value chosen in the first grid search and a second
grid search applied to obtain a finer estimate of the minimizing p value. This
value and the corresponding 8’s obtained from Eq. (9-66a) constitute the point
estimates, and the asymptotic standard errors can be obtained from Eq. (9-67).

MA(1) Disturbance
Instead of the AR(1) disturbance process assumed in Eq. (9-55), let us now
consider an MA(1) process. As has been shown, this is likely to occur in a simple
Koyck scheme or in an adaptive expectations model. In each of these cases there
is the further significant feature that the parameter of the MA(1) process is also
the coefficient of the lagged dependent variable. The model to be considered is
thus
Y,=a+AY,_,
+ BX,+(u,—Au,) [AL <1 (9-68)
where it is assumed that

u ~ N(0, 621)
Utilizing the existence of the common parameter, this relation may be rewritten
as
LZ, SOA NZ. sete piAy (9-69)
where
Z,=¥,- 4,
Successive substitution for the Z variable in Eq. (9-69) gives
Z,=a(1+A+V4---4+2X7')
+ B(X,+AX,_) + PX. + + + NIX) + ZX
or Y,=a(1+A+--- +7!) + BX* + ZN + u, (9-70)
where now
XP = X, A NX,_4 + VXLg Hee HAY,
which may be computed recursively, for any given A, as
XP =X tN Aeon with X* = X,
Relation (9-70) has a well-behaved disturbance term suitable for ML (or equiva-
lently OLS) estimation, with Z, treated as a nuisance parameter. The data matrix
for OLS estimation would be

1 AtEN
1+A XEN
Oe 1+A+2”

The appropriate procedure is then a grid search over the interval 0 < A < 1. For
LAGGED VARIABLES 369

each value of A, X(A) is computed, OLS applied to Eq. (9-70), and the set of
parameters is chosen which minimizes the residual sum of squares.
The asymptotic standard errors may be obtained from the information matrix
in the usual way. The log likelihood for Eq. (9-70) may be written
1
nL = -— 7in(27)— sino; ~ 392 oti (9-71)
where
u, = Y, — aW, — BX* — ZN
and W.=1+A4---4)01
The unknown parameters in Eq. (9-71) are a, 8, A, Zy, and 0°.It may be oun
that the expected values of the cross ee order partial derivatives involvin07
g
are all zero. Thus inferences about 07 may be made independently of the other
parameters. The ML estimator is
a2_ dt;ae
0,
n

with asymptotic variance 20,1/n. The information matrix for the remaining four
parameters ist

a YW?
t [Link]* DWV, IWR
: B S ae axe xe, exe (0-72)
r a, yy LV, x

a DM!
where W, and X* have already been defined and
Ou,
EX
=a{1+2A+---+(¢-1)N-?] + B[X,_, + 20X,_,
4+-
+ (¢— 1)N~?X,] + tZ,rN-!
For a penultimate problem we return to Eq. (9-27), which represents a
combination of adaptive expectations and partial adjustment. The equation is
MS a A ag) (Nak) YBN Ages
P(A, )(L AZ)X, + (0, = ABU)
Defining Z, = Y, — u,, this may be rewritten as
Zim gia Ng Yo et By kee NoLyay (9-73)
where & = a(l —A,)(1 -—A,)
Bo =P = A) —A3)
Baas
bal ey voeo.

+ See Problem 9-7.


370 ECONOMETRIC METHODS

Successive substitution for Z and transformation back to Y gives


Y,=a{1 +A, +-°°: +Ay']
tA [YA + AVE. + + NYE]
NGXan
+)Be [XP cst AG yee (9-74)
The disturbance term in Eq. (9-74) is well-behaved. The “variables” in square
brackets are all dependent on A,. Thus a grid search over 0 <A, < 1 and the
choice of the error minimizing version of Eq. (9-74) will yield point estimates of
all the parameters.
Finally we take a look at the estimation problems of Eq. (9-14) where it was
assumed that Y, responded to two separate Koyck lags with different parameters.
The equation was

Y, = p* + (a; + @,)¥,_, — a02¥,-2 + BX, — a, BX,_; +. YZ, — 0 YZ,-1 + 2,


with
Ore ae (a, ot Oy) Uy + QQ ,U,_2

The disturbance series {v,} follows an MA(2) process. Ignoring this complication
for the moment and assuming the v’s to be independently and identically
distributed normal variables, the application of unrestricted OLS to Eq. (9-14)
would not yield the ML estimators since the seven coefficients are functions of
only five parameters. However, the relation may be rewritten as
Y* = p* + BX* + yZ* + v; (9-75)
where
Ve Ye (a op) Xe aaae
XPS X05 X55
LZ, = ZZ pl

The transformed variables in Eq. (9-75) depend on the a,, a, parameters. Given
any pair of a,, a, values and assuming the v’s to be independently distributed,
OLS could then be applied to Eq. (9-75) to yield estimates of u*, 8B, y, and the
residual sum of squares. The indicated estimation procedure would be a two-
dimensional grid search over a,, a, pairs, each parameter being constrained to the
(0, 1) interval.
Alternatively, if one makes the explicit assumption that the v, follow an
MA(2) process and if u ~ N(0, 021), the variance matrix for the v’s is given by
Opt) Ope Uae. vee 0
8, 8 § 8 0 te 0
s O54 f(s bona Oro vee 0
E (vy)
i=07 een ae (9-76)
LAGGED VARIABLES 371

Table 9-1 Lagged variable models


a ee RE Pvt wey Set ds ad ee
Model Assumption Estimators
ae ae
l ¥,= n+ D(L)X, + u, D(L) a polynomial in the OLS (ML) possibly plagued
lag operator: {u,} white with imprecision due to
noise collinearity
2. .¥,= e+ D(L)X, + u, Almon approximation OLS (ML)
to D(L)
3. YS et DL) Xu, Koyck approximation ML (grid search)
to D(L)
Be Y mrt PaAp tr bal)44,tot, {u,) white noise OLS asymptotically
normal and efficient
53 ¥n= By By A, + B3¥,-p-t-a; {u,} follows AR(1) process OLS now inconsistent;
consistent estimators
via IV or ML
6. ¥Y,=a+A¥Y,_, + BX,+(u,—dAu,_,) (u,) white noise ML (grid search)
7. Y, = a(1 — A,(1 — Az) + (Ay + AQ)Y,_,; Combination of ML (grid search)
~A\AZ¥,_2 + B(L—A,)\(1—2,)X, adaptive expectations
itt es AG en) and partial adjustment
Sang pete (oy Ae ey) Yk | A aaa dna 5 Two explanatory variables ML (grid search) or GLS
BX, 05B X42) yZ, with separate Koyck depending on treatment
HONOL5 ae lags of {v,)
ee ee ee eee
yIf the u, are independently and identically distributed as N(0, 0), then {u,} is said to be a white noise
series.

where
8) = 1+ (a, +.a,)° + aa?
Ota (a, iz a,)(1 Ee Ay)
6) = aa,
The appropriate estimation procedure for Eq. (9-75) is then a combination of
GLS and a two-dimensional grid search over a,,a,. For each a,, a, pair the
variance matrix in Eq. (9-76) is computed and then GLS applied to Eq. (9-75).
One chooses the set of parameters that minimizes the weighted sum of squares
e’Q~ 'e, where e is the vector of residuals computed from Eq. (9-75) by using the
GLS estimates and Q is the matrix in Eq. (9-76).
Various models have been considered in this section, and it may be helpful to
summarize them briefly in Table 9-1.

9-3 TIME-SERIES METHODS

The models summarized in Table 9-1 incorporated various theoretical specifica-


tions. The first model embodied the least a priori specification, but direct
estimation was liable to be somewhat imprecise, which led to the development of
372 ECONOMETRIC METHODS

Almon approximations. The Koyck hypothesis in model 3 is a very strong


assumption. Models 4 to 6 are versions of adaptive expectations and partial
adjustment depending on the treatment of the disturbance term. Model 7 is a
combination of adaptive expectations and partial adjustment, while model 8
incorporates two explanatory variables with separate Koyck lags.
In recent years time-series methods of estimating a lagged relationship such
as Eq. (9-1)
Yipee D(L)X, + U,

have come to be more extensively employed. As seen in Sec. 9-1, this relation may
be formulated equivalently as Eq. (9-5),

Y=pt
B(L) Xa
A(L)
t

where D(L), B(L), and A(L) are all polynomials in the lag operator, but the
orders of B(L) and A(L) are expected to be small relative to the order of D(L).
The relation (9-5) is known as a transfer function in the time-series literature.}
There are four main characteristics which distinguish time-series estimation
methods from the various estimation procedures described in Sec. 9-2.

1. Before estimating the transfer function, the “input” series (X,} and the
“output” series {Y,} are subjected to sufficient differencing to render both
resultant series stationary.
2. The orders of the A(L), B(L) polynomials are determined empirically from
the data by an identification process and without imposing any a priori
theoretical specifications, such as a set of declining exponential coefficients.
3. The disturbance term in the transfer function is estimated as a general
ARMA process, as described in Sec. 8-5, rather than as a low-order AR or
MA process as in some of the models in Sec. 9-2.
4. The transfer function approach has been most extensively developed for the
single-input case (that is, one explanatory variable with various lagged values),
and there is no firm agreement yet on the appropriate extension to cope with
two or more inputs, each with aset of lags.

To get a grasp of the methodology we need to discuss each of these four


points in greater detail.

Stationarity
The simplest example of a stationary process is the white noise series {€,}, where
the e’s are independently and identically distributed as N(0, 67). It follows from

+ The basic reference is G. E. P. Box and G. M. Jenkins, Time Series Analysis: Forecasting and
Control, revised edition, Holden-Day, San Francisco, 1976, especially Chaps. 10 and 11.
LAGGED VARIABLES 373

the definition that


E(e,)=0 forall
var(e,) = E(e?)=02
t forall
Y, = cov(é,, &,_ = E(¢,e,,)=0 foralltands
+0
=e {1] fors = 0
pe, ter Xo 0 s+ 0
for
Thus the mean and the variance of the series are constant, finite, and independent
of the time subscript, as are the covariances and the autocorrelations. These
conditions constitute a definition of second-order, or weak, Pa A series is
said to be strictly stationary if the joint probability distribution of Xe aS
the same as the joint distribution of X, epee Ape fOLzalli ty aot eT:asince a
multivariate normal distribution is eo tees specified by the first- and second-
order moments, the {¢,} series is also strictly stationary.
Now consider
X,= oX,_, + &, (9-77)
where {e,) is white noise. This process may also be expressed as

p(L)X, e? (1 = oL) X, mae


giving
X= é, + pe.) + a ge:
Thus
E(X,)=0 — forall+
and
var(X,) = E(X7) =o2(1+ ¢+¢4+---)
This last expression only converges if |¢| < 1. We have already seen in Sec. 8-5
that if |o| < 1,

and the autocorrelation function is given by


pra
Thus Eq. (9-77) is a stationary process if |¢| < 1. This condition is also stated
equivalently as the root of p(L), or the zero of the polynomial g(L), lying outside
the unit circle. This root is obtained by setting
g(L)=1-4L=0
and solving for LZ to find L = 1/9. Clearly, the condition |¢| < 1 implies
Vel 1:
If |¢| > 1, the root of y(L) lies inside the unit circle and Eq. (9-77) is an
explosive series.t Now consider the in-between case where ¢ = 1. Relation OWT

+ See Problem 9-8.


374 ECONOMETRIC METHODS

then defines the random walk


KX rey

Clearly, var(X,) still explodes and X, is not a stationary series. However, AX, =
(1 — L)X, is a stationary series since it is equal to €,. Thus first differencing the
random walk series produces a stationary series, but no finite number of differences
of Eq. (9-77) can produce astationary series if |p| > 1.
Extending the model to a second-order scheme gives
X, = 9 X,_1 + Oy.X-2 + & (9-78)
or p(L)X, =, (9-79)
where p(L)=1-6,L—-¢L’
By analogy with the first-order case we seek conditions on the roots of »(L)
which might distinguish between the stationary case, the explosive case, and the
intermediate case, where differencing might produce a stationary series. The
polynomial may be factorized as
p(L) = (1-¢,L)Q - © L)
and so the roots of the polynomial are c,;' and c; '. From Eq. (9-79)
Xone,
=. See
a
(l—¢,L)d=oL£)*
The term 1/(1 — c,L)(1 — cL) may be expanded in partial fractions as
Ree
Tt ao
Se aeeen
Vea ber) aa py ka)
where d = c,/(c, — ¢), as may be verified by multiplying out. Thus
d head
A Tay eae oe eens
= d(e, + ¢,@,_, + ¢78,25 + --+)
+ (1 — d)(e, + cye,_, + che, +-:-)
and the variance of X, will only be finite and constant if |c,| and |c,| are both
less than unity, that is, if the roots of p(L) lie outside the unit circle. The condition
on the roots may be stated equivalently in terms of the $,, ¢, parameters of Eq.
(9-78) as}
Ip.|
<1
g, + o, < 1

g, — 9, < 1

+ G. E. P. Box and G. M. Jenkins, Time Series Analysis: Forecasting and Control, revised edition,
Holden-Day, San Francisco, 1976, p. 58.
LAGGED VARIABLES 375

If the polynomial factorizes as

PUL) (Leela L)
Eq. (9-79) becomes
(VP GLE BE) XS ep) AX, = &, (9-80)
Even if |c,| < 1, the X, series is nonstationary since the other root lies on the unit
circle. However, it is clear from Eq. (9-80) that the A X, series is stationary as long
as |c,| < 1. If a third-degree polynomial factorizes as
2
p(L)'= (l= eh) L)
then second differencing the X, series will yield a stationary series as long as
leq] <1.
So far we have just considered AR processes of the form y(L)X, = €,, where
€, 1s white noise, and have seen that the condition for stationarity can be
expressed in terms of the roots of »(L). The same conditions hold when the
disturbance of the right-hand side follows an MA scheme, for if we write
p(L)X, = O(L)e, (9-81)
where 6(L) is a finite MA operator,
6(L)=1-6,L—-6,L?—---- we Ogee
then 6(L)e, is a stationary series. It has zero mean, a constant variance, and an
autocorrelation function which is nonzero for the first g lags and zero thereafter.
Thus the stationarity of the X, series still depends on the roots of g(L). The
general form of Eq. (9-81) is

(1 iby lis Gyle = joa $,L? )(1 a LOX,

=(1-0L-6,L?-----6L*)e, (9-82)
This is an autoregressive, integrated, moving average, ARIMA(p, d, q) scheme,
where p is the order of the AR polynomial, d is the degree of differencing required
to yield a stationary series (or equivalently, the number of unit roots in p(L)),
and gq is the order of the MA polynomial. The term integrated refers to the reverse
of the differencing operation since the differenced series have to be summed (or
integrated) to retrieve the original series. It did not arise in Sec. 8-5 where ARMA
processes were introduced to model a disturbance series which was already
stationary.
The general ARIMA model of Eq. (9-82) has been found to be a very flexible
tool for the univariate modeling and forecasting of a wide variety of homogeneous
nonstationary series—series that are not explosive, but which may display drift or
apparent short-run trends as well as various irregular oscillations. The univariate
modeling procedure consists of first determining the amount of differencing
required to produce approximate stationarity. Typically it appears that, if
differencing is required, first or at most second differences suffice. Defining

x, = @ aa exe
376 ECONOMETRIC METHODS

the second stage is making a judgment about the orders p and q in

(1 — ¢,L — 1? - +++ — @,L?)x,=(1- 62 -6,L*—---— 6,L")e,


This is done by comparing the pattern of the estimated autocorrelation coefficients
of the {x,} series with the theoretical patterns corresponding to various small
values of p and q. An initial estimate of the }, 8 parameters is then derived, which
serves as the first round in a nonlinear iterative estimation process. Finally
various diagnostic checks are applied to the fitted model.
Econometricians are naturally more interested in the estimation of transfer
functions than in univariate time-series modeling. However, the latter turns out to
be an essential component of the former. Returning to the transfer function (9-5),
we may write it explicitly as

(Taphe ag? Ss 8 2 al EV ei 8p ek da Deer


or A(L)y, = B(L)x,_, + 4,

which is a transfer function of order (r, s, b), where b > 0 represents any delay in
the transmission of an effect from X to Y. The X and Yseries are appropriately
differenced to achieve (near) stationarity and are also expressed as deviations
from the sample means, if necessary. The problem now is the determination of
the values of r, s, and 6 and the estimation of the consequent a and 6 parameters.
The Box-Jenkins starting point is the calculation of the covariances (current and
lagged) between x and y and the autocovariances of the x series. The solution of a
set of simultaneous equations yields estimates of the 6 coefficients.{ From the
resultant 6 coefficients rough guesses are made of the values of r, s, and b on the
basis of a comparison between the pattern of the 6 coefficients and the theoretical
patterns for various values of r, s, and b. From the 6’s initial estimates of the a’s
and £’s can be derived and an iterative estimation process carried out, with
interaction between the estimation of the transfer function weights and the fitting
of an ARIMA scheme to the disturbance term. If the original disturbance was a
white noise series, any differencing will have produced an MA process in the
transformed disturbances, and if the original disturbance was complicated, the
transformed disturbance will normally be more complicated.
Box and Jenkins also suggest that the efficiency of the above process could be
improved if an ARIMA model was first fitted to the x, series. Denote such a

+ It is assumed that the same degree of differencing has been applied to each series. However, Box
and Jenkins state, “the procedures outlined can equally well be used when different degrees of
differencing are employed for input and output” (op. cit., ftn., p. 378). Consider
Y,=a+t BX,+ u,
First differencing both Y and X gives
AY, = BAX, + Au,
so that the original f coefficient is retained while the intercept disappears. If different degrees of
differencing are applied to each variable, one would no longer be estimating the original f coefficient.
+ These are the coefficients of the various lagged values of X, defined earlier in Eq. (9-1).
LAGGED VARIABLES 377

model by

$(L)x,
=6(L)n,
where 7, is approximately a white noise series. Now multiply through the model

y, = D(L)x, + u,
by 0. '(L)o,(L). The result is
y* = D(L)n, + 2, (9-83)
where

y* = 0-'(L)o,(L)y, and v, = 6, '(L)o,(L)u,


Equation (9-83) still preserves the 6 coefficients of the original equation, but the N
variable is approximately white noise and so its lagged covariances will be
approximately zero. This leads to a considerable simplification in obtaining the
original 6 estimates, since a series of single equations is solved rather than a set of
simultaneous equations. The process leading to Eq. (9-83) is termed prewhitening
the input series. The expression 6, '(L)¢,(L) is termed a filter, and the same
filter is applied to both the input and the output series.
There is an analogy between these procedures and the transformation con-
ventionally applied in econometrics. The model (9-1) with various lagged values
of just a single input may be written in the usual matrix form as

y=Xd+u (9-84)
As seen in Chap. 8, a nonspherical variance matrix for the disturbance term leads
to GLS estimation procedures. The GLS procedure is equivalent to premultiply-
ing Eq. (9-84) by a transformation matrix T and applying OLS to the transformed
data Ty and TX. The matrix T is chosen according to the assumed properties of
the disturbance term so as to make Tu a white noise series. The time-series
approach concentrates first of all on the properties of y and X in Eq. (9-84) and
not on the nature of u. A common differencing procedure is applied to Y, and X,,
followed by a common filter derived from the ARIMA model fitted to X,. The
D(L) polynomial containing the “long” series of 5 coefficients is finally repre-
sented by the ratio of two low-order polynomials which are estimated along with
an ARIMA model for the disturbance term.
It is impossible to give here a detailed operational description of the time-series
procedures.+ However, it is clear that a considerable amount of “judgment’’ is
required at various stages in choosing between different ARIMA and different
transfer function models. Time-series analysts also stress that long runs of
observations, preferably in excess of 100, are desirable, which requires the

+ Reference should be made to G. E. P. Box and G. M. Jenkins, Time Series Analysis: Forecasting
and Control, revised edition, Holden-Day, San Francisco, 1976, or to G. W. J. Granger and P.
Newbold, Forecasting Economic Time Series, Academic Press, New York, 1977. A lucid introduction
to a wide range of time series topics is provided by C. Chatfield, The Analysis of Time Series: Theory
and Practice, Chapman and Hall, London, 1975.
378 ECONOMETRIC METHODS

assumption that the underlying economic structure has been stable for that length
of time. The estimation of even the single-input case is fairly complicated. The
model is perhaps most appropriate to a “black-box” situation where interest
centers on a single input variable which can be controlled in any desired manner
but the researcher has no clearly articulated theory of the relation between the
input and the output.
The approach described above cannot be simply extended to multiple-input
models, since the covariances between the output and any input are contaminated
by the effects of the other inputs, unless the inputs are orthogonal. Spectral
methods are a possibility, but are not yet well developed for this case. A
somewhat different approach for dealing with two or more inputs has recently
been suggested by Liu and Hanssens.f Their approach is a modification of the
corner method for ARMA identification proposed by Beguin, Gourieroux, and
Monfort.£ Much work is proceeding in this field, and it is too soon to assess the
likely practical significance of the methods currently under development.
A final time-series approach that may be noted for the two-variable case is
the prewhitening of both series. Letting y, and x, denote appropriately differenced
series as usual, a separate ARMA model is fitted to each series, denoted by

$,(L)y, = 9,(L)a,,
(9-85)
and ,(L)x, mT 6.(L)i,,

where a, and #,, denote estimated residuals which are approximately white noise
series. Thus y, is prewhitened by the filter 6, \(L)$,(L) to yield a#,,, and x, is
prewhitened by its filter to yield @,,. It is argued that this approach is useful in
cases where there is doubt about the direction of causation. Does x cause y so that
one expects nonzero correlations between y and earlier values of x, or is it the
other way around, or is there joint causation and feedback? The suggested
procedure is to compute the cross correlations at various lags, positive and
negative, between a, and @,,. Inspection of these correlations should lead to a
decision about causation. For example, if causation is thought to run from x to y,
a transfer function model is estimated for d@,, on a,,,xt say,

i,, = T(L)a,, + noise (9-86)


where the parameters of this transfer function are indicated by

(L)=y+yL+yl?+--:
to emphasize that they are not the original structural coefficients D( L) connecting

7 L. M. Liu and D. M. Hanssens, “Identification of Multiple-Input Transfer Function Models,”


Communications in Statistics, 1982.
£J. M. Beguin, C. Gourieroux, and A. Monfort, “Identification of a Mixed Autoregressive- Moving
Average Process: The Corner Method,” in O. D. Anderson, Ed., Time Series Analysis, North-Holland,
Amsterdam, 1980.
LAGGED VARIABLES 379

y and x. Finally substituting for u,,and a,, from Eq. (9-85) gives

6 '(L)9,(L)y, = P(L)6-(L)$,(L)x,+ noise (9-87)


which is a relation connecting y and x. The resultant estimates of the structural
coefficients are obtained by equating coefficients of the powers of L in

D(L) = 6,(L)$, '\(L)P(L)6"(L)$,(L) (9-88)


This is a rather different procedure than just prewhitening the input, applying
the same filter to the output, and then estimating D(L) directly as in Eq. (9-83).
In principle the D(L) polynomial is recoverable and capable of being estimated
by either method. For example, suppose y, and x, simply follow different AR(1)
processes,
(1-aL)y=u,, and (1 ~ agL)x=,uy
Substituting for y, and x, in y = D(L)x, gives

uy, = (1ae a,L)D(L)( = CES us


which is the implied transfer function between the separate white noise processes.
Substituting now for u,, and u,, gives

(i= aL )y= = yh) DL) = 0,1) (=a, L)x,


which gets us back to

y= D(L)x,
However, the above is in terms of the true coefficients and has also ignored the
noise terms. In practice the bivariate prewhitening approach requires the estima-
tion of more parameters and greater manipulations of those estimated parameters
than does the univariate prewhitening method. It would be interesting to see
comparative case studies of the results yielded by the two approaches, but there
do not yet appear to be any. An extensive application of the bivariate pre-
whitening approach to various time series of money and interest rates yielded “a
surprising, probably disconcerting, lack of relationship among several variables.” +
Pierce’s main conclusion was, “Extensions of time series modeling procedures of
Box and Jenkins reveal that numerous economic variables which are generally
regarded as being strongly interrelated may with equal validity, based on recent
empirical evidence, be regarded as independent or only weakly related.” A further
study by Haugh and Box illustrated the same approach to the study of the
connection between the GNP X and the unemployment rate Y in the United
Kingdom.} Each series was first differenced to yield x, = X,— X,_, and y, =

7 D. A. Pierce, “Relationships—and the Lack Thereof—Between Economic Time Series, with


Special Reference to Money and Interest Rates,” Journal of the American Statistical Association, vol.
72, 1977, pp. 11-22.
+L. D. Haugh and G. E. P. Box, “Identification of Dynamic Regression (Distributed Lag) Models
Connecting Two Time Series,” Journal of the American Statistical Association, vol. 72, 1977, pp.
121-130.
380
21GB. 7-6 sor) suoyejess09
28 L(1)
BeT (y) SI-— tv1— €l— Z1- Il Ol 6 8 L 9 ¢ €- 7 I-
et 610-
+ 900 910- L770 p00 so0o- 100 970 FrI0 O10 £00 3800 ZWO- 910- 610—

3eT -(7)
0 I Z € v ¢ 9 iB 8 6 Ol Il ral €I vl SI
vd 6£€0- v70- 970- €00 €00- 600 910- 810 +270 3800 600 100 670 810 OK) 70-
4nXn

“TL ‘q y8nezyy
pure‘°O “q ‘gq ‘xog UONROYNUEPT,,
Jo omMeUXG Uorssa1Zay painquysiq)
(eT sfopoy] SUNIBUUOD
OM], PWT], ,“SITI2G DUNO,
fo ay]
uDIIUaU JDIIISIIDIS ‘UONDIIOSSP
JOA ‘ZL ALLO“LT
LAGGED VARIABLES 381

Table 9-3 Summary statistics for r;


Sa

Number of coefficients exceeding Mean abéolyte


ee akRs MGS SE et SEL) er
Two standard errors One standard error coefficient
2 ws SE
EE eee
Negative lags 2 7 ; 0.13
Positive lags 2 8 0.14
ee

Y, — Y,_,. The separate ARMA models were estimated from 56 deseasonalized


quarterly observations as
(L063)y, Uy,

and x=. 0°60 44 xt

The first concern was the direction of causation. Table 9-2 shows the various
lagged cross correlations. A positive lag is here defined as the y series lagging
behind the x series. The asymptotic standard error for r is 0.13. Table 9-3 presents
three summary statistics computed by the author from the data in Table 9-2. The
data in these two tables hardly seem to give any clear indication of the direction
of causation. However the authors of the paper state, “It is concluded that any
feedback effect is of secondary importance, as evidenced by the small cross
correlations at negative lags... . This direction of causation from x to y agrees
with that considered by Bray.”+ Another time series analyst might well interpret
these cross correlations differently and fit a different transfer function to the
residuals.
There is as yet no clear consensus on the relative roles of time-series
techniques and the more orthodox econometric methods. Some mistakenly view
them as competitive rather than complementary. Each is still an “art,” as distinct
from a “science,” in that time-series practitioners have to make various subjective
judgments in the course of their analyses just as econometricians conventionally
“choose” between different regressions and specifications. Investigators with
strong prior beliefs can usually see “patterns” in the data that may be invisible to
more sceptical colleagues.

PROBLEMS

9-1 Deduce the 6 coefficients implied for D(L) = B(L)/A(L) where


ACL) == 6,L.= a,f7
B(L) = Bo + BL
Derive an expression for the mean lag in this process.

+ Haugh and G. E. P. Box, op. cit., p. 127.


£ Years ago in Ireland “reading the tea cups” was a favorite social pastime before the advent of
the ubiquitous tea bag. On draining the tea cup the haphazard pattern of the remaining leaves could
be interpreted by the skilled “reader” as full of meaning and significance. Nowadays a different class
of professionals apply similar “skills” to the interpretation of computer printouts.
382 ECONOMETRIC METHODS

9-2 A model is specified as


YooYout ea, ott
Une Obra

with
e ~ N(0, 0/1)
The 6 parameter is estimated by § = SYdeo 1/2 ).= jasnowe
ge nt OCbetoa): a
(a) plim d = 8 + +55,
286 where $ eee

(b)plim("54? = 0,[1 + a(a — a*)]


where

MS (1 — 3°) . =a ‘
OY,
7 1+ 26¢ ond ef : ed
9-3 Show that a second-degree approximation for the Almon lag implies the restrictions

On ~e 36, oe 38, Fa6, — 0

Oh = Buh ar haw = hag = 0


and hence verify the R, matrix shown in Eq. (9-37).
9-4 Verify that the information matrix for the model of Eq. (9-44) is given by Eq. (9-46).
9-5 Prove Eq. (9-56). [Hint: Use the lag operator to express Eq. (9-47) as

Y, = constant + B,(X, + B3X,-; + B3X,-2 + ---) + (u, + B3u,_, + B3u,-2 + ---)]


9-6 Derive the information matrix (9-67). [Hint: Write e, in the alternative forms

6 = — Bi = py Boy —1Ba ey

&— Uy Pur)
where

Vee pen Lee xX, = X,— pX;,_| by = Bi Pak Ps tp]


and then find all the second-order partial derivatives of
n n l
Ing — 3 In(277) - 5In 0, - Fo2 et

9-7 Derive the information matrix (9-72) and show also that in this model the estimator of 0, is
asymptotically independent of the remaining estimators.
9-8 Consider X, = 2 X,_, + e, where (e,) is a white noise series. Draw some sets of e’s from a table of
random normal deviates and compute the corresponding sample realizations of the process for
t= 1,..., 10, starting each realization off by setting Xy = 0. Satisfy yourself that X, can become
“very large” in both positive and negative directions.
9-9 If u,=(1-0,L—0,L? —--- — 6,L7)e, and {e,) is white noise, derive the autocorrelation
function of the {uw,) series.
9-10 In the rational lag equation
3L
iia ee
L092 20274
determine:
(a) The total multiplier
(b) The mean lag
(c) The coefficients of Xp forj = 0,1, 2,3.
(UL, 1980)
LAGGED VARIABLES 383

9-11 The model generating the (y,) series is presumed to be


Vet Osyee atone: ja| <1
with
UN—SpUy— yb > je| <1
and the e’s independently and identically distributed with zero mean and constant variance 02. Show
that
n oe
p= re 1
Tete.
La 24i— 1

where u, = y, — ay,_, and a is the OLS estimator of a, is not a consistent estimator of p.


(UL, 1973)
9-12 A simple stochastic version of the permanent income model is
Ze panies
where z, is observed income, X, 1S permanent income which is unobserved, and u , 1S a serially random
transitory element. Let x, evolve according to x, = x,_, + v,. Assume u and v are independent.
(a) What constraints does this model place on the autocorrelation function of (z, — z,_ 1)?
(b) Discuss how you would estimate o? and o2 from data on z.
(c) What would positive sample autocorrelation at lag 1 in z,—z,_, suggest about the
plausibility of the model?
Suppose now that u and v are not assumed to be independent but have covariance a,,,:
(d) Are the parameters 02, «2, and a,,, identified?
(e) What would positive autocorrelation in z, — z,_, imply about the sign of o,,,? About the
magnitude of a, relative to o2?
(University of Washington, 1980)
CHAPTER

TEN
A SMORGASBORD OF FURTHER TOPICS

Chaps. 5 to 9 have presented the “standard fare” of the single-equation linear


model. This chapter outlines a number of additional topics, some of which are
“golden oldies” that have been around for some time, while others have come into
prominence more recently. Some enthusiasts may wish to study all the topics;
other readers may be interested in some topics but not in others. To a large degree
the sections stand alone and can be read independently.

10-1 RECURSIVE RESIDUALS

As shown in Chap. 5, the vector of OLS residuals is given by


e = Mu
where
M = I — X(X’X) _‘X’
which is a symmetric idempotent matrix of rank n — k. If the u’s are indepen-
dently and identically distributed, it then follows that
E(ee’) = 02M
Thus the calculated residuals will, in general, display heteroscedasticity and
nonzero covariances, even when homoscedasticity and zero covariances hold for
the true disturbances. This leads to the difficulties in testing for heteroscedasticity
and autocorrelation already discussed in Secs. 8-4 and 8-5.
384
A SMORGASBORD OF FURTHER TOPICS 385

Recursive residuals are a set of residuals which, if the disturbances are


independently and identically distributed, will themselves be independently and
identically distributed, thus greatly facilitating tests of the null hypothesis.+ We
postulate the usual linear model
y=Xb+u
(10-1)
with u ~ N(0, 071)
and X a nonstochastic matrix of order n x k. Let x ; denote the k Xx | vector of
observations on the k explanatory variables at sample point j.& Thus
/


/
xX, ——

Let X,_, denote the (r — 1) X k matrix consisting of the first r — 1 rows of X.


Provided r — 1 = k, this matrix may be used to estimate B. Denote the resultant
estimator by b,_,, that is,
, —ly,
b,_ a (X p=} X,_1) eye

where y,_, denotes the subvector consisting of the first r — 1 elements of y. Using
b,_,;
r one may “forecast” y, at sample point r, corresponding to the vector x, of
explanatory variables at that point. The forecast error is
ae xpTae
and, as shown in Sec. 5-4, the variance of this forecast error is
, , al
o7(1 a (XC Xs) x,)

Define the recursive residual w, as


= yy" b
A
We = ee ee (10-2)
ve fs x (ex |

Clearly, under assumption (10-1)


w, ~ N(0, 07)
since it is a linear function of ncrmal variables and the OLS forecast is unbiased.
A sequence of recursive residuals may be generated as follows.

1. Choose a base of k observations. For the moment let this be the first k
observations in the sample, whether it be composed of time-series or cross-
+ Recursive residuals are a member of the general class of LUS residuals (linear unbiased with a
scalar variance matrix). Another important set is the BLUS residuals due to Theil. See H. Theil,
Principles of Econometrics, Wiley, New York, 1971, Chap. 5.
+ As is customary, the first element in each x vector will be unity to accommodate the intercept
term.
386 ECONOMETRIC METHODS

section data. Compute the vector b, and the recursive residual

Vari — Xs 1D,
Weil >
Vl 2 Kan (XX) Uae

2. Update, or extend, the base to include the first k + 1 observations; compute


b,,, and hence w, ,5.
3. Repeat step 2, adding one new observation point at a time.

There is thus a sequence of n — k recursive residuals as defined in Eq. (10-2)


for r = k + 1,..., n. The practical importance of recursive residuals is due to the
fact that, under assumption (10-1), the vector of residuals defined by Eq. (10-2) is
multivariate normal with zero mean vector and scalar variance matrix, that is,

w ~ N(0,07I,_;) (10-3)
Since we have already seen that each w, is normal with zero mean and variance
o*, the proof of Eq. (10-3) just requires the establishment of zero covariances. The
numerator in Eq. (10-2) may be written

vr xb, | er x (X/OHK ee Nee


Thus
E{(y, 2 xb, 1 )(y, 7: x’,b,_ ,)}

= E{|u, es Ke (X#up Xan (oka toeee wi Ke X ane NG


(10-4)

We may assume that r < s without any loss of generality. Thus

E(u,u,) = 0

E(u,u,)

E(u,u,)
E(u,_,u,) = fe =0

E(u,_,u,)

E(u,w,¢)/=1[0 7)conOuie a0F a0)


J
rthposition

a
Wy
E(u,_w,_,) = E [std ithe eat paras 27)
a

ea ipee 0: ey)
A SMORGASBORD OF FURTHER TOPICS 387

Multiplying out the right-hand side of Eq. (10-4), remembering that the expecta-
tion of a scalar can also be written as the expectation of the transpose of the
scalar, and using the above results, easily establishes
E(ww,)=0 forallr,s;r#s
and so Eq. (10-3) is proved.
The computation of the recursive residuals might be achieved by using the
conventional OLS formula repeatedly to compute each b vector in the sequence
b,,b,,),--., b,. However, the calculations are simplified by using the following
recursion formulas:

(SEX (KS Keen)ie (X,_.X,_1) XX), (X"


r= Xela
(10-5)
1 + x (X)_ |X yak
and b,= b_, + (X,X,) 'x,(y,
r b,_, —xib,_,) (10-6)
Since

it follows that
XX, = X_|X,_,) + XX,
Eq. (10-5) may then be checked by multiplying the left-hand side by X’.X,, the
right-hand side by X,,_,X,_, + x,x’,, and seeing that both reduce to the identity
matrix.t Relation (10-6) may be simply derived since
(X/X,Jb,= X,y
Be eee,
= XPRpX Dey ax
(XX) ky, Xeb, 7)
Finally, relations (10-5) and (10-6) may be used to derive the following: §
RSS,
= RSS,_, rN + w? r=k+ Vn (10-7)
where
RSS, - (y, aiX,b, )’(y, re X,b, )

These theoretical results on recursive residuals have a number of important


practical applications. First of all they provide an alternative derivation of the test

+ See R. L. Brown, J. Durbin, and J. M. Evans, “Techniques for Testing the Constancy of
Regression Relationships over Time,” Journal of the Royal Statistical Society, ser. B, vol. 37, 1975, pp.
149-192, for a statement of these formulas and some notes on their history. A useful survey of
recursion formulas for various models is to be found in W. C. Riddell, “Recursive Estimation
Algorithms for Economic Research,” Annals of Economic and Social Measurement, vol. 4, 1975, pp.
397-406.
+ See Problems 10-1 and 10-2.
§ See Problem 10-3.
388 ECONOMETRIC METHODS

for structural change in the case where the second sample contains fewer than k
observations.} Based only on a heuristic proof, it was asserted in Eq. (6-27) that
under the hypothesis of no structural change
- (ee, — ee) /n»
—F ny, 1 =i)
eie,/(n, ak)
where e,e, denotes the residual sum of squares from a regression fitted to all
n, + n, observations and e'e, is the residual sum of squares from a regression
fitted to the first n, observations. From Eq. (10-7) it follows that for a regression
with n observations,
n

RSS ioe
r=k+1

since RSS, = 0, as a regression with k parameters fitted to k observation points


will have zero residuals. Thus
ny
, pees 2.
Ci cian x WwW,
r=k+1
ny t+n
, = 2
Cees = DL Ww,
r=k+1

and so the Fstatistic defined above becomes


+ 2
Ee "Ww, /N,

ea (ier)
Since under the null hypothesis the w, are independently and identically distrib-
uted normal variables, the F statistic is seen to be the ratio of two independent x?
variables, each divided by the appropriate number of degrees of freedom, and so’
it has the F(n,, n, — k) distribution.
A second useful application of recursive residuals lies in testing for hetero-
scedasticity.t If the alternative hypothesis to homoscedasticity is that 0,° varies
with Xjns the procedure would be as follows.

1. Order the data according to the values of X; and choose a base of at least k
points from among the central observations.
2. From that base compute a vector w, of recursive residuals corresponding to
the first m observations, and another vector w, of recursive residuals corre-
sponding to the last m observations.§ Since the smallest feasible base is of size
k, the maximum value of m is (n — k)/2.

+ See A. C. Harvey, “An Alternative Proof and Generalization of a Test for Structural Change,”
The American Statistician, vol. 30, 1976, pp. 122-123.
$¢A. C. Harvey and G. D. A. Phillips, “A Comparison of the Power of Some Tests for
Heteroscedasticity in the General Linear Model,” Journal of Econometrics, vol. 2, 1974, pp. 307-316.
§ Notice that there is no problem in computing recursive residuals backward or forward in a
sample from any suitably chosen base, or indeed in adding “new” observations in any order.
A SMORGASBORD OF FURTHER TOPICS 389

3. Under the null hypothesis it follows directly from the properties of recursive
residuals that the test statistic

F= WW, ~ F(m,m)
of (10-8)
ww)

Some sampling experiments by Harvey and Phillips indicate that the power of the
test in Eq. (10-8) compares favorably with that of the Goldfeld-Quandt test
described in Sec. 8-4. They recommend setting m at approximately n/3. An
advantage of the recursive residuals test over that of Goldfeld and Quandt is the
greater flexibility of the former. If, for example, one now wished to test whether
o* varies with some other variable X;, one could simply regroup the existing
recursive residuals according to low and high values of X, and compute Eq. (10-8)
afresh, whereas the Goldfeld-Quandt test would require the computation of two
new regressions.
A third application of recursive residuals is in testing for autocorrelation.} In
a time-series application one may take the first k observations as the base. From
the resultant n — k recursive residuals the conventional von Neumann ratio ist

& = va Kk+2(% = w,_1) /(n ete 1)

s* Trees (%, — WY /(n — k)


where w = L_,,,w,/(n — k). This is the ratio of the mean-square successive
difference to the variance. An exact test against serial correlation would be
provided by referring the calculated value of 5*/s? to the significance points of
the von Neumann ratio.§ These critical values, however, were derived for the
general case where the expected value of the series being tested is some unknown
constant. In this application the w’s are known to have zero mean. Incorporating
this information, Press and Brooks have computed significance points for a
modified von Neumann ratiof

(10-9)
Ss
z ei (ate)
These points are tabulated in App. B-7. The von Neumann ratio is arithmetically
closely related to the Durbin-Watson statistic, which could, of course, be com-
puted from the recursive residuals. The crucial point, however, is that the
multivariate normal distribution for w specified in Eq. (10-3) satisfies the assump-
tions underlying the derivation of the von Neumann (Press and Brooks) signifi-

7G. D. A. Phillips and A. C. Harvey, “A Simple Test for Serial Correlation in Regression
Analysis,” Journal of the American Statistical Association, vol. 69, 1974, pp. 935-939.
+J. von Neumann, “Distribution of the Ratio of the Mean Square Successive Difference to the
Variance,” Annals of Mathematical Statistics, vol. 12, 1941, pp. 367-395.
§ B. I. Hart, “Significance Levels for the Ratio of the Mean Square Successive Difference to the
Variance,” Annals of Mathematical Statistics, vol. 13, 1942, pp. 445-447.
4S. J. Press and R. B. Brooks, “Testing for Serial Correlation in Regression,” Report no. 6911,
Center for Mathematical Studies in Business and Economics, University of Chicago, Chicago, 1969.
390 ECONOMETRIC METHODS

cance points so that an exact test is available, thus avoiding the inconclusive zone
associated with the Durbin-Watson statistic calculated from the OLS residuals.
Some sampling experiments by Phillips and Harvey suggest that the power of this
test may be increased by forming the initial base from a mixture of the first and
last observations.
Fourth, recursive residuals provide a test of some possible forms of
misspecification.; Since, under the null hypothesis, the recursive residuals are
independently and identically distributed normal variables with zero expectation,
the mean of the residuals divided by its estimated standard error will follow a ¢
distribution. Formally
Ww
Ge ee ae 1) (i0-10)

where
ae Linkt
homed
and
n wir 2
Le w )
=
tke |
As an illustration of the use of this test in specification analysis suppose the
postulated model is a linear relation between Y and X. If the true relation is
convex (concave) and the data are ordered by the size of X, the recursive residuals
would be expected to be mainly positive (negative) and the computed f statistic
will tend to be large in absolute value. In a multivariate situation this specification
test could still be carried out for any single explanatory variable, if it were thought
that the other explanatory variables were correctly specified, but this type of a
priori knowledge is seldom available. Several specification errors might have a
self-canceling effect on the recursive residuals, so this test is not likely to be very
effective in multivariate situations.
Finally Brown, Durbin, and Evans describet an important application of
recursive residuals in testing for structural change over time. The null hypothesis
of no structural change for the model y = XB + wis specified as

Hy) By Bye BB
re =97 =o?

where B, denotes the vector of coefficients ruling in period ¢ and o/ the dis-
turbance variance in that period. It is clear that the null hypothesis would be
violated if the B vectors remained constant but o? varies. This would be the classic

y+ A. C. Harvey and P. Collier, “Testing for Functional Misspecification in Regression Analysis,”


Journal of Econometrics, vol. 6, 1977, pp. 103-119.
+R. L. Brown, J. Durbin, and J. M. Evans, “Techniques for Testing the Constancy of Regression
Relationships over Time,” Journal of the Royal Statistical Society, ser. B, vol. 37, 1975, pp. 149-192.
A SMORGASBORD OF FURTHER TOPICS 391

case of heteroscedasticity, which might be tested by some of the procedures


already outlined. The main concern in problems of structural change, however, is
variation in the B’s.
The authors suggest a pair of tests, namely, the cusum test and the cusum of
Squares test. The first test statistic is the cusum quantity

W.= Viw/é r=k+1,...,0 (10-11)


k+1

where
RSS
67 = ——
Ok
W, is seen to be a cumulative sum, and it should be plotted against r. As long as
the B vectors are constant, E(W,) = 0, but if the B’s change W,, will tend to
diverge from the zero mean value line. For a forward recursion the significance of
the departure of W, from the zero line may be assessed by reference to a pair of
straight lines which pass through the points
{k, tavn—k} and ({n, +3a¥vn—k}
where a is a parameter depending on the significance level a chosen for the test.
The correspondence for some conventional significance levels is
a= 0.01 a = 1.143
a = 0.05 a = 0.948
a = 0.10 a = 0.850
The lines are shown in Fig. 10-1.
The equation of the upper line in Fig. 10-1 may be determined from
Wi Wisk | avn Kk
t—k n—k

Figure 10-1 Cusum plot.


392 ECONOMETRIC METHODS

or
Zatzkh)
= avn—-k + ————
Vik
and the equation of the lower line is given by its negative.
The second test statistic is based on cumulative sums of the squared residuals,
namely,
oe 2

ieee
I eT ae (10-12)
ae
The mean value line giving the expected value of the test statistic under the null
hypothesis 1s
hak
Ee Dacarks
which goes from zero at r = k to unity at r = n. The significance of the departure
of s, from its expected value may be assessed by reference to a pair of lines drawn
parallel to the E(s,) line at a distance cy above and below. Values of cy for
various sample sizes and levels of significance are tabulated in App. B-8. Refer-
ence should be made to the Brown, Durbin, and Evans article for practical
illustrations of the technique and for interpretations of various plots. The basic
idea is that instability of the parameters would be indicated if the plot of W, or s,
crossed the significance lines described above. There is some evidence that the
cusum test is less powerful than the cusum of squares test. Some Monte Carlo
experiments by Garbade also suggest that the latter may not be very powerful in
comparison with tests based on variable parameter models.t However, the
explanatory variable in his experiments was random over time, and it would be
interesting to see if the same result was obtained with an autoregressive explana-
tory variable.

10-2 SPLINE FUNCTIONS

In an interesting study Poirier and Garber examined the determinants of profit


rates in the aerospace industry over the period 1951—1971.4 They were particu-
larly interested in the behavior of profit rates, ceteris paribus, in three distinct
periods, 1951-1954 (Korean war), 1954-1965 (peace), and 1965-1971 (Vietnam
war). To cover the ceteris paribus proviso, they included eleven explanatory
variables, apart from time, in their equation. They treated time by means of spline
functions, and to illustrate the basic idea we will assume that the profit rate has
been adjusted for the effects of the eleven variables and look at the behavior of
the net, or adjusted, profit rate over time. Assuming a linear time trend, the

+ K. Garbade, “Two Methods for Examining the Stability of Regression Coefficients,” Journal of
the American Statistical Association, vol. 72, 1977, pp. 54-63.
t See D. J. Poirier and S. G. Garber, “The Determinants of Aerospace Profit Rates, 1951-1971,”
Southern Economic Journal, vol. 41, 1974, pp. 228-238; or D. J. Poirier, The Econometrics of Structural
Change, North-Holland, Amsterdam, 1974, Chap. 2.
A SMORGASBORD OF FURTHER TOPICS 393

postulated model would be


Period 1 DOr ie dt ie, SG
Period 2 y= a, + Bot + u, Gt p (10-13)
Period 3 Vi ORE Bat ett, b<t
In this example we might take the origin of time to be 1950. Measuring in years
then gives a = 4 (i.e., 1954) and b = 15 (1965). The data might be split into three
distinct subsets and three separate time trends estimated. The result, in general,
would look like Fig. 10-2a. There is nothing in the unrestricted estimation process
to ensure that the functions meet at the join points ¢ = a and ¢ = b. Fig. 10-25
illustrates a /inear spline, or piecewise linear, function, which eliminates instanta-
neous jumps or discontinuities in the function at the join points or knots.
The linear spline function may be fitted in two alternative fashions. One is to
define the following variables:
Wi,=t
5 age i if fd
t—a ifa<t
aes ‘a ii)
= t—b gp
and reparameterize the function as
Y, = a, + b\w, + d,w,, + b,w,, + u, (10-14)
Comparing Eqs. (10-13) and (10-14) it is easy to see that

B, = 6,
B, = 6, + 46, a, =a, — d,a (10-15)
B, = 6, + 6, + 6, a, = a, — 5b

y y
{ t

: a b mae Oi a b ie

(a) (b)

Figure 10-2
394 ECONOMETRIC METHODS

Fitting Eq. (10-14) directly by OLS will yield estimated functions which meet at
the knots, and the estimated a and 8 parameters of those functions can be
determined from Eqs. (10-15). Tests on a’s and B’s imply equivalent tests on the
6’s. Thus testing the significance of 6,(= 8,) is asking whether there is a positive
(or negative) trend in the first period. Testing the significance of 5, is asking
whether the trend slope in the second period differs significantly from that in the
first, and similarly, testing the significance of 6, amounts to asking whether the
trend slope in the third period differs from that in the second. Setting up the null
hypothesis
5, 0
at bs
| ; |
0|
is equivalent to postulating that the B’s and the a’s are the same in all three
periods, that is, that the data may be adequately described bya single linear
trend. This test may be carried out most simply by fitting
y, =at bw, + u,

as the restricted model, the full spline function (10-14) as the unrestricted model,
and calculating the test statistic defined in Eq. (6-8).
An alternative estimation procedure is restricted least squares. Returning to
Eqs. (10-13), the restrictions implied by the join points are
a, + B\a=a,+
B,a
a, + B,b =a, + B3b
which may be set up in the conventional framework as
R B r

ay
B,
1 oe —a 0 Ones “|
0 0 beet o b, -|° (10-16)
a3
B;
Thus the model

ae |
rz,|
an 3 Q,

: eG pel) A feesape By
is! bel ean Y a,
ys aga 2), Polit ar
Y |= ee 10-
¥3 ee ete eon sl | iia ete a,
| |
| oe b+ 1 B,
| |
|
A SMORGASBORD OF FURTHER TOPICS 395

where the empty cells in the data matrix are all zero, is fitted subject to the
restrictions in Eq. (10-16). The appropriate formula is given in Eq. (6-5). The
estimates of the a and 8 parameters will be identical to those derived from the
estimated coefficients of the spline function in Eq. (10-14).+
This simplified example used time as an explanatory variable. The procedure
works equally well for any explanatory variable x with known join points, or
knots, at x,, x,, and so on. A possible disadvantage of the linear spline is that
while the function itself is continuous at the knots, there is a discontinuity or
jump in the first derivative. This may be overcome by the introduction of
quadratic or cubic splines. To illustrate a cubic spline function, suppose we have a
two-variable relation with known knots at x, and x,. Within each subset y is
expressed as a third-degree polynomial in x, namely,

Vo A Bh Bx + Bae aloe Lele,3 (10-18)


where the subsets are defined by

i=] Ks xe

i= 2 My a es

i=3 Xho

The restrictions implied by continuity at the knots are then


2 As 2 3
Caer Pypeaee Pigk ete igX =" Oar Poi
Xe typhy Gee Poa,
2 aie 2 3
>, + BoX_ + Boxy + By3xX5 = 3, + B3,X, + Bs2x5 + B33x;
We further impose continuity of the first derivatives of the cubic spline function,
which implies

By, + 2By2Xq + 3Bi3x7 = Boy + 2ByxXq + 3B 3x2

By, + 2ByX, + 3By3x5 = Bs, + 2B32x, + 3B33x5


In addition, continuity of the second derivatives implies

2B12 + 6B13X4 = 2B) + 6B23x,

2B. + 6By3X, = 2832 + 6B33x,


The cubic spline merely allows discontinuities in the third derivatives at the join
points. Thus the cubic spline may be estimated by fitting Eq. (10-18) and
estimating the twelve parameters subject to the six restrictions set out above.

+ See Problem 10-4.


+See A. Buse and L. Lim, “Cubic Splines as a Special Case of Restricted Least Squares,” Journal
of the American Statistical Association, vol. 72, 1977, pp. 64-72, which develops the restricted least
squares approach; and D. J. Poirier, op. cit., Chap. 3, for an alternative estimation procedure.
396 ECONOMETRIC METHODS

Both previous examples have been in terms of spline functions on a single


explanatory variable. There are now several examples of applications of bilinear
splines, where linear splines are specified for two variables with main effects and
interaction effects at a two-dimensional grid of specified knots.t

10-3 POOLING OF TIME-SERIES AND CROSS-SECTION DATA

In many problems the investigator may have access to observations on the


behavior of a “panel” of decision units at a number of different (and usually
successive) time periods. We will assume there are p distinct decision units or
groups indexed by i = 1,..., p and m successive time periods indexed by ¢ =
1,..., m, giving a total of nm = pm sample points. The variables are denoted by
Y,, = value of the dependent variable for unit / in period ¢
P= le, Pb = eee
X jit = Value of jth explanatory variable for unitiin periodt j= 2,...,k
The linear hypothesis would then be
Y,, = a + B,X;, + By X3;, + a +B, Xia Ui, (10-19)
where, for the moment, we assume a common set of parameters for all units in all
time periods. To illustrate some of the many possible applications of the model
consider some examples.

1. The panel consists of, say, 1000 households whose savings behavior Y,, is
monitored along with various explanatory variables X,,,, such as income,
family size, and composition over a number of time periods.
2. The panel consists of a set of firms, and the object of study is the size and
timing of their investment expenditures Y,, as a function of the group of.
explanatory variables thought to influence investment.
3. The panel might consist of the 50 states of the United States, and the focus of
investigation are the determinants of the unemployment rate Y,, across states
and over time.
4. The panel consists of the OECD countries, and Y,, indicates the per capita
consumption of gasoline in country 7 in year t. The relevant question is
whether the usual economic variables such as income and relative prices can
adequately explain the variation in Y,,.

The most common way of organizing the data in Eq. (10-19) is by decision
units. Thus let
Yi ei
ae F Xx Xyi X31 Xx i :
y; = {= poe mee oe u,= :
No 2im 3im
im
ki
Ues

7 See D. J. Poirier, The Econometrics of Structural Change, North-Holland, Amsterdam, 1974,


Chap. 4.
A SMORGASBORD OF FURTHER TOPICS 397

denote the data and the disturbances relevant to the ith unit. The data may be
“stacked” to form
Y; X, u,
Tele ols ce | (10-20
where y isn X 1, Xisn X (k — 1), and uisn X 1. The model in Eq. (10-19) may
be expressed as
y= [ix] 8 +u (10-21)
where i is an n X 1 vector of units, a is a scalar, andB=(f, 8B; --- B,J’.
A variety of models has been proposed for time-series and cross-section data,
and most have been fitted to some data set or another. These models may all be
derived from Eq. (10-21) by varying the assumptions made about the systematic
part of the equation and/or the assumptions made about the disturbance vector.
A possible taxonomy of models is indicated in Table 10-1. The meaning of
various terms in the table may not be clear at first sight but will become so as the
models are explained.
Model I(a) is perfectly straightforward. The systematic part of Eq. (10-21)
postulates a common intercept and a common set of slope coefficients for all units
at all time periods. The disturbance assumption is
u,, ~ iid(0,02) for alli,t
where iid means independently and identically distributed. Thus there is no serial
correlation in the disturbances for any individual unit, there is no dependence
between the disturbances for different units, either contemporaneous or lagged,
and the disturbance has a constant variance at all points. The appropriate
estimation method is OLS applied to the stacked data of Eqs. (10-20). If, in
addition, the u;, are assumed to be normally distributed, all the finite sample
inference procedures of Chaps. 5 and 6 are valid.
Model I(b) allows a richer specification for the disturbance term. There are,
in fact, several versions of model I(b) depending upon the precise assumptions

Table 10-1 Taxonomy of time-series, cross-section models

Assumptions about
Vector of slope
Intercept coefficients Disturbance term
Model a B Ui,

I(a) Common for all 7,t Common for all i, ¢ E(uv’) = oI,
I(b) Common for all i, ¢ Common for all i, ¢ E(uu’) = V
I(a) Varying over i Common for all i, ¢ Fixed effects model
II(b) Varying over i Common for all 7,t Random effects model
III(a) Varying over i, t Common for all i, ¢ Fixed effects model
III(b) Varying over i, f Common for all i, ¢ Random effects model
IV Varying over i Varying over i E(uu’) = 671 or E(uu’) = V
398 ECONOMETRIC METHODS

made about var(u). Suppose, for instance, one postulates


E(u2)=o6,
it
forallt;i=1,...,p
E(uj,uj,) =4;; for allt andi + /

E(u,,u,,)it4js = 0 for alli,


j, andr +5
These assumptions allow for heteroscedasticity of the disturbance term across
units and for nonzero contemporaneous covariances between the disturbances in
different units but rule out lagged correlations within and between disturbances.
The resultant variance matrix is

9,1, 9,1, a 91,1,

E(uu’) Vee 9,01, OI, Ws 07,1, (10-22)


eee ; a i cu eMosn elrowe orl

The application of GLS to Eq. (10-21) using Eq. (10-22) would now yield the
[Link].e. of B,

by = (XV (X) X’V"“y


The o;;, however, are unknown. They may be estimated by the following proce-
dure.
Fit Eq. (10-21) by OLS and partition the residual vector into the subvectors e;
(i = 1,..., p) relating to decision units. Then calculate
ee,
Teak
Substitution of the s;, in Eq. (10-22) gives a V matrix which may be used to
compute the feasible GLS estimator. The usual inference procedures now apply
asymptotically. Another version of model I(b) could be produced by adding an
assumption of autocorrelated disturbances within each decision unit.
Model II relaxes the assumption of a common intercept but retains the
assumption of a common vector of slope coefficients for all decision units. The
matrix formulation of this model is then

y: ai:Ec
v> i, o 0 X,
ae Lin cons’oop Oe Boe heme aee (10-23)
.
Y,
0 0 In
x,P || 8

or
y=Za+xXBP+u (10-24)

+ See Problem 10-5 and also J. Kmenta, Elements of Econometrics, Macmillan, New York, 1971,
pp. 512-514, for a discussion of this case.
A SMORGASBORD OF FURTHER TOPICS 399

where the definition of Z is obvious from the comparison of Eqs. (10-23) and
(10-24). Define the matrix B as

B = 2(Z'Z) ‘2’
It is easily seen that B is ann X n matrix given by

J, 0 0
B=—|0 J, 0
T7Ua |WR eetcecahtyy
Ge EF Heer
0 oO J,
where

is an m X m matrix consisting entirely of ones. From the definition of J,


Ye

1 Y,
Sl es

Y,
where

iameay
i m ~ ij
1

Thus premultiplication of any n X 1 vector by B will replace each observation for


any decision unit by the sample mean of that variable for the decision unit. If we
then define
P=I1,-B
premultiplication by P will replace the original observations by the deviations
from their unit sample means. It is also clear that P is a symmetric idempotent
matrix, which is orthogonal to Z, that is,
PZ =0
Premultiplication of Eq. (10-24) by P then gives
Py = (PX)B + Pu (10-25)
Thus estimation of the B vector may be achieved by applying OLS to the data
expressed in terms of deviations from group (unit) means. The resultant estimator
is
b = (X’PX) 'X’Py (10-26)
This is, of course, exactly the same vector as results from the application of OLS
to Eq. (10-24). The normal equations are
Z'Za + Z’'Xb = Z’y
X’Za + X’Xb = X’y
400 ECONOMETRIC METHODS

Solving the first equation for a,

a= (Z’Z) '(Z'y — Z’Xb) (10-27)


Substituting in the second normal equation and solving for b gives, after some
manipulation,
b = (X’PX) 'X’Py
as before. From the definition of Z it may be seen that (Z’Z)' is a p X p
diagonal matrix, namely,

(Z'Z)
| =diag{m—! m=! --- m")
and premultiplying an n X 1 vector by Z’ serves to sum elements within each
group. Thus Eq. (10-27) implies

NaNO
Oye ts Vip Yh J, ee (10-28)
This model, which is designated as model II(a), is usually known as the fixed
effects model. The fixed effects are the intercepts a,, one for each group. It is
usually assumed that the u vector in Eq. (10-24) is homoscedastic and nonauto-
correlated so that OLS provides b.l.u.e.’s, though GLS estimators could be
constructed on the lines of model I(b). The b vector of Eq. (10-26) is also
sometimes referred to as the “within” estimator, since it is based on the
within-group deviations (Y,, — Y,) and (X;;, — X;;). Equations (10-21) and (10-24)
have already appeared in Sec. 6-2 on Tests of Structural Change. The exposition
in that section was solely in terms of a time-series application where the “groups”
referred to p different subperiods, not necessarily all of the same length. Equation
(10-21) is a restricted version of Eq. (10-24), and tests of the restrictions may be
made in the context of OLS estimates as in Sec. 6-2, or in the context of GLS
estimates as in Sec. 8-6.
Model II(b) is the random effects, or error component, model. Instead of
assuming a set of given (unknown) constants a,,..., a, for the p groups, a single
intercept a is postulated, and the differential intercepts are merged with the
disturbance term. The model is now formulated as in Eq. (10-21), namely,

y = [ix] 8 +u
but the assumptions about u are
Uy = a; + bj,
where the a; are drawn at random from M(0, 62) and the e,, are drawn at random
from N(0, 07). The a, are now increments (positive or negative) to the common
intercept a. To derive the variance matrix of u we note that for the ith group we
may write
Us Gl as.
A SMORGASBORD OF FURTHER TOPICS 401

It then follows directly that


E(ua=
,)075, #671,a =m f= 1,...,p
0, + 02 02 a.
= oa, o; + a2 a2
oa, oa a, a o,

p p
=o;|p 1 e|=0,7A
p ]
where
62
0, =90,+06,2 and (ise
0,

Since E(u,u’,) = 0,
A 0 0
V=E(u’)=o67/0 A 0
FO tay BHA
= ole @A
The matrix A may also be expressed as

Alto Lia pin


This facilitates finding the inverse. Let

A-'=),J, + Ajl,, (10-29)


where A, and A, are constants to be determined. Multiplying out and noting that
JZ =m4J,,,
AA~' = (1 — p)AaI,, + [(1 — p)A, + mpa, + pag],
Equating the right-hand side to I, gives
a
x Saar Mics1 (10-30)
which, on substitution in Eq. (10-29), gives A~'. The GLS estimator of model
II(a) might then be obtained from Eq. (10-21) using

A7! 0 ae 0

Vir OS Age lise pa (10-31)


0, suis mieasste shTontoMsiTs! a

The difficulty, however, is that V~' involves the unknown oa? and o2. Before
dealing with this problem it may be shown that the GLS estimates can also be
402 ECONOMETRIC METHODS

achieved by applying OLS to suitably transformed variables. From Eq. (10-29) we


may write
Mica r, oes r,
ADs r, AG NS r,
: x, Beige x ee eee i, oe

Letz=[z, z, -:: Z,,]/ denote a column vector of m observations on some


variable.7 It then follows that

z’A~'z =A,(Zz) + A,Lz?


ue [Es2 Seeara
= a, (Xz) : i
using Eq. (10-30)
2
=
=r, 2nd
|2: MP. 52|

The form of this expression suggests defining a quasi deviation as


Geran CL
and asking whether a constant c can be found such that
2
Lz*so = Lz ee Daas
mp =2

Now Lz? = Lz? — (2cm — c*m)z?


Solving for c gives

veer oe
Leper imp

Tee os (10 49)


~ Ve o2 + mo? a)
recalling that p = 0, /(o, + 0,’). It is customary to take the negative sign in Eq.
(10-32), and thus we have

ZA z= > (z,- czy (10-33)


t=]

wheret
j
Coates \aio (10-34)
0; + mo?

+z can thus represent the sample observations on the dependent or an independent variable for
any given unit.
+ This result is stated without proof in J. A. Hausman, “Specification Tests in Econometrics,”
Econometrica, vol. 46, 1978, p. 1262.
A SMORGASBORD OF FURTHER TOPICS 403

The estimation of Eq. (10-21) by GLS using the V~! defined in Eq. (10-31) then
gives

|= Ut, SV Tex te Sy (10-35)


From the definition of V~' in Eq. (10-31) the elements in the matrix and vector
on the right-hand side of Eq. (10-35) are of the form
al
ZA 2;
where z; and z, represent m X 1 vectors. Thus the GLS estimator defined in Eq.
(10-35) is equivalent to applying OLS to the quasi deviations

vir Vier

Dears Ae ON, bai le A Pte ails oot fee ee, oak

Either procedure requires an estimate of the variances appearing in V~! or in c.


These may be developed as follows. The disturbance in this model is

uj, = a; + E,,
Averaging over ¢ for unit i gives
S| I Ri ates I
Averaging then over i gives
Sl RI +€
and the usual decomposition of sums of squares gives
N
Y(u,, * iz)” a Lui i it,)” a X(a = Sl
2
Beak

The resulting analysis of variance is shown in Table 10-2.


Looking first of all at the within-group mean square,
Na Bo
Vitus 4) = Dee. ie)
Pot aie

Recall that if x,, x.,..., x,, are drawn at random from x ~ N(0, o?),+

p{PRa |a
m— | eo

Thus z| Ge a} =(m— 1)o2


t=]

and BYMs Be, = p(m-— 1)o2


1

+ The assumption of normality is not required for this result, only that x ~ iid(0, 02).
404 ECONOMETRIC METHODS

Table 10-2 ANOVA of the disturbance term

Density Expected
Source Sum of squares function Mean square mean square

Between groups- yaa ~ i)’ p-1 rata ani a; + moZ


i,t

2
Within groups D(H - a;)” p(m— 1) a= Q€

Total ¥(44- i)’ pm— |

giving o, as the expected, within-group mean square. Similarly,

¥ (a, - 7) =Y(e,- a) + le é)’ + 2Y (a; — aye —@)


Beat i,t

But |¥ (a- a)'|=(p~ 193


i=]

and e{¥ (4-8'}=(e- 0%


i=] te

since the é, are drawn at random from N(0, o7/m). Thus

AS tne tene
re
= =\2 2 2

and the expected between-group mean square is mo? + 62. The u,, in Table 10-2
are, of course, unobserved, but we can estimate the relevant disturbance variances
by substituting estimated u’s in these formulas.
The estimation procedures may be summarized as follows.

OLS on transformed data

1. Fit the basic model, Eq. (10-21), by OLS and obtain the n X 1 vector & of
OLS residuals. Compute also the mean residual a, for each unit, and note
that
a= 0.
2. Compute

3. Compute the quasi deviations y,, = Y,,— cY,, and so on, and apply OLS.
A SMORGASBORD OF FURTHER TOPICS 405

The direct application of the GLS formula requires estimates of p = 0,/o, and of
o,,. These are also obtained from the OLS residuals, The steps are as follows.

1. As in OLS procedure.
2. Compute

s=2=— 1 _ YF(a,- a)
Peey sy
t= 4,

eeDepea
] Oe
t ares)hehe
A ne aa? Be?

Ss, =s2+s
2
b=3
3. Using 6 and s2, compute V and the GLS estimator defined in Eq. (10-35).

The estimation of the variance component from the OLS residuals is not to be
recommended when lagged values of Y appear in the X matrix. Since p = 02/
(0, + 9,’) is constrained to be in the (0, 1) interval, a grid search over this interval
for the ML estimator is a feasible procedure.
The final question with respect to model II is the choice between fitting either
the fixed effects or the random effects model. The choice basically has to be made
by the researcher based on the institutional realities relevant to the problem being
studied. Returning to the examples given at the beginning of this section, suppose
certain monetary /fiscal policies are set in place in an attempt to reduce unem-
ployment rates across the country, and after some time an analysis is made of the
experience of the various states. As a result of historical developments, the states
have variable mixtures of industrial, commercial, private, and public structures.
One would thus expect differential effects across states, which would be modeled
appropriately by the fixed effects assumption. On the other hand, if we look at the
per capita consumption of gasoline in the OECD countries, we will certainly
observe very different levels of the dependent variable in different countries.
However, it is also true that for tax and other reasons the real price of gasoline
has historically been very different in different countries. For sound economic
reasons this may be expected to have Jong-run effects on the size of automobiles
and on per capita gasoline consumption. Inserting dummy variables to allow
different intercepts across countries removes this variation from the data, and the
“effects” of the explanatory variables are estimated solely from the within

+ See P. Balestra and M. Nerlove, “Pooling Cross-Section and Time Series Data in the Estimation
of a Dynamic Model: The Demand for Natural Gas,” Econometrica, vol. 34, 1966, pp. 585-612; and
G. S. Maddala, “The Use of Variance Components Models in Pooling Cross-Section and Time-Series
Data,” Econometrica, vol. 39, 1971, pp. 341-358.
406 ECONOMETRIC METHODS

estimator, Eq. (10-26), which is based on the within-country variation and is not
influenced by the between-country variation. Thus a fixed effects model would be
liable to underestimate the price elasticity. The random effects model would be
equally inappropriate since it would attribute significant variations in consump-
tion to unidentified stochastic factors rather than to price. In this case a more
sensible estimator of long-run price and income effects would be obtained by
computing the between estimator based on country (group) means. Averaging Eq.
(10-19) over groups gives

Y, =a Bi.X5, + o> teBe r= 1,...,P

This is the same as transforming the original data by premultiplication by the B


matrix defined earlier and computing the OLS estimatort

fp]= (li. XBL, xii, xy (10-36)


The random effects model would seem appropriate when the decision units (say,
households or firms) have been drawn from some population of such units.
Conditional on the explanatory variables, there will be an average level of
response in the population, and individual levels will vary around that average as
a consequence of unidentified stochastic factors.
Looking at the statistic defined in Eq. (10-34), which produces the quasi
deviations underlying the GLS (random effects) model, we see

1. Asm — oo, c > 1, and the GLS (random effects) estimator of B tends to the
fixed effects estimator of B.
2. As a, becomes very large relative to 07, c > 1, and again the random effects
and fixed effects estimators of B will tend to coincide.
3. As o2 — 0, c — 0, and the random effects estimator would tend to the OLS
estimator (X’X) 'X’y.

Returning to the taxonomy in Table 10-1, model III allows the intercept to
vary over units and time periods, while retaining the assumption of a common B
vector for all 7, t. This again may be estimated bya fixed effects or random effects
approach. The former extends Eq. (10-24) to include dummy variables for the
time periods, taking care to use only m — 1 such dummies in order to avoid a
singular data matrix. The random effects model postulates the disturbance to be

Ui Oped fi

where the y’s are assigned at random to the time periods from some postulated
distribution. Just as a, is assumed common to the ith unit for all time periods, so

+ For a lucid and practical discussion of these issues see J. M. Griffin, Energy Conservation in the
OECD: 1980 to 2000. Ballinger, Mass., 1979, Chap. 2.
A SMORGASBORD OF FURTHER TOPICS 407

is y, assumed common to all units in the ¢th time period. The extensions from
model II are relatively straightforward, and we will not go into them here.+
Model IV allows both the intercept and the B vector, or some components of
it, to vary across units. This model has already been studied in Sec. 6-2 under the
simplest possible assumptions about the disturbance term and in Sec. 8-6 in the
context of the SURE model. The random effects version of model IV might also
be extended to allow for time-specific as well as unit-specific error components. In
testing for the stability of the B vector (whether across unit or over time) it is then
especially important to use the procedures of Sec. 8-6 with an appropriately
specified variance matrix for the disturbance.

10-4 VARIABLE-PARAMETER MODELS

This topic has already appeared in several places. Sec. 6-2 on structural change
investigated variations in some or all of the parameters of a relation, but it was
known apriori at which point possible structural breaks might have occurred
(peacetime, wartime, and so on). Section 10-2 on spline functions showed how
different functions might be fitted so as to meet at the known join points. Section
10-3 on time-series and cross-section data considered many possible variations in
parameters, but again, as in Sec. 6-2, there were obvious points at which such
changes might be expected. Only in Sec. 10-1 on recursive residuals was there
some discussion of the case where the B vector might change at unknown points.
We must now consider cases where there is no a priori information on the
observational points at which structural changes might have taken place, and in
this brief section we will consider just two possible approaches. The approach of
switching regressions is based on the assumption that there is a known (small)
number of different regimes, but the switching points are unknown. The other
approach is based on the assumption of continuous parameter variation.

Switching Regimes
The simplest case of switching regimes is based on the assumption of just two
different regimes. The switch may depend on time or on a “threshold” value for
some variable, or it may be triggered stochastically. For instance, wage and price
decisions may be different in periods of low inflation and in periods of high
inflation. The pioneering treatment of switching regimes is due to Quandt.§ To
+ Reference may be made to the articles by Maddala and by Balestra and Nerlove already cited,
and also to T. D. Wallace and A. Hussain, “The Use of Error Component Models in Combining
Cross-Section with Time-Series Data,” Econometrica, vol. 37, 1969, pp. 55-72; and Y. Mundlak, “On
the Pooling of Time-Series and Cross-Section Data,” Econometrica, vol. 46, 1978, pp. 69-86.
£See Problem 10-7 and B. H. Baltagi, “An Experimental Study of Alternative Testing and
Estimation Procedures in a Two-Way Error Component Model,” Journal of Econometrics, vol. 17,
1981, pp. 21-49.
§ See S. M. Goldfeld and R. Quandt, Studies in Nonlinear Estimation, Ballinger, Mass., 1976,
Chap. 1, and references therein.
408 ECONOMETRIC METHODS

illustrate the approach suppose we have ¢ = 1,..., 2 sample observations and the
hypothesis is that
Regime 1: y, = a, + B,x, + u,, holds fort < 1*
Regime 2: y, =a, + B,x, + up, holds fort > ¢*
where ¢* is unknown. Assuming the u’s to be normally and independently
distributed with zero means and variances 0) and 93, the log likelihood is
t* ron t*

In’ => In2m ee tnoee Ino?


1 ED > 1 n 5

pee ye) = Bae 7 DEE oe)


Lora 207 p=r* +1
(10-37)
ML estimates of a@,, 8,, and o7 (i = 1,2) would be given by two separate OLS
regressions for any assumed value of ¢*. On replacing these parameters by their
ML estimates the last two terms in Eq. (10-37) become

26? 26? 2
and so
* ere
Inf= 2
in de — ene
ae 2
ee (10-38)
An estimate of the switch point r* could then be made by evaluating Eq. (10-38)
for all possible values of t* and choosing the one that maximizes the likelihood.
With n sample observations and two variables the possible range for ¢* is from
t* = 3 to * = n — 3, implying the calculation of n — 5 pairs of regressions..
Riddell, however, has recently pointed out that the computational burden is
considerably reduced by making use of recursive residuals.+ Consider the set of
forward recursive residuals w,, w,,... . From Eq. (10-7) we have

RSS, = RSS,_,
+ w?
ie

Thus RSS2= 5)
t=3
and

620) = —
2 RSS.

In a similar fashion 6;(t*) can be constructed from the set of backward recursive
residuals. Thus just two passes of a recursive residuals program will generate all
the data required to find the maximum of Eq. (10-38).

+ W. C. Riddell, “Estimating Switching Regressions: A Computational Note,” Journal of Statistical


Computation and Simulation, vol. 10, 1980, pp. 95-101.
A SMORGASBORD OF FURTHER TOPICS 409

The null hypothesis that no switch occurred may be examined by means of


the likelihood ratio statistic.} Let

pe)
L(&)
where L(&) is the unrestricted maximum of the likelihood function over
the
entire parameter space. In this example it is the antilogarithm of the maximum of
Eq. (10-38) since it is assumed to be known that there is at most one switch point
and the restriction of a single regression (no switch point) has not been imposed.
L(@) is the maximum of the likelihood function over the subspace w C Q2 to
which one is restricted by the hypothesis. In this problem it is the maximum value
of the likelihood for a single regression. Under the hypothesis of no switch

In L(6) = — Finda ~ 5 — sine?


where

é and B being the OLS coefficients. Thus


(ar vaea Nisha
A=
(67)"7
The conditions required for —2 In A to follow an approximate x? distribution are
not fulfilled since the likelihood function is only defined for integral values of 1*.
However, the graph of A (or In A) against ¢ can be instructive, as shown in Brown,
Durbin, and Evans, especially when considered in conjunction with other tests.t
For a discussion of procedures when the switch is triggered in various other
deterministic or stochastic fashions the reader should consult Goldfeld and
Quandt.§ A special case of switching regimes arises in the context of disequi-
librium models where in some periods we have observations on the demand
function and in others on the supply function.

+ For a brief account of likelihood ratio tests see P. G. Hoel, Introduction to Mathematical
Statistics, 4th edition, Wiley, New York, 1971, pp. 211-217.
+R. L. Brown, J. Durbin, and J. M. Evans, “Techniques for Testing the Constancy of Regression
Relationships over Time,” Journal of the Royal Statistical Society, ser. B, vol. 37, 1975, pp. 149-192,
especially p. 161.
§ S. M. Goldfeld and R. Quandt, Studies in Nonlinear Estimation, Ballinger, Mass., 1976.
A treatment of disequilibrium models is beyond the scope of this book. Some important
references are R. C. Fair and D. M. Jaffee, ““Methods of Estimation for Markets in Disequilibrium,”
Econometrica, vol. 40, 1972, pp. 497-514; R. C. Fair and H. H. Kelejian, “Methods of Estimation for
Markets in Disequilibrium: A Further Study,” Econometrica, vol. 42, 1974, pp. 177-190; T. Amemiya,
“A Note on a Fair and Jaffee Model,” Econometrica, vol. 42, 1974, pp. 759-762; G. S. Maddala and
F. D. Nelson, “Maximum Likelihood Methods for Models of Markets in Disequilibrium,”
Econometrica, vol. 42, 1974, pp. 1013-1030; S. M. Goldfeld and R. E. Quandt, “Estimation in a
Disequilibrium Model and the Value of Information,” Journal of Econometrics, vol. 3, 1975, pp.
325-348.
410 ECONOMETRIC METHODS

Continuous Parameter Variation


The capacity of econometric theorists to “invent” new varieties of models with
continuous parameter variation tends to exceed the willingness and sometimes
even the computational ability of researchers to apply them to real-world situa-
tions. We will illustrate two main approaches, namely, random coefficient models
and adaptive regression models, and in each case emphasize one or two major
publications without attempting to give a comprehensive coverage of all recent
theoretical developments.

Random coefficient models. The traditional single-equation model y = XB +


u
puts the ignorance or uncertainty into the disturbance vector u, while the B vector
is assumed to be fixed at all sample points. An alternative assumption is to make
the B vector stochastic and write the model as

os (B, ff 01;) + (B, 45 0 ;)Xp, Grea saatiate (B, a5 On;) Xe; Vo Allee nN

(10-39)
The ’s in Eq. (10-39) are unknown constants common to all sample points.
The
v,; are stochastic variables which determine the coefficient vector for the
jth
sample point. The n sample points might, for example, be a cross
section of
households where important explanatory variables may be unobserved,
and their
influence affects slope coefficients as well as the disturbance term.
The reaction of
mortgage debt to, say, the measured rate of interest may well
depend on the
unobserved age of the head of household. There is no need to
insert the usual
equation disturbance term in Eq. (10-39) since it will merge with
v,,. Equation
(10-39) may be rewritten as

Y=
XiBi Wendel lau (10-40)
where uz = XV,

x LL Seat
y= [ei P.)|
Assumptions about the vy, are required to make the model
operational. A simple
set of assumptions is

E(v,) =0 J=1,...57
areas0 0
E(vv') = 0 a, ae Osli—A alee (10-41)
ed) a,
E(vv/) = 0 f=. nti sy
The stochastic elements in the coefficients
are thus assumed to have zero means
A SMORGASBORD OF FURTHER TOPICS 411

and to be uncorrelated between sample points and also between different coeffi-
cients for any given sample point. The last assumption is possibly the least
plausible. If, for example, age has an effect on the reaction of mortgage debt to
the rate of interest, it may have a related effect on the response to income. These,
however, are the assumptions of the original Hildreth-Houck random coefficient
model.
From Eqs. (10-41) the disturbances in Eq. (10-40) have the following proper-
ties:
E(u;)=0 fecal
E(u;) = E(x'vvix, jae leees ai
= x)AX;
E(u,u;) = 0 LJ Sahat ea

Since A is diagonal, var(u;) simplifies to


k

3; E(u?) a x Xia,
i=]

= Xa (10-42)
Sane
where x; = [1 XD)2 Sa: Xa2

and a’ =[a, Qa eeOl


Thus Eq. (10-40) constitutes a model with a heteroscedastic disturbance term, the
variance at each sample point being the same linear combination of the squares of
the explanatory variables at that point. Collecting the n variances in Eq. (10-42)
gives

where X denotes the matrix obtained from X by squaring each element. The form
of this relation suggests that if estimates of the left-hand vector could be obtained,
a regression on X could yield an estimate of a. Looking at the residuals obtained
from the OLS fit to Eq. (10-40), e = y — Xb, we know from Chap. 5 that

E(ee’) = ME(uu’)M (10-43)

where M =I-— X(X’X) 'X’



+ C. Hildreth and J. P. Houck, “Some Estimates for a Linear Model with Random Coefficients,
Journal of the American Statistical Association, vol. 63, 1968, pp. 584-595.
412 ECONOMETRIC METHODS

It follows from Eq. (10-43) that

Ee?
E(é) =| 2* |= Mo?
Ee?
Thus E(é) = MXa (10-44)
Equation (10-44) leads to the following procedure for constructing a feasible GLS
estimator.

— Fit OLS to Eq. (10-40) and square each residual to obtain the vector é.
2. Regress é on MX, which can be constructed from the original data matrix X,
to obtain an estimated vector &.
~ Substitute & in Eq. (10-42) to obtain estimates o of the variances of the w’s.
4. Using the s? obtain the GLS estimate of B in Eq. (10-40).

As Hildreth and Houck pointed out, there is no constraint on step 2 of this


process that ensures that the @’s are all nonnegative. They suggest setting any
negative @’s to zero. They also suggest a number of other methods of obtaining
consistent estimators of B, but it is difficult to know how to choose between them,
and their small sample properties are unknown. A small sample test of the
significance of the a vector (that is, whether the v’s have nonzero variances) might
be based on the OLS regression of é on MX, but the precise significance levels are
unknown.
The Hildreth-Houck model is applicable to a sample where there is just one
observation per unit. The Swamy model is designed for cross-section
time-series
data.j The data for the ‘th unit or group are modeled by

¥, = SV (Bae,
ctu ale eee (10-45)
There are p separate units with m sample observations on each. The
X, are all of
order m X k and rank k. The B vector of k coefficients is common
to all units. The
¥, vectors model the stochastic variation of the coefficient vector across
units. For
all i, 7 = 1,..., p it is assumed that

LE 0 E(u’) o,,1,, ifi=j/


" lL le
as 0 ifi +j
2. Ev,=0 (10-46)
: A ifi =j
= Bow) ={4 ifi+j
4. v, and u, are independent

+P. A. V. B. Swamy, “Efficient Inference in a Random Coefficient Regression Model,”


Econometrica, vol. 38, 1970, pp. 311-323.
A SMORGASBORD OF FURTHER TOPICS 413

The model would be written in full as

Yi XxX, X 0 0 vy Bi
Y2 X, V> U5
Sele Pree ceo ui 1 Ot ae leila (10-47)
_ ae

Yp Xx, fic lve u,


The disturbance vector is the sum of the last two terms in this expression. Using
the assumptions in Eqs. (10-46), the variance matrix for the composite dis-
turbance term is
NeAKG Pols 0 oe 0
Wie 0 NAX tf 205,1, ose 0
0 0 X, AX’, + 6,,I,,
(10-48)
A feasible GLS estimator of B could be constructed by first obtaining estimates of
A and the o,, in V. These estimates are obtained as follows.

1. Compute the OLS vectors for each unit separately, that is,

b, = (X,X,)'Xiy,
and the vectors of OLS residuals e; = y, — X;b,.
2. An unbiased estimator of o,,; is given by
_ ee;
Ce iran ke
3. An unbiased estimator of A is given by
A 5 Thee “Fi
A =—*2 -— ¥5,,(XX,
Dis 1 P 2d . ( )

where
P ee
SiGe L bb - > ebb!
i=] i=l i=1

4. Substitution in Eq. (10-48) gives V which may then be used to derive the
feasible GLS estimator of B in Eq. (10-47) as
b, = (X’V-'X)
UX’ ly
where y and X denote the stacked vector and matrix in Eq. (10-47). The
estimated variance matrix is
est var(by) = OVX)
and the conventional tests on b, would be valid asymptotically.

Before fitting Eq. (10-47) by the above procedure it is desirable to test


whether the coefficient vectors are truly different across units. Letting B; = B+ y;
414 ECONOMETRIC METHODS

denote the k xX 1 vector of coefficients for the ith unit, we set up the null
hypothesis
Hot BiB = Peak
This hypothesis may be tested by computing the test statistic defined in Eq. (8-91)
for the SURE model, where the = in that formula is the variance matrix of the u’s
defined in Eq. (10-46), line 1. The R matrix would be set up by reformulating the
null hypothesis as
B, = B,

B, = B;

B, = 8,

However, as shown in Sec. 6-1, the same test statistic can be derived from the
residual sums of squares from the restricted and unrestricted versions of the
model. Under the null hypothesis the restricted model is

Yi x, u;

: Mealae
¥; X u

and the unrestricted model is

Yp X, B, u,

The assumptions about the u; in Eq. (10-46), line 1, give

Cr
E(uu’) = +2 @I m

Dp

Thus GLS estimates of each model are achieved by applying OLS


to transformed
variables, where the transformation is to divide the observations
for the ith unit
by /9;; (¢ = 1,..., p). The residual sum of squares from the restric
ted model is
then

/ l / l /
Cree Be fa im ae

where
h l
b = (=|>—xX’xX.
cae $1
. x/x,] y —X’y.
a xy, i
(10-49)
and all summations are over i = 1,..., p. The residu
al sum of squares from the
A SMORGASBORD OF FURTHER TOPICS 415

unrestricted model is
1 1
ee = Lyi cae Lyi Xib, (10-50)
i

where b, = (X)X,) 'X4y, (10-51)


Thus
l 1
e,e, — e’e = L— y;X,b; — L—y,;X,b
9;UW Gj;

= 5 1 (b, - b)Xy,
= 5+(b,- byX;X,b,
Finally it may be shown thatt
P
1 (10-52)
ee, —ee= > a) — b)’X’X,(b, — b)
j=] 0

where b and b, are defined in Eqs. (10-49) and (10-51). If, in addition to the
assumptions already made, the w’s are normally distributed, then under the null
hypothesis,
pylecesunre
ere AEG 6) Kp) as F[k(p— fe1), p(m—k)]
i

This development has, however, used the unknown o,,. Replacing them by the
estimated values s,,, the same test statistic can be computed, but it will now just
have asymptotic validity. This model has been extended to include lagged
variables and more complicated assumptions about the vy, vectors.{

Adaptive regression models. A different form of modeling variable-parameter


schemes is the adaptive regression model associated mainly with the names of
Cooley and Prescott.§ The model is designed for application to time-series data.
We will illustrate the basic idea first of all with reference to the intercept term.

+ See Problem 10-9.


Springer-
+See P. A. V. B. Swamy, Statistical Inference in Random Coefficient Regression Models,
earity in Random
Verlag, New York, 1971; P. A. V. B. Swamy, “Criteria, Constraints and Multicollin
2, 1973, pp.
Coefficient Regression Models,” Annals of Economic and Social Measurement, vol.
Swamy, “Linear Models with Random Coefficient s,” in P. Zarembka, Ed.,
429-450; and P. A. V. B.
Frontiers in Econometrics, Academic Press, New York, 1974.
§T. F. Cooley and E. C. Prescott, “An Adaptive Regression Model,” International Economic
Review, vol. 14, 1973, pp. 364-371; T. F. Cooley and E. C. Prescott, “Tests of an Adaptive Regression
and E. C.
Model,” Review of Economics and Statistics, vol. 55, 1973, pp. 248-256; and T. F. Cooley
A Theory
Prescott, “Systematic (Non-Random) Variation Models: Varying Parameter Regression:
and Some Applications,” Annals of Economic and Social Measuremen t, vol. 2, 1973, pp. 463-474.
416 ECONOMETRIC METHODS

Consider the relation


ire ait Bx, + U, (10-53)
The additive disturbance u, shifts the function up or down period by period.
Cooley and Prescott make the additional assumption that the intercept term is
subject to change according to
Oat fen (10-54)
Assume the u’s and v’s to be independently distributed with zero means and
variances 0,7 and 0, and assume also that u, and v, are independent for all ¢, s.
The model is similar to the conventional regression with fixed parameters and an
autoregressive disturbance process. The difference is that autoregressive shocks
are subject to exponential decay while the effects of the v’s persist. An AR(1)
disturbance process would give
y, = a + BX, + (6, + 98,5 4 pe,5 Fo + p''e, + p'uy)
while the adaptive model gives
yg tf Pat (o,_| TyrU pep
tat ores + 09) + u,
where a, is the intercept in the immediate presample period. One might estimate
the adaptive model in the form just given, in which case the parameters would be
a, 8, 67, and o,. Cooley and Prescott, with an emphasis on forecasting, express
the model in terms of «,,, ,, the intercept in the first postsample period. From Eq.
(10-54)

Oy a Ope Dre:

Thus Eq. (10-53) may be written as


Vr = O41 + Bx, + w, (10-55)
where Ww, =U, — yv,
s=t
Estimation of Eq. (10-55) is simplified by reparameterizing the disturbance
variances as
O71) 0% Non yor) a Osecavae il (10-56)
The larger y, the greater is the importance of the “permanent” compon
ent v in
the shift of the function relative to that of the transient component u. Using
Eqs.
(10-56) and the assumptions previously made about u and v gives the
variance
matrix of w as
E(ww’)
= 07
where

n al Sey) <a
iO 0 em Sines eara
Onna ome Cneee prac BLOM, cel Soe Peg
ie oe oe ; ‘ Sia

1 Loa hese
l
A SMORGASBORD OF FURTHER TOPICS 417

If the w’s and v’s are normally distributed, the log likelihood is
n n 1 1 pis:
InL= —- 5nd = 5 Ino® = 5 In|2| ar
aa a XB) Q iy = XB)

where XB represents the systematic part, a, ,, + BX,, of Eq. (10-55). If y, and


hence Q, were known, the ML estimates of B and o? could be computed from

B = (xQ-'x) 'x'Q-'y
and 6° = (y — XB)'@"'(y — XB)
However, y is unknown, but it is confined to the interval (0, 1), which suggests a
grid search. Substituting B and 6? for B and o? in the log likelihood gives the
concentrated function
In L = constant — sin se 5In|o (10-58)

Maximizing Eq. (10-58) over y yields ¥, which then gives (2. The feasible GLS
estimators are then
b, = (X'0-'x) 'x'O-'y
and s? = (y — Xb,)’07'(y — Xb,)
The asymptotic distribution of by is normal with mean B and variance matrix
o2(X’Q~'X)~!. The asymptotic variance matrix for (y, 0”) is more complicated
and is given in the first of the Cooley-Prescott papers.
The idea of adaptive coefficients can obviously be extended to slopes as well
as intercepts. This is done in the third of the Cooley-Prescott papers. To illustrate
the treatment consider the three-variable model
Bi,
Y, ee [1 X5, X3,] Bo,
Bs,
The assumptions now are
Bi, = Bi + u; (10-59)
Bei eno f= 1.2°3
PAP
it ee try

where the superscript p denotes the permanent part of a coefficient. The Cooley-
Prescott assumptions about u,, and v,, are
~ N(0,(1 — y)o73,
ee Cat iene P=Aeekn (10-60)
v, ~ N(0, yo?S,)

where u, and v, are 3 x 1 vectors. In addition the u, and the vy, are serially
independent, and u, and v, are independent for all s, ¢. The new feature is the
appearance of the 3 X 3 variance matrices >,, and &,,. For the estimation method
to work, these matrices have to be known up to scale factors. Thus they can be
418 ECONOMETRIC METHODS

normalized by setting, say, the element in the top left-hand position to unity.
Writing 2, as
LP OSe rn
r,=|9 o% 9%
0 033 033
implies that u,, is independent of u,, and u;, for all ¢, that the random
components in 8, and £, have variances proportional to 0} and o%4, respectively,
and that these same random components have a covariance proportional to 055. If
one has no reason to expect a nonzero covariance, 044 is set at zero and an
becomes diagonal with only two elements to specify. If one assumes that the
intercept is the only coefficient subject to transitory changes and that the
permanent changes are independent, the matrices become

=,={10 0 0 x, =|9-9%- 0
O00 OP ene
Finally if one assumes the slope coefficients to be constant, the matrices reduce to

jis ie LO)
2,=2,=10 0 0
Om Ose
and Pie Bian
Bf, ae Bite + Ui,
which is simply another way of writing the adaptive intercept case already
studied. The intercept is then B? (= a,), the equation disturbance is
u,,, and
Re
t petty at Dia rs
For the general case of variability in all coefficients Cooley and Prescot
t
Suggest that unless there is special a priori knowledge, one assumes
the matrices
2,, and 2, to be equal. In one practical application the diagonal element
s in this
common matrix were set equal to the estimated sampling varianc
es of the
parameters computed under the assumption of parameter constan
cy. The authors,
however, report that losses in efficiency are surprisingly small,
even for sizable
errors, in specifying 2, and ,.
The general model may now be sketched briefly:
Y= xB feee
eyall
where x, is the k X 1 vector of explanatory variables at time
t, including unity in
the first position to take care of the intercept. The variable-p
arameter assump-
tions are
B=BP+u, BP=BP,+y, ¢t=1,...,n
It then follows that
n+]

Br =p a Ds V,
s=t+]
A SMORGASBORD OF FURTHER TOPICS 419

and so
TXB
aS ,
ie (10-61)
Rael
where wW=x'u,—x, ), Y,
S=tt|

and emphasis is placed on estimating the permanent coefficients for the first
postsample period. The variance matrix for the disturbance term in Eq. (10-61) is
E(ww’) = o?[(1 — y)R + yQ] = 0’ (10-62)
where Ris a diagonal matrix with
ry = Xj2,X;
and Q is defined by
qi; = min(n —i + 1,n—j + 1)x,2,x,
Given =, and &,,, Q depends only on y. Thus a grid search over the (0, 1) interval
will yield a ¥ which in turn gives () and the estimators
b?,, = (XQ7'xX) 'xO-'y
2 e (y a Xb’, ,)'Q7'(y a Xb/?, ,)
and
n

The grid search for 7 is in terms of the concentrated likelihood function Eq.
(10-58), with Q now defined in Eq. (10-62).

10-5 QUALITATIVE DEPENDENT VARIABLES

We saw in Sec. 6-3 on dummy variables that there was no essential difficulty in
the incorporation of qualitative variables in the X matrix. It is, however, quite a
different matter when the dependent variable is qualitative or categorical in
nature. We may distinguish three main cases.

1. Dichotomous, binary, or quantal responses. These can be characterized by a


variable Y which takes on the value one or zero according to which of two
possible results occurs. For example, Y; = | or 0 if individual i dies (lives); if
person i goes to college (does not go to college); if family i goes abroad on
vacation (does not go abroad); and so on.
2. Polytomous responses. This case refers to more than two possible responses.
Thus a family may have
eno vacation
ea vacation in the United States
ea vacation in Europe
ea vacation elsewhere
3. Limited dependent variable. This includes both cases | and 2 as special cases,
but is also more general. One may have a quantitative dependent variable
420 ECONOMETRIC METHODS

which is subject to some limit, whether upper or lower, or both. This is also
referred to as the case of censored, or truncated, variables.

Space forbids a treatment of all three cases. We will concentrate on the basic
ideas underlying the binary case, which are also the foundation for any treatment
of the more complicated cases.

A Single Dichotomous Variable


Suppose the members of a union who have been on strike in a wage dispute are
now being asked to vote on a specific wage increase wy. Let us assume that each
worker has a reservation wage increase and there is some distribution f(w) of this
figure over the population of workers. The response of an individual worker is
denoted by the dichotomous variable
y= ] if the worker accepts the offer
0 if the worker rejects the offer
A worker accepts the offer if it exceeds his or her reservation figure. Thus the
proportion of the population accepting the specific offer wy is given by

t% = [Hoy dw (10-63)
Management would clearly like to know as much as possible about the distribu-
tion f(w). What value of wy, for example, would be required in Eq. (10-63) to
yield a probability in excess of, say, 0.5? If the distribution f(w) remained
constant over a sequence of contracts and various wage increases were subjected
to ballots, estimation of the parameters of f(w) would be a possibility. Alterna-
tively, at a given period in time, one might imagine a government mediator
sampling various groups of workers with a variety of hypothetical wage increases .
in an attempt to chart the f(w) distribution.
The main use of this type of analysis has not been in economics but in
bioassay.f Applications in economics are, however, increasing with the ever
expanding supply of micropanel data. In bioassay a specific dosage z, of, say, a
poison is administered to each member of a population (insect, animal, human).
The responses of the individual members are presumed independent of each other.
For a great variety of reasons the tolerance to the poison varies from individual to
individual and may be described by some distribution f(z). If the tolerance is less
than the dosage, the individual succumbs to the poison. Thus the proportion
of
the population dying at dosage z, is

% = [Ve dz

Finney suggests that the distribution of tolerances is often skew and


approxi-

+ Two basic references are D. J. Finney, Probit Analysis, 3d edition,


Cambridge University Press,
New York, 1971; and D. R. Cox, The Analysis of Binary Data, Methuen
, London, 1970.
A SMORGASBORD OF FURTHER TOPICS 421

mately log normal. Thus a transformation of tolerances (and dosage) by


x=Inz
will render f(x) approximately normal. The dosage-response curve would then be
represented by the cumulative normal distribution as shown in Fig. 10-3, where
x ~ N(u, 07). At dosage x, the proportion dying is read off from the curve as 7,
at x, the proportion is 7, and so forth. The practical problem is now the
estimation of » and o7. Suppose, to this end, an experimenter selects a set of
dosages x,, X7,..., X,. The ith dose is administered to n, individuals and the
proportion p, dying is measured. On the assumption of a normal distribution for
the tolerances the sample proportions will be scattered around the cumulative
curve in Fig. 10-3. The use of p, to estimate » and o* is difficult since p is a
nonlinear function of x. The probit transformation linearizes the relationship and
makes the estimation of p and o° relatively straightforward. First define
x—h
oO
Thus y ~ N(0,
1)
and any dosage x, can also be expressed in terms of y. The probability of death
with dosage x, 1s now given by
To = F(y) (10-64)

where F(-) is the cumulative standard normal distribution and yy = (Xo — )/9.
This is shown in Fig. 10-4, which is simply a repeat of Fig. 10-3 with the
horizontal axis translated to y. Inverting Eq. (10-64) gives

(10-65)

PX

Figure 10-3
422 ECONOMETRIC METHODS

™ = Fo) —--———s :

Figure 10-4

Given a value of y, one can read off the corresponding 7. Conversely, given 77, one
can read off the corresponding value of y. The y variable is defined as the normal
equivalent deviate (n.e.d.) or by the somewhat unattractive term “normit.” A
probit is defined as
Probit = y + 5
From Eq. (10-65) there is an exact linear relationship between the n.e.d. and
dosage or, equivalently, between probit and dosage. The n.e.d. will be negative
whenever 7 < 0.5, whereas the probit will almost never be negative.f Fisher and
Yates give a table transforming percentages to probits.+
In a typical experiment dosages X 1, Xz,..., X, are administered to n,,n5,...,
n, Subjects, respectively. The resultant proportions P\, P2,--+, Pg are measured.
The estimation procedure then follows directly from Eq. (10-65).

1. Convert the sample proportions P\> Pz». Pz into n.e.d.’s and plot against
dosage x.
2. If the scatter in step 1 is approximately linear, then fit the regression
NED = a+ bx (10-66)
where§
=p
a = estimate of ——
oO

:
I estimate 1
of —
oO

+ The number 5 was chosen to eliminate negative probits, since negative


standard normal deviates
with absolute values approaching 5 will almost never be found. As
Finney explains, “At a time when
most biologists lacked even simple calculating machines, and
many had little skill in Statistical
arithmetic, avoidance of negative quantities was an appreciable practical
advantage.” D. J. Finney,
op. cit., p. 23.
+R. A. Fisher and F. Yates, Statistical Tables for Biological, Agricult
ural and Medical Research, 6th
edition, Oliver and Boyd, Edinburgh, 1963, Table IX,
§ If probits are used instead of n.e.d.’s, a is an estimate of 5 — b/o.
A SMORGASBORD OF FURTHER TOPICS 423

A simple OLS regression would be unbiased but inefficient since it ignores the
properties of the error structure. A GLS estimator may be obtained by taking
account of the likely nature of the errors. Write the sample proportions as
Dp, a8; Vices £

Thus}

je binomial ;5
n(1— 7
t

The exact relationship is

=I
F-'(a,) ne te
Seen as Naess,

The computed n.e.d.’s are given by

F-'(p,) = F(a, + €;)


Applying a first-order Taylor expansion

F-"(p,) = Fn) +2, 4


Fim

ap; Pi=7;

Thus
ax Le 1
Ep Vas als ea, (10-67)
GG o&

dF!
where ie
ap; Dim

Returning to
ae | 2
Pi == F(y,)j= beac
See dy

aia ae deal 7, ordinate of the standard normal curve at y,


Oe 22
Thust
di al
dp; Z;

+ The binomial distribution applies since each individual in the ith group is subjected to dosage x;
and hence to a probability 7, of death or whatever. Moreover, individual responses are assumed to be
independent of one another.
+ p = F(y) is a monotonic function, and so is its inverse. We have
dF dF!
dp =—eo dy and ly EE Ip
dy = ——d

Thus

aE ay os A
dp lp dF /dy
424 ECONOMETRIC METHODS

ane at (l—q.
7
eee (10-68)
The regression equation (10-67) thus has a heteroscedastic disturbance given by
Eq. (10-68). Feasible GLS estimators would be achieved by computing a weighted
regression of the empirical n.e.d.’s on dosage x using n,;Z?/p,(1 — p,) as weights.
The next extension to consider is where the stimulus or dosage is not a single
variable but some linear combination of variables. Thus the ith level of the
stimulus might be denoted byt

= eeD
where x; is a column vector of k variables and B is a k X 1 vector of coefficients
presumed constant over all individuals. For example, in the question of whether
or not to purchase a new car in a given year the x vector would include such
variables as income, the relative prices of cars and gasoline, the age of the present
car, and so forth. We still assume that each individual has a threshold level for car
purchase, and we postulate a distribution f(s) over the population, where s
indicates the threshold or minimum stimulus required to trigger a new car
purchase. Thus the probability of a car purchase at stimulus level 5; 1S

7, = ff) ds
If the f(s) distribution were normal with mean p and variance o7, then

7,= F|
=}
Sit:

where F(-) again indicates the cumulative standard normal distribution. The
observed sample proportions p, are transformed into n.e.d.’s, and the appropriate
regression is

ae FED | = F~'(;) + u;

or y, = i
Ss. =
+agyeas
U, LL= — og
x’B

The relationship actually estimated is then

y= Xi Pear a ian ee ee (10-69)


where
h
B *
Zi[(B, — 2)
=>=—
Brame Bel)
Since p, is still a binomial variable with mean 7, and variance 7,(1
— 7,)/n,, the
disturbance term u, will have the same properties as above.
Thus GLS may be
applied to Eq. (10-69) with the correction for heteroscedasticity
implied by Eq.
(10-68).

+ We are now using s (rather than x) to indicate stimulu


s or dosage, since we wish to use x; to
indicate a vector of explanatory variables, in conformity with
the notation in regression analysis.
A SMORGASBORD OF FURTHER TOPICS 425

An obvious problem with the application of this model in economics is the


difficulty of ensuring that n, individuals are subjected to a given stimulus x’,B. The
B vector, of course, is unknown, but an appropriate method available with large
data sets is to classify units into subsets with given values of explanatory variables
such as income, age of car, and'so on. A second problem is that we may have little
justification for the normality assumption underlying the n.e.d. or probit ap-
proach. This may be explored by making different assumptions about the rela-
tionship between the probabilities 7, and the stimulus level s; = x’,B.
The simplest alternative assumption is that of a linear relationship, namely,

7, = x'B
or p= x Btu;
If this is estimated by OLS, or by GLS taking account of the heteroscedasticity in
u, it may give a reasonable fit to “middle-range” data, but it is doomed to run
into difficulty for extreme values of x’,B since there is nothing in either procedure
to prevent estimated probabilities turning out to be negative or in excess of unity.
The more common and more sensible procedure is to model the probabilities
a, by some distribution function other than the cumulative normal. Perhaps the
most frequently used is the Jogistic.} This may be formulated as
Xj
a, = partgaUsoabona ia (10-70)
1+e%8 1 +e 7x8
Clearly, 7 is constrained to the (0, 1) interval. It increases monotonically with the
stimulus x’B, it equals 0.5 when x’B = 0, and it has a shape similar to that of the
cumulative normal.t It is, however, simpler to work with than the cumulative
normal.
It follows directly from Eq. (10-70) that
T.
| =X;xB (10-71)
10-71

that is, the logarithm of the odds ratio or /ogit is an exact linear function of the
x’s. As before, the observed sample proportions p; = 7, + ¢; follow the binomial
distribution
7m a
[Doe binomial, nN:
1

We seek a relationship between the observed logits and the true logits. Letting

fp) = In ap
|
+ The classic reference is D. McFadden, “Conditional Logit Analysis of Qualitative Choice
Chap. 4.
Behavior,” in P. Zarembka, Ed., Frontiers in Econometrics, Academic Press, New York, 1974,
London, 1970, p. 28, Table 2.1.
+ See D. R. Cox, The Analysis of Binary Data, Methuen,
426 ECONOMETRIC METHODS

a first-order Taylor expansion around z7, gives

f( 2) = f(7;,) PIER oa
1 |pj=,

and |of l
=——_
Op; Pi=7, m(1 — 7)
Thus

Vf ]=x’ B+ u, (10-72)
Dep; ‘

h
where a
fae aan

so that
l
E(u;)=0
j= and — var(u,))=——_—_
Sa (10-73 )

The appropriate estimation procedure is then as follows.

1. Compute the observed logits In[ p,/(1 — p;)| from the sample proportions.
2. Carry out a GLS regression of Eq. (10-72) using the disturbance variances
obtained from Eq. (10-73) by replacing the unknown m1, by p,.

So far in both the probit and the logit approaches we have assumed that there
were several observations at each level of the stimulus so that sample proport
ions
could be computed. In some cases this may be infeasible and we Just have a
single
observation, y = 1 or y = 0, at each xB. The scatter would then look like
Figs
10-5.

f ee

1 2
Z
ye
S
Ov Ko?
a oS
7 oS
YES
eo
x
4
4
Yi
4
Ze
Vz
7
ee
—©-_-©-@ Ss @)— Agee > x8
ya

Figure 10-5
A SMORGASBORD OF FURTHER TOPICS 427

Fitting a linear regression of y on x’B is unlikely to approximate the true


probabilities over the middle range and gives nonsense results at the extremities.
The logistic assumption, however, allows the derivation of a fairly simple ML
estimator, which does not violate the constraints on the probability number.
For each of n individuals in the sample we now observe a k X | vector x; of
stimulus variables and a response variable y, (i = 1,..., ). The scalar stimulus
experienced by an individual is given by s; = x’,B, and y, is a dummy variable,
taking the value unity when a response is observed and the value zero when there
is no response. The probability of a response is assumed to be logistic, that is,

n= Pr(y, = 1) =
Ss.
e I

ee ‘ (10-74)
and 1 — a, = Pr(y, = 0) eae
Suppose that r responses and n — r nonresponses occur in a sample. Let us
reorder the sample observations so that the responses come first and the nonre-
sponses last. The log likelihood is then

int—»> Inn >) inl)


i=1 i=rt+l
From Eqs. (10-74)
Oln 7. ei
a Es .=(1—q. :
op ( 1 Bie aS ( 17;)X;

d dln(1 — 7, ) eo . a
oF ap leer! ot
eww“ - tS —_— ._ => —-T7: 3

Thus
dln L
a = y TX;
@ i=r+l

Mm
v
M-~ sas en
I i=1

The ML estimates of B must then satisfy the equation


Te n

wx = Do ax, (10-75)
al a
The left-hand side is the sum of the x vectors just for the individuals displaying a
response. The right-hand side is nonlinear in B, and an iterative nonlinear
program is required for the estimation of B. The asymptotic standard errors may
be obtained as follows. From Eggs. (10-74)
On,
i=" e*
Os; (1 +e%)
= (1-7)
Thus
Oa SC 7, )X,
428 ECONOMETRIC METHODS

and
2 n n
Ga
dB de
op’ i ae ay ts
2B ~ $3 AG os 7;)X ;X’;
i=] i=]

If the x’s are treated as nonstochastic, the information matrix is


n

R(B) = my 7, (1 — 7, )X ;X,
i=]

and the asymptotic variance matrix is R~ '(B).+

10-6 ERRORS IN VARIABLES

So far we have implicitly assumed that the X variables have been measured
without error and that the only form of error in the equation has been in the
disturbance term u. The latter has generally been thought of as representing the
influence of various explanatory variables that have not actually been included in
the relation. It could, of course, also have a component representing measurement
error in the dependent variable Y, and the previous results would still be valid.
We now pose the question of what happens if the X variables are subject to
measurement error. We assume that the B vector represents the coefficients of the
correctly measured X variables. Thus the model is assumed to be

y=XBP+u (10-76)
where X is the n X k matrix of the true (but unobserved) values of the explana-
tory variables. The matrix of observed values is

X=X+V (10-77) ©
where V is the n X k matrix of measurement errors. If some variables are
measured without error, the appropriate columns of V are zero vectors. Combin-
ing Eqs. (10-76) and (10-77) gives the following relation between the observed
variables:
y = XB + (u — VB) (10-78)
The OLS estimator of B in Eq. (10-78) is then

b =B + (X’X) 'X’(u— Vp)


Conventional assumptions about the error terms are as follows.

} For extensions of the material in this section the reader should


refer to the works of Cox and
Finney already cited; also to M. Nerlove and S. J. Press, Univaria
te and Multivariate Log-Linear and
Logistic Models, Rand Corporation, R-1306-EDA/NIH, 1973;
G. G. Judge et al., The Theory and
Practice of Econometrics, Wiley, New York, 1980, Chap. 14; and
T. Amemya, “Qualitative Response
Models: A Survey,” Journal of Economic Literature, vol. 19, 1981, pp.
1483-1536.
A SMORGASBORD OF FURTHER TOPICS 429

1. The
~
measurement errors in X are uncorrelated in the limit with the true values
X. Thus
Aas
plim(—-Xv] =0

and so

plim(—-x’x) = plim(—-X°X] + plim|“vy


n n n
=2+Q0
2. The equation disturbance (plus any measurement error in Y) is uncorrelated
in the limit both with X and V, that is,

plim(—-V'u) =0 and plim( Xu) =0

With these assumptions


plimb = B — (= + Q) ‘QB (10-79)
and so OLS estimates are inconsistent. The inconsistency is due to the correlation
between the data matrix X and the composite disturbance term (u — VB) in Eq.
(10-78).
As an illustration of the result consider the two-variable model
Y=a
+ px, pu,
where Keke,
Then

1 “Di,
z= plim{—X’X} = plim 1 i
—yX, —LX?
n

1 p
eer ae
where p and o7 denote, respectively, the mean and the variance of X. Further
1 0 0
sate villi) anes 1
Q plim|; Vvv] plim 0 “Zo?

mAliOey.0
DynOiviNor

since there is no error in the dummy variable for the intercept term. Substitution
in Eq. (10-79) gives
i Aili Get 1 —po,B
aly 5 (OMe fe 0,8
430 ECONOMETRIC METHODS

from which

lim(b)= B - =
ORs
Soe
B
po ae o* +0, l +107 /a7
Errors of measurement in X thus bias the estimate of 8 downward. The per-
centage bias is approximately given by the error variance as a percentage of the
variance of the X values. The estimate of the intercept is also inconsistent, and
this result extends to the multivariate case: even if some explanatory variables are
measured correctly, all coefficients will in general be inconsistent.
The measurement error in the X variables thus poses a possibly serious
estimation problem, and alternative estimators are required. There are two main
types of estimator described in the literature. One is based on instrumental
variables of various kinds and the other on ML methods, buttressed with fairly
strong assumptions about the covariance matrix of the measurement errors.
Before describing the estimators it is worth emphasizing the possibility that in
certain circumstances economic agents may react to the measured values rather
than the true values of economic variables. Firms may base investment decisions
on some extrapolation of national income trends and in so doing will use the
latest national income statistics complete with such errors as they contain. If
decision makers respond to measured data, then the measurement error is
irrelevant and our previous techniques will be valid.

Instrumental Variable Estimators


The IV method requires a matrix Z of variables which are correlated with the true
X but uncorrelated in the limit with the measurement errors V. The IV estimator
is
bry = (ZX) Zy (10-80)
which will then be consistent and have asymptotic variance matrix

asy var(b);y = 02(Z/X) 'Z/Z(X’Z) '


To illustrate some of the instrumental variables that have been suggested,
consider first the two-variable model, which may be written
Y,BX,
=a + (u,+
— Bv,)
Suppose there is an even number of sample observations. Define Z as
1 l ] Poy ]
Z' =
sj heel tae
where the elements in the second row are plus or minus 1 according to whether
the corresponding value of X is above or below the median YX value. Application
of Eq. (10-80) then gives
#3 te we
ay n nX n Y
= nS — n> = ae
by 0 3(% - X) a i)
A SMORGASBORD OF FURTHER TOPICS 431

where x; and X, denote the means of the values above and below the median and
Y, and Y, the means of the corresponding Y values. The estimator of the slope is
Bo
OO eX
Deal
and the intercept is estimated by
aw = x a bX

This procedure amounts to partitioning the data into two subsets by the median
X value and passing astraight line through the mean points (X,, Y,) and (X3, Y,).
If n is odd, one should omit the central observation before beginning the
computations. This estimator was first proposed by Wald.+ Under fairly general
conditions the Wald estimator is consistent but likely to have a large sampling
variance. Bartlett has shown that the efficiency may be increased by dividing the
X values into approximately three equally sized groups, the first containing
the n/3 smallest X values and the third the 7/3 greatest X values.t Omitting the
central n/3 observation, the slope is estimated by

IV
Yi
xX, ae 1
and the intercept as usual by ayy = Y — bX.
Extension of the grouping methods of Wald and Bartlett to more than one
explanatory variable is cumbersome and tedious. A somewhat different IV
estimator suggested by Durbin does not have this drawback.§ The suggestion is to
rank the X values in ascending order and then define the Z matrix as
TESRba 16gray
Pea yaa yee eter
where the second row indicates the rank values of the X’s.{] Substitution in Eq.
(10-80) then gives the estimate of the slope as
ey
by
= alt ale (
10-81
8 )

where y, = Y, — Y and x, = X, — X. The estimate of the intercept turns out to be


YEG XU,
ary = a (10-82)

+A. Wald, “The Fitting of Straight Lines if Both Variables Are Subject to Error,” Annals of
Mathematical Statistics, vol. 11, 1940, pp. 284-300.
+M. S. Bartlett, “Fitting a Straight Line when Both Variables Are Subject to Error,” Biometrics,
vol. 5, 1949, pp. 207-212. It is easily seen that this is equivalent to making the second row in Z’
consist of equal numbers of zeros and plus and minus ones according to the ranks of the X values.
§ J. M. Durbin, “Errors in Variables,” Review of the International Statistical Institute, vol. 22, 1954,
DD yo aoe
4 With this formulation plim((1/n)Z’Z) would not exist as required for the consistency of the IV
estimator. However, if the second row is replaced by 1/n,2/n,..., 1, the condition will be satisfied
and the same estimates as in Eqs. (10-81) and (10-82) will result.
432 ECONOMETRIC METHODS

This procedure can easily be extended to replace additional explanatory variables


by their ranks. Asymptotic standard errors may be estimated by the usual IV
formula. It is likely that the instrumental variables for these grouping schemes
will not be highly correlated with the X variables. Thus the IV estimators will
probably have fairly large standard errors compared with those of OLS, which is
the price that has to be paid for consistency.
To illustrate ML methods, which depend on some specific prior knowledge of
the disturbance variances, consider again the two-variable model

Y=a+ BX +u
pte pet a t=l--yn (10-83)
with X, = X, + v,
where X denotes the observed value and X the true unobserved value. The u term
is an amalgam of the conventional disturbance term and any measurement error
in Y. Thus the model might be written equivalently as an exact relation between
two variables, both subject to error, that is,

Y,=a+ BX,
5 : (10-84)
with eee tas and X, = X, + 0,
The errors u, and v, are assumed to follow normal distributions with the following
properties:

E(u,)=E(v,)=0 Elu?)=02 E(v?)=062 forallz


E(u,ue)
= E(0,0,)=0) 9 sr (10-85)
E(u,v,)=0 foralls,t
Thus the errors are taken to be serially and mutually independent. The relation
between the errors and the true X, Y values depends on the nature of these latter )
variables. We will distinguish two cases.

Case 10-1. X,, X,,..., X, are a set of given numbers. This case has two possible
interpretations. One is that the set of X’s can be held fixed in repeated sampling.
This situation would be of little interest, even in the experimental sciences, for if
the X’s are truly unobservable, how can the experimenter know that they have
been held constant in repeated trials. The more useful interpretation, especially in
the social sciences, is the one treating the X’s as fixed amounts for making
inferences conditional on the set of X’s underlying the sample observations.

Case 10-2. The X’s are random drawings from a normal distribution with mean p
and variance o*. This is hardly a plausible description of the generating mecha-
nism of most economic variables, but this case leads to the simplest estimating
equations and there are interesting parallels between the estimators in the two
cases.
If the X’s are fixed, then so are the Y’s, and the assumptions already made in
Eq. (10-85) would ensure zero covariances between errors and true values.
A SMORGASBORD OF FURTHER TOPICS 433

Specifically
E(X,u,) = E(X,v,) = E(¥,u,) = E(¥,v,)=0 forall (10-86)
If, however, the assumptions of Case 10-2 apply and the X’s and hence the Y’s
are random variables, the conditions in Eq. (10-86) would constitute an additional
set of assumptions.

Estimation of Case 10-2. Given the assumptions listed above, the observed X, Y
values would come from a bivariate normal distribution which is fully determined
by the following five parameters:

E(X)= E(X)= p
E(Y) = E(Y) =a + Bp
var(X) = 07 + 02 (10-87)
var(Y) = of + of = B’o* + 0,
cov( X, Y) = cov( X,Y) = Bo?
The ML estimates of the parameters on the left-hand side of Eqs. (10-87) are
given by the corresponding sample statistics, and we then hope to solve the
resultant equations for estimates of the parameters of the model. The estimating
equations for a, B,... are

m,, = 6° + 62 (10-88)

where the m’s indicate second-order moments of the sample data, that is,

may 5 D(X — XH¥)


and so on. The dilemma with Eqs. (10-88) is that there are six unknowns but only
five equations. Only p is identifiable and estimable. There is no hope of estimating
the other parameters unless additional information can be brought to bear. Three
possible sources of additional information are conventionally considered.

1. Knowledge of o2. It is becoming more common for economic statisticians to


indicate the approximate degree of error in major statistical series. Thus in
some circumstances it may be possible to gauge the probable error in the
explanatory variable and to replace o, by an estimate s;. The third and fifth
equations in Eq. (10-88) then give
m
(eae (10-89)
Myx, —~ Sp
434 ECONOMETRIC METHODS

Thus the sample variance in X is reduced by the estimated error variance


before dividing into the covariance term. If there were zero measurement
error in X, Eq. (10-89) reduces to the slope of the OLS regression of Y on X.
The first and second equations in Eqs. (10-88) give

é= Y—BX (10-90)
2. Knowledge of 62. This is perhaps a less likely situation than prior knowledge
of 0, since «2 incorporates both the measurement error in Y and also the
conventional equation error. If, however, we have a prior estimate s?, the
fourth and fifth equations in Eq. (10-88) yield

p= —— (10-91)

If s;, were zero, this estimate becomes the reciprocal of the slope in the OLS
regression of X on Y.
3. Knowledge of the ratio \ = 62/02. After some manipulation the last three
equations of Eq. (10-88) now give

mB? =) (m,, aa Am,,)B aa Am, ra 0 (10-92)

with roots

(my, —Am,,) + f(m,, — Am)? + 42,


ieSE (10-93)
xy

The sign of ® must be the same as that of m,,. This will be so only if the
numerator of Eq. (10-93) is positive, and that in turn will be so only if the
positive sign before the square root is taken. Thus the estimator is

i (m, =Am,,) + He aa NGe ce 4\m<,


pe (10-94)
xy,

Estimation of Case 10-1. We now assume that there is a set of unknown values
X,, X,,..., X, underlying the sample data, and we wish to make inferences
conditional on this set. We still retain assumptions (10-84) and (10- 85). The log
likelihood function is

=
In L = constant — 7n ino, - 7n ino, Be es
X (x, — X,)
=, \2

~, \2
= LS) fea eae (10-95)
20,7 i=]

The major aoe is ae the likelihood function now contains n + 4 parame-


ters, namely, a, B, 0,6 a and the n values of X. Straightforward maximization
of
A SMORGASBORD OF FURTHER TOPICS 435

Eq. (10-95) leads to unacceptable results.} The situation cannot be rescued by


increasing the sample size since this automatically increases the number of
unknown X’’s. It ea ROWo ee be improved by the use of prior knowledge,
typically that \= 07/0, is known. Making this substitution in the log likelihood
and carrying through the maximization process gives exactly the same quadratic
in B as already derived in Eq. (10-92) for Case 10-2. Thus the f defined in Eq.
(10-94) is the ML estimator for Case 10-1.
The range of A is zero to infinity. The extremes correspond to the two simple
regressions in Case 10-2, information 1 and 2. The estimator defined in Eq.
(10-94), which is based on a known A, will lie between the two OLS regression
lines. This estimator is a consistent estimator of 8. Kendall and Stuart show that
a consistent estimator of 02 is provided by

ioe 2n rv
i mew : at ey™ = 2Bm =r B>m.,,) (10-96)

The hypothesis Hj: $8 = 0 may be tested by computing the sample correlation


coefficient

and using the result that, under the null hypothesis,

TV ee
~ t(n
— 2)
vl-r?

The computation of a confidence interval for B is somewhat more complicated.§


Define the angle 6 by B = tan6, or 6= arctan B. The 95 percent confidence
interval for 6 is given by

2 1/2

Oe aC 2!oinis leas esciis, saa Nae ie 3eee aT (10-97)


A . M,,My, A: Myx,

(1 = 2)|(m,., — i) + 4m?,|

The corresponding limits for B are the tangents of these angles. The assumptions
required for the development of Eq. (10-97) render this essentially a large sample
method, and, of course, all the above rests on exact knowledge of A, which is not
often likely to be forthcoming. The technique may be extended to a multivariate
regression if the investigator has knowledge of the ratios of all the error variances.
Details are given in the Kendall and Stuart treatise.

+See M. G. Kendall and A. Stuart, The Advanced Theory of Statistics, vol. 2, Griffin, London,
1961, pp. 383 ff.
+ M. G. Kendall and A. Stuart, op. cit., pp. 385-386.
§ M. G. Kendall and A. Stuart, op. cit., pp. 388-391.
436 ECONOMETRIC METHODS

PROBLEMS

10-1 Prove Eq. (10-5) by the method suggested in the text. [Hint: Remember that expressions such as
x’,(X’,_,X,_,)'x, are scalars and may be moved back and forth in matrix formulas, that is,
cAB = AcB = ABc, where c is a scalar and A and B are matrices.]
10-2 Relation (10-5) is a special case of a general result given by Plackett.+ His problem and method
of proof may be stated as follows:
First sample data y}, X,(" X k)

Additional data Yo, X,(m X k)

Complete sample -[3'| X = e

The problem is to find the simplest computational way of updating least-squares statistics from the
first sample to the complete sample.

Method: Define

2 =i ,
R, = X,(X,X,) X45

and R = X,(X’X) 'X’


Prove that

R,R=R,-R
and hence that

CyeRat R) = I,,
Then show that

(I, + Ri) K2(X,X;)_) =X, (Xx)!


and thus that

(XX)
, zi
"= (XX)
, ik
— (KX)
, al
XS [L, + Ry] "X2(XX,)
, = , re

Finally show that this result yields Eq. (10-5) when X, is just a row vector of observations on one
additional sample point.
10-3 For the recursive residuals defined in Sec. 10-1, prove

RSS, = RSS,_, + w
[Hint: Express y, — X,b, as y, — X,b,_, — X,(b, — b,_,). Partition

18 Yr-1 = De 1
yn [75| and x= ("|

and using Eq. (10-6) show that

RSS, G; — X,b,_ 1)’(Y, “7 x bp) Fr KE (XIX Se (Geex x’, nm ey

Applying the partitioning again and using Eq. (10-5) gives the desired result.]
10-4 Take a simple time series and verify that the restricted estimation of Eq. (10-17) yields the
same
point estimates of the a and £ parameters as those derived from the estimated coefficients of the spline
function (10-14).

7+ R. L. Plackett, “Some Theorems in Least Squares,” Biometrika, vol. 37, 1950, pp. 149-157.
A SMORGASBORD OF FURTHER TOPICS 437

10-5 For the disturbance term in Eq. (10-21) make the following assumptions:

E(u? = 6;

E(uj,4 jr) aay)

ui = Pei fe | leh <1


€;, = iid(0, o2)
Derive var(u) and discuss how a feasible GLS estimator of the parameters of Eq. (10-21) might be
constructed.
10-6 Show that in Table 10-2
Pp n
|
E, ———— =)? ) =02
uj, —u) + m(p—
—— 1) 2
eet | Te
and hence show that it is a biased estimator of 62 = 62 + o2.
10-7 For a two-way error component model assume
Uj, = ey + Ay t+ Ei Dt ason OR = Anon i

where p; is a unit-specific time-invariant effect, A, is a period-specific unit-invariant effect, and, ¢;, is a


random disturbance at observation j, t.
The p,, A,, and e,, are random variables having zero means, independent among themselves and
with each other, with variances o,°, ox, and o2, respectively. Show that

V = E(u’) = 0? [pA + wB+ (1 —p—w)Ipm|


where A=1eJ,

B=J,@I,,
G,
2 Oo
2
CeO 22 ra 2
ania 2 eee
ee eeew N

and J,, is an m X m matrix of ones.


10-8 For the disturbance u,, defined in Problem 10-7 develop the ANOVA table similar to Table 10-2.
Hence indicate possible estimators of 07, of, and o,’.
10-9 Establish the result stated in Eq. (10-52).
10-10 Prove Eq. (10-57).
10-11 Derive formulas (10-81) and (10-82).
10-12 Consider the following regression model for a sample of panel data:
Yj, = a + a, Xj; + a Xi; + a3 X34); + 8);

i = 1,2,..., n (panel members), j = 1,2,..., ¢ (time periods), and the X’s are exogenous variables.
The ¢;; are assumed to be normally and independently distributed with zero mean and constant
variance for all i, /.
(a) If X3,, is not observed and an investigator regresses Y,; on just X,;; and X,,; with a constant
term in the regression, what is the bias in the least-squares estimate of a5? If the algebraic sign of the
simple correlation coefficient for X3;; and X3,; were known, is this sufficient information to determine
the algebraic sign of the bias? If not, explain what information is required to determine the algebraic
sign of the bias.
(b) If the unobserved independent variable X%; ; is assumed to satisfy X3;; = X3; for all j and is
assumed to be nonstochastic, explain how to obtain estimates of a, and a, and their associated
standard errors.
(University of Chicago, 1977)
438 ECONOMETRIC METHODS

10-13 Consider the following errors in variables model:


Vig = a + Dx* Ds)
ay NG ok =ale
Xi tare Vin = Vit + ej,

Ee;, = Ee,, = E(xie,,) = E(xke,,) = E(yre,,) = 0

E(yireis) = E(eréis) = 0
E(e) =o? E(e7) =o2 E(en8:2) = p02 E(e€€:2) = p02
E(x#?)=02 = E(x4x%)=p,02 —foralli;t = 1,2;5 = 1,2
Xjz» Vix ate Observed for i = 1,..., N; t= 1,2. Let b be the IV estimate of b from a cross-section
regression using data from the second time period and x,, as the instrument. Let b be the IV estimate
of b using the same cross section hut with y,, as the instrument. Show that if p,, p,, and p, are all
positive, then plim(b) < b < plim(4).
(UL, 1981)
CHAPTER
ELEVEN
SIMULTANEOUS EQUATION SYSTEMS

So far our interest has centered mainly on the inference problems associated with
a single equation, although there was some discussion of groups of equations in
Chap. 8. Economists, of course, often focus on a single equation, such as an
aggregate consumption function, a demand function for gasoline, a wage-change
equation, and so forth. However, economic theory teaches that such equations are
embedded in a system or subset of related equations. Thus one must examine
whether the presence of these related equations has any implications for the
estimation of the focus equation. More importantly, the estimation of a complete
system of equations is often an important practical problem, whether the objec-
tive is to test economic theories about the nature of the system or to use the
complete system to make joint predictions of a set of related variables.

11-1 SOME ILLUSTRATIVE SIMULTANEOUS SYSTEMS

In this section we will consider a few very simplified systems in order to illustrate
the main problems that arise, and then in subsequent sections we will give a more
general and formal treatment.
Consider first an even simpler income determination model than the one
outlined in Chap. 1. This one consists solely of a consumption function and the
national income identity, namely,
C=a+
t
BY, + u, (11-1)
Y= C4
t (11-2)
439
440 ECONOMETRIC METHODS

where C= aggregate consumption expenditure


Y= national income
Z =nonconsumption expenditure
u=a Stochastic disturbance term

We regard the model as explaining the values taken by C, and Y, conditional


on Z,. Thus C and Yare classified as endogenous variables and Z as an
exogenous variable. We will make two assumptions, namely:

1. u~ NO, 021)
2. Z and u are independent, which will be satisfied if either Z is a set of fixed
numbers or Z is a random variable distributed independently of u. Z could be
taken as representing autonomous investment and government spending
controlled by some central authority. The model does not discuss the determi-
nants of Z.

The reduced form of the model ist

Ci a
Tee B Gee
ate (11-3)
a ]
Y-Toptpepot (11-4)

where v, = u,/(1 — B), so that


62
v~N Dein a
(lB)
It is immediately obvious from Eq. (11-4) that v,, and hence u,, influences Y,. In
fact,
o2
ye el ;
plim(— E¥,2,] = plim(—0? =
(lis A)
Thus the application of OLS to the consumption function (11-1) would yield
inconsistent estimates.{ The nature of the inconsistency is illustrated diagram-
matically in Fig. 11-1. The line a + BY shows the relation between C and Yif the
disturbance u were zero. The line Y — Z’ illustrates the identity (11-2) for a
specified Z’. The equilibrium of the system would then be indicated by the point
P,. Imagine now that Z is held constant at Z’ and that the disturbance takes on
various positive and negative values in some finite range.§ The economy would

+ As shown in Chap. 1, the reduced form is obtained by solving the model so as to express each
current endogenous variable solely in terms of exogenous variables and lagged endogenous variables.
¥ If necessary, review the discussion of consistency in Sec. 7-2 and illustrations of inconsistency in
the presence of lagged variables in Sec. 9-2 and in the presence of errors of measurement in Sec. 10-6.
§ The range is, of course, infinite for a normally distributed disturbance, but the finite range is a
convenient assumption to keep the diagram simple.
SIMULTANEOUS EQUATION SYSTEMS 441

a+ BY

Figure 11-1

then trace out points in successive periods in the range P, to P, along the Y — Z’
line. If Z never changed from Z’, these would be the only points ever observed for
this economy, no matter how many observations were taken. The estimated
regression of C on Y would coincide with the line Y — Z’, and the estimated
marginal propensity to consume would be unity, no matter what the true B happened
to be. Now suppose that over a large number of time periods Z ranges between Z’
and Z”. Observations on C and Y would then fill in the parallelogram P, P, P; P,.
The least-squares regression of C on Y minimizes the sum of squares of the
residuals measured in the vertical (that is, C) direction. Thus in the limit the OLS
line will tend to pass through the points P,, P,;. The estimated slope will now be
less than unity but will still be greater than the true B, so that the asymptotic bias
is positive.

Instrumental Variable Estimation


We saw in Sec. 9-2 that the use of suitable instrumental variables can produce
consistent estimators. The obvious instrument in the present model is the Z
variable which, by assumption, is independent of u, and by Eq. (11-4) will be
correlated with Y. Applying the IV estimator defined in Eq. (9-60) to this model
gives
aw — CG—byY (11-5)

DoCz
and bry = Dyz (11-6)

+ See Problem 11-1.


442 ECONOMETRIC METHODS

where c, y, and z denote deviations from the sample means. From Eqs. (11-3) and
(11-4) we may derive
eee
Yen = 7B
2 + }izo

]
Lyz = Lz? + Lz
=f,
Thus, provided
a OL =
plim(—:20 =0 and plim( 52" =M,,

a finite number,

plim(d,;y) = B
and hence plim(a,y) = a

Indirect Least-Squares Estimation


The above development already contains a clue to a second estimation principle,
that of indirect least squares (ILS). Looking at the reduced-form equations it is
clear that they satisfy the assumptions under which OLS estimators are consistent
(and indeed best linear unbiased) so that
Bez
=a is a consistent estimator of B
22 1-8
Lyz
and <2 is a consistent estimator of ———,
De hip
which suggests taking the ratio
DCZy ae :
Bie = a en = as an estimate of B
LZ Le
The principle of ILS is to estimate reduced-form coefficients by OLS and then to
compute structural coefficients by an appropriate transformation of the estimated
reduced-form coefficients. We see immediately that in this case
Sez
bis a Yyz Ny

Two-Stage Least-Squares Estimation


A third estimation principle is that of two-stage least squares (2SLS). It starts
from the problem of Y, and u, in Eq. (11-1) being correlated. The first stage is to
regress Y on the exogenous variables in the model, which in this case are Z, and a
dummy variable that is always unity to allow for the intercept term. This
reduced-form regression yields an estimated Y series, which it is hoped will
display less correlation with the wu series than does the original Y series. We may
SIMULTANEOUS EQUATION SYSTEMS 443

write Eq. (11-4) in deviation form as


y, = 5z, + 0,
where 6 = 1/(1 — B), and we have also omitted © as it does not affect the
subsequent derivation. The regression values are then given by
Be VESEY Z
j= b,— (2s,

6+ ae
Sz

Thus Lyu = 6Lzu +228 - Lizu


Z
On the assumptions made earlier,
A alos eUEL.
plim|= E20} ~ plim(—Szu] = 0

so that in the limit Y is uncorrelated with u. In the second stage C is regressed on


Y to estimate a and £, that is, Eq. (11-1) is reformulated as

C,=a+ BY,+[u,+ B(Y,- ¥,)]


with C, as the dependent variable and x as the explanatory variable. The
disturbance term is shown in square brackets. From the OLS regression of Y on Z
it follows that Y, will have zero correlation in the sample with the residual
Y, — Y,, and we have just shown that Y, is uncorrelated in the limit with u,. Thus
Y, is uncorrelated in the limit with the combined disturbance term [u, + B(Y, —
Y,)], and the 2SLS estimators will be consistent. The 2SLS estimate of the slope 8
is
SCR ipOleez, Mike gC?
Pass ry? §Ez2 Ez? Lyz | ENZ
Thus we see that in this case all three principles of estimation, IV, ILS, and 2SLS,
would yield identical consistent estimates.
The two-equation model of Eqs. (11-1) and (11-2) is the simplest possible
simultaneous equation model, consisting of just one stochastic behavioral equa-
tion and an identity, but that is enough to generate a dependence between the
explanatory variable and the disturbance in the structural relation, rendering OLS
inconsistent. More complicated models may be expected to generate further
problems in addition to those already encountered.
Consider next a two-equation model in which both equations are stochastic
behavioral relations. With a slight change of notation we write
Vit Biya t+ Yn =
Petia oon (11-7)
Bo Vie + Yar + Yor = Yar
In this and subsequent models lowercase letters denote the actual values of the
444 ECONOMETRIC METHODS

variables and not deviations from sample means. We will reserve the letter y for
endogenous variables so that y,, denotes the rth observation on the ith endoge-
nous variable. Likewise x,, will denote the th observation on the / th exogenous
variable. The structural parameters B and y also have two subscripts, the first
indicating the equation and the second the variable to which it is attached.
Model (11-7) would be a conventional demand-and-supply model if y,
denotes price, y, denotes quantity, and we impose the restrictions

Bi, > 0 B,,


<0
so that the first equation represents a downward sloping demand curve and the
second an upward sloping supply curve. We would also want to impose an
additional restriction y,, < 0 to ensure a positive intercept for the demand
function. If the disturbances in period t were both zero (u,, = 0 = u,), the model
would be represented by the D, S lines in Fig. 11-2, and we would observe the
equilibrium price and quantity indicated by y*, y3. Nonzero disturbances shift
the D, S curves up or down from the position shown in Fig. 11-2. Thus a set of
random disturbances would generate a two-dimensional scatter of observations
clustered around the y*, yx point.
A fundamentally new problem now arises. Given this two-dimensional scatter
in price-quantity space, demand analysts might fit a regression and think they
were estimating a demand function. Supply analysts might fit a regression to the
same data and presume they were estimating a supply function. “General
equilibrium” economists, wishing to estimate both functions, would presumably
be halted on their way to the computer by the thought, “How can we estimate
two separate functions from one two-dimensional scatter?” The new problem is
labeled the identification problem. It is concerned with the question of whether any
specific equation in a model can in fact be estimated. It is not a question of the
method of estimation nor of sample size, but of whether meaningful estimates of
structural coefficients can be obtained. On the assumptions made so far neither -
equation in Eggs. (11-7) is identified. A regression fitted to the scatter in y,, y,
space is not an estimate of either the demand or the supply function.

y\

S: Bai¥, + Yo + ¥21 = 0

yf H—-————— ——
|
| Diet Bynys
vig
|
||
O | WY
V3 2 Figure 11-2
SIMULTANEOUS EQUATION SYSTEMS 445

The identification problem may be investigated by looking at the relation


between the structural and the reduced forms of the model. The reduced-form
equations corresponding to Eq. (11-7) are
1
iltoae oreldberrar a Bi2Ya1) as (u,, i Bite.) |
: (11-8)
ae A LCBairn = Yo) a (=B,)%), a Uy,)|

where A = | — £,,8,,. The first term on the right-hand side of each equation is a
constant. Thus we may write the reduced form more simply as
(11-9)
Vara Fy TOiy

y21 = by 7 U2,

where

hy =
See inee
ma
yt

f
2
ees
(11-10)
7
OTS
— Mir Bio
A
t

Ome
~ ai
_ B A
t
+Ure
If we postulate that

stu)=| 50) |= E(u

E(u,)

, Gije 3Cq2
lag 1 ie ss
=
then
E(v,)=0

var(v,) = E(v3,) = tu7


CBG
Pi ta Pa te
Bo

9 pe
B3101, + Aa 61
— e281ae
a E(v3,) = Eire
var(v,)

and
— B01; — B29 tall 5 B 2B) Ov
cov(v;, v>) as E(v,0,) am A

+ Here there are no lagged endogenous variables and the only exogenous variable is the dummy
variable x,;, = 1 for all t, which is required to take care of the intercept term in the structural
equations.
446 ECONOMETRIC METHODS

It also follows from Eqs. (11-9) that

E(y,) ="

E(y)
= Mo
var( y,) = var(v,) (11-11)
var( y,) = var(v,)

cov( y, Y2) = cov(v,, v2)


Sample data on y,, y, can only yield estimates of the five parameters in Eqs.
(11-11). These in turn are functions of the seven parameters of the structural
model, namely, 8,5, Bs), Y11> Yo1> %11> 922» and o,. On the assumptions made so far
the structural parameters are unidentifiable.
As a numerical illustration of this situation suppose the true structure
corresponding to Eqs. (11-7) is
y, + 2y, —-10 =u,

= BY ty) 2= Uy (11-12)
01}, = 9, = 1 01, = 0.5
Equations (11-7) define a model, and a structure like Eqs. (11-12) is obtained from
a model by assigning specific numerical values to the 8 and y parameters and also
to the variances and the covariance of the u’s. Solving this structure for Eqs.
(11-9) gives
yy =2+0,

Va Aas
u, — 2u
where v, Leen

ee 3u, LS
+u 2

Thus E() =m, =2


E(y,) =p, =4
3
var( y,) = var(v,) = i (11-13)
13
var( y,) = var(v,) = 49

—1.5
cov( y1, ¥2) = cov(0,, 02) = o—

The true structure (11-12) is, of course, known only to the “deity” who sets the
economic system in motion. Now suppose that one of the deity’s vice-presidents
tinkers with the institutions in an attempt to confuse the econometricians of the
world and concocts a new structure by the following rule, where (1) and (2)
SIMULTANEOUS EQUATION SYSTEMS 447

indicate the first and second equations in Eqs. (11-12),


New first equation = 4(1) + 1(2)
New second equation = — (1) + 3(2)
This yields the structure
y, + 9y, — 38= ut (11-14)
— 10y, +y, + 16 = uh
where
ux = 4u, + up,
ry as

The new structure obeys the same a priori constraints on signs as Eqs. (11-12).
Solving this structure for Eqs. (11-9) gives
y=a2t+vy
yy = 44 v3
where
uy = 945 uy
— 205
LU aeOV Se, CMEep
10g, 2us) Buy,
DoT eNG ITCtie reg Ae Ae
Thus the five parameters of the reduced form E(y,), E(y2), var(y,), var(y2), and
cov(y,, ¥>) are identical for the two different structures and indeed for all
structures derived by taking linear combinations of the original structural equa-
tions.
It is instructive to see what type of further information might help identify
one or both equations of this model. There are three basic possibilities, namely,
(1) restrictions on the 8 and y parameters, (2) restrictions on the 2 matrix, and (3)
respecifications of the model to incorporate additional variables. To illustrate the
first category, suppose the supply function is presumed to go through the origin.
The a priori restriction is thus
Yr,
=9
This reduces the number of structural parameters to six, but the number of
reduced-form parameters is five, as before, so that it is still not clear that any
structural parameters can be identified. However, making the substitution y,, = 0
in Egs. (11-10) gives
isp!
be An

oo Buti
oe A

me
so that be
es By
448 ECONOMETRIC METHODS

showing that 8,, can be determined from a knowledge of the reduced-form


parameters and also suggesting a possible estimator as £,, = —y,/y,. This
restriction would enable the supply function to be identified, but the demand
equation remains unidentified. Linear combinations of the demand and sup-
ply equations would be statistically indistinguishable from the original demand
equation. However, any linear combination that assigns a nonzero weight to the
demand function will fail, with probability 1, to have a zero intercept and thus
will not look like the new supply function.
Now suppose we return to Eqs. (11-7) and impose the restriction
var(u,) = 0,, = 0
This also implies that o,, = 0. Looking at Eqs. (11-11) we now find

= Bin095
var( y,) oi A2

0.
var( y,) = a

bee — B 12%
cov(y;, 2) = = aie
so that

phe )ae3 _ ~cow(


ys,y»)
var( y>) var( y)

ea — var(y,)
cov( y;, V2)

and thus the slope of the demand function is identified. Taking expectations of
the demand function in Eqs. (11-7) gives
Yin = Shake
and substitution for », and pw, from Eqs. (11-10) verifies that this relation holds.
Thus y,, and £,, can both be expressed in terms of the parameters in Eqs. (11-11),
and the demand equation is identified. This case is pictured in Fig.11-3. The
combination of o,, =0 and o,, + 0 generates a set of observations on the
demand function.
A less extreme version of this case would occur if o,, were “small” as
compared with o,,. The scatter of observations would then tend to be con-
centrated around the demand function rather than lying exactly on it. However,
knowledge about the relative sizes of disturbance variances is not likely to be
generally available, though a possible reason for a large o,, might be the omission
of important explanatory variables from the supply function in Eqs. (11-7). The
appropriate remedy is the respecification of the supply function to include such
variables. In practice the demand function should also be looked at since the
simple two-variable model of Eqs. (11-7) is hardly a realistic specification with
which to commence empirical work.
SIMULTANEOUS EQUATION SYSTEMS 449

O >y, Figure 11-3

Consider now arespecification of Eqs. (11-7) which is, say,

Vi + Bid. + Wx + Yi2%2 =U, (11-15)

Bait Yo + Yai%y + ¥23%3 + Y2qX4 =U


where we still retain the restrictions B,, > 0 and B,, < 0 to conform with the
demand-and-supply analogy. The variable x, could be taken as a dummy with a
value of unity in all periods to cater for the intercept term, x, might represent
income, which is expected to influence demand, and x, and x, would represent
variables influencing supply. The reduced form of this model is
xy
yi Wi (yar Bove vin Piss (Biv iloo v,
yy2 es (Bai¥i1 — Yar) BuYi2 —Yo3 — Y24 Eee
ee 2

where A = 1 — £,,8,, and the v’s are given in Eqs. (11-10). Let us denote the
reduced-form coefficients by 1;, (i= 1,2; j=1,..., 4). It is clear that the
structural coefficients can be obtained from the reduced-form coefficients. For
example,
Bo, = eh)
21 TT

pean aseit, Aycan


is 7793 4
and having found the £’s, the y’s can be obtained from 7,, and 7,,. Leaving the
disturbance parameters aside, there are eight reduced-form coefficients and just
seven structural coefficients. The imbalance is reflected in the existence of two
alternative (but equivalent) expressions for B,,. This indicates, however, that we
may expect the ILS technique to run into trouble here since the estimated
reduced-form coefficients will in general not satisfy the equality 73/73 = 74/24
that holds for the true coefficients.
450 ECONOMETRIC METHODS

Further investigation of identification and estimation problems by way of


specific models of increasing size and complexity would be inefficient. We now
move to a more general and more formal treatment, which can then be specialized
to deal with particular cases.

11-2 THE IDENTIFICATION PROBLEM

Let us assume a linear model containing G structural relations. The /th relation at
time ¢ may be written

Bia Yie + +> + BigVer + Yar t °F Vix = (11-16)


b= Lee Ge tee
where the y,, denote endogenous variables at time ¢, and the x,, indicate exogenous
variables (current or lagged) and may also include lagged endogenous variables.
The latter two groups constitute the class of predetermined variables. The model
may then be regarded as a theory explaining the determination of the G jointly
dependent variables y,, (i = 1,..., G; t = 1,..., 2) in terms of the predetermined
variables x,, (i = 1,..., K; t= 1,..., m) and the disturbances u,, (i = 1,..., G;
t = 1,..., n). The underlying theory will in general specify that some of the B, y
coefficients are zero. If it did not, all the equations in the model would look alike
statistically, as in Eqs. (11-7), and no equation could be identified. As mentioned
earlier, the lowercase letters denote actual values of the variables and not
deviations from arithmetic means, and setting one of the x variables at unity
caters for a constant term in any equation that requires it.
The model may be written in matrix form as
By,+ Ix,=u, . t=1,...,n (11-17)
where B is a G X G matrix of coefficients of current endogenous variables, I is a
G X K matrix of coefficients of predetermined variables, and y,, x,, and u, are
column vectors of G, K, and G elements, respectively,

By, By as Big YA WA ee Yik


B=], By Bog T=] 21 Y22 Y2kK
Ber Be2 Boe YGIaNG2 YGK

Vit Xi, Ur,

V21 Xo; Uy,


J, a x, a U, i

YGt XKt UG,

It is plausible to assume that the B matrix is nonsingular since, if it were not, one

+ Notice that for the moment we‘have not normalized the structural equations
by setting any of the
8 coefficients at unity.
SIMULTANEOUS EQUATION SYSTEMS 451

or more of the structural relations would merely be a linear combination of other


structural relations, thus being redundant, or, if the rows of the I matrix did not
obey the same linear restrictions as the rows of B, the G structural equations
would be inconsistent. Assuming, therefore, that B~' exists, the reduced form of
the model is
y, = IIx, +), b=) (11-18)

where

IlI=-B'T and y,=B'u, (11-19)


The II matrix is of order G X K and thus contains GK elements. The B and
matrices contain at most G? + GK elements. There is thus an infinity of B and T
structures corresponding to any given II matrix.
The identification problem arises because the most that can be determined
from observational data on y, and x, (¢t = 1,..., 2) is a knowledge of the elements
of II and the elements of the variance-covariance matrix of the v’s. This may be
seen in a number of ways. The reduced form Eqs. (11-18) show explicitly that the
model provides an explanation of y, conditional on x, and on the disturbance
vector v,. From Eggs. (11-19) it is clear that the stochastic properties of v, depend
on the assumed stochastic properties of the structural disturbance vector u,.
Assuming E(u,) = 9 for all ¢ then givest

E(y,|x,) = Ix,
Thus the mean of the conditional distribution of y,, given x,, depends solely on
the IT matrix. A finite sample of observations (y,,x,; t = 1,..., 1) will yield some
estimate ne which will deviate from the true II due to the fluctuations of random
sampling. Suppose, however, that we dispense with sampling problems by assum-
ing that an infinitely large sample of observations can be made available. In
general the true II may then be determined with any desired degree of precision.
This is all that can be afforded by the sample data. Thus knowledge of B and
can only come from knowledge of II.
To see the same point in a likelihood context, let us assume
u,~ N(O, =)

and also that the u, vectors are serially independent. It then follows from Eqs.
(11-19) that
v,~ N(O, &)
where C= BoISBa) (11-20)
and the y, are serially independent. From the reduced-form equation (11-18)

P(y,|X,) oanpP(v,) > (27) °/7|Q| -'exp(— $v/Q~'v,)

+ When x, contains lagged y values, this expectation has to be read as conditional on these lagged
endogenous values.
452 ECONOMETRIC METHODS

Thus the likelihood of the sample y’s conditional on the x’s is


oon
L = p(Y1+Y24-+++ Yul X) = 29200 -*Pexo|—5 t=1
Y va i
=H 2 ere ive
= (29) "19 "Aexo|—3 t=13 (,~ Hs,) (o—1)| (11-21)
Alternatively one might set up the likelihood in terms of the structural equations
(11-17). This gives

P(yIx,) = p(u,) eS
= p(u,) - ||BI|
where ||B|| denotes the absolute value of the determinant of B. The likelihood of
the sample y’s conditions on the x’s is then

L =(20) aoe
"°"|BII"|2|
n St
Perel— 3Dw
l

t=1
[SS
1,

= (29) <n -"°"Bi\"| = "Pexo]— 3X


ly
t=1
(By + eit
Px) (By,+ T)|
(11-22)
Comparing Eqs. (11-21) and (11-22) it is easily seen, using Eq. (11-20), that

(y, — IIx,)’Q~"(y, — IIx,) = (By, + I'x,)’2""(By, + T'x,)


and |Q|~ "7? = |[BI"|2|>
"7?
so that Eqs. (11-21) and (11-22) are equivalent. Leaving aside the variance
matrices 2 and Q, each of which contains G(G + 1)/2 parameters, there are >
G* + GK parameters in Eq. (11-22) and just GK in Eq. (11-21). The likelihood
function is thus completely specified by the GK parameters in II. Identification of
structural parameters in B and I thus depends on the addition of further
information to the model specified in Eq. (11-16). Such information usually takes
the form of restrictions on various elements of B and I and, less frequently, on
the elements of 2.

Restrictions on the Structural Coefficients


We will consider the identification of the first equation in the system. The
methods derived can then be applied to any structural equation. Let us rewrite the
structural form of the model (11-17) as

Az, = (|B ri =U, (11-23)


where A = [B_ I] is the G x (G + K) matrix of all structural coefficients and z,
is a (G+ K) X1 vector of observations on all variables at time ¢. The first
SIMULTANEOUS EQUATION SYSTEMS 453

structural equation may then be written as


O12, = Uy,
where a, denotes the first row of A.
Economic theory typically places restrictions on the elements of a,. The most
common restrictions are exclusion restrictions, which specify that certain variables
do not appear in certain equations. Suppose, for example, that y, does not appear
in the first equation. The appropriate restriction is then
By, = 0
which may be expressed as a linear restriction on the elements of a,, namely,
0
0
1
Bambu baer tine ovielioyo

0
There may also be linear homogeneous restrictions involving two or more
elements of a,. The specification that, say, the coefficients of y, and y, are equal
would be expressed as

a
eee ee are rea re . =)0)

0
If these were the only a priori restrictions on «,, they may be expressed in the
form
a,® = 0 (11-24)

where

0 1
(ee |
® =) 1 0
00 00
The ® matrix has G + K rows and a column for each a priori restriction on the
first equation.
In addition to the restrictions embodied in Eq. (11-24) there will also be
restrictions on a, arising from the relations between structural and reduced-form
coefficients. From Eqs. (11-19) we may write
BIL+T=0
or AW =0
454 ECONOMETRIC METHODS

where W= LT

The restrictions on the coefficients of the first structural equation are thus
aW =0 (11-25)
Combining Eqs. (11-24) and (11-25) gives
a[W o]=0 (11-26)
There are G + K unknowns in a,. The matrix [W 9] is of order (G + K) X (K
+ R), where R is the number of columns in ®. On the assumption that II is
known all the elements in [W ©] are known. Thus Eq. (11-26) constitutes a set
of K + R equations in G+ K unknowns. Identification of the first equation
requires that the rank of [W ®] be G+ K — 1, for then all solutions to Eq.
(11-26) would lie on a single ray through the origin. This suffices to determine the
coefficients of the first equation uniquely, for in specifying the general model in
Eq. (11-17) a B or y coefficient was attached to each variable in every equation.
Normalizing the first equation by setting one coefficient at unity (say, B,, = 1)
will now give a single point on the solution ray, and this determines a, uniquely.
e[W ®]=G+K-1 (11-27)
is clearly a necessary and sufficient condition for the identifiability of the first
equation. The condition for the identification of the ith structural equation is
e[W ®]=G+K-1
where ®, is the matrix embodying the a priori restrictions on the ith equation. The
basic difficulty with the rank condition, as stated in Eq. (11-27), is that it is not a
convenient one to apply since it requires the construction of the II matrix, which
is complicated even in small models. We will give below an equivalent condition ~
in terms of structural parameters which is easier to apply. However, condition
(11-27) does yield necessary conditions for identification which are very simple to
apply. Since[W ®] has K + R columns, a necessary condition for Eq. (11-27) to
hold is that
IM drdk = Gar ix = Il
or Kea Geael (11-28)

that is,

The number of a priori restrictions should not be less than the number of
equations in the model less 1.

When the restrictions are solely exclusion restrictions, the necessary condition is
restated as:

The number of variables excluded from the equation must be at least as great as
the number of equations in the model less 1.
SIMULTANEOUS EQUATION SYSTEMS 455

Finally, an alternative form of this last condition may be derived by letting


g = number of current endogenous variables included in equation
k = number of predetermined variables included in equation
Then
R=(G-—g)+(K-k)
and the necessary condition becomes
(G-—g)+(K-k)=G-1
or Ke Rie Cal
that is,

The number of predetermined variables excluded from the equation must be at


least as great as the number of endogenous variables included less 1.

The necessary condition is referred to as the order condition for identifiabil-


ity. In large models this is often the only condition that can be applied since
application of the rank condition becomes difficult, if not impossible.
The rank condition (11-27) may be restated ast
ep[W ®]=G+K-1 if and only if o(A®)=G-— 1 (11-29)
Note carefully that [W @®] is a matrix consisting of the two indicated sub-
matrices, while A® is the product of two matrices. The second form of this
condition only involves the structural coefficients and thus affords an easier
application. When the restrictions are all exclusion restrictions, the first row of
A® is a zero vector and the remaining G — | rows consist of the coefficients in the
other structural equations of the variables which do not appear in the first
equation.
If equality holds in Eq. (11-28), that is, R = G — 1, so that the number of
restrictions on the first equation is just equal to the number of structural
equations less 1, the matrix A® is then of order G x (G — 1). However, the first
row of this matrix is zero by virtue of a,® = 0. This leaves a square matrix of
order G — 1 which, apart from some freakish conjunction of coefficients, will be
nonsingular. The first equation is then said to be exactly identified or just
identified. Suppose instead that R > G — 1. Then A® has G or more columns.
There are now more restrictions than strictly required for identification, and in
general there will be more than one square submatrix of order G — | to satisfy
the rank condition. The equation is then said to be overidentified.
A direct proof of the rank condition in terms of the A® matrix may be
obtained from an alternative approach to the identification problem. We saw in
one of the examples how taking linear combinations of the equations in a given

+ See F. M. Fisher, The Identification Problem in Econometrics, McGraw-Hill, New York, 1966,
Chap. 2; or for a shorter proof, R. W. Farebrother, “A Short Proof of the Basic Lemma of the Linear
Identification Problem,” International Economic Review, vol. 12, 1971, pp. 515-516.
456 ECONOMETRIC METHODS

structure could yield a new structure which satisfied the same a priori constraints
as the original structure and had identical reduced-form coefficients. Let
A=[B TI]
denote an original set of structural coefficients (that is, with specific numerical
values), and let FA denote a new structure obtained from A by premultiplication
with an arbitrary G X G nonsingular transformation matrix F. The new structure
is said to be admissible, or equivalently F is said to be an admissible transforma-
tion matrix, if FA satisfies all a priori restrictions on A.f Identifiability of the first
equation then requires that the first equation of every admissible structure be
some scalar multiple of the true first equation. The first row of A may be
expressed as
a, =e,A
where e, is a 1 X G row vector with unity in the first position and zero elsewhere.
Thus the a priori restrictions on the first equation may be written
e,(A®) = 0
The first row of coefficients in the transformed structure may be written as f,A,
where f, denotes the first row of F. For an admissible structure this must obey the
same restrictions as «,, and so we must have
f,(A®) = 0
Identifiability requires that f,A be a scalar multiple of e,A, that is, that f, be a
scalar multiple of e,, which gives the condition that p(A®) = G — 1. If all the
equations of a model are identified, the only admissible transformation matrices
are diagonal matrices.

Examples. To illustrate the application of the conditions for identifiability we


shall work with the two-equation system
Bi die + Biz¥ar + WX + Yi2X21 = Uj,
Bo Vie + Boo Yor + YarX1y + Yx2X2, = Uy,
As it stands, both equations are unidentifiable since no a priori restrictions have
yet been imposed. Each example will postulate a different set of restrictions.

Example 11-1 Suppose the a priori restrictions are


Yn = 0 Yor 9
For the first equation ® is then a four-element column vector
0
A
a ako
1

7 The general definition of admissibility also requires that the variance matrix of the transformed
disturbances satisfy all the a priori restrictions on the original variance matrix, but we are restricting
consideration here to the structural coefficients.
SIMULTANEOUS EQUATION SYSTEMS 457

Yi2 0
and A® = ne] = ie

Thus p(A®) = 1 = G — 1, and the first equation is identified, provided, of


course, that y,, * 0. If y,, were zero, the variable x, would not appear in
either equation, and so the fact that it was absent from the first would be of
no help in identifying that equation. In a similar fashion, the restriction on
the second equation gives
0
0
Smal
0
Yu
ea |0 |
and p(A®)=1=G-1
Alternatively the equations
a[W ]=0
in the parameters of the first equation give
™, ™M O
Bybee
[Bu Bo Yn Yl 1 ee [0 0 0]

0 1gach
that is,
Bum, + Bim, + W11 = 9
By M2 + ByyM + Y2 = 9
Yi2 = 0
If we normalize by setting, say, 8,, = 1, these give

pis 2
9
ig, 22
and Nitta aah
eaterare
2
which shows explicitly how the parameters of the first equation may be
derived uniquely from those of the reduced form. The parameters of the
second equation may be obtained in a similar fashion.

Example 11-2 The restrictions are


Yo =0 Yx = 0
For the first equation

-—-

oO
458 ECONOMETRIC METHODS

Be i
and A® = a

which has zero rank. Thus the first equation is not identifiable; nor is the
second, for this is the case we alluded to in Example 11-1, where x, appears
in neither equation.

Example 11-3 The restrictions are

ib at Yi2 = 0 Yx.
=0
This example might be treated in two ways. In one approach we note that the
restrictions y,, = 0 = y>, mean that x, does not appear in the model at all.
Thus the model could be reduced to one with just a single exogenous variable,
in which case the only restriction is y,, = 0, and that suffices to identify the
first equation, but leaves the second unidentified. Alternatively, retaining the
dimensions of the original model, the restrictions on the first equation give

0 O
ATEN ; 205 ad)
® = 1 0 with A® ey |

OF al
Thus p(A®) = 1 = G — 1, and so the first equation is identified. For the
second equation

0
® eer= O0 with A® ae |
]
so that this equation is not identified. Alternatively, for the second equation .
a,[W ®|=0
gives

™m, M2 O

[Br By» Yr Yr] Tr Tn 0) 8 [0 0 0]


] Omen
0 ae
This appears to give three equations in four unknowns. Setting B,, = 1 would
then determine the remaining parameters of the second equation. However,
the restrictions y,, = 0 = y2) imply 7,, = 0 = 75). Thus the second and third
columns in[W_ 9] are identical, and so we only have two equations plus a
normalization rule, which are insufficient to identify the second equation.

Example 11-4 The restrictions are

Yi =e ¥o
—0
SIMULTANEOUS EQUATION SYSTEMS 459

For the first equation


0. 0
OA Ag
Bealby +0
reer:
d =
0 0
a mM Ee a
so p(A®) = | and the first equation is identified, while the second is not.
a[W ®]=0
gives Bum, + Bim, + 1 = 9
Bit + Brym + Y= 0
% = 0
Viz — 9
which, on setting B,, = 1, gives
po=
T
= sel a2
7
i <b) 2
This does not imply a contradiction, for both expressions for B,, will yield an
identical value. The prior specifications and the normalization rule in this
example give the model
Vit te BirVay = Uj,

Bo Vie + Yar + YaiXie + Y22%01 = Ur,


The matrix of reduced-form coefficients is

= ie a oe |Pte fara
7, 72 A= Yq aay)

where A = 1 — B,,f>,. Although II is a 2 X 2 matrix, its rank is only 1. This


is an example of overidentification. Only one prior restriction is needed to
identify the first equation, but we have two. The consequence isarestriction
on the reduced-form coefficients. Notice also that even in the overidentified
case p(A®) cannot exceed G — 1. A® has G rows, but the first row is always
zero for homogeneous restrictions, so p(A®) < G— 1 even in cases of
overidentification where A® has G or more columns. If II is replaced in an
actual two-equation problem by II, the matrix of estimated reduced-form
coefficients, then p(I1) will almost certainly be 2 and not 1, so that estimating
Bo by —%,,/%, or by — #,2/%) would yield two different values. ILS is thus
not a suitable estimation method for overidentified equations, since it fails to
yield unique estimates.

Example 11-5 The restrictions are


Yin 0 Yn = 0 By, + 21 = 9 =9
Yo.
460 ECONOMETRIC METHODS

This is Example 11-3 with the additional specification £,, + Y2; = 0. In


Example 11-3 the first equation was identifiable and the second not. Leaving
x, out of the model, we now have for the second equation
|
®=1/0
1

and av =|F]
0

so p(A®) = 1 and the second equation is now identified.

In all the above examples readers should check for themselves that the
necessary condition (or order condition, as it is often called) would correctly
indicate the presence or absence of identification. This need not always be the
case. For example, if 8,, in Example 11-5 were zero, the rank condition would fail
even though there is one restriction on the second equation.

Treatment of Identities
Identities themselves do not raise any identification problems since in general the
coefficients are known and indeed are usually unity. The general model
By, + Ix, =u,
may, however, be formulated in two alternative fashions. In one version all
identities appear explicitly in the model. In the alternative version the identities
may be substituted in other structural equations, thus effectively reducing the size
of the model. The identification rules may be applied to either version. Solving .
out the identities will not change any conclusions about the identifiability of any
behavioral or other structural equation whether in its original or revised form.
As an illustration consider the simple supply-and-demand model

qP=at+aptu

Gq = By + Bip. + Baw tu
q?=¢°
where q? = quantity demanded
q° = quantity supplied
P price
w = an index of weather conditions
This is a model containing three endogenous variables q”, g°, and p (G = 3) and
two exogenous variables w and z (a dummy variable) set at unity to take care of
the intercept term in the first two equations. Rearranging the model in more
SIMULTANEOUS EQUATION SYSTEMS 461

suitable form we have

D
1 OS aeO. 0 — Qo ei uy
0 losweBin Aa! 1 Bo i! seg
ee | 0 0 LO) w 0
Z
For the first equation
0 0
A® = ae Bo
al 0
and p(A®) = 2 = G — | so that the equation is identified. Notice that when we
have exclusion restrictions, the A® matrix can be written down directly by taking
the columns of the A matrix which contain zeros in the row corresponding to the
equation under study. For the second equation
1
A® = |0
1
which only has rank unity, and so the second equation is not identified.
If we rewrite the model without the identity, it becomes a two-equation model
in two endogenous variables q and p,
T= Ay 1 aptr uy

g=B)+B\p+
Bw u,
where now G = 2, and the first equation is again just identified because it has one
restriction on its coefficients while the second equation is not identified because
there are no restrictions on its coefficients.

Inhomogeneous Linear Restrictions


The linear restrictions embodied in Eq. (11-24) are all homogeneous, that is,
specified coefficients or linear combinations of coefficients are set equal to zero.
Many restrictions indicated by economic theory occur naturally in a nonhomoge-
neous form, an illustration being the specification, say, that the elasticities in a
production function sum to unity. Such restrictions, however, have no meaning
until a normalization rule has been imposed. Thus if we have the restriction
By sia ieee 1

it can be written as

Bitty vein ©
plus the normalization rule B,, = 1. Thus inhomogeneous [Link] be
recast in homogeneous form before normalization and the previous procedures
still apply.
462 ECONOMETRIC METHODS

Restrictions across Equations


So far we have only considered linear restrictions within a structural equation.
There are cases, however, where theory suggests restrictions across equations,
some examples of which have already been encountered in Sec. 8-6. These can
also serve to ensure identifiability, as is shown in the following simplified
examples.

Example 11-6 Consider the model


Yi + Bry. + MX) =
BoyV + Y2 + Ya1%X1 = 42
Without further restrictions neither equation is identified. The imposition of
cross-equation restrictions requires that each equation be normalized, other-
wise the restriction is ambiguous. Suppose there is a theoretical basis for
postulating
Vike Yor =e
Identifiability in the presence of the restriction may be examined either by
looking at the relationship between structural and reduced-form parameters
or by investigating the set of admissible transformed structures that satisfy
the restriction. The reduced-form equations are
mail
Dita aad sf Bi2)x, + 0,
y
yD eal + Boy) + U2

where

Noel Oss Bir Boy


The reduced form yields only two parameters and, even with the restriction,
there are still three structural parameters. It is clear that neither equation is
identified.7

Example 11-7 Consider


Vt VX = Uy
Ba, + Y2 + YX) = U2

+ The argument to the contrary in G. S. Maddala, Econometrics, McGraw-Hill, New York, 1977, p.
230, is incorrect. Maddala investigates identifiability via transformation matrices. However, he

aia
essentially postulates a transformation matrix

and then finds that the restriction implies A = 0, which leads him to conclude that both equations are
identified. But F has already assumed that the second equation is identified, which is an invalid
assumption. The identifiability of both equations has to be considered jointly.
SIMULTANEOUS EQUATION SYSTEMS 463

Postulating the transformation matrix

ie ye
GRE a
the transformed structure is

(fu + fi2Bo) + fi2y2 + (fyi + fi2%o) 4 =r

by + fo Boi) V1 + for
Yo + (firn + foyYo1) X = uy
The requirement that the transformed structure satisfies the same a priori
constraints as the original structure, namely, that y, does not appear in the
first equation, gives
fio = 9
The normalized transformed structure is then

Veo Yi te
(2 + frBo havin + fora x ia
|
Vt Io = iS)
fo fo
If we now impose the cross-equation constraint y,, + y2, = 0 on the original
structure, the same condition on the transformed structure gives
ae fallFa =

or fay = 9
giving
fa, = 9
so that all admissible transformation matrices are diagonal and both equa-
tions are identified.
Alternatively the reduced form of the model is
Vi re
Vola (Bayi a Yo1) x 105. ¥11 (Ba am 1)x, + Uv,
The parameter y,, can be obtained from the first reduced-form coefficient and
8, can be derived from the second reduced-form coefficient, so that both
equations are identified.

Restrictions on the Variance Matrix

So far the only explicit assumption about the disturbances has been that of serial
independence, but we have made no explicit assumptions about contemporaneous
correlations between disturbances in different structural equations. Let
z= E(u’)
> is then a G X G matrix, the terms on the principal diagonal indicating the
464 ECONOMETRIC METHODS

variances (assumed constant) of the disturbances in the G structural equations


and the off-diagonal terms indicating the covariances between pairs of dis-
turbances. If specific restrictions can be placed on some of these elements, they
constitute an additional source of identifying power.
Let us examine first of all restrictions on covariances. Consider the model
Vit Vie 1
BoY) + V2 + YaiX1 = U2
As is easily seen, the first equation of this model is identifiable and the second is
not. We shall, however, examine the identifiability of the model again by
considering admissible transformation matrices, as this approach facilitates the
study of restrictions on variances and covariances.
Using
is fir fo
e- | ‘
the transformed first equation becomes
(fn + fia Ba) yi + fir
yy + (fivu + firYa1) 1 = fii + firte
If the coefficients of the transformed equation are to obey the same restrictions as
those of the original equation, we must have

fir + fi2Ba = 1
fiz = 9
giving f;, = 1 and f,, = 0. The only restriction on the second equation is the
normalization condition, which is held in abeyance. Thus admissible transforma-
tion matrices are given by
] 0
he B al
showing that the first equation is identified and the second not.
Suppose we can now postulate
O11 0
> —

0) 05>

The vector of disturbances in the transformed structure is Fu,, and so the


variance-covariance matrix for the disturbances of the transformed structure is
WV = E(Fuw,F’)
= | 3s)1”
This must obey the restriction that the covariance between the two transformed
disturbances is zero, that is,

~ wale Sf
f,2f5,= 0

Ori far sah


SIMULTANEOUS EQUATION SYSTEMS 465

that is,
fo), = 0
which gives
fy, = 9
The value of f,, is then settled by the normalization condition that the coefficient
of y, in the second equation must be unity. The coefficients of the transformed

giving the coefficient of y, in the second equation as f,,. Thus f,, = 1, and the

rls
only admissible transformation matrix is

so that both equations are identified.


As a further illustration consider the model
Yt n%) =
ByVy, + Y2 + YX, = Uy
B31¥1 + Byr
V2 + V3 + Y31X1 = U3
Without further restrictions only the first equation is identifiable. If, however, we
assume
o;, O 0
Dee): Ure y0555 0.0
0 0 653
the second and third equations become identifiable. Consider
16 “00 ] On Oly,
FA=|fi fo fs || Bu PO ays
fa fo fs || Bs B21 Ys
The normalization condition on y, in the second equation and on y, in the third
give
ha + fo3B32 = 1
fag = 1

+ It is convenient algebraically, but not necessary, to impose the normalization condition on the
coefficients of y, and y, in the first and second equations, respectively, of the transformed structure.
The absence of y, from the first equation gives |, = 0. The zero covariance term then gives f,, = 0
and so the class of admissible transformation matrices is

ea
r-(4 |
which secures the identification of both equations.
466 ECONOMETRIC METHODS

and the exclusion of y, from the second equation gives


hoz = 9
which also implies
fro = 1
Thus F is now
1 Or 0
F=|fr | 9
fy fro 1
We have not yet considered the effect of the zero covariance restrictions 0), = 0)3
= 0), = 0. These must be satisfied by the transformed structure. Hence
f,2, = 0
f,2f, = 0
f, =f, = 0
The first of these gives

a, 0 0 hry
[1 0° O} 0) <o55% 0 1 | =,0;,=0
0 Ona to

so that
hay = 0

and in a similar fashion the second and third conditions gives f;,; = 0 and f;, = 0.
Thus the only admissible transformation matrix is
Ee0a 0
F = \0ipele0
Oe Oe
and all three equations are identified.
The above model has two special features, namely, a triangular B matrix and
a diagonal = matrix. The presence of these two features defines a recursive system.
All the equations of the recursive system are identified and, as we shall see below,
simple estimation procedures are available for this model.
Zero covariances can aid identification and not necessarily just in recursive
systems. For example, in
Yt Byy2 =
Bo¥, + Yo + Yn1%1 = U2
the first equation is identified and the second is not. However, the additional
specification o,, = 0 would serve to identify the second equation as readers can
easily prove for themselves. There is no simple necessary and sufficient condition
for the zero covariance case as there was for restrictions on the B and y
parameters, so each case must be examined from first principles.
SIMULTANEOUS EQUATION SYSTEMS 467

The discussion has dealt only with models which are linear in variables and
parameters. Many realistic models, however, may be nonlinear in variables
and/or a priori restrictions. Identification theory for such models is difficult and
has only been partially developed. Owing to the unsatisfactory state of the theory
it will not be summarized here. Interested readers should consult Fisher.

11-3 ESTIMATION OF SIMULTANEOUS EQUATION MODELS

Whether we wish to estimate an equation which is one of a set of equations


constituting a complete model or whether we wish to estimate all the equations of
a model we are in a situation where OLS and the variants of OLS that we have
considered so far in the context of a single-equation model are, in general,
unsatisfactory estimating techniques. If OLS is applied to an equation in a model,
there will usually be more than one current endogenous variable in the relation,
and whichever variable one selects as the “dependent” variable, the remaining
endogenous variable(s) will generally be correlated with the disturbance term in
the equation so that OLS estimates will be biased and inconsistent. Only in the
case of recursive models will OLS be an optimal estimating technique.
In the more general simultaneous case, where the special assumptions of a
recursive system are not fulfilled, the main estimating techniques are indirect least
squares (ILS), two-stage least squares (2SLS), both of which may be interpreted
as IV estimators, limited-information maximum likelihood (LIML), three-stage
least squares (3SLS), and full-information maximum likelihood (FIML). ILS,
2SLS, and LIML are essentially single-equation methods, in which attention is
focused on one equation at a time without using all the information contained in
the detailed specification of the rest of the model. 3SLS and FIML are system
methods, where all the equations of the fully specified structural model are
estimated simultaneously.

Recursive Systems
As we have seen already, the two crucial features of a recursive system are a
triangular B matrix and a diagonal 2 matrix. As an illustration consider the
model
Viet WX = Ue
Boi Vis + Yar + Yai, = Ure
with the specification
o,, O
E(uu’) = = =
0 05>

+ F. M. Fisher, The Identification Problem in Econometrics, McGraw-Hill, New York, 1966,


Chap. 5.
468 ECONOMETRIC METHODS

To explore the connection between the y’s and the u’s we look at the reduced-form
equations which are
Vie = TVM1% + U1

VD aa (BY ial Yo) x; a (uy, = Bo, U1,)


The first equation is the same in each case. Since the exogenous variable x is by
assumption uncorrelated with the u’s, the first equation may be estimated
consistently by OLS. The second reduced-form equation shows y,, to be a
function of both u,, and u,,. Thus it would be inappropriate to estimate the
second structural equation by an OLS regression of y, on y, and x. However, y,, 18
uncorrelated with w,, since it is a function only of u,,, which has zero correlation
with w,,. Thus an OLS regression of y, on y, and x will yield consistent estimates
of the second structural equation.
More generally, the disturbance vector in the reduced form of a model is
a

(11-30)

When Bis lower triangular, then so is B~'. Thus Eq. (11-30) gives
Vir =f (uy)
Yor =f (Us Uo,)
Vou = F (tags Uays M34)
0
Yor = f(Uygs ays <+s MGs)
The assumption of a diagonal = matrix then ensures that y,, is uncorrelated with .
u5,, that y,, is uncorrelated with u,,, and so forth. Thus the second structural
equation may be estimated consistently by an OLS regression with y, as the
dependent variable, the third with y, as the dependent variable, and so on.
It is also easy to show that if the u’s are normally distributed, OLS yields ML
estimates. As was shown in the previous section, the likelihood of the sample y’s,
conditional on the x’s, for the model
By, + Ix, =u,

is given by
aids ut 1 n

L= (20) "Bil" -(2 “Pexp|- aie v2",


t=1

For recursive systems |B| is unity and = and ="! are both diagonal. Thus finding
the B and f to minimize L is equivalent to finding the B and f to minimize
n

Lu, 2 'u,
t=1
SIMULTANEOUS EQUATION SYSTEMS 469

For a three-equation system this sum of squares is

]
91)
n : 1

S= a [u,, Uy, U3,| 0 6 0 Uy,


t=1 22
]
0 0 a U3,
0. So)

‘ Ur, Tee us
= — $+ SH
Re oolioe . -927-, 9 O33
Thus the partial derivatives of In L with respect to the coefficients of the ith
structural equation are simply the partial derivatives of
2
se
ga Oi

Setting these partial derivatives to zero gives the OLS equations for the ith
structural equation. Thus under the special assumptions of the recursive model
the OLS estimators of the structural equations will have the desirable properties
of consistency, asymptotic normality, and efficiency. They will also have the usual
small sample properties.

Indirect Least Squares


As indicated in Sec. 11-1, ILS is a feasible estimation technique for an equation
which is just identified. The first step consists of estimating the matrix of
reduced-form coefficients by the application of OLS to each of the reduced-form
equations. The estimates of the structural coefficients are then obtained from the
algebraic relations existing between structural and reduced-form coefficients.
The structural model at time period ¢ has been written as
By,+ Ix, =u, (11-31)
where

Vit X11

y21 X91
Y= Ne ands ky =

YGt X kt

are, respectively, the G X 1 vector of observations on the jointly dependent


endogenous variables at time ¢ and the K X 1 vector of observations on the

+ For a proof that the usual small sample inference procedures apply see E. Malinvaud, Statistical
Methods of Econometrics, 2nd edition, North-Holland, Amsterdam, 1970, pp. 679-681.
470 ECONOMETRIC METHODS

predetermined variables at time ¢. Let us define Y and X as


/ iy

OY ees aaa
/ /

er kop a =) he ae
Y= ; xs ;
/ /
7 Yn a i xX, oot

so that Y is the n X G matrix of the sample observations on the endogenous


variables and X is the n X K matrix of sample observations on the predetermined
variables. From Eq. (11-31) we then have
YB’ + XI’ =U (11-32)
where U is the n X G matrix of all the sample disturbances. The reduced form
may then be written
Y=xXII’+V (11-33)
where
I’ = -I(B’) (11-34)
and V =U(B)|
The matrix of reduced-form coefficients defined in Eq. (11-34) is simply the
transpose of the matrix previously defined in Eqs. (11-19). The estimation of IT’ is
accomplished by applying OLS to Eq. (11-33) giving
P’ = (X’X) 'X’Y (i235)
This yields the set of estimated reduced-form coefficients for the first stage of ILS.
Let us denote the equation we are interested in estimating by
y=Y,6P+X,y+u (11-36)
where y=n X 1 vector of observations on the dependent (endogenous) variable in
the equation
Y, =n X(g—1) matrix of observations on the other g—1 current endoge-
nous variables in the equation
X,=n Xk matrix of observations on the k predetermined variables in the
equation
u=n X1 vector of disturbances in the equation.

Rewriting Eq. (11-36) gives


1
—B| =u
[y Y, X,]
may,
or, more fully,
SIMULTANEOUS EQUATION SYSTEMS 471

where Y, and X, are matrices of observations on G — g endogenous and K — k


predetermined variables which are excluded from the equation.
The relations between structural and reduced-form equations are given in Eq.
(11-34), which may be rewritten as
VB’ = =."
The relations holding for the coefficients of the structural equation (11-36) are
then

| -B =|5
0 (11-37)
0
J J y
KX Ge Goo) Kx 1

Substituting in this from Eq. (11-35) gives the ILS coefficients as the vectors b and
c obtained by solving

: c
(XX) “X’Y =p} = i (11-38)
0
The crucial question is whether there are unique solution vectors b and c.
Rewriting Eq. (11-38) as

l c
(XX) 'X’[y Y, Y,] Fi |0|
0
gives

(X’X)'X’y = (X’X) 'X’Y,b — lo (11-39)


Premultiplying by (X’X), partitioning X as [X, X,], and rearranging gives the
pair of equations

(XY,)b+ (Xi X,)e = Xiy (11-40)


(X,Y,)b+ (XX,)e= Xyy (11-41)
Together these constitute K equations in (g — 1) +k unknowns. Since the
necessary condition for exact identification is
Khe
we have the same number of equations as unknowns so that, in general, Eqs.
(11-40) and (11-41) solve uniquely for the ILS estimates b and ec.
These equations also indicate how the ILS estimates may be interpreted as IV
estimates. Returning to the structural equation
y=Y,B+Xy+t+u
the inconsistency of OLS arises from the correlations between Y, and u. The x
variables, however, are uncorrelated with u, and in the exactly identified case X,
472 ECONOMETRIC METHODS

will have the same number of columns as Yj. This suggests using
[X, X,]
as the set of instruments for [Y, X,]. The resultant IV estimates are given by

X5Y, X)X, || diy i zi


MY, XX, | civ Xiy

which are identical with Eqs. (11-40) and (11-41). Notice that the ordering of the
instrumental variables is unimportant. We can just as well take
x= [X, X,]

as the matrix of instrumental variables. Rewriting Eq. (11-36) as


y=Zd+u
where

z=
1
(Yee
1
xsl] and ||

the IV estimator of 8 is
b z,
diy = | I= (X’Z,) 'X’y (11-42)
Civ
which is easily seen to be identical to the ILS estimator defined in Eqs. (11-40)
and (11-41).

Two-Stage Least Squares


In practice ILS is not a widely used technique since it is rare for an equation to be .
exactly identified. 2SLS is perhaps the most important and widely used proce-
dure. It is applicable to equations which are overidentified or exactly identified.
Moreover, it turns out that in the case of an exactly identified equation the 2SLS
estimates are identical with the ILS estimates given by Eqs. (11-40) and (11-41).
Consider again the estimation of the equation
y=Y,P+Xyt+u
where the necessary condition for identification requires that
Kies
As we have seen, the trouble about applying OLS directly to this equation is that
the embedding of the equation in a simultaneous equation model makes the
variables in Y, correlated with u. The 2SLS technique consists of replacing Y, by a
computed matrix Y,, which hopefully is purged of the stochastic element, and
then performing an OLS regression of y on Y, and X,.
The matrix Y, is computed in the first stage by regressing each variable in Y,
on all the predetermined variables in the complete model and replacing the actual
SIMULTANEOUS EQUATION SYSTEMS 473

observations on the y variables by the corresponding regression values. Thus

Wi = X(X’X). XY, (11-43)


In the second stage the regression of y on Y, and X, yields the estimating
equations

YY, YX, b Yiy


(11-44)
X.Y, XxX, c Xiy
where [>] now denotes the 2SLS estimator of er For the actual estimation
there is no need to compute the regression values in Y, explicitly. An alternative
form of Eq. (11-44) can be derived which involves only the matrices of actual
observations. The matrix Y, can be written as
Y,=Y,+YV,
where Y, is given by Eq. (11-43) and V, is the n X (g — 1) matrix of OLS
residuals. The usual properties of OLS residuals give
VV = 0
and X’V, = 0

Thus VV VY, =v)


iF ay,

= Y/X(X’X) 'X’Y,
and VX, = (ViANV,)X,
= YX,
Thus the equations for the 2SLS estimator can now be written

YX(KR) WY VEX, |b] | VEXOWX) Xy (11.45)


X.Y, x’ X,|} ¢ X’y
Yet another form of the 2SLS equations, which is useful for further theoretical
developments, is

YY; - Viv, YX b = (Vea)y (11-46)


xy, XX, c XY

The equivalence between Eqs. (11-45) and (11-46) may be proved by the reader as
an exercise.

Example 11-8 The first structural equation in a three-equation model is


Vie = BroVae + VWiX1e + Vi2X%ae + Ue
There are four predetermined variables in the complete model, and the X’X
474 ECONOMETRIC METHODS

matrix is
—_

X’X =
SY >
nT
oo
Se)"=) oO
ohowore

In addition we are given


ie
Plo eee
F 7 x= |? 0 2 1
The necessary condition for identification is satisfied since K — k = 2 and
g — 1 =1 so that the equation is overidentified. To estimate the parameters
by 2SLS we need to establish a correspondence between the data in this
problem and the vectors and matrices in Eq. (11-44). Thus
| | esa [ims]
Yan Y,= | ¥2 X,=|%1 Xe X,=|%3 %4
| | | peal ree
YX=[1 0 2 1] Y{X,=[l 0]
2
; 10 O F 3 , Z

]
and so

WVX(X°K)
©XY) = [1 te eet |

2
Y’X(X’X) 'X’y =[0.1 0 05 0.5] ; OG
1
The 2SLS equations are then
1:65, 1 Ocha. 2
Li 910% 0) ine raeaen ee
0 0 © Sre5 3
with solution
bi 1.6667
Crp P= 11020333
Cp 0.6000
Example 11-9 For a model
Vie = Bio aetna
Yor = Bar Wie + Yo2%or + 23X31 + Ur,
SIMULTANEOUS EQUATION SYSTEMS 475

The sample matrices are}


l = 0 10
XX=/0 20 0 XY = 140 20
Os 0) 10 20.30
We will illustrate the application of 2SLS: and ILS, as appropriate, to this
model and also look at the estimation of the reduced-form coefficients.
The first equation is overidentified and is estimated by 2SLS. The
correspondence between the variables in the equation and the matrix expres-
sions in Eq. (11-44) is given by

y= |" Y, = |¥2 X, =] *1 X,=|X2 Xs;


| oe!
Thus
10 5
X’Y, = | 20 XY, = 10 X’y = | 40 Xiy =5 X/X, = 1
30 20
P00. Ot ia0
Y/X(X’X)'X’Y, = [10 20 30]/0 0.05 0 || 20
Ory Ocho
10
=[10 1 3]]20] =210
30
5
Y/X(X’X) 'X’y =[10 1 3]
0 = 150
20

Ps el-[4
The 2SLS equations are thus

with solution

loaBey
The second equation is just identified and thus may be estimated by
2SLS or ILS. For the 2SLS approach

| | ae |
y=|%2 Y= ("1 X,=|*%2 %3} X,=|*1

+In this and the previous example the X’X matrices are assumed to be diagonal to keep the
arithmetic simple. In realistic situations orthogonal variables are, of course, very rare.
476 ECONOMETRIC METHODS

Thus
5 40 10
XVA="40 V5 ie X’y = | 20
20 30
7 4a 20 ele 4
A= ba ag |0 10
lowe 0 5
Y/X(X’X) 'X’Y, =[5 40 20]]}0 0.05 0 || 40
0 0 0.1} 20
5
=[5 2 | 0|- 5
20
10
Y/X(X’X)'X’y =[5 2 2]] 20] = 150
30
The 2SLS equations are thus
145 40 20}] 5), 150

20 QO: 10 |Hhes 30

with solution
by, D
> |= \'—3

To obtain the ILS estimates of the second equation we need to specify


the additional matrices appearing in Eq. (11-42). These are
XY, = 5 x,X, =[0 0] X,y = 10
The ILS equations are then
5 207 0455, 10
AN 20) aeOllie eleeeloe
20oO 0 10]\c,, 30
with the same solution vector as 2SLS. This is an illustration of a general
result that 2SLS and ILS estimates, where the latter exist, are identical. The
general result will be proved below, but in the meantime we continue with the
numerical example.
The reduced-form coefficients, estimated by OLS, are

P’ = (X’X) ‘X’Y
LO 0 ar ah!
= Obs OOS eer 40 20
OnRO OF 120> 30
Selo
=r 2 1
ee
SIMULTANEOUS EQUATION SYSTEMS 477

giving
Viper ekap eka Oe
and Vor LU take, teks es,
The reduced-form matrix may also be estimated by substituting the estimated
structural coefficients B and f in Eq. (11-19),
i= -B-'f
However, care must be taken in making this substitution since Eq. (11-19)
was derived from the structural equations specified as By, + Tx, = u,, whereas
the equations of this model have been specified with just a single endogenous
variable on the left-hand side of each equation. The 2SLS estimates of the
structure are
Valea 1) 21me 4tx1, + uy,
Vp = Vi, — 3g, — Xz, + Ud,
Rearranging with all variables on the left-hand side gives
Xt
| eee Bole are On| ee =e
—2 1 y20 0 Se ery Uy,
3t

Thus

fe Lt eet “N44 0 0
ae 1 Ore ar ei
A SPss as
10> %35
These are somewhat different than the OLS estimates. The reason is that the
OLS estimates are unrestricted and thus fail to satisfy the restrictions placed
on the reduced-form parameters by the overidentification in the system. With
two endogenous and three predetermined variables there are six reduced-form
coefficients, which are functions of just five structural coefficients. The true
reduced-form matrix is
1 Yu Bir¥x. Bi z¥x3
II =

(1 i BB) BY Y22 Y23


in which the second and third columns are linearly dependent.

Interpretation of Two-Stage Least Squares as an Instrumental Variable


Estimator
The structural equation to be estimated may be written as
y=YP+Xytu=Zd+u (11-47)
where

Z,=([Y, X,] and 6= a


478 ECONOMETRIC METHODS

Let us recapitulate the discussion of IV estimates in Chap. 9, with the vector of


unknown parameters now indicated by 8 rather than B and with the matrix of
explanatory variables simply indicated by Z. The equation to be estimated is
y = Z8 + u, the problem being that
eA etees:
plim(—-Z'u) + 0

which is the difficulty with Z, in Eq. (11-47). Provided a matrix W can be found
such that

le plim(- ww] = oe, a finite symmetric positive definite matrix


ala
2. plim(—W’Z] = 2 a finite nonsingular matrix

S plim(—-W'u] =0

the IV estimator

dy = (WZ) ‘Wy (11-48)


will be consistent and will have an asymptotic variance matrix estimated by

asy var(dyy) = s2(W’Z) '(W’W)(Z’/W) | (11-49)


ea (y eS Zd 1 )'(y = Zdy )
where
n

In the present case let us set

Le Tila)

and W= [Y, Xx,|

so that Y, is the set of instruments for Y,. The IV estimator defined in Eq. (11-48)
is then

VV eye x! oy i Yiy
xy 11-50
(11-50)
xiY, Xx XxX, Civ
but we have already seen that Y/Y, = Y/Y, and Y/X, = Y/X,. Thus Eqs. (11-50)
and (11-44) are identical, so that 2SLS is in fact an IV estimator with Y, as the
instruments for Y,.
The consistency of the 2SLS (IV) estimator requires the three conditions on
W, stated above, to be fulfilled. We will assume that

ale!
plim|Ww) and plim|,W72}
SIMULTANEOUS EQUATION SYSTEMS 479

are both finite.t The third condition is

1 plim(+ Yu]
plim(—-W'] = ¢ =0
ek ore le,
plim| Xiu}

Insofar as X, contains exogenous variables, whether current or lagged, these are,


by assumption, uncorrelated in the limit with the equation disturbance. The same
result will also hold for any lagged endogenous variables in X, provided the
disturbance term is serially uncorrelated. The remaining term is
its © aia :
plim{--Yu] = plim(—¥;X(X°X) Xu)
ee e AY ALSIRAON RS OAR ialhe
= plim(—Y;x . plim|7 xX x| : plim|5 xX u]

=0
since the first two terms are finite and the last is the zero vector.
It was also shown in Sec. 9-2 that the IV estimators are asymptotically
normally distributed with an asymptotic variance matrix estimated by Eq. (11-49).
Substituting for W and Z and using the fact that Y/Y, = Y;Y, and Y;X, = Y;X,
gives

b NY UN |
asy var =a in
c KYA DKK)
Se WX(XX)
/ , =A
XV Gn YX |
r , =
nn
XY, xX,
where
2 = = Vib — X,0)(y — Yb = Xo) (11-52)
n

which is a consistent estimator of «7. Some authors prefer to use the number of
degrees of freedom n — g — k + 1 as the divisor in s” rather than n. This is also a
consistent estimator of 0,7. The 2SLS estimators are thus consistent and asymptot-
ically normally distributed with estimated variance matrix given in Eq. (11-51).
A problem sometimes arises in the application of 2SLS to equations in
medium-size or large-size econometric models. The difficulty is that the number of
predetermined variables in such a model may become large in relation to the
number of observation points. Suppose, to consider a special case, that the
number of predetermined variables becomes as great as the number of observa-
tions, K = n. The X matrix is then square and, in the absence of any exact linear

+ The detailed conditions for this to be true are set out in H. Theil, Principles of Econometrics,
Wiley, New York, 1971, pp. 484-488.
480 ECONOMETRIC METHODS

relations between the predetermined variables, nonsingular. Formula (11-43) thus


reduces to
Y, = X(X’X) ‘X’Y,
Sere)

and 2SLS is equivalent to OLS. The 2SLS estimates would, of course, no longer
be consistent, since the matrix of instrumental variables is now W = [Y, X,] and
ial
plim(—-Yiu +0 so that plim(--W') +0

as was required for consistency.


When K > n, the X’X matrix is of order K X K and of rank n. Thus it is
singular, and the inverse (X’X)' does not exist. This has often led to the
conclusion that the 2SLS estimator will not exist, since Eq. (11-45), for example,
involves (X’X)~!. Fisher and Wadycki have pointed out that this is not necessarily
the case.} They argue that the Y, matrix will be unique in spite of the multiplicity
of solutions for the reduced-form coefficients. Consider, for instance, the first
variable in Y, and denote the n X 1 vector of observations on that variable by y,.
Letting p denote the K xX 1 vector of OLS reduced-form coefficients for that
variable, the usual formula gives
(X’X)p = X’y, (11-53)
Since X’X is of order K X K with rank n < K, Eq. (11-53) has an infinity of
solutions. Letting p, and p, be any two solution vectors, we have

(X’X)p, = X’y,
(X’X)p, = X’y,

Thus (X’X)(p, — p,) = 0


Premultiplying by (p, — p,)’ gives

(p; — P.)’(X’X)(p, — p,) = 0


Thus X(p; — Pp) = 0
so that ¥, = Xp, = Xp,
Moreover Eq. (11-53) may be rewritten as

X’(Xp — y,) = 0
J J
KXn nx1

Since X’ has rank n (< K), the only solution vector is Xp — y, = 0 so that

7+W. D. Fisher and W. J. Wadycki, “Estimating a Structural Equation in a Large System,”


Econometrica, vol. 39, 1971, pp. 461-465.
SIMULTANEOUS EQUATION SYSTEMS 481

¥, = y,. The same result will hold for each variable in Y, so that once again
Y, = Y,, and 2SLS would be equivalent to OLS.
Various suggestions have been made for dealing with the problem of an
excess of predetermined variables. Kloek and Mennes suggested replacing X, in
the first-stage regressions by a smaller number of principal components.} Let F
denote the n X / matrix of the / chosen principal components and then define

Z=[X, F]
This Z matrix takes the place of the X matrix in Eq. (11-45), and the 2SLS
estimates based on the principal components approach would then be given by

WD) LN Ne Xl bee I VL2Z)) Ly


; ; : (11-54)
XY, XX, || Cec X‘y
Various problems arise with this approach. The first concerns the number / of
principal components to be used. Kloek and Mennes state that identification
requires
i oe]
but it is difficult to see the reason for this condition since the problem is to find a
suitable matrix Z for the first-stage regressions in which Y, is replaced by an
estimated matrix Y,. It is, of course, true that identification of the structural
equation requires that the number of columns in X,, namely, K — k, should be at
least equal to g — 1, but there is no reason to carry this condition over to the
choice of variables used in computing Y..
A second problem concerns the criterion to be used in selecting principal
components. One possibility is to choose the components with the greatest
eigenvalues, that is, the components which account for the greatest variance of the
variables in X,. Some of these components, however, may be highly correlated
with variables in X,, thus providing little additional assistance in explaining Y,
and possibly also causing Z’Z to be nearly singular, so that numerical difficulties
arise in computing the inverse. Kloek and Mennes have suggested components
which have the /east correlation with the X, matrix. Both approaches involve
substantial computation and also imply different sets of principal components for
different structural equations.
The last difficulty is avoided by calculating, once and for all, principal
components of the complete set of predetermined variables and using a subset of
these in the first-stage regressions for each structural equation. In a very interest-
ing study Klein estimated a revised version of the Klein-Goldberger model of the
U. S. economy by using just the principal components corresponding (1) to
the four largest and (2) to the eight largest eigenvalues of X’X in the first stage of
the 2SLS procedure. Comparing the predictions of GNP in the sample period

+T. Kloek and L. B. M. Mennes, “Simultaneous Equation Estimation Based on Principal


Components of Predetermined Variables,” Econometrica, vol. 28, 1960, pp. 45-61. For a review of
principal components see App. A-10.
482 ECONOMETRIC METHODS

from these two estimators with OLS and FIML, Klein found 2SLS based on just
four principal components to give the smallest absolute percentage error followed
by the other 2SLS estimator, OLS, and FIML in that order.+
An alternative approach based on instrumental variables has been suggested
by Brundy and Jorgenson to bypass the substantial computation involved in
calculating the reduced-form coefficients required for Y,.¢ Let
E(Y,) = XII,
where II, is the K X (g — 1) submatrix of reduced-form coefficients relevant to
the variables in Y,. The Brundy-Jorgenson suggestion is as follows.

1. Define a matrix of instrumental variables as

w, = [xfl, X,] (11-55)


where IT, is any consistent, estimator of IT 1°
2. Then compute the structural coefficient estimator from the IV formula as

a= [>] = (Wiz, "wy (11-56)


where Dye N, ox

The regular 2SLS estimator satisfies these conditions, for ne = (X’X) 'X’Y,
is a consistent estimator of II, and W, then becomes [Y, X,]. The novelty of the
Brundy-Jorgenson approach is to avoid computing reduced-form coefficients and
to derive an appropriate I, by first obtaining B and fas consistent estimators of
B and [ and then using

bat
from which the relevant submatrix II, can be extracted and XII, computed for
insertion in Eq. (11-55). Thus even if one is interested in Just a single structural
equation, this approach requires the initial computation of consistent estimators
of all structural coefficients. On the other hand, if one is estimating all
the
equations of a model, the single [I matrix is used to provide all relevant IT,
submatrices.
Several suggestions are offered for initial consistent estimation of the B and
r
matrices, all of them essentially IV estimators. Considering Eq. (11-47) again,
the
matrix of right-hand side variables is

Z, = [Y, X,]
where Y, is n X (g — 1) and X, isn X k. Define

Wi = [xt X,]
7 L. R. Klein, “Estimation of Interdependent Systems in Macroeconometrics,” Econometrica, vol.
STNO6SS ppl 92»
¢J. M. Brundy and D. W. Jorgenson, “Efficient Estima
tion of Simultaneous Equations by
Instrumental Variables,” Review of Economics and Statistics,
vol. 53, 1971, pp. 207-224.
SIMULTANEOUS EQUATION SYSTEMS 483

where X7 is the matrix of any g — 1 predetermined variables which do not appear


in the first structural equation. These variables could be chosen from the
predetermined variables appearing in the structural equations for Y,, as suggested
by Fisher.j The resultant IV estimator of 8 is (W#*’Z,)” 'Wi’y. Repeating this
procedure for each structural equation yields the preliminary consistent estima-
tors B and f for insertion in I!= —B~'f, and the computations outlined in Eqs.
(11-55) and (11-56) would then yield the final estimator. Another possibility is to
define

= [F, X,]
where F, is a subset of g — 1 principal components of X. This differs, of course,
from the Kloek and Mennes procedure, where the principal components were
used in quasireduced-form estimation to compute Y,. Here the principal compo-
nents are used as instrumental variables ina first-round estimation of structural
coefficients. The Brundy-Jorgenson estimator is known as the limited-information
instrumental variables efficient (LIVE) estimator. The asymptotic variance-covari-
ance matrix for d is estimated by

asy var(d) = s2(W/W,)| (11-57)


where W, is defined in Eq. (11-55) and

2 — LY = ZA)'y ~ Za)
n

The LIVE estimates can thus be computed even where the 2SLS estimates cannot,
but the actual point estimates will, of course, vary with the variables chosen as
instruments.

Limited-Information Maximum Likelihood (Least Variance Ratio)


Estimators
This alternative approach to the estimation of a structural equation preceded the
development of 2SLS, which has largely replaced it on grounds of greater
simplicity. Consider again the structural equation
y=YB+Xyt+u
and rewrite it as

Y,By - Xv =u (11-58)
where

Vly Yj ovand Pa (11-59)

+ F. M. Fisher, “Dynamic Structure and Estimation in Economy-Wide Econometric Models,” in


J. Duesenberry, G. Fromm, L. R. Klein, and E. Kuh, Eds., The Brookings Quarterly Econometric
Model of the United States, Rand-McNally, Skokie, IL, 1965, pp. 589-636.
484 ECONOMETRIC METHODS

Let us suppose that the endogenous variables have been so numbered that Y,
constitutes the first g such variables and likewise that X, refers to the first k
predetermined variables. The likelihood function for the endogenous variables in
Y, will involve the parameters in the first g rows of the reduced-form matrix II.
Let these rows be partitioned into the two submatrices [II,, II,,] which are of
order g X k and g X (K — k), respectively. We know that
BIT
= -T
The first row of each side of this equation may be written

[Bx 0,JII=[-y' 0]
where 0, indicates a row vector of G — g zeros and 0, a row vector of K — k
zeros. Using the partitioning of II then gives

Bl ante (11-60)
BxIT,, = 0, (11-61)
Eq. (11-61) constitutes K — k homogeneous equations in the g elements of By.
However, one of the 8’s has been set at unity so that we merely need to determine
the ratios of the elements in B,. This can be done uniquely if the rank of II,> is
g — |. Even in the overidentified case where K — k > g — 1 and II,, thus has g
rows and at least g columns, the rank of II,, cannot exceed g —) a eats ts
obvious intuitively since Eq. (11-61) is just a subset of equations from BII = —T,
which gives the relations between the true structural coefficients and the true
reduced-form coefficients. However, the true II,, is unknown, and when it is
replaced in Eq. (11-61) by, say, the ML estimate Ihe this matrix in the
overidentified case will almost certainly have rank g so that one cannot solve
for
nonzero B,, except by arbitrarily dropping one of the equations.
The limited-information maximum likelihood (LIML) approach is to maxi-
mize the likelihood function for the g endogenous variables in Y, subject
to the
restriction that e(II,5) = g — 1. This approach was developed by Anderson
and
Rubin.¢ The application of the method requires one to know, in additio
n to the
specification of the equation being estimated, merely the predetermined
variables
appearing in the other equations of the model, as in 2SLS. The
mathematical
development of the LIML estimator is complicated and lengthy,
but it may be
shown that it reduces to the choice of the elements of B, to minimize

jo B’Wars By
Zi Br Wa By (1 #62)

7 See W. C. Hood and T. C. Koopmans, Studies in Econome


tric Method, Wiley, New York, 1953,
pp. 185-186.
¢T. W. Anderson and H. Rubin, “Estimation of the Paramet
ers of a Single Equation in a
Complete System of Stochastic Equations,” Annals of Mathematical
Statistics, vol. 20, pp. 46-63, 1949.
SIMULTANEOUS EQUATION SYSTEMS 485

where Wy‘, and W,, are certain matrices of residuals.+ The explanation of these
residuals is given in the following account of least variance ratio (LVR) estima-
tors.
Rewrite Eq. (11-58) as
z=X,ytu
where z= YP
so that the z vector is a linear combination of the endogenous variables appearing
in the equation, the coefficients of the combination being the unknown B
parameters. If z is regressed on X,, the residual sum of squares is

v'z — 2'X,(X,X,) /Xiz = BYLY,By — BYYAX,(X,X,) XV, By = BKWats Bs


where Ws, = Y¥xY, — WX OX GXG ie Xe (11-63)
Similarly, if z is regressed on all the predetermined variables, X = [X, X,], the
residual sum of squares is
BxWaaB
where Waa = YY, — Y{X(X’X) 'X’Y, (11-64)
The second residual sum of squares will be no greater than the first since the
second regression includes all the explanatory variables in the first regression X,
plus the set X,. However, the specification of the structural equation asserts that z
depends on X, but not on X,. Thus the LVR principle suggests that the estimate
of B, should be chosen to keep this reduction in the residual sum of squares as
small as possible, that is, to minimize the ratio

l
_ BiWesBs
BxWa Ba
which is the same criterion as that for the LIML estimator. Differentiating / with
respect to B, and setting the result equal to the zero vector gives
(Wiis — 7Wys)By = 0 (11-65)
This set of equations will only have a nontrivial solution if the determinantal
equation
|Waa a LWy al =)
is satisfied. This gives a polynomial in /, which must be solved for the smallest
root /. This root is substituted back on Eq. (11-65) and the estimator B, obtained

+ T. W. Anderson and H. Rubin, op. cit.; see also W. C. Hood and T. C. Koopmans, op. cit.,
Chap. 6. Hood and Koopmans arrive at Eq. (11-62) by a different method from the original approach
of Anderson and Rubin, who maximized the likelihood function subject to appropriate constraints by
using Lagrange multipliers. Hood and Koopmans start with the likelihood function for the complete
model of G equations for all G endogenous variables and then, by a series of stepwise maximizations,
eliminate from the likelihood function all parameters other than those of the equation to be estimated.
Finally, even y is eliminated and the concentrated likelihood function expressed in term of By.
486 ECONOMETRIC METHODS

from

(Wit — 7Wy.)B, = 0 (11-66)


by setting the first element of 8, equal to unity. Defining

a= Y, By
and regressing Z on X, gives

9 = (XiX,) Xi¥sB, (11-67)


Equations (11-66) and (11-67) define the LIML estimates of the structural
equation. The LIML estimators have the same asymptotic variance-covariance
matrix as 2SLS. The estimates of the asymptotic variances, however, will differ
since s* is computed from the estimated structural coefficients, which will be
different in the two cases.

Three-Stage Least Squares and Full-Information Maximum Likelihood


The estimators considered so far, namely ILS, 2SLS, LIVE, and LIML, are all
essentially limited-information estimators in that in the estimation of any struc-
tural equation complete information on all the other structural equations in the
model is not taken into account.j In principle information on the complete
structure, if correct, will yield estimators with greater asymptotic efficiency than
that attainable by limited-information methods. There are two main full-informa-
tion methods, namely, three-stage least squares, (3SLS) and full-information
maximum likelihood (FIML).
The initial development of 3SLS is due to Zellner and Theil.¢ Consider again
the general linear model containing G jointly dependent endogenous variables
and K predetermined variables. The ith equation may be written
y, = Y,B; + X,7, + u, (11-68)
where y, is an n X | vector of sample observations on the dependent variable in
the ith equation, Y, is an n X g, matrix of observations on the other endogenous
variables in the equation, X,; is an n X k, matrix of observations on the prede-
termined variables in the equation, B; and y, are vectors of structural paramete
rs,
and u, is a vector of disturbances. Rewrite Eq. (1 1-68) as

y,= 28, + u,; (11-69)


where L, = |X) Xela and aoe F
If Eq. (11-69) is premultiplied by X, the n x K matrix of all the predete
rmined
variables in the model, then

Xy, = XZ;8,+ Xu, i= 1,...,G (11-70)


y An exception is the LIVE estimator where initial estimat
es of the B and T matrices are made to
derive an estimate of IT.
¥ A. Zellner and H. Theil, “Three Stage Least Squares:
Simultaneous Estimation of Simultaneous
Equations,” Econometrica, vol. 30, 1962, pp. 54-78.
SIMULTANEOUS EQUATION SYSTEMS 487

The variance-covariance matrix of the disturbance term in Eq. (11-70) is


E(X’u,u’,X) = 0,,X’X (11-71)
on the assumption that E(u,u’,) = o,,1. Considering Eq. (11-70) as a relationship
between a dependent variable X’y, and explanatory variables X’Z,, the nonspheri-
cal disturbance matrix in Eq. (11-71) suggests using generalized least squares. The
GLS estimator of 8, is then
d, = [Z;X(X’x)'x’Z,]'Z/x(X’x)'x’y, (11-72)
Equation (11-72) is simply another way of writing the 2SLS estimator of Eg.
(11-69), as may be verified by substituting for Z,, multiplying out, and comparing
with the original expression for the 2SLS estimator in Eq. (11-45).
We may note in passing that Eq. (11-72) affords a simple demonstration of
the equivalence of 2SLS and ILS in the case of a just identified equation. The
order condition for exact identification of the ith structural equation is
K = k= ige = | or Act cei
Wi
Thus Z; is of order n X K so that X’Z; is of order K X K and may be assumed to
be nonsingular. In this special case Eq. (11-72) gives

d, = (X7Z,)”'(X’X)(Z;X)~'(Z;X)(X’X)'X'y, = (XZ) 'Xy,


which, from Eq. (11-42), is seen to be the ILS estimator for the ith structural
equation.
We also know from the discussion of GLS estimators in Chap. 8 that it is
possible to interpret the GLS estimator as equivalent to the estimator given by the
application of OLS to suitably transformed data. The present case may be so
interpreted, and this leads to a considerable simplification in the presentation of
the 3SLS estimator. Consider again Eq. (11-70) whose disturbance has a variance
matrix given by o,,X’X. Since X’X is positive definite, we know from Chap. 4 that
its inverse is also positive definite and that a nonsingular matrix P exists such that
(X’X)| = PP’ (eas)
from which it follows that
P’X’XP = I (11-74)
Premultiplying Eq. (11-70) by P’ gives
P’X’y, = P’X’Z,5, + P’X’u,
or w, = WS, + ¥; (11-75)
where w, = P’X’,
W, = PX’Z,
vy, = PX,
The variance matrix for the disturbance term in Eq. (11-75) is
E(v,v/) = E(P’X’u,u,, XP)
= 0,,P’X’XP
= 6,1 (11-76)
488 ECONOMETRIC METHODS

The application of OLS to Eq. (11-75) then gives

d, = (W/W,) 'W/w, (11-77)


which is easily seen to reduce to the 2SLS estimator in Eq. (11-72).
Collecting all G structural equations gives

" Ww, 0 Ol 5 Poly


Ww, 1 Vv,

= (PO ew, Cet Vg7 op ae tm Clisis)


Wo 0 W; a Vc

or, more compactly,

w= Wd+ Vv (11-79)
where the definition of the symbols in Eq. (11-79) is obvious from the comparison
with Eq. (11-78). The variance matrix for the v vector is

Ol opt --- ol
V= E(w’) =| 1 oyI +--+ gl |= Tel (11-80)
ani a Jo nee ay

The variance terms in Eq. (11-80) follow directly from Eq. (11-76). The typical
covariance term is

E(vy/) = E(P’X'uu’,XP) = o,1


Thus the basic assumption is that each structural equation has a homoscedastic
nonautocorrelated error term and that the disturbances in different structural
equations may be contemporaneously correlated. Provided that at least some 0;;
are nonzero, the arguments underlying the Zellner SURE estimator, already -
considered in Chap. 8, would suggest that any of the G equations defined by Eq.
(11-75) would be more efficiently estimated as a member of the complete set
defined in Eqs. (11-78) and (11-79). 3SLS is, in fact, simply the SURE estimator
of 8 in Eq. (11-79). The only difficulty is that the = matrix in Eq. (11-80) is
unknown. The Zellner-Theil suggestion is to estimate first each structural equa-
tion by 2SLS, giving the residual vectors
Oy Lid ae AG
where d; is the 2SLS estimator of 6,. The elements of = are then estimated by

Si
a7a,
ae
ie
for alli, j

giving

V=Se1
The 3SLS estimator of 8 is then

dssts = (WV
'W) (WV! w (11-81)
SIMULTANEOUS EQUATION SYSTEMS 489

with asymptotic variance matrix estimated by

asy var(d3sis) = (WV 'W)


Substituting in Eq. (11-81) for the elements of w and W, the 3SLS estimator may
be expressed in terms of the original data as

EUZX(N XR):NZ PO ASTRA RE XZ a SSIOTER WN Ze


Geeetlus eo X(N XS)” XZ pads PGR XX) XMZon us? RUZ
SZEX(X'X) 'X’Z, sZX(XX) XZ, SOZLX(XX) XZ
s s/Z,X(X'X) 'X’y,
e

Sel SID, X(X'K): X'y, (11-82)


~S

G
wel ZeX(X’X) Xy J
vt

where the s‘/ denote the elements in S~!.


A crucial question concerns the conditions under which 3SLS will be asymp-
totically more efficient than 2SLS. A necessary condition for the superior efficiency
of a full-information, or complete-system, method of estimation over a limited-
information method is that the specification of the complete model should be
correct. In many systems this is a formidable requirement, and the larger and
more detailed the system, the more difficult does it become. Even granted a
correct full-system specification, there are two conditions under which 2SLS and
3SLS will give identical point estimates with identical asymptotic sampling
variances. The first is
o,,=0
ij for alli + j

that is, the contemporaneous correlations between the disturbances in different


structural equations are all zero. The equivalence follows directly from the result
for the SURE model that a diagonal 2 matrix gives equality between the SURE
and OLS coefficients.} It may also be seen directly by substituting s’/ = 0 in Eq.
(11-82).£ The other condition under which one would find equivalence of 2SLS
and 3SLS estimators is all equations being exactly identified. We have already
seen that the order condition for exact identification of the ith equation leads to
the result that X’Z, is of order K X K and may be assumed to be nonsingular. The

+ See Problem 8-2.


. Notice that Eq. (11-82) refers to the feasible 3SLS estimator where =~! has been replaced by
=~ |. If 3 is diagonal, then so is S~' and s‘/ = 0 for all i + j. However, even if this condition is
satisfied, the s,; (and hence the s'/) estimated from 2SLS residuals will in general not vanish.
490 ECONOMETRIC METHODS

P matrix defined in Eq. (11-73) is also square of order K and nonsingular. Thus
W, = P’X’Z,
is K X K and nonsingular. This result holds for all i = 1,..., G. Thus the
block-diagonal W matrix defined in Eqs. (11-78) and (11-79) is nonsingular and
each component submatrix is nonsingular. The 3SLS estimator defined in Eq.
(11-81) may then be written

dogs = W'V(W’) 'WV-'!w


=W 'w
Under the same assumption the 2SLS estimator for the ith equation defined in
Eq. (11-77) reduces to

d, = W, l ‘Ww,
Thus the collection of 2SLS estimators for the complete system may be written

“t
d,
[wrt 0 Daal Mb)
dosts = <1 a 0 wW,' 0 : catia
dq 0 0 ars WwW, Wo

which is identical with d4,,


s.
So far we have assumed that all the structural equations in the model are
identified. Before attempting to apply 3SLS in practice one must omit all
unidentified equations and also all identities, since the latter have zero dis-
turbances which would render the = matrix singular. Suppose that there remain G
identified equations of which G, are exactly identified and G xx Overidentified.
Zellner and Theil have shown that the 3SLS estimator of the G x equations,-
treated as a complete group, is the same as that obtained from the application of
3SLS to the complete system of G equations. Thus it is computationally efficient
to obtain the 3SLS estimates in two steps. First compute the 3SLS estimates of
the overidentified equations. The 3SLS estimates of the Just identified equations
are then obtained by adding to the relevant 2SLS estimates a linear combination
of the 3SLS estimates of the overidentified equations.}

Full-Information Maximum Likelihood


As with 3SLS this is a complete system method of estimation. It is comput
a-
tionally more expensive than 3SLS as it involves the solution of
nonlinear
equations. We will merely sketch the outlines of the approach. Consider
again the
linear simultaneous equation model in G current endogenous variables
By, + Ix, =u, LN ei

+ The precise formula is given in A. Zellner and H. Theil, “Three Stage


Least Squares: Simulta-
neous Estimation of Simultaneous Equations,” Econometrica, vol. 30,
1962, p. 67.
SIMULTANEOUS EQUATION SYSTEMS 491

with BU)SSOF er Sa Ee
E(u) ==
If it is assumed that the G disturbances follow a multivariate normal distribution
>
we may write

any
Gey
1
P| 7l ue
"y-1
u,]

Assuming, in addition, that the u vectors are serially uncorrelated, the likelihood
for the n vectors u,,U,,..., u, is then

AUT reece a BAC)

= (27) "(det =) "exp| - s v2"


N|—
=|

The likelihood for y,, y,,..., y,, is


—nG/2
PAY ¥o.--+s
Ye)= 27)” detB|"n (det =) ~””
18/2,

xX exp -5 » (By, + Ix,)’27'(By, + rx) (11-83)


t=1

If we write

By, + Ix, = [B ma = Az,


t

the exponent in the likelihood in Eq. (11-83) can be written

l
|
vz A'S'Az, = — 5tr(ZA’E- 147’)
7—

— 5tr(B>'AZ/ZA’)
where

yx)
Z=(Y XJ=/y x,
yc xy,
is the n X (G+ K) matrix of observations on all the endogenous and prede-
termined variables. Defining

M = loz
n
tr(=~'AZ’ZA’) = ntr(=~'AMA’)
492 ECONOMETRIC METHODS

Table 11-1 Estimation methods in the models of Project Link

Total Number of
number of stochastic Estimation
Country Data} equations equations method

Australia Q 82 42 OLS
Austria A 128 54 OLS
Belgium Q 25 19 OLS
Canada A 183 44 OLS
Finland Q 144 60 OLS
France A 32 19 OLS
West Germany A 137 51 FIML
Italy Q 104 53 OLS
Japan Q 78 43 OLS
Netherlands A 87 13 LIML and 2SLS
Sweden A 133 75 OLS
United Kingdom Q 226 106 OLS
United States Q 207 70 OLS
Developing America A 12 11 OLS
Developing South and East Asia A 14 13 OLS
Developing Middle East and Libya A 10 9 OLS
Developing Africa less Libya A 11 10 . OLS

+Q—dquarterly data; A —annual data.


Source: J. Waelbroeck, The Models of Project Link, North-Holland, Amsterdam, 1976.

and so the logarithm of the likelihood in Eq. (11-83) may be written

L(A, 2) = constant + nIn|detB| — 5Indet =- 5tr(2-'AMA’) (11-84)


The FIML estimator results from the maximization of L(A, 2) with respect to
the °
elements of A and 2. The equations are nonlinear and computationally expensive,
though less so with each advance in computer technology. The asymptotic
variance matrix of the FIML estimator, however, turns out to be identical
with
that for 3SLS, thus indicating the asymptotic efficiency of the latter method.+
This feature, combined with its less severe computational problems, leads
some authors to recommend 3SLS over FIML. Most practical applications of
3SLS or FIML occur, not surprisingly, with fairly small models. What is perhaps
surprising is the continued dominance of OLS over all other methods, especial
ly
in the estimation of major econometric models. A recent study by
Waelbroeckt
documents the main features of the various countrywide econome
tric models in
Project Link. Of the 17 models summarized, OLS is the estimating
method in 15,
FIML is used in only one model, and a combination of LIML and
2SLS in the
remaining model. Details are given in Table 11-1.

+ See H. Theil, Principles of Econometrics, Wiley, New York, 1971, pp.


524-527.
+J. Waelbroeck, The Models of Project Link, North-Holland, Amsterd
am, 1976.
SIMULTANEOUS EQUATION SYSTEMS 493

PROBLEMS

11-1 For the model defined by Eqs. (11-1) and (11-2) show that

where 5 is the slope of OLS regression of C on Y and

m., = tim|+2(z = z)|


Zz Pp n t

11-2 Prove the equivalence between the alternative expressions for the 2SLS estimator in Egs. (11-45)
and (11-46).
11-3 The structure of the Klein model is

C =a + a (W, + W,;) + a,II + a3I1_, + uy

T=
By + Bl + B,T_, + B3K_, + uy

Wp= Yo FY AT = We) 2 (Y+ PH We) ch yt


$ ay

Y=C+/+G

i= Y= W,—7
KG Kee tel

The six endogenous variables are Y (output), C (consumption), / (net investment), W, (private wages),
II (profits), and K (capital stock at year-end). The four exogenous variables are G (government
nonwage expenditure), W, (public wages), T (business taxes), and f (time).
Examine the rank condition for the identifiability of the consumption function.
11-4 Tintner’s model of the U.S. meat market is specified as follows:

V(t) = a + a y(t) + ax, (t) + u(t) demand

it) = Bo + Biya(t) + Bax2(4) + B3x3(t) + u2(t) supply


(a) Determine the identification status of each equation.
(b) Suppose it is known a priori that B,/8, = k where k is a known number. Determine the
identification status of each equation under this specification.
(c) Suppose the model stated at the outset is changed by specifying that a, = B, = B; = 0.
What prior restrictions (if any) on the disturbance variance-covariance matrix would lead to the
identification of both equations?
(University of Michigan, 1981)
11-5 In the model

Vir + BiaYar + WX ie = Me

Yar + Bo Vie + Y22%00 + Y23%X3¢ = Yay

the y’s are endogenous, the x’s exogenous, and u/ = [u,v ,] is a vector of serially independent
normal random disturbances with mean zero vector and the same nonsingular covariance matrix for
494 ECONOMETRIC METHODS

each ¢. Given the following sample second moment matrix:

calculate the LIML and 2SLS estimates of 8, and yj).


(University of Michigan, 1981)
11-6 An investigator has specified the following two models and proposes to use them in some
empirical work with macroeconomic time series data.
Model 1: C, = Oy, + Agm,_, + Uy,
i, = By, + Bor, + ur,
Viet

Jointly dependent variables: Gicthig Mp


Predetermined variables: Tarren
Model 2: MeN 2M.)
+ Oi;
r, = 5\m, + d,m,_; + 63y, + 0,
Jointly dependent variables: Met
Predetermined variables: aay;
(a) Assess the identifiability of the parameters that appear as coefficients in the above two
models (treating the two models separately).
(b) Obtain the reduced-form equation for y, in model | and the reduced-form equation for r, in
model 2.
(c) Assess the identifiability of the two-equation model comprising the reduced-form equation
for y, in model | (an IS curve) and the reduced-form equation for r, in model 2 (an LM curve).
(Yale University, 1980)
11-7 Suppose the following sample second moment matrix (based on 36 observations) has been
obtained for the variables in the Tintner meat model of Problem 11-4:

yy 1) xy xX X3

yy 10 0 l 0 al
V2 0 10 =i = 0
xy ] sail! ] 0 0
X5 0 cal 0 l 0
x3 =I 0 0 0 l

(a) Estimate the parameters a, and a, by 2SLS and test the hypothesis
a, = 0 against the
alternative a, + 0.
(6) Repeat part (a) using IV estimates of a, and a5 obtained with x5,
as an instrument for Yo,
and x,, as its own instrument.
(Yale University, 1980)
SIMULTANEOUS EQUATION SYSTEMS 495

11-8 (a) Assess the identification of the parameters of the following five-equation system:

Vie + BirVar + Bra Yar + WiZue + YNaZae = Me


Yor + Bo3¥3¢ + Bos Ys, + Y22Z01 = U4
Yar + ¥31211 + 3323" = U3,
+ Yaa2ar = Ugy
Barir+ Bas Yar + Yar + YanZ21
2 y3, + Vsr — 2, = 0
(b) How are your conclusions altered if y;; = 0? Comment.
(c) Briefly explain how you would estimate the parameters of this model. What can be said
about the parameters of the second equation?
(WER 979)
11-9 The model given by

Vie = BiaYar + Yuzu + Yi2Za1 + ei (1)


Yor = Bar Vir + Y2323¢ + €2, (2)
generates the following matrix of second moments:

Calculate:
(a) Least-squares estimates of the unrestricted reduced-form parameters
(6) ILS estimates of the parameters of Eq. (1)
(c) 2SLS estimates of the parameters of Eq. (2)
(d) The restricted reduced form derived from parts (6) and (c)
(e) A consistent estimate of E(€)7&,) = 0)
(WEF 1973)
11-10 Let the model be

Viet BizYar + YX + Mia %3e = He


Bar Vie + Yar + Yr2%a1 + Y23%31 = U2, ()
and suppose the observations on the variables are
tase
ea STG pditea(sicgatte 2, My
X= 4! 4-3
Wo. 12420 it ee aa (2)
Dele
Ae Sees 10
(a) Examine the rank and order conditions for identification on the basis of Eqs. (1) and
suggest a suitable estimation procedure for both equations.
(6) In the light of Eqs. (2) investigate whether the answer under part (a) needs modification
and, if so, in what way. Interpret by reformulating the model in Egs. (1).
(UL, 1971)
11-11 Let the model be

Vie + Biz¥or + VNi2%20 + M13%3e = 4


Bay Vir + Yar + Yaite+ Yra%ar = Ure
496 ECONOMETRIC METHODS

If the second moment matrices of a sample of 100 observations are

vy -|
80.0 -4.0 rox 2.0 10 -3.0 7
-40 50 -0.5 15 05 -1.0
3.0520 0 0
0 DD 10 0
X’X =
0 0 12 OREO)
0 0 0 0.5
find the 2SLS estimates of the coefficients of the first equation and their standard errors.
(UL, 1970)
11-12 In the following market model

Supply = =9,=ByP,
+ Yio + 41,
Demand = Q, = ByP,+ Y20 + Y21Z11 + Y22Z2¢ + U2,
quantity Q, and price P, are endogenous, while income Z,, and the price of some other good Z,, are
exogenous. If the supply function is estimated directly by least squares, will the resulting estimate of
8, be biased? If so, in which direction will the bias occur?
(UL, 1972)
11-13 If

Vie = BizYar + MX 1e+ Ni2X21 + Uy


Yar = Bar Vie + Y23%31 + Ur,
10 O 0 10 20
and xX’X = Oy S 0 X’Y =/ 20 10
OR ORO 30. =—.20
estimate the parameters in the model and comment on your results. If £5, is known to be equal to 0.6,
would you modify your estimation procedure, and if so how?
11-14 The X’X matrix for all the exogenous variables in a model is

7 0 oan
wee 10 en)
ce Son Se 4
l 0 bet
Only the first of these exogenous variables has a nonzero coefficient in a structural equation to be -
estimated by 2SLS. This equation includes two endogenous variables, and the least-squares estimates
of the reduced-form coefficients for these two variables are
iP pas |
Lie eal el
Taking the first endogenous variable as the dependent variable, state and solve the equation for the
2SLS estimates.
11-15 For the model

Vie = BizYar + YX + Uy
Yay = Bar Vir + Yo2%21 + ¥23X3, + U,
you are given the following information:

1. The least-squares estimates of the reduced-form coefficients are

SO Al
KD MO Ss
2. The estimates of variance of the errors of the coefficients in the first reduced-form equation
are |,
Ws), Ofte
3. The corresponding covariances are estimated to be all zero,
4. The estimated variance of the error on the first reduced-form equation is 2.0.
SIMULTANEOUS EQUATION SYSTEMS 497

Use this information to reconstruct the 2SLS equations for the estimates
of the coefficients of the
first structural equation, and compute these estimates.
(UL, 1969)
11-16
Vit = Bia Yar + Big ¥3e + Yx1, + uy,
is One equation in a three-equation model which contains three other exogenou
s variables x5,, x3,, and
X4,- Observations give the following matrices:

20 15 ae) 2 2 4 2) ; : ;
YY =] 15 COP = 457 ¥ X= 0 4 Wh 5)|| 2,05 O10 eee
4S 70 ON 2 ee 10 0 OF Os
Obtain 2SLS estimates of the parameters of the equation and estimate their standard
errors (on the
assumption that the sample consisted of 30 observation points).
(UL, 1968)
CHAPTER

TWELVE
ECONOMETRICS IN PRACTICE:
PROBLEMS AND PERSPECTIVES

A careful study of the material covered in the previous eleven chapters would not,
unfortunately, equip the reader to conduct a successful piece of applied econo-
metric research, since that involves many more problems than those already
discussed. We will tentatively explore some of these issues in the present chapter,
but the reality should be faced at the outset that it is not feasible to write a.
comprehensive manual that would prepare applied econometricians for all the
problems that can arise in a wide variety of research projects. Successful econo-
metric modeling is not a collection of mechanistic and routine procedures but
more of an art requiring wide-ranging knowledge and judgment. Such an art is
best learned by practice, hopefully with talented supervisors and colleagues, and
by study of “best practice” examples. It is, however, not always easy to find the
latter. Indeed a very instructive book might be written under the title, How NOT
to Do Econometrics, with every chapter illustrated by one or more published
articles. The author of such a book would have to time its publication carefully in
relation to his own impending demise or retirement from contact with his
professional colleagues: he might also face the difficult problem of choosing some
of his own previous work for inclusion.
There is a widespread view that econometrics has in some sense not lived up
to its early promise, and there is much scepticism about the value of the plethora
of empirical results embedded in the literature. This state of affairs should not be
too surprising. There is, after all, a sound proposition in economics that the use of
a good or service tends to expand to the point at which price and marginal utility
498
ECONOMETRICS IN PRACTICE: PROBLEMS AND PERSPECTIVES 499

are equated. If computers are essentially free goods, if “researchers” can plug into
a data bank without any understanding of where the series come from or how
they were constructed, if they can press buttons to implement computer programs
whose contents they dimly comprehend, it then follows, as night follows the day,
that some work of zero worth will emerge. Indeed, given the uncertainty inherent
in the research process, compounded by the fallibility of the researcher, some
outputs may err on the wrong side of zero and be positively dangerous. Our twin
defenses against further encroachments by a flow of dubious work lie in improv-
ing still further the quality of the editorial screening process and raising also the
quality of the training given to would-be practitioners.

The Origins and Objectives of an Econometric Research Project


The origins and objectives of econometric research projects are as many and
various as the persons and groups who undertake them. It may be a lone graduate
student scratching his head or rummaging in his supervisor’s “bottom drawer”
for a thesis project; it may be an academic intrigued by a theoretical debate in the
literature or impelled by some idea of her own; it may be a public or commercial
research group building or expanding an econometric model to be used for
short-term forecasting. In my own experience the writing of this book was
delayed for several years by two separate phone calls. One led to a year’s work
estimating demand functions for oil for the major industrial countries of the
world. A subsequent, but not unrelated, call led to two years’ intense activity as
the econometric consultant on the construction of a world model of energy
demands and supplies. It is platitudinous but important to say that in all cases
one should be as clear and precise as possible about the objectives of the research,
since these condition the design and layout of the project, though of course they
may have to be revised as the project proceeds.

Data and Model Specification


These two topics are inextricably linked. The model specification will have strong
implications for the data required and, conversely, data limitations may constrain
the feasible specification. As an illustration suppose an objective is the estimation
of a demand function for crude oil in the United Kingdom that might then be
used to forecast demand, conditional on various assumptions about the future
paths of income and relative prices. The first step is to investigate the range of
possible strategies. Should one estimate an aggregative demand function for
crude, using some measure of “income” and some relative price? If so, should the
income measure be real GDP, or an index of industrial production, or what?
What, in turn, are the appropriate price series from which an index of relative
prices should be constructed? Or, alternatively, should one use a more disaggre-
gated approach looking at the final demands for specific products from the
refining process, such as gasoline, jet fuel, heating oil, and so on, and should one
disaggregate also by consuming sector, whether residential, commercial, in-
dustrial, public utilities, and so forth? The disaggregated approach would also
500 ECONOMETRIC METHODS

require the modeling of the refining decision and of the relationship between the
price of crude and the prices of refined products. In this decision process it is of
great importance to have as much knowledge as possible of what may be called
the “institutional realities” of the situation, specifically in this case such things as
the nature of the refining process and the constraints on the refining decision, the
quantitative importance of various groups of consumers, and the crucial factors in
their decision processes. An econometrician coming cold to the study would run
the risk of very slow progress with much searching through inappropriate formu-
lations. In my own experience collaboration with an experienced oil specialist
greatly improved the research efficiency.
Knowledge of the “institutional realities” is, of course, valuable in all areas.
In a study of cost-output relationships in coal mining this author felt it necessary
to don a safety helmet and get to the coal face in the narrow and twisting seams
of the Lancashire coal field in order to see at first hand the nature of the
production process before sitting down to peruse the statistics at the regional
headquarters of the National Coal Board. Similarly in studies of scale, costs, and
profitability in road passenger transport and of cost-output variations in a
multiple-product firm the author spent time at each firm talking to accountants
and managers to study their accounting and decision processes before extracting
the relevant data by hand from the firm’s records.} To take a final data problem,
monetary theory postulates the demand for money to be positively related to
income and negatively related to the rate of interest. Each of the three nouns in
this proposition raises formidable problems of definition and measurement. There
are numerous definitions of money and almost continual evolution of payments
technology, there are many interest rates, and even income is not unambiguous.
When appropriate data series have been identified, the next decision in
time-series contexts is what data period (hourly, weekly, monthly, quarterly,
annual, or whatever) to use. Again if we had institutional information about
decision processes (who decides when about what) we could make the appropriate’
choice. If, for example, production decisions are revised at the start of each
month, a model of the production decision employing monthly data would have
the best chance of capturing the essential features of the process. Quarterly or
annual data would in this case involve an inappropriate aggregation over time,
thus making it difficult, if not impossible, to determine the lag structure. Often,
however, there is little firm information about decision procedures, and the main
choice between quarterly and annual data is based largely on a mixture of
empirical considerations and the objectives of the modeling process. As Table
11-1 shows, the macroeconometric models for the developed economies are split
roughly evenly between those based on quarterly and those based on annual data.
By far the most difficult problem of all is the initial specification of the
model, be it a single equation or a set of equations. By specification we mean the

+ For these and other studies see J. Johnston, Statistical Cost Analysis, McGraw-H
ill, New York
1960.
ECONOMETRICS IN PRACTICE: PROBLEMS AND PERSPECTIVES 501

following:

1. The listing of explanatory variables, including lagged values, in each equation


2. The functional form relating these variables to the dependent variable
3. The stochastic properties of the disturbance term or terms

Economic theory is mostly about equilibrium situations and contains little in


the way of systematically developed dynamic theory. Thus it cannot be expected
to yield strong insights about lag structure. Nor can it be expected to indicate the
correct functional form. Thus items 1 and 2 inevitably lead to a certain amount of
interaction between theory and data. This interaction also impinges on item 3,
which essentially consists of assumptions about unobservable variables. However,
each specification under items 1 and 2 provides estimates of the unobservables,
and the interaction between specification and data usually continues until the
researcher feels that a “reasonable” set of results under items 1, 2, and 3 has been
obtained.

Data Mining and Specification Searches}


This interactive process has been labeled data mining or, more recently and less
pejoratively, specification searches. At one extreme it is alleged that data mining
invalidates all the conventional significance levels or, even more strongly, that the
final results are quite valueless, since the researcher has gone on a “fishing
expedition” or beaten the data set into submission until they finally yielded the
desired conclusion. At the other extreme it is suggested that if the set of models
includes the “true” model, that model will have the smallest residual variance and
hence the highest true R*, so that searching for the best fit to the sample data is a
reasonable and sensible procedure.
Let us take a look at the data mining problem by means of two hypothetical
examples.

Example 1. A researcher’s objective is to explain the variation in a variable y. He


has 10 candidate explanatory variables x,,..., x, 9 on the basis of his a priori
theory. The underlying theory can only be characterized as “lousy” for

1. It specifies that only three of the possible 10 variables actually influence y,


but it does not know which three, and, more seriously,
2. The theory is totally in error for, in fact, none of the 10 variables has any
effect on y.

The first defect of the theory actually appears as an advantage to our


researcher for his computer cannot handle more than three explanatory variables

+ Our brief discussion cannot hope to do adequate justice to this topic. The interested reader will
find much nourishment in the elegant, entertaining, and enlightening E. E. Leamer, Specification
Searches, Wiley, New York, 1978.
502 ECONOMETRIC METHODS

at a time. Thus he computes all '°C, = 120 possible multiple regressions and the
attendant F statistics for the overall fit. The true value of all 120 population F
statistics is of course zero, but the reader would not be surprised to find that our
researcher discovers some significant sample regressions.+ His theoretical and
institutional knowledge enables him to write a plausible commentary on these
regressions and perhaps select one as the seemingly best theory for the explana-
tion of y. Sending the write-up to an editor, who likes to publish “significant”
results, guarantees another “scientific” paper and a further small step by the
author up the academic ladder.
Another variant of Example 12-1 is a theory that only identifies the three
candidate variables, none of which, in fact, has any relevance to y. A series of
investigators drawing different sets of sample data from y, x,, x», x; fail to find a
significant regression, consigning their computer printout to the waste paper
basket or filing cabinet, according to temperament. In either case their profes-
sional colleagues are unaware of this accumulation of “negative” results, and so
testing of the theory continues. Working at any conventional level of significance,
it is only a matter of time until a set of sample data is drawn that yields a
“significant” result, which will, of course, have a good chance of being published.

The moral of Example 12-1 is clear. In an area where theory is poor and
provides little guidance on specification to the researcher, data mining is a highly
dangerous activity. Combined with the propensity of editors to publish only
significant results, it can in extreme cases result in the publication of falsehoods
and the suppression of truth. However, take heart, faint reader, the above surely
cannot be a description of economics, the queen of the social sciences, richly
endowed with well articulated theory. Consider then Example 12-2.

Example 2. The minister of petroleum in the mythical oil-rich country of Sandia


desires to know the demand function for crude oil so that he may better inject
some good sense and realism into the next round of cartel discussions. Having

+ The 120 models may be represented by


y = constant + B)x; + Bx; + Byx, +u for alli, j,k;itj*#k
In each case the null hypothesis is

Hy: B= B = By
Working at the 5 percent level of significance, the probability of accepting the null hypothesis
for any
specific model is 0.95. Assuming independence of the models, the probability of accepting
the null
hypothesis for all the models considered is (0.95)'7° = 0.0021. Thus the chance that the researcher
finds ar least one “significant” regression is 0.998. Working at the more stringent
| percent level of
significance, the probability of finding at least one significant regression is still as
high as 0.70. The
models will not all be independent of each other because of overlapping explanator
y variables, so
these startling probabilities need not be taken too seriously, but they do indicate
the nature of the
potential problem associated with data mining.
+A small but constructive step toward addressing this problem was taken a few
years ago by the
editors of the Journal of Political Economy, who initiated a section for the publication
of “confirma-
tions and contradictions.”
ECONOMETRICS IN PRACTICE: PROBLEMS AND PERSPECTIVES 503

ample resources at his disposal, he commissions four separate econometricians to


estimate and deliver such a demand function. All four have access to the same
data base, namely, all the statistics ever published in the world plus the internal
files of the Sandia ministry of petroleum, but they are to work completely
independently.
Econometrician A is an able econometrician, trained in a good graduate
school and already prewarned of the sin of data mining. After much cogitation,
research, and study of the institutional realities, he specifies his demand function,
estimates it by the appropriate procedures, and sends off the results to Sandia.
Econometrician B is a better econometrician, who went to a better graduate
school than A. He too is not about to engage in data mining, so he formulates his
specification, which happens to differ somewhat from that of A, estimates the
equation, and dispatches the results. Econometrician C is a clever chap from
Cambridge (either one). He follows the same procedure as A and B, but as one
might expect, his a priori specification is, in truth, superior to theirs. Econometri-
cian D is a data miner from Dublin, who unhappily never had the good fortune to
go to graduate school nor even to attend a lecture on statistics, but nonetheless
has a certain degree of native intelligence. As benefits an Irishman, he has been
warned about so many sins that he has completely forgotten the sin of data
mining. His first attempt at the problem just happens to be the specification used
by A. However, D does not much like the results and respecifies, just happening
now to arrive at the specification used by B. The results of that are still not quite
to his pleasing, so he respecifies once more and now happens to hit on the
specification used by C. The results of that please him and are sent off to Sandia,
but he does not confuse the minister by including the results of his earlier and, to
him, unsatisfactory specifications.
Suppose we are privileged to have one further piece of information, which is
that the C/D specification is the true and correct one. The statistical purist would
presumably congratulate C and criticize D. However, their standard errors,
confidence intervals, and associated F statistics are identical. Classical inference
establishes the properties of estimators and tests of hypotheses by examining what
might be expected to happen in repeated sampling from a given population or
model. If the model has not been correctly specified, the tests are strictly invalid
and the various probability statements are not correct. In the present hypothetical
example inferences based on the A or B specifications would, strictly speaking,
not be correct, while those based on the C/D specification are. Data mining has
only enabled D to make good the defects in his education and has been beneficial
rather than damaging. Finally we may observe that most classical procedures are
fairly robust to specification errors, and in practice, probably no finite model will
ever be the “true and correct” model so we should not be slavish devotees of
spuriously precise significance levels.t+

+ The development of econometric theory has been heavily influenced by the early work of the
Cowles Commission, which emphasized problems of equation error to the almost total exclusion of
problems of measurement error. Little is known about significance levels or the relative properties of
different estimators when these problems jointly coexist, as indeed they do in practice.
504 ECONOMETRIC METHODS

We conclude that the circumstances of Example 12-2 are closer to those of


real economic research than those of Example 12-1, that interaction between
theory and data is both inevitable and, indeed, desirable, and we turn now to a
discussion of some specific guides to that respecification.

Criteria for Model Selection

Residual variance (R”) criterion. Most of the operational criteria have been
developed in the context of a single equation model. The first is the residual
variance, or R*, criterion. Suppose there are just two competing models for the
explanation of y, namely,
y=X,B, + u, and y = X,B,
+ u,
where X, is nonstochastic, of order n X k;, and of full column rank. Suppose that,
in fact, the first model is correct. If the second model is fitted, the vector of OLS
residuals is

e, = Moy
= M,(X,B, + u,)
where M, =1-—X,(XX,)
'X
Thus the residual sum of squares is

ese, = B,X{M,X,B, + 2B; X,{M.u, + u,Mu,


Taking expectations
E(eje,)
= B{X\M,X,B, + (n — k,)o? (12-1)
Since M, is idempotent, the quadratic form on the right-hand side of Eq. (12-1) is
positive semidefinite. Defining sj = eSe,/(n — k,), it then follows that
Bea) a2
If the first (and correct) model is fitted, we know from Sec. 5-3 that

E(s}) = of
where

(n-—k,)sf = Ce. VV y'X,(X,X,)_Xiy


Thust+

E(s?) < E(o3) (12-2)


Notice that the inequality is in terms of the expected values of the residual
variances. In practice we can only compare the estimated residual variances. It is,
of course, possible for sj to be less than s?, even though, in fact, E(s2) < Ess).
Thus minimum residual variance cannot be taken as asingle overriding criterion

+ This argument is due to H. Theil, Principles of Econometrics, Wiley, New York,


1971, p. 543.
ECONOMETRICS IN PRACTICE: PROBLEMS AND PERSPECTIVES 505

for equation selection. Since there is a monotonic negative relationship between


R’ and s’, the same comments apply to a maximum R? criterion, but the fact that
the degree of fit does not discriminate perfectly between the true and competing
models does not mean that evidence on fit is to be ignored.

Criteria for individual coefficients. There are two important criteria under this
heading. Economic theory is rich in qualitative predictions about the direction of
various effects. Thus one looks for agreement between a priori expectation and
the signs of estimated coefficients. Second, one looks for correctly signed coeffi-
cients which have reasonable statistical significance. The latter criterion should
not be applied too stringently since we have seen, for example, that collinearity
among the regressors can inflate estimated standard errors. The R? criterion also
has implications for the significance level of individual coefficients. As shown in
Problem 5-12, R* only increases with the addition of an extra regressor if the F or,
equivalently, the ¢ statistic for that variable exceeds unity, which corresponds to
the use of a significance level of about 30 percent rather than the conventional 5
or | percent level.
The previous remark is in the context of a fixed sample size. However, any
substantial increase in sample size has implications for significance levels. As seen
in Chap. 5, the test of the hypothesis that a subvector of g elements in B is the
zero vector is given by

_ (€xex — e’e)/q if
Ae e'e/(n — k) Aga ie)
where e’e is the residual sum of squares from the unrestricted model and eye,
that from the restricted model, the relevant g variables having been omitted. This
statistic is written equivalently as
Re SoRA Usk
1a Re q
Thus even though R* — R% may be very small, the test statistic can become
arbitrarily large with increasing sample size. Using a given significance level, the
null hypothesis is more and more likely to be rejected as n increases. This point
has been emphasized by Leamer, who, along with others, argues that the signifi-
cance level for this kind of test should be adjusted downward for larger samples.+

Well-behaved disturbances. As seen in earlier chapters, a homoscedastic nonauto-


correlated disturbance term is a wonderfully powerful assumption from astatisti-
cal point of view. Its presence underlies the derivation of a battery of statistical
tests, while its absence seriously distorts some of these tests and calls for revised
procedures. Thus it is essential to examine the properties of the disturbance term
in order to assess the validity of the statistical tests being applied. However, the

+ Recall that r(r) = VF (1, 7) . See App. A-7.


+E. E. Leamer, Specification Searches, Wiley, New York, 1978, pp. 88-89.
506 ECONOMETRIC METHODS

same property is often implicitly, and occasionally explicitly, taken as a desirable


feature of a well-specified economic relationship. The purpose of such arelation is
to model the behavior of some group of economic agents. Can it be a good model
if the net effect of the omitted variables displays some systematic autocorrelated
pattern? A misspecification of functional form can also lead to nonrandom
disturbances. Thus statistical and economic considerations alike lead one to look
for relations with well-behaved disturbances. However, as noted in Chap. 8,
discrepancies between decision periods and data periods may well produce
autocorrelated disturbances in a properly specified economic model. It is also the
case that efficient estimation of fairly complex dynamic regressions may require
an autoregressive specification for the disturbance term, but this is a by-product
of considerations of statistical efficiency: the original relationship is desired to
have a nonautocorrelated disturbance term. This point is emphasized by Hendry
and Mizon.j Suppose one-period lags on both variables are sufficient to give a
white noise disturbance. The original (general) dynamic relationship between y
and x may then be written as

Vo Beh Yo vient v, (12-3)


where |8,| < 1 and {v,} is white noise. Using the lag operator, this may be
rewritten as

Cle BL) y, = (¥% P ¥,L)x, + 0,


If it then were true that the parameters satisfied a restriction
ile — Bi
the relation would become

(1 — BL) y, = yo(1 = Bre, (12-4)


which gives
Ye = Vox oe

with ie On (12-5)
If the restriction were valid, estimation of Eqs. (12-5) would involve just three
parameters, namely 8,, yy, and o,, whereas estimation of Eq. (12-3) involves four
parameters. However, Eq. (12-3) has, in fact, to be estimated to test the restric-
tion. The payoff is improved statistical efficiency of the parameter estimates if the
restriction is upheld. Comparing Eqs. (12-3) and (12-4) the restriction implies that
{y,} and {x,} have a common factor with root B,.t There may be no economic
rationale for the restriction or common factor. If so, it is likely to be rejected and
the “general” equation (12-3) cannot then legitimately be reduced to the “simpler”
form in Eq. (12-5). Sargan’s COMFAC program tests for the existence of

t D. F. Hendry and G. E. Mizon, “Serial Correlation as a Conveni


ent Simplification, Not a
Nuisance: A Comment on a Study of the Demand for Money by the Bank
of England,” Economic
Journal, vol. 88, 1978, pp. 549-563.,,
t Strictly speaking the root of the polynomial 1 — B,L = 0 is 178).
ECONOMETRICS IN PRACTICE: PROBLEMS AND PERSPECTIVES 507

common factors in dynamic regressions.} Suppose a general dynamic relation

B(L)y, = y(L)x, + 2, (12-6)


can be found where {v,} is white noise. If B(L) and y(L) have, say, two common
roots, there exists a quadratic in L, say 6(L), which is common to both B(L) and
y(L). Thus we can write
B(L)=S(L)B*(L) ands y(L
= 8(L))
y*(L)
so that Eq. (12-6) becomes

Pe yea ay(Lyx e,
=,»,
5(L)u
(12-7)
which involves considerably fewer parameters than Eq. (12-6). Suppose, for
example, that a relationship
Vy as Biy-| ee By y,—> a Yo*X, at YixX1-1 a Y2X 1-2 a ¥3%;-3 ntsv, (12-8)

was estimated and a common polynomial 6(L) = (1 — L)(1 — pL) found. The
relation (12-8) may then be written
(1—L)(1 — pL)y, = (1 — L)(1 = pL) (vs + IL)x, + ©,
which only involves three parameters instead of six. For estimation purposes it
may be put in the form
DV = Ve AX tay (NXeae ta
(12-9)
t

with (1 = pL )u, =»,


that is, a simple relationship between first differences with an AR(1) process in
the disturbance. The Sargan-Hendry message is that researchers should not begin
with simplified specifications such as Eqs. (12-5) or Eqs. (12-9), but should instead
commence with a general model containing sufficient lags to yield a white noise
disturbance and then test to see how far it can be legitimately simplified.

Stability of the relationship. A very important indicator of the quality of a


functional specification is the stability of the parameters over various data sets.
This may be examined in two alternative fashions. One is a straightforward test
for structural change as outlined in some detail in Chap. 6. This presupposes
sufficient observations in each subset of data to permit estimation of all parame-
ters. When that is not the case, the Chow forecasting test may be applied. This
has already been set out in Example 6-5 of Sec. 6-2 and was derived in Sec. 10-1
by using recursive residuals. However, it is often derived in an alternative fashion
as follows.
Suppose the usual linear model has been fitted to n observations of k
variables. The OLS coefficient vector is
b =8 + (X’X) ‘Xu

+J. D. Sargan and J. D. Sylwestrowicz, “COMFAC: Algorithm for Wald Tests of Common
Factors in Lag Polynomials,” User’s Manual, London School of Economics, London, 1976.
508 ECONOMETRIC METHODS

and it is assumed, as usual, that u ~ N(0, o7I,,). Now suppose a new set of m
(< k) observations on these same variables becomes available. On the assumption
that the original model still holds, the new observations may be characterized by
Yo= XoB+ uy
where E(u,u,) = o7I,,. The m observations are insufficient to allow reestimation
of the model, but one may forecast the yy vector by
Jo = Xob
The vector of forecast errors is

€©) = Yo — Jo = XoB + uy — X,[B G (x’X) 'x’ul


uy — Xo(X’X) 'X’u
It then follows directly that

where V =I,, + Xo(X’X)” 'X4


ev le
Thus at = x?(m)
0

Since e’e/o* has an independent x?(n — k) distribution, it follows that under the
hypothesis of parameter constancy
/ / il / =,

= en + Xo(XX) Xo] co/m ~ F(m,n—k) (12-10)


e’e/(n — k)
The hypothesis of a stable relationship would be rejected if the F statistic in Eq:
(12-10) exceeded some preselected critical value. Chow demonstrates the equality
of Eq. (12-10) with the alternative expression in Eq. (6-27).} In an interesting and —
important study Jorgenson, Hunter, and Nadiri have used measures of fit,
Durbin-Watson statistics, and tests of structural change to assess and compare
different investment equations. When the regressors are stochastic, the test
Statistic in Eq. (12-10) will only be approximately distributed as F. Hendry
Suggests using an asymptotically equivalent test which neglects the variation due
to estimating the B vector. Under the hypothesis of parameter constancy§
eoe, D
Sona Ace) (12-11)

7G. C. Chow, “Tests of Equality between Sets of Coefficients in Two Linear Regressions,”
Econometrica, vol. 28, 1960, pp. 591-605.
+D. W. Jorgenson, J. Hunter, and M. I. Nadiri, “A Comparison of Alternative Econometric
Models of Quarterly Investment Behavior,” and “The Predictive Performa
nce of Econometric Models
of Quarterly Investment Behavior,” Econometrica, vol. 38, 1970, pp. 187-224.
§D. F. Hendry, “Predictive Failure and Econometric Modelling
in Macro-Economics: The
Transactions Demand for Money,” London School of Economics, London,
September 1978.
ECONOMETRICS IN PRACTICE: PROBLEMS AND PERSPECTIVES 509

where
nee
er eay
The test of forecast errors in Eq. (12-10) may be extended to deal with joint
forecasts from the reduced form of a simultaneous equation model.}
An aspect of prediction which is frequently ignored, and unjustly so, is the
longer-term implications of the dynamic regression that has been estimated. For
example, return again to our hypothetical demand function for oil, and ask
questions such as:

e What does it imply about the long-run elasticities?


e What does it imply about the length of the long run, how long is it estimated to
take for a full adjustment to a “shock”?
els the reaction path plausible or is it the result of a statistical straitjacket
imposed on the data?

The model’s answers to questions such as these have to be put up against the
intuition and good sense of the researchers themselves and, more importantly, the
intuition and good sense of informed critics. This may seem very “unscientific”
and perhaps it is, but it is nonetheless very important and in the next two sections
we present a brief discussion of some ways in which it is attempted.

The Cairncross Test

We have suggested that there are various aspects of any specification which are
important, namely,

Residual variance (or fit)


Signs and precision of specific coefficients
Properties of the disturbance
Parameter stability (predictive performance)
aS

In comparing different specifications there is no serious problem if the


indicators more or less all point in the same direction, as was the case in the
Jorgenson, Hunter, and Nadiri study of the investment equation. Where contrary
indicators emerge, the choice between specifications has to rest on the relative
importance of various factors to the decision maker. As in the choice of a
husband or a place to live, a specification is a “package deal.” No one has yet
found a way to piece together the perfect package, though we continually try to
improve, as is evidenced by the statistics on divorce, population mobility, and the
flood of computer printout. In the case of economic specifications the choice can
be based less on purely subjective personal considerations and more on the
accumulated knowledge and experience of the critic.

+ See P. H. Dhrymes et al., “Criteria for Evaluation of Econometric Models,” Annals of Economic
and Social Measurement, vol. 1, 1972, pp. 307-308.
510 ECONOMETRIC METHODS

This point was brought home to me forcefully and convincingly a few years
ago when I was working on an energy research project. Each month a report on
the econometric activity had to be presented to a steering committee in London,
presided over by Sir Alex Cairncross.¢ Cairncross had (and still has) a healthy
scepticism of econometrics, no doubt partly due to his days at the U.K. Treasury
when, to quote, “the young men might present me with thirty different equations
to ‘explain’ British imports, so that, at the end of the day, neither they nor I knew
what determined British imports.” Each month the econometric output was
subjected to his shrewd, informed, and penetrating scrutiny. Eventually, however,
there came a monthly report which secured the approbation, “I wouldn’t mind
getting on a plane and taking this to Riyadh.” Presumably I had been engaged in
some successful data mining or, perhaps, had been “learning by doing,” so I
suggested to him jokingly that the Cairncross test would appear in the next
edition of Econometric Methods. The two-step Cairncross test is thus as follows.

1. Compute your R*, Durbin-Watson statistic, assorted t, F, and x? statistics for


the best specification you can manage.
2. Send the resultant report to Sir Alex Cairncross with the question “Would
you be prepared to take this to Riyadh?”

The suggestion is, of course, not entirely frivolous. Researchers circulating


their discussion papers are carrying out informal Cairncross tests. For Cairncross
Substitute the expert of your choice and for Riyadh substitute Washington, the
editorial offices of the American Economic Review, or some other preferred
location.
Economists, however, are not alone in facing difficult choice problems in
which all the elements cannot be fully quantified and brought together in a single
equation. Circumstances comparable to those of the Cairncross test arise in a
broad spectrum of commercial, industrial, and governmental decisions, where the
best possible research still does not eliminate the need for some personal element
based on judgment and experience.

The Bayesian Approacht


A Bayesian would criticize the Cairncross test on the grounds that the opinions,
judgment, and, possibly, prejudices of the expert have only been introduced in
some implicit, informal, and nonreproducible fashion. Bayesians also tend to
make a more general criticism of classical inference in that its procedures are

¥ Sir Alex Cairncross, a very distinguished British economist, was for many
years economic advisor
to Her Majesty’s Government and subsequently Master of an Oxford College.
I hesitate to give his
present address lest he be deluged with manuscripts from aspiring econometricians,
but I suspect he is
mostly to be found in his Scottish retreat north of the Solway Firth, enjoying
the Scotsman’s favorite
view “looking down upon England.”
+ The Bayesian approach requires a book of its own. The premier
references are A. Zellner, An
Introduction to Bayesian Inference in Econometrics, Wiley,
New York, 1971; and E. E. Leamer,
Specification Searches, Wiley, New York, 1978.
ECONOMETRICS IN PRACTICE: PROBLEMS AND PERSPECTIVES 511

justified in terms of sampling distributions, which picture the behavior of estima-


tors in repeated sets of sample data. Such repeated samples are never drawn. We
typically have one sample and have to do the best we can with that. Moreover the
possible losses from incorrect conclusions are not usually considered in the choice
of inference procedures.
In principle the Bayesian approach can cope with these problems in one
integrated framework. The underlying principle is both simple and beautiful, but
there are problems in the way of practical applications, especially in the area of
large and complex models. The approach may be illustrated in two steps.
Suppose, first of all, that there is no uncertainty about the form of the relevant
model but only about its parameters. To be more specific, let us assume that we
have n observations drawn at random from

p(y) = (2003)
'””en] 3 ian? — 1) i
where o, is known, but the mean p is unknown. This gives the vector

Yin wires ra ool!


The probability density function (pdf) for y is then

Plate)= (2993)"exp ~5 Low! (12-12)


O5 i=1

This is the likelihood for the sample observations, conditional on the parameters pu.
and o,, but the latter has been omitted from the left-hand side since it is assumed
known.
The first crucial element in the Bayesian approach is to postulate the
existence of prior information about p. This may come from theoretical sources,
from previous empirical studies, hunch, judgment, or what have you. Such
information cannot be exact, so it is formulated in a stochastic fashion. It is
theoretically convenient to model this information in a way that is compatible
with the likelihood in Eq. (12-12). This leads to the concept of the conjugate prior.
The prior pdf for p is thus taken to be normal and written

a(t) = na) L exp|-(a


]
~ m)} 0
(12-13)
where m and o° are specified numerically.+
From elementary probability theory for any two events A and B we can write

Pr( A, B) = Pr( A) - Pr( B|A) = Pr(B) - Pr( A|B)

+ We are using p(-) to indicate a pdf for sample data and 7(-) to indicate a pdf for parameters.
This practice was suggested in K. M. Gaver and M. S. Geisel, “Discriminating Among Alternative
Models: Bayesian and non-Bayesian Methods,” in P. Zarembka, Ed., Frontiers in Econometrics,
Academic Press, New York, 1974, Chap. 2.
512 ECONOMETRIC METHODS

from which

PE ry
Pr( B) - Pr( A|B)
Pr( B| A) = ———_——— (12-14)
12-14

Letting A represent the sample vector y, B the unknown parameter p, and


replacing probabilities by pdf’s, we have

7(uIy) =
mu): p(y|e) (12-15)
P(y)
In Eq. (12-15) the expression 7(u|y) represents the posterior pdf for m, and
comparison with 7() indicates the change in the researcher’s beliefs about p
brought about by the sample information in y. The denominator in Eq. (12-15) is
given by

P(y) = Jp(y|u)7(H) dp
For given y, m, oj, and o” this reduces to a constant. Thus Eq. (12-15) can be
rewritten as

m(uly) & m(m) > p(y|u)


Substituting from Eqs. (12-12) and (12-13),

(ply) x e9|
-1) ee
Pineal

As shown by Zellner, this pdf can be simplified tot

al |
m(ply) & o|- |
207o3/n Ore}. oo/n
where ji = Ly,/n. Thus the posterior pdf for p is also normal with mean

E(p) =
fi(ag/n)a' + m(o?)
a
| (12-16)
(o5/n) + (07)
and

1
var(p) =
(og /n)' + (0?)
Formula (12-16) shows that the posterior mean is a weighted average
of the
sample mean and the prior mean, the weights being the reciprocals
of the
respective variances. Strong prior information (low 0°) gives the prior mean
a
large role to play in determining the posterior mean, and conversely,
strong
sample information (large n and/or low 06.) gives the sample mean a dominat
ing
role. The importance of the posterior mean rests on a basic result in
Bayesian

7 A. Zellner, An Introduction to Bayesian Inference in Econometrics,


Wiley, New York, LOT eS 0)
ECONOMETRICS IN PRACTICE: PROBLEMS AND PERSPECTIVES 513

statistics that if one assumes a quadratic loss function for errors in estimating p,
the estimate which minimizes the expected loss is the posterior mean.+
A parallel result holds in the linear regression case when the form of the
model is assumed known, but a multivariate Bayesian prior distribution is
specified for the B vector.t Assuming a multivariate normal prior, this involves
specifying the mean vector by and all the elements in the variance-covariance
matrix, indicated, say, by o7N, '. Assuming the usual linear model
y=XB+u_ with u~ N(0,071)
the mean of the posterior distribution is§

bs» = (Ny + X’X) '(Nyby + X’Xb) (12-17)


where b = (X’X)_'X’y is the OLS estimate of B from the sample data. This is the
linear regression equivalent of Eq. (12-16). The posterior mean vector is seen to be
a matrix weighted combination of the prior vector b, and the OLS vector b, with
weights proportional to the inverses of the respective variance matrices. Two
immediate problems arise with any attempt to implement Eq. (12-17) in practice.
The first relates to the problem of specifying numerically the elements of the
variance matrix N, '. One may have some intuition about the mean vector b,, but
it is difficult to see the source of numerical information about variances and
covariances. The second problem, as Leamer emphasizes, is that if b, and N, are
specified, the posterior is then a function of these specific values. Other investiga-
tors might specify different parameters for the prior distribution with different
implications for the posterior distribution. What is important is to make clear the
mapping from priors to posteriors and to investigate, if possible, the implications
of various classes of prior distribution for posterior distributions.
The second step in the Bayesian approach relaxes the assumption that the
form of the model is known and that the only uncertainty relates to the parameter
values. The extension allows uncertainty about both models and parameters. To
simplify the exposition let us suppose that there are just two competing models,
both in the linear regression format. They are specified as
M,: y=X,6, +4,
M,: y=X,B,+u,

where X, is nonstochastic of order n X k;, of full column rank, u; ~ N(0, o/I,,),


and there are n sample observations. There are two possible situations. One is
where the models (or hypotheses) are nested, which is the case when X,, say,
includes all the variables in X, plus some others. The nonnested case occurs when
some (or all) variables in X, do not appear in X, and vice versa. Classical
inference procedures apply ina straightforward fashion to the nested case, but

+ A. Zellner, op. cit., p. 24.


+ In practice, there is no justification for assuming the disturbance variance o,? to be known. Thus
the prior distribution should incorporate B and a.
§ E. E. Leamer, Specification Searches, Wiley, New York, 1978, p. 78.
514 ECONOMETRIC METHODS

have no simple treatment for the nonnested case, where Bayesian procedures
admit, in principle, of a very simple solution.
For any model the marginal density of the observations M,, sometimes
referred to as the predictive pdf, is given byt

p(yiM,) = [r(y1B,, M,)7(BM,) ap; (12-18)

In Eq. (12-18) p(y|B;, M;) is the likelihood for the sample observations, condi-
tional on the model M, and its parameters B,, and 7(B,|M,) is the prior density for
the parameters, given that the model is M,. Equation (12-18) says that if M, is the
true model, the marginal pdf for the sample observations is found by taking a
weighted average of the sample likelihoods, where the weights are the elements of
the prior distribution for the parameters, given the model. Now suppose that
associated with each model there is a nonnegative fraction P(M,) indicating the
prior subjective probability that M, is the true model and, in this case, such that

P(M,) + P(M,) =1
The unconditional pdf for the sample observations is then

Ply) = P(M,) p(y|M,) cls P(M,) p(y|M,)


and the application of Bayes’s rule to revise the prior probabilities of the models
gives

P(M,) p(y|M,)
P(M\ly) = (12-19)
P(y)
In comparing two models there are just two possible losses, one if M, is chosen
when M;j is the true model and the other if M, is chosen when M , applies. If these
losses were equal, the decision rule that minimizes the posterior expected loss is as ©
follows. Choose M, if

P(M,|ly)
P(My\y)
= P(M,) p(y|M,)
P(M,) p(y|M,)
(12-20)
is greater than 1. Equation (12-20) defines the posterior odds ratio, which is seen to
be equal to the prior odds ratio multiplied by the ratio of the marginal densities
(weighted likelihood functions). If there are more than two models, posterior odds
as defined in Eq. (12-20) can be computed for any pair.
It is clear from Eq. (12-18) that the formidable task in the computation of
Eq.
(12-20) is the evaluation of the marginal pdf’s. If a multivariate normal prior
is
assumed for B,, given M,, that is,

7(b;|M,) 1s N(b;+,Njx')
then Leamer has shown that p(y|M,) varies inversely with a quadratic form Q,

+ We are assuming unrealistically, but for simplicity, that there is no uncertainty about
the
disturbance variances.
ECONOMETRICS IN PRACTICE: PROBLEMS AND PERSPECTIVES 515

which may be expressed in alternative ways ast

Q; ir (y ~ X;b; )'(y i Xb; ) a (b; = b;)/(Nix r N, ')(b; a b; «) (12-21)

Or

Cr= ys X;b; x)'(y = X;b; «) = (b; a bj «)’N,(N¥* at N,) 'N,(b, r b; «)


(12-22)

where N, = X/X, and b, = (X,X,)~'X‘y,. The first term in Eq. (12-21) is the
residual sum of squares from the OLS fit of model M,. This has to be increased by
a factor depending on the discrepancy between the OLS vector and the prior
mean vector. Alternatively, the first term in Eq. (12-22) is the error sum of squares
if the coefficient vector were set equal to the prior mean vector. This is adjusted
downward by a term which is again dependent on the discrepancy between the
sample and the prior coefficient vectors. Thus apart from the prior odds ratio, the
choice between models would depend on these adjusted sums of squares, which
are a mixture of sample and prior information.
Readers must judge for themselves whether a criterion such as Eq. (12-20),
for all its elegance and simplicity, is a valid guide for choice. Suppose that just
two crude and simple models are being compared. M,, say, is a “Keynesian”
reduced-form equation relating GNP to “exogenous expenditures,” while M, is a
“Friedmanian” equation relating GNP to “money.” If the Ghost of Keynes could
be contacted, he would presumably offer a prior-odds ratio P(M,)/P(M,),
dramatically different from that forthcoming from Professor Friedman. How can
the protagonist of one theory begin to specify the prior pdf’s for the parameters
of the opposing theory, which he basically regards as false? Must then a Bayesian
researcher be certified ideologically pure and unbiased before being allowed to
specify prior odds and prior densities for model parameters?
A partial resolution to the problem of excessive dependence on priors, which
are, perhaps, spuriously precise, idiosyncratic, or just personal to one investigator,
is provided by some recent work by Chamberlin and Leamer.{ It is assumed that
in a single equation there are one or more “focus” variables, whose coefficients
are of crucial interest. The equation may also contain other “doubtful” variables.
The investigator specifies a prior zero mean vector for the doubtful variables.
However, he does not have to specify the elements of the prior variance matrix,
merely that it belongs to the class of positive definite or semidefinite matrices.
Leamer’s SEARCH program computes bounds on the focus coefficients, so that
the researcher can study the robustness of these coefficients under a variety of
specifications. This approach is appealing and seems likely to be developed and
considerably extended. It adds yet another dimension to the array of information
that we can obtain on any specific problem. How to weigh and interpret the

+E. E. Leamer, op. cit., p. 109.


+E. E. Leamer, op. cit., pp. 182-201. See also T. F. Cooley and S. F. LeRoy, “Tdentification and
Estimation of Money Demand,” American Economic Review, vol. 71, 1981, pp. 825-844, for a very
interesting detailed application of the procedure.
516 ECONOMETRIC METHODS

jigsaw of computation and information will still depend on the vital spark of
human imagination and powers of judgment.
The position can best be summarized by a quotation from the late Jacob
Bronowski.+ Though writing of the physical world, his comments are very
apposite to the economic and social world that we study.

The world is not a fixed, solid array of objects, out there, for it cannot be
fully separated from our perception of it. It shifts under our gaze, it interacts
with us, and the knowledge that it yields has to be interpreted by us. There is
no way of exchanging information that does not demand an act of judgment.

Science is a very human form of knowledge. We are always at the brink of the
known, we always feel forward for what is to be hoped. Every judgment in
science stands on the edge of error, and is personal.

+ J. Bronowski, The Ascent of Man, Little, Brown, Boston, 1973,


pp. 364 and 374.
APPENDIX

A
MATHEMATICAL
AND STATISTICAL APPENDICES

A-1 FUNCTIONS AND DERIVATIVES

The purpose of this section is merely to remind the reader of various notational
conventions for functions and derivatives. It is not intended to review the basic
rules of differentiation.} If y is a function of x, the relationship may be denoted
variously as

y=y(x) y=f(x) y=ealx) y= F(x)


and so on. Once a specific functional form for the relationship has been assumed,
one can determine the shape of the function by studying the behavior of y in
response to variation in x. Starting from an initial value, say, x), and moving to
xX, = X) + Ax will trace a movement in the dependent variable from yy to y,. The
ratio
AV eit 6
Ne Ka,
measures the change in y per unit change in x. Taking the limit of this ratio as
Ax — 0 gives the derivative of y with respect to x, written variously as

® = 7(x)= lim (32)

The derivative measures the slope of the function at a specific point and is, in
general, a function of x, as is emphasized by the f’(x) notation. Thus it may itself

+ For a lucid introduction to the calculus and other mathematical topics of special relevance to
economists see A. C. Chiang, Fundamental Methods of Mathematical Economics, 2d edition, McGraw-
Hill, 1974.
517
518 ECONOMETRIC METHODS

be differentiated with respect to x, giving the second-order derivative, denoted by


d? y
OF 4, SESH)
axe
When y is a function of several variables, say,

Vesa OX ore |
where the x’s are capable of moving independently of one another, then one may
study the change in y in response to the change in any one of the independent
variables (or arguments) of the function, the other independent variables being
held constant at any arbitrary set of values. This gives rise to the partial
derivatives, denoted by

In this notation f,, for example, would indicate the rate of change of y with
respect to x,. Once again further partial differentiation may be carried out,
yielding the second-order partial derivatives

d*y a ae =
Ox;0x, Ox,0x, “4
Alternatively, if the independent variables are given separate labels, as in

Vi flues)
the partial derivatives may be denoted by

Ces Cvs Oy
One Ap ie
Onde ae
though, even here, one may see f, used for f, and Jo lOEp.

A-2 EXPONENTIAL AND LOGARITHMIC FUNCTIONS

Consider the function

y= b* b>0O (A-1)
This is called an exponential function since the variable x appears as the exponent
of the constant, or base, b. We rule out negative values for b, since if x were, say,
one-half, » would be the square root of a negative number, which is imaginary. If
x denoted time ¢ measured at equal intervals, then

y, = b' and Vt =b
Veet

Thus y, denotes a series which is growing (b > 1) or declining (0 < b <


1) ata
constant rate. If we set b = 1 + r, then r denotes the proportionate rate of
change
in y per unit period of time.
MATHEMATICAL AND STATISTICAL APPENDICES 519

(a)

Figure A-1 (a) Exponential function; (b) logarithmic function.

The logarithm of a number to a given base is defined as the power to which


the base must be raised to give the number. Thus in Eq. (A-1) x is the logarithm
of y to base 5, written

x = log, b y A-2

This is the inverse of the exponential function. The first expresses y as a function
of x and the second expresses x as a function of y. Typical graphs for b > 1 are
shown in Fig. A-1. If the graph in Fig. A-15 were superimposed on Fig. A-la with
the y axis on the y axis and the x axis on the x axis, the curves would coincide.
Numerical calculations are facilitated by the tables of common logarithms, which
are taken to the base 10. Thus, for example, log,,100 = 2 since 100 = (10)?. In
practice the subscript 10 is rarely shown explicitly. For mathematical purposes it
is usually much more convenient to work with natural logarithms, which are taken
to base e. This is the mathematical constant defined byt

e= lim ( +2 = 2.41828

This has the remarkable property that if


x
aa ae.

then

ayes aroracal = ("

dx dx?

that is, all derivatives are equal to the original function. The function is written in

+ See also the footnote on p. 68.


520 ECONOMETRIC METHODS

alternative forms as
y=e* or y= exp{x}
and the inverse logarithmic function is written ast
x = log, y or x=Iny
The general exponential function is written as
y = Ae™ or y= Aexp{cx}
which has the effect of stretching or contracting the typical exponential shape in
Fig. A-la vertically and horizontally.
If the inverse function exists, as it does when y = f(x) is monotonic (that is,
to each value of x there corresponds a unique value of y and vice versa), then
dx 1
dy dy/dx
Ify = e*, then dy/dx = e* = y, and so for the inverse function, x = In y, dx /dy
= 1/y. Thus we have the two standard forms:

Wiaes
yH=e ay meeahs
Be

e
y=Inx Deo
ei

Suppose we have y = log x. What then is dy/dx? We may write


x = 10”
Thus
In x
¥=Ti9 ~ nx: loge

since it may easily be shown that In 10 - log e = 1. It then follows that

d(logx) 1
ie ie log e (A-3)
Finally we may note a frequently used connection between logarithms and
elasticities. Ify = f(x) and a change Ax is imposed leading to a change Ay, then

BE ee ge
perl iy Ax yy
measures the proportionate change in y per unit proportionate change in
x. The
elasticity of y with respect to x is defined as the limiting value of this ratio
as
Ax — 0, that is,

(Point) elasticity of y with respect to x = o ee


x y

+ In general In denotes a logarithm to base e and log a logarithm to base


10.
MATHEMATICAL AND STATISTICAL APPENDICES 521

It may be shown that


d d(l d(l
ax” ye d(lnx)- d(loex)
To show the first part of the identity let
z=Iny y=f(x) and x=e” sothatw=Inx
Then

oe usin ON eae
ven 5 dw dw/dx — %
= elasticity of y with respect to x
The second part of the identity shows that the same relation holds if logarithms
are taken to base 10, since there is a proportionate relationship between loga-
rithms to the two bases.
It follows from Eq. (A-4) that a functional form which implies a linear
relation between the logs of the variables is a constant elasticity function. For
instance,
y = Ax*
gives
log y = log A + a(log x)
so that a is the elasticity of y with respect to x. A simple way to fix the meaning of
an elasticity is that it measures the percentage change in y produced by a / percent
change in x.
The elasticity concept extends to functions of several variables. Thus
y = AxFz¥
is a constant elasticity function, where a, 8, and y are the partial elasticities with
respect to the arguments x, v, and z.

A-3 OPERATIONS WITH SUMMATION SIGNS

The Greek capital sigma is used to indicate summation. Thus

De Ng
Mae cick X,)
i=l
This sum is variously denoted by

(LX
i=n n n

Seat
i=1
Oeee ima ys orjust!
i=] 1

so long as no ambiguity is involved in any particular application.


522 ECONOMETRIC METHODS

If each value of X is multiplied by a constant a, the sum is


n n

(aX, + aX, tie aX, = aX, = a, (A-5)


i=] i=]
Thus a constant appearing after a summation sign may be moved in front of it to
multiply the sum. If each X; in Eq. (A-5) were equal to unity, YX, = n and
n n

Dea dian
j= i=

Thus the summation of a constant over n points is n times the constant.


The arithmetic mean of the X’s is

x=
UX; t
n
It follows directly from the definition that

2 (Xin Xia Xp
i=]
so that the algebraic sum of deviations around an arithmetic mean is zero. The
sum of squared deviations from the arithmetic mean is

E(=
Y
i=]
¥(aes
-¥ axea yy i=]

I DEX DX SE ae

-Dx- (Ex)
or alternatively,
n 2

DEN NEO Neer Xe


i=]
that is, the sum of squared deviations about the sample mean can be
expressed as
the sum of the squared values of the original variables less a correction factor
for
the mean. Similarly, one may derive

SoKay
= aeee
i=]1

= OGY) as
l
a)
Suppose a variable has two subscripts, say,

Xj P= NV2p sy Ps ely eae


MATHEMATICAL AND STATISTICAL APPENDICES 523

This is illustrated in the following table:

| Xi, X2 Xin,
2 X91, X225-++5 Man,
Class :

Pp Xp X,2 SIA sIs Xpn,

As an example, X might measure personal income and the sample data consist of
n, observations from social group 1, n, observations from social group 2, and so
forth. Total income in the sample is defined by
peak
» » xX, or, more simply, ae
Hes | Tey

The total number of sample observations is

and the overall mean income is then

ae 2, Xi
n
The mean income for the ith group, or class, is
nj
af ie j
x; n;
The sum of squared deviations about the overall mean is

E(%)- ¥) = E[(x,- %)+(%-¥)]


iJ

=(%,-
my %) +0(¥-X)
Al: +20(%,-
; ¥)(¥-%)
The last term may be written
P

SE(4 HI he eae | (x= X))


.—

since the factor (X, — X) does not involve the j subscript and so may be moved in
front of the summation over /. But
524 ECONOMETRIC METHODS

for each 7, and so the whole term vanishes. The middle term may be written

(xem) |><!=ei=1 j=leee


fey

since (X, — x )* is a constant for each element in the ith group, so that the sum
overj is simply n,;(X; — X). Thus

(x,-¥) =L(%,-
us
RY +Da(%-¥) Ey) I

This decomposition is often written as


Total sum of squares = within-group sum of squares
+ between-group sum of squares

A-4 RANDOM VARIABLES AND PROBABILITY DISTRIBUTIONS}

A discrete random variable X consists of a set of possible values Xo a


and associated positive fractions (probabilities) p,, p5,..., Pp, such that
k
2 Diaeae |
i=]
The two most important features of the probability distribution are the mean and
the variance. The mean, often denoted by p, is defined as
k
B= E(X)= Dee Te, (A-6)
i=]
which is just a weighted average of the x values, the weights being the respective
probabilities. E is the expectation operator, and it may also be applied to various
functions of X. For example, E(X*) indicates the expected value of X2. The
possible values for X* are x?, x3,...,x2, which occur with probabilities
P\> Pr»---, Py. Thus
k

EXE) = DE xp,
t=]

The second most important feature of the probability distribution is the


variance

} For the statistical paragraphs in this appendix, two of the most lucid
texts at an introductory
level are P. J. Hoel, Introduction to: Mathematical Statistics, 4th
edition, Wiley, New York, 1971, and
L. D. Taylor, Probability and Mathematical Statistics, Harper and Row, New
York, 1974.
MATHEMATICAL AND STATISTICAL APPENDICES 525

or expected squared deviation about the mean. This is usually denoted by o”. Thus

a? = E((X - n)’)2 (A-7)


Evaluating this from first principles,

2 : | 2
BCX)= Cas By
i=l

= Exjp; — 2uLx,p, + weXp,


y ap: - (Lx,p,)”

= BCX) Eee)
This result may also be obtained by squaring the expression in Eq. (A-7) and
applying the expectation operator to each term in turn. Thus

E({(X— w)°} = E{X? — 2X+ y?}


= E(X7)— 2pE(X) + E(w)
= EM) ME0)
since E(u*) indicates the expectation of a constant, which is simply the constant.
When the random variable is continuous, the discrete probabilities are
replaced by a continuous probability density function (pdf), usually denoted by
p(x) or f(x). An example is shown in Fig. A-2.
The pdf has the properties that
f(x) =0- “for-allx

[f(x) ax =]

and [Pes ax = Pra


<x < 6]

The mean and the variance are defined as before, but integrals now replace
summation signs.

Figure A-2
526 ECONOMETRIC METHODS

We are often interested in the joint variation of a pair of random variables.


Let the variables X, Y have a bivariate pdf denoted by f(x, y). Then
f (Sayre; stor allay

[{fC y) dx dy = ]

and

[f/f 0% y)ax dy = Prla< X<b,e<y<d]

Given the joint density, a marginal density is obtained for each variable by
integrating over the range of the other variable. Thus

Marginal pdf for X -[- f(x, y) dy = f(x)

and

Marginal pdf for y-[~ f(x, y) dx = f(y)

A conditional pdf for Y, given X, is defined as

Tey)
(
EO)
{(x)
aes (A-8)
and similarly, a conditional pdf for X, given Y, is defined as

Ae)
(
Lip =
/0)
Two variables are said to be statistically independent, or independently distrib-
uted, if the marginal and conditional densities are the same. Thus the joint
density can be written as the product of the marginal densities

f(x, vy) =f(x)- f(y) | (A-9)


The mean and variance for each variable may be obtained from the marginal
densities. Thus

Be = E(X) = ffxfy),
dxdy= fxf(x) ax
oy = var(X) = f(x — w,)’f(x) ax
and similarly for the mean and the variance of Y. A new statistic for the bivariat
e
case is the covariance. It is defined as

Sry = cov( X,Y) = El(x— wy -n,)) = ffx — wy )(y - wy) f(x, ») de dy


and measures the linear association between the two variables. For indepe
ndently
MATHEMATICAL AND STATISTICAL APPENDICES 527

distributed variables the covariance is zero since

Lf — wr = wy) fCe y) dx dy = [x= BIC) axf(y = Hy) f(y) &y


= 0
In general, the converse of this proposition is not true, that is, a zero covariance
does not necessarily imply independence. An important exception to the proposi-
tion, however, exists in the case of normally distributed variables, as is shown in
App. A-5.

A-5 NORMAL PROBABILITY DISTRIBUTION

The pdf for the univariate normal distribution is

f(x) = 1
ep|- 1
(x-0) 2 (A-10)
This defines a two-parameter family of distributions, the parameters being the
mean p and the variance o*. The bell-shaped curve reaches its maximum at x =
and is symmetrical about that point. A special member of the family is the
standard normal distribution, which has zero mean and unit variance. An area
under any specific normal distribution may be expressed as an equivalent area
under the standard distribution by defining
eee eg
ea
Clearly, E(z) = 0 and var(z) = 1, so that
l 2
f(z)z)= as eRe A-11
(A-11)
Then [fo dx = [°f(z) a

where z; = (x; — »)/o. The areas under Eq. (A-11) are tabulated in App. B-1.
Three very important results about the normal distribution are as follows.

1. Linear combinations of normally distributed variables are themselves normally


distribufed.+ For example, if x denotes an n X 1 vector of variables, which
follow the multivariate normal distribution
x ~ N(p, 2)
and if a vector y is defined by y = Dx where D is an m X n matrix of rank

+ See L. D. Taylor, Probability and Mathematical Statistics, Harper and Row, New York, 1974, pp.
154-160.
528 ECONOMETRIC METHODS

m <n, then

y ~ N(Dp, DID’)
2. Central limit theorem.} If (qv,= )15-a phere: a independent random
variables with means (j1,, “,... ) and variances (0/7, o7,... ), then

ys (ea ile)
lim
n— oo
Pr Ae
= 3
Ss
= a ee
exe" dz) eC Ae)
0;
i=]

Notice first of all that nothing is assumed about the specific forms of the
various pdf’s other than the existence of means and variances. The remark-
able result embodied in Eq. (A-12) is that the limiting or asymptotic distribu-
tion of the quantity U(x;— p;)/\Zo? iis the standard normal distribution. A
special case of the result may help to make its meaning clearer. Suppose the
means and variances are all identical. The statistic in Eq. (A-12) then reduces
to
n _ _*¥-p

Date np o/Vvn
1=1

and the theorem states that x is asymptotically normally distributed with


mean p. and variance o7/n.
3. Zero covariance between two normally distributed variables implies statistical
independence. The bivariate normal distribution is

f(x») =mae l
2770,0,¥ 1 — p°
1
||x — p,\?
|
/ pe) Ox

a r a | l e o e l a l a = ) | |
“2lir
where p = 0,,/0,0,. When the covariance d,, 18 Zero, the joint pdf simplifi
es

fon [aeeeal-4(*58)]ataoe ill


which is the product of two separate normal pdf's. Thus X and Y are
independently distributed.

7S. S. Wilks, Mathematical Statistics, Wiley, New York, 1962, pp. 257-258.
MATHEMATICAL AND STATISTICAL APPENDICES 529

A-6 LAGRANGE MULTIPLIERS AND CONSTRAINED


OPTIMIZATION

Frequently in economics one has to find the maximum or minimum of a function


subject to some constraint, or set of constraints, on the independent variables.
Thus one may have to maximize profits subject to the constraint of the produc-
tion function, or maximize utility subject to the constraint of the budget equation.
In the context of the regression model, as in Chap. 6, one may need to minimize a
residual sum of squares subject to a set of constraints on the regression coefficients.
Such constrained optimization problems may be tackled in two alternative
fashions. The first is to substitute the constraints into the objective function, thus
reducing the number of independent variables and find the stationary values of
the resultant, unrestricted function. The second is to use the method of Lagrange
multipliers. We will illustrate with a simple example.
Suppose the problem is to find the minimum value of y = f(x, z) = x? + z?
subject to x + 2z = 10. Substituting the constraint in the objective function
means that the latter may be expressed as a function of just a single variable,
either x or z. For example, replacing x by 10 — 2z gives

y = (10 — 2z)* +z? = 100 — 40z + 52?


Differentiating
2
ee Agel ie eat
dz dz?
Thus a minimum value occurs at z = 4 (x = 2), and that minimum value is
y = 20.
Alternatively, define a new, or augmented, objective function as
p= x? +2z7—X(x
+ 2z — 10)
where A is a Lagrange multiplier, whose value is as yet unknown. So long as the
constraint is satisfied, the term A(x + 2z — 10) vanishes, irrespective of the value
of A, and ¢ will have the same stationary value as y. To find the stationary value
of @ we must take the three partial derivatives and equate to zero. Thus
Op _ mes
ay ox Aa)

OTE eal
dz

=x 422-10=0

The third equation ensures that the constraint is satisfied. Eliminating A from the
first two gives 2x = z, which on substitution in the third gives x = 2(z = 4) and,
as before,

Pmin a Yimin = 20
530 ECONOMETRIC METHODS

The solution value for A is A = 2x = z = 4, which in this case has no specific


significance. There are problems, however, where A may have a meaningful
economic interpretation.
The technique extends to handle more than one constraint, as is evidenced by
the examples in Chaps. 2 and 6.

A-7 RELATIONS BETWEEN THE NORMAL, x’, t AND F


DISTRIBUTIONS

Let z ~ N(0,1) be a standard normal variable. If n random values z,, z,..., Z,


are drawn from this distribution, squared, and summed, the resultant statistic is
said to have a x? distribution with n degrees of freedom,

(2? +22 +--+ +27) ~x?(n)


The precise mathematical form of the x’ distribution need not concern us here.
The important point is that it constitutes a one-parameter family of distributions,
and the parameter is conventionally labeled the degrees of freedom of the
distribution. As the degrees of freedom tend to infinity, the x? distribution
approaches the normal density. Critical values of the x* distribution are given in
App. B-3.
The ¢ distribution may be defined in terms of a normal and an independent
x° variable. Let

z= N(0,1) **and= © "42 (@)


where z and v are independently distributed. Then

ae (A-13)
vo
has Student’s ¢ distribution with v degrees of freedom. The ¢ distribution, like x’,
is a one-parameter family. It is symmetrical about zero and tends asymptotically
to the standard normal distribution. Its critical values are given in App. B-2.
The F distribution is defined in terms of two independent x? variables. Let u
and v be independently distributed x? variables with vy, and pv, degrees of
freedom, respectively. Then the statistic

u/V,
Fete
4s, 2
(A-14)

has the F distribution with (v,, v,) degrees of freedom. Critical values are given in
App. B-4. In using the table note carefully that v, refers to the degrees of freedom
attaching to the expression in the numerator and v, to the expression in the
denominator.
MATHEMATICAL AND STATISTICAL APPENDICES 531

If we square the expression for 1, the result may be written

t
Oe 27
v/v
where z? eine the square of a standard normal variable, has the x?(1) distribu-
tion. Tus t? = F(1,v), that is, the square of a ¢ variable with v degrees of
freedom i Syan F variable with (1, v) degrees of freedom.
The x* variable was formed from the sum of squares of a standard normal
variable. Suppose, however, that the z variables are still independent but distrib-
uted as

Z; ia N(u,;, 1)

The statistic z> + z3 + --- + z? now has the noncentral x? distribution with n
degrees of freedom. The previous distribution is sometimes referred to as the
central x* distribution. Corresponding to a noncentral x? distribution, there are
noncentral ¢ and F distributions, the former arising when the v variable in Eq.
(A-13) is noncentral and the latter when u in Eq. (A-14) is noncentral, but v is
central.+

A-8 EXPECTATIONS IN BIVARIATE DISTRIBUTIONS

Let X and Y be two variables with a bivariate pdf denoted by f(x, y). Let g(x, y)
be some function of the variables. The problem is to evaluate E{g(x, y)}. By
definition

E(g(x, y)) = [fax »)f(x, y) dx dy

= [falx, »)f(b) f(y) ax ay (A-15)


where f(x|y) denotes the conditional distribution of X given Y and f(y) denotes
the marginal distribution of Y. Rearranging

E{g(x, y)} = {|fol, y)f(xly) dx} f(y) ay (A-16)


The term inside the square brackets gives the expected value of g(x, y) in the
conditional distribution f(x|y), and we will denote the operation by £,,,. This
conditional expectation is a function of Y, and it is then averaged over the

+ For references to some tables for noncentral distributions see B. W. Lindgren, Statistical Theory,
2d edition, Macmillan, New York, 1968, p. 383.
+ We exclude functions g(-) which may have some values undefined such as x/0 or 0/0.
532 ECONOMETRIC METHODS

marginal distribution f(y). Thus we may write

E{g(x, y)} = E,{E,,,8(x, y)}


Example Consider the simple bivariate distribution
f(x, y)

y=2 y=4 f(x)

x=1 0.2 0.4 0.6


x=2 0.3 0.1 0.4
f(y) 0.5 0.5 1.0

Let g(x, y) = x/). The straightforward application of Eq. (A-15) would then
give

E(= = 4(0.2) + 4(0.4) + 3(0.3) + 3(0.1) = 0.55


Using Eq. (A-16) we would first find E(x/y) within each of the two columns
of the table, which contain the conditional distributions f(x|y). Thus
a ILff Oe 2 {0:3
E(5b =2)=5(55) +3lgs) - 98
and
x 1 /0.4 2D (eOs1
e(Sp=4)-4(53) + alos) 702
Then averaging these expectations over f(y) gives, as before,

E(=] = 0,8(0.5) + 0.3(0.5) = 0.55


Clearly the process is symmetrical and we could average first of all over
each conditional distribution f( y|x) and then average the results over Toe:
The procedure, however, would break down for the function g(x, y) = x /y if
zero is a possible value for y.

This result has useful applications in regression theory, but it has to be


handled carefully. As a simple illustration consider the model

Var Pp tas
where the u’s are well-behaved, the x’s are stochastic and distributed indepen-
dently of the u’s so that, in particular,
E(x
=,u,
E(x,)E(u)
,)=0 forallz
The OLS estimator of B is

hale pee
xs DE
MATHEMATICAL AND STATISTICAL APPENDICES 533

To examine the bias of b we need to evaluate E(Xx,u,/Xx7). Applying the


theorem

Sat] {Bue S|
xu Dx.
Be a Ee Ne)
so that 5 is still unbiased when x is stochastic, provided it is independent of wu.
The sampling avariance is given by

var(b) Sep
= E{(b — B)’} heeeies oH|aaa
sh

Now

ie yay : a os
Sepang? Dn,
Thus

]
var(b) = ott
t

which is the one-dimensional version of the general result given in Eq. (7-26).
Now consider the case where x is stochastic but no longer independent of u.
Suppose, for example, that x, u follow a bivariate normal distribution

f(x, u) =

-2»(22#\(4)+ (#)]
tomcat HET
with marginal densities

and

f(u).= ee
—exe|5)

The conditional density for u, given x, is

f(ulx) me=
an eat eeae
‘ zs i= Ze ep shat 262(1 amexa pea
0”) | 0, ( )

(A-17)
Thus
E(xu|x) = xE(u|x) =x-
po,
- (otSic)
534 ECONOMETRIC METHODS

since Eq. (A-17) shows f(u|x) to be normal about mean po,(x — p,)/o,. It then
follows that

go ee ue al
exe Bele eee

— PO,
l ee —
= z.|6, oe Dexa K)|

oP On,, E De em eo
0. Ling |

On dividing the top and bottom by n, the term in brackets is approximately the
ratio of the sample variance of the x observations to their sum of squares. This
expectation does not vanish, and so b is a biased estimator, both for finite samples
and also asymptotically. This example is a legitimate application of Eq. (A-16)
since f(x, uw) is a well-defined bivariate distribution.
Now consider the model

aa By, F Uu, (A-18)

The OLS estimator is

With the w’s well-behaved there is no problem about assuming


E(y,_,\u,)=0 — forall?
An application of Eq. (A-16) would then appear to give

Be = EBay | = 0

ye fal e ye
However, the OLS 5is well known to be biased in finite samples.} The source of
the error is that (y,_,,u,) does not have a well-defined bivariate pdf, which
renders the application of Eq. (A-16) invalid. Given some starting value Yo, Once a
u vector is drawn, the y vector is exactly determined by Eq. (A-18). The stochastic
behavior of y is completely determined by u, so there is, in effect, only one
stochastic variable. Thus E(y,_ ,u,/Xy/_,} has to be evaluated solely over the u
distribution, and the two-step procedure of Eq. (A-16) does not apply.
A more complicated version of the same error can arise in IV estimation.
Consider

Vr a By,_ sie YX; ts U,

Suppose {x,} is taken to be nonstochastic and x,_, is used as an instrument for

+ J. S. White, “Asymptotic Expansions for the Mean and Variance of the Serial Correlation
Coefficient,” Biometrika, vol. 48, 1961, pp. 85-94.
MATHEMATICAL AND STATISTICAL APPENDICES 535

¥,—;. The W’Z matrix of Sec. 9-2 will then include terms in Lx, y,.,/and
2x,-1);—-1. These are stochastic simply because wu is stochastic and, as in the
simple example given above, a two-stage evaluation via E.,,, followed by E,, is
invalid.

A-9 CHANGE OF VARIABLES IN DENSITY FUNCTIONS}

The basic idea may be simply illustrated for the univariate case. Suppose u is a
random variable with density function p(w), and suppose that a new variable y is
defined by the relation y = f(u). The y variable must then also have a density
function for y in terms of the density function for u and the relation y =f(u).
Suppose the relation between y and u is monotonically increasing, as shown in
Fig. A-3. Whenever u lies in the interval Au, y will be in the corresponding
interval A y. Thus
Pr{y lies in Ay} = Pr{u lies in Au}
or p(y’) Ay
= p(w’) Au
where u’ and y’ denote appropriate values of u and y in the intervals Au and Ay,
and p(y) indicates the postulated density function for y. Taking limits as Au goes
to zero gives

p(y) = p(w) (A-19)


If y were a decreasing function of u, the derivative in Eq. (A-19) would be
negative, thus giving an impossible negative value for the density function. Thus
the absolute value of the derivative must be taken and the result reformulated to

7 o Figure A-3

+A detailed treatment of this topic is given in L. D. Taylor, Probability and Mathematical


Statistics, Harper and Row, New York, 1974, Chap. 10.
536 ECONOMETRIC METHODS

read

du
P(y) =p(u)- dy (A-20)
If y = f(u) were not a monotonic function, Eq. (A-20) would require amendment,
but we are only concerned here with monotonic functions.
In the multivariate case u and y now indicate vectors of, say, n variables each.
Under suitable conditions a result similar to Eq. (A-20) still holds, namely,
ou
p(y) = p(u) dy (A-21)
where |du/dy| indicates the absolute value of the determinant formed from the
matrix of partial derivatives

du, ee
| du, eedu,

dtr, tk ws du,
ORO
i ae
eeieon
er ae

dy, ay, OY,

A-10 PRINCIPAL COMPONENTS

Suppose we have a matrix X of n observations on k variables,

where the observations have been expressed as deviations from the sample means,
for we are concerned with studying the variation in the data.
The nature of principal components may be approached in a number of ways.
One is to ask how many dimensions there are or how much independence there
really is in the set of k variables. More explicitly we consider the transformation
of the X’s to a new set of variables which will be pairwise uncorrelated and of
which the first will have the maximum possible variance, the second the maximum
possible variance among those uncorrelated with the first, and so forth. Let
21p— yy, ¥ OaiXae Mee Gye, belt
denote the first new variable. In matrix form
z, = Xa, (A-22)
where z, is an n-element vector and a, a k-element vector. The sum of squares of
z, 1S
Z\Z, = a',X'Xa, (A-23)
MATHEMATICAL AND STATISTICAL APPENDICES 537

We wish to choose a, to maximize z/z,, but clearly some constraint must be


imposed on aj, otherwise zz, could be made infinitely large. So let us normalize
by setting

The problem now is to maximize Eq. (A-23) subject Eq. (A-24). Define
o = a X’'Xa, — A,(aia, — 1)
where A, is a Lagrange multiplier. Thus

a = 2X’'Xa, — 2,2,
Setting
Op
dan
gives
(X’X)a, =A,a, (A-25)
Thus a, is an eigenvector of X’X corresponding to the root A,. From Eqs. (A-23)
and (A-25) we see that
Z\z, = A,a\a, =A,
and so we must choose A, as the largest eigenvalue of X’X. The X’X matrix, in the
absence of perfect collinearity, will be positive definite and thus have positive
eigenvalues. The first principal component of X is then z,.
Now define z, = Xa,. We wish to choose a, to maximize a’, X’Xa, subject to
a’,a, = | and aja, = 0. The reason for the second condition is that z, is to be
uncorrelated with z,. The covariation between them is given by
a) X’Xa, = A, aa,
=0 if and only if aja, = 0
Define
@ = a,X’Ka, — A,(aa, — 1) — w(a\az)
where A, and p are Lagrange multipliers.

SOON
da,
a Eh as = ia 0
Premultiply by a’
2a) X’Xa, — p = 0
But from
(X’X)a, = A,a,
a’,(X’X)a, = A,a5a, = 0
Thus
w=0
538 ECONOMETRIC METHODS

and we have

and A, should obviously be chosen as the second largest latent root of X’X.
We can proceed in this way for each of the k roots of X’K and assemble the
resultant vectors in the orthogonal matrix

Asay a, 2"? Paz] (A-27)


The k principal components of X are then given by the n X k matrix Z,
Z=XA (A-28)
Moreover,

Na) 0
ZZ=AXXA=A=]0 A, <*> 0 (A-29)
geet .

showing that the principal components are indeed pairwise uncorrelated and that
their variances are given by
z.z,=A, ies lee (A-30)
If the rank of X were r < k, k — r eigenvalues would be zero and the variation in
the X’s could be completely expressed in terms of r independent variables. Even
if X has full column rank, some of the A’s may be fairly close to zero so that a
small number of principal components account for a substantial proportion of the
variance of the X’s. The total variation in the X’s is given by

ae T Dene tebe Pare Die, 7 tr(X’X)


t t t
but

tr(A’X’KA) = tr(X’XAA’)
tr(X’X)
since AA’ = I, and so from Eq. (A-29)
k n k
DY Lo xf = (XX)= VA,
=212, +--+ + 2,2,
i=1 t=1 i=]

Thus

Ale
AN ean
represent the proportionate contributions of each principal component to the
total variation of the X’s, and since the components are orthogonal, these
contributions sum to unity.
It is sometimes difficult to attach a concrete meaning to specific principal
components. Occasionally a suggestion may be found in the correlations of a
component with various X’s. To find the correlation between, say, the first
MATHEMATICAL AND STATISTICAL APPENDICES 539

principal component and the X variables, we proceed as follows. The vector X’Z,
gives the cross products between z, and each X variable. But
X’z, = X’Xa, = X,a,
Thus the correlation between X, and z, is
an Aiaiy
hms a
2
Ay Di
t=1

aA,
eea ee Le re ke (A-31)
2
ys Xir
Gal

where a;, is the ith element in the vector a,. In general, the correlation between_X,
and z;, is
a;, d,
rj =———= — ii, j=l,...,k (A-32)
| LuXit; t

These correlation coefficients may also be used to show how the variations in
each X variable may be decomposed into the contribution due to each compo-
nent. From
Z=XA
we have
L — XG

and X’ = AZ’
since A is orthogonal.
So X’X = AZ’ZA’
= AAA’
from Eq. (A-29), and so
n k
xe Dag A Lele i (A-33)
t=1 j=l
Dividing both sides of Eq. (A-33) by X,x7, gives
ie aid . 4i2rdr ee did 4 (A-34)
Lixe DXi DXi
where the terms on the right-hand side are the squares of the correlation
coefficients defined in Eq. (A-32). Thus the proportions of the variation in X;
associated with the various principal components are given by
22 aed, 2
Vito Vid.+ ++ Tik
540 ECONOMETRIC METHODS

and since the components are uncorrelated, these proportions sum to unity, as is
shown by Eq. (A-34).
A note of warning should be inserted here. The development so far has
proceeded on the implicit assumption that the X variables are all measured in the
same units. If not, it is difficult to attach a meaning to concepts such as the total
variation of the X’s and the partitioning of that total variation into the contribu-
tion due to each component. It is still, of course, possible to compute the
eigenvalues and eigenvectors of X’X even if the dimensions of the variables are
not all the same and the correlations in Eq. (A-32) and the partitioning in Eq.
(A-34) would still be meaningful even though the partitioning of the total
variation in the X’s would not. As an alternative, analyses are sometimes carried
out after all the X variables have been standardized, that is, each deviation from
the sample mean is divided by Vn times the sample standard deviation of that
variable. X’X is now the matrix of zero-order correlation coefficients of the X
variables. The analysis can proceed from X’X as before. Now tr(X’X) = k, and
from the development following Eq. (A-29),

ED Neenem,
The eigenvalues and eigenvectors will in general be different from those yielded
by unstandardized variables. We leave it as an exercise for the reader to establish
whether the correlation coefficients in Eq. (A-32) are affected by the standardi-
zation of the X variables.
Empirically, then, one may compute the principal components for a given X
matrix and see how much of the variation of the X’s is accounted for by various
components. Frequently the intercorrelation of economic and social data means
that a small number of components will account for a large proportion of the
total variation, and it is desirable to have a test for Judging the number of
components to retain for further analysis. Suppose that we have computed the:
roots A,,A,,..., A, and that the first r roots Aj, Ag,---,A, (7 < k) seem
both
sufficiently large and sufficiently different to be retained. The question then
is
whether the remaining k — r roots and their associated vectors and the compo-
nents are sufficiently alike for one to conclude that the true values are equal. A
very approximate test is based on

Sif Nee HaNec op actA \ee


p= (NEN ee ND) ee (A-35)
The proposed test is to consider nlog,.p to follow a x? distribution
with 4
(k —r— 1)(k —r + 2) degrees of freedom, if the null hypothesis of
equality of
the remaining latent roots is true.t+ One hopes in practical applica
tions that the
number r of significantly different components to be retained is substan
tially less
than the number of variables k from which the components have been
computed.

+ See M. G. Kendall and A. Stuart, The Advanced


Theory of Statistics, vol. 3, Griffin, London,
1966, pp. 292-293, for details and qualifications.
MATHEMATICAL AND STATISTICAL APPENDICES 541

A somewhat similar result is achieved by factor analysis in which the X


variables are specified ab initio to be linear combinations of a small number of
independent standard normal variables (factors) plus an independent normal
error term. From the principal component analysis we have

Z = XA
and hence

xX = ZA’ (A-36)

Equations (A-36) express the X’s as exact linear combinations of the components
with coefficients given by the elements of A. If, however, we retain less than k
principal components, Eqs. (A-36) would have to be replaced by

X = Z*A*’ + U (A-37)
where Z* and A* denote the submatrices of Z and A giving the retained
components and the corresponding eigenvectors, and U is a matrix of errors.
Principal components is obviously a possible estimation method in factor analy-
sis, but slight modifications are required to the A* coefficients to conform to the
imposed assumption that the factors should have unit variance. Without addi-
tional restrictions z;z; = A,, as we have seen in Eq. (A-29). When the A*
coefficients have been adjusted, they are referred to as factor loadings. However,
several other estimation methods are used in factor analysis, and we do not
propose to discuss them here. An interesting application of factor analysis is
given by Adelman and Morris, who find that 66 percent of the variance of the
GNP per capita in 74 underdeveloped countries associated with just four factors,
which have in turn been based on a complex of more than 20 social and political
variables.t
Table A-1 shows another example in which a small number of components
effectively account for the variation in a set of data. The basic data are 11 series
of average quarterly interest rates in the United Kingdom from the first quarter of
1963 to the first quarter of 1969. They include various national and local
government rates as well as commercial rates, such as those on Building Society
deposits. The series were standardized and the second row of the table gives the
values of A,/LA for the first four principal components. The first principal
component, which turned out to be effectively a simple arithmetic average of the
standardized series, accounts for over 83 percent of the total variance and the first
three components account for almost 97 percent. The last seven components
account for less than 2 percent of the total variation.

+ See J. T. Scott, Jr., “Factor Analysis and Regression,” Econometrica, vol. 34, 1966, pp. 552-562;
M. G. Kendall and A. Stuart, op. cit., pp. 306-311; and H. H. Hyman, Modern Factor Analysis,
University of Chicago Press, Chicago, 1960.
+ Adelman and C. T. Morris, “Factor Analysis of the Interrelationship between Social and
Political Variables and Per Capita Gross National Product,” Quarterly Journal of Economics, vol. 79,
1965, pp. 555-578.
542 ECONOMETRIC METHODS

Table A-1 Contribution of principal components to the total variation of


11 interest rates}

Component l 2 3 4

Contribution 0.8368 0.0831 0.0482 0.0156


Cumulative contribution 0.8368 0.9199 0.9681 0.9837

Source: L. D. D. Price and P. Burman, Bank of England.

An important but as yet unresolved question concerns the use of principal


components in conventional econometric regression problems. There appear to be
at least two possibilities that are worth distinguishing. In both we are still
assuming that some Y variable is to be explained in terms of a set of X variables.
In the first problem, however, the number of variables that might possibly be
included in the X matrix on theoretical or other grounds is so large and possibly
so intercorrelated that conventional estimation procedures would be dubious for
lack of degrees of freedom aggravated by multicollinearity. An obvious approach
is then to apply principal component analysis to the X variables to see whether a
small number of components might account for a sufficiently large proportion of
the total variation of the X’s and then to use these components as explanatory
variables in a conventional regression with Y as the dependent variable. Some
discussion of this topic is given in the treatment of 2SLS in Chap. 11. A possible
variant on this approach is to retain a small number of specific important X
variables in the final regression along with principal components determined from
the other X variables. This seems valid and useful, as far as it goes, and if some
economic or social significance can be attached to specific components, so
much |
the better. The second suggested use is more doubtful and requires more examina-
tion than it has yet received.t It concerns the case where multicollinearity
rather
than an excessive number of the X variables is the problem. As is well known,
least-squares estimation of the coefficients of the X variables becomes
very
imprecise. Kendall’s suggestion is to compute the principal components
of the Y
variables, discard those with low eigenvalues, regress Y on the retained principal
components, and transform back from the regression coefficients on
the principal
components to obtain estimates of the coefficients of the X variables.
Suppose, for
example, that there are five X variables and we retain just two principal
compo-
nents,

Zy = Cars Fido) Xo 42 as5\X5

22 = Qj2X) + AnX_ + ++ + Asx,

7 For an illustration see G. B. Pidot, Jr., “A Principal Compon


ents Analysis of the Determinants
of Local Government Fiscal Patterns,” Review of Economics and
Statistics, vol. 51, 1969, pp. 176-188.
£ See M. G. Kendall, 4 Course in Multivariate Analysis, Griffin,
London, 1957, pp. 70-74.
MATHEMATICAL AND STATISTICAL APPENDICES 543

The regression of Y on z,, z, is then


Y= bz) + by25 + e€

7 b,(a,,x, Posie et as)Xs) af by(a,x, Tha ese As)X55) ae

az (b,a), a bya17) x; sc Vea Ae (d,as, ie byas) Xs ake (A-38)


If one retained all five principal components, the coefficients of the x’s in Eq.
(A-38) would be identical with those given by a direct regression of Y on the x’s.
How should we decide on the number of components to retain? Purely subjective
decision on the size of the latent roots, as in Kendall’s illustrative example, is
hardly satisfactory. Should one use the test based on Eq. (A-35) or a conventional
analysis of variance test on the regression? The procedure would give a nonsense
result in the case of perfectly collinear x variables. For example, suppose
Xx, = 2x, and let
pele 2
Rois
The eigenvalues are A, = 5 and A, = 0. ForA, = 5,

7a) | ee ee
ae |e ans
and for A, = 0,

ey Ree
PEE AGE aS
The second principal component does not exist, for

Z,= poe x, =0
eee
since x, = 2x,. However, the first component does exist, for
l 2
a Rear + —x, = 75x,
v5 v5
and so the coefficient of z, in the regression with Y as the dependent variable can
be computed as

= — = =
oy
wat A, (oma

Y=5,z, +e
544 ECONOMETRIC METHODS

and apparently the relative influence of x, and x, has been determined in a


perfectly collinear case where such a determination is impossible. The coefficients
on x, and x, from the principal component regression are seen to reflect simply
the fact that x, = 2x, and are unrelated to the true but unknown parameters.
Nevertheless the question remains whether or not the approach might work
reasonably well inaless than perfectly collinear case.}

7 For a further contribution which shows that the principa


l component approach can be an
improvement over OLS in certain circumstances see B.
T. McCallum, “Artificial Orthogonalization in
Regress ion Analysis,” Review of Economics and Statistics, vol. 52,
1970, pp. 110-113. The basic point
is that the principal component estimators will be biased
but will have smaller variances than the
unbiased OLS estimators. Thus under certain conditions, the
principal components estimators may
have smaller mean-square errors than OLS estimators.
en =

APPENDIX

—_—_--
Oreeooooo
— o—!_—O

STATISTICAL TABLES

B-1 Areas of a Standard Normal Distribution


B-2 Student’s ¢ Distribution
x’ Distribution
B-4 F Distribution
B-5 Durbin-Watson Statistic (Savin-White Tables)
Wallis Statistic for Fourth-Order Autocorrelation
B-7 The Modified Von Neumann Ratio (Press and Brooks)
B-8 Cusum of Squares Test (Brown, Durbin, and Evans)

545
ba ieee
9 we tiie so
ihtapdl , re ee

reat ve iti gid



- ai Pie ‘ 14%) Ae

sf"
ee } ray wr 7

7 , . oy fae ie tr

is
; 1s iy
CA
A i

a+ j
a) es
i

/!
STATISTICAL TABLES 547

Table B-1 Areas of a standard normal distribution

An entry in the table is the proportion


under the entire curve which is between
z =0 and a positive value of z. Areas for
negative values of z are obtained by
symmetry. 0z

i < . _ - joe |os|oe| a Ey ae

0.0 -0000 -0040 -0080 -0120


0.1 -0398 -0438 -0478 -0517
0.2 -0793 -0832 -0871 -0910
0.3 -1179 1217 1255 1293
0.4 -1554 01591 -1628 -1664

0.5 -1915 -1950 1985 -2019


0.6 2257 +2291 +2324 +2357
0.7 -2580 -2611 -2642 -2673
0.8 -2881 +2910 2939 -2967
0.9 3159 -3186 -3212 -3238

1.0 3413 -3438 -3461 -3485


1.1 +3643 23665 -3686 -3708
1.2 +3849 +3869 -3888 -3907
1.3 -4032 -4049 -4066 -4082
1.4 -4192 -4207 4222 -4236

5 -4322 24345 4357 -4370


6 4452 -4463 4474 4484
7 4554 4564 -4573 -4582
8 -4641 -4649 -4656 -4664
9 4713 -4719 -4726 -4732

2.0 4772 -4778 -4783 -4788


2.1 -4821 -4826 -4830 -4834
2.2 -4861 -4864 -4868 4871
2.3 -4893 -4896 -4898 -4901
2.4 -4918 -4920 -4922 -4925

2.5 -4938 -4940 -4941 -4943


2.6 -4953 4955 -4956 4957
7 -4965 -4966 -4967 -4968
2.8 4974 4975 4976 4977
2.9 -4981 4982 -4982 -4983

3.0 -4987 -4987 -4987 -4988

Reprinted from P. G. Hoel, Introduction to Mathematical Statistics, 4th ed.,


New York, Wiley, 1971, by permission of the publishers.
548 ECONOMETRIC METHODS

Table B-2 Student’s ¢ distribution


The first column lists the number of
degrees of freedom (v). The headings of
the other columns give probabilities (P)
for ¢ to exceed the entry value. Use
symmetry for negative ¢ values.

-05

Reprinted from P. B. Hoel, Introduction to Mathematical Statistics,


4th ed., New York, Wiley, 1971, by permission of the publishers.
*wopaal
SZL°9Z

ooo°z<e
L1Z°97

8LS°0¢

607°¢¢
So08°7e

086°77
682°00
7£6°8E
999°LE

Z209°SY
£96°97
picyy

Z268°0S
T61°9¢
980°SI
LLO*<ET
Z18°9T
SLv°8I
060°02

ZL0°SZ 889°LZ

saaibap
£L8°97 TVT°6Z

889°67
8L2°80
10°0 so9°9
O12°6

Jo
BE9
Ive°tl 99917
602°€Z

It
7S0°7Z

L89°¢Ce
6S2°8Z

$66°0¢
££9°6Z

Oz20°S¢
IETS
778°L
899°TT
8Be°el 89T°8T
6L9°61 £7O°9S
896°8¢
6S9°LE 99S9S8°77
OvI°7y
£6990

Jaquunu
SL9°61 819°2Z

yo
Z17°S L¢8°6 17
ZZ9°91 T9T°IZ
£¢0°ST OLZ°0” 617°St
796°LY

962°9Z
1Z 970°

ie OW
166°S OLO°TTL90°7T
L0S°SI CS9°LE Leeiy
LS9°20

ayj
C9IE*7Z
S$89°¢Z
966°7Z

L8S°LZ

pyt°Os
698°8Z

“O-D ‘OUy
178°¢ ST8°L
8817°6 69°71 616°91
LO<"81

U st
1L9°Z¢
CLI“SE

URTV] SUTYsTqng
NZ6°CE
SsT7°9¢
see"se
e107 ellen

aJaym
S79°O1
LIO°ZI

L86°ST
789°
162°9

‘aouelseA
SLO°LIZ18°61
697S°8T
790°1Z 69L°0C S19°6ZLO0°Z¢78S" L80°6¢
90L°Z

9£2°6
s09°)

686°S 6LL°L

O<O" 79S"

Ee
Z02°71 1

LO¢*2Z
CUS°ST686°SZ
702°LZ
Z17°8Z <18°0¢961°SE£ISSE
TvL°g9e
9I6°LE9S2°07
T

c7°¢
c79°7
1 279°

IL1°9Z sL9°0¢Z16°ZE
Lz0°v<

UQIM y1UN
S86°91
Z18°ST

$917°0Z

006°¢Z
8<0°SZ
IsT’st
612°¢

682°L

£08°6
8SS°8

6217°82
TO0<°Lz
TT

£SS°62S6L°Ie 6£1°SE
osz°9¢

a3eIAap
T¢e9°t

TI<*61

S19
IZ
O9L°72

“YIOX
YLO°I 790°9 £82°8 9S59°01 668°Z1
Tee°k 61ST cZe°LI
c22°91
T10°vT 8I7°st 689°1Z

ogs"e¢

JeuJOU
s99°¢
807°2 8L8°7 02S°618Z°IT T1S°6I
109°0ZSLL°7Z

MON
8S8°¢Z

810°9Z
6£6°7Z

ZLI°8Z

612°O¢
960°LZ

907°6Z

19%°Z<
l6c"l¢
6£2°CT
6£o°VT

B<e"8T
Le¢"6l
Ove*Z

Beers

98¢°1 £0o°8 Lee*nz9£E°9T


9ILE°BT
9EE°LT

Yip] “po
pasnse
8ee°9l
8eerLT
17°01

ssv°0992°C
Looe
Icey
B1E"S
900°9
perl 70S°6
Ove

YoADasay ‘siay4Os4
L¢ee°0z
LESES
7£0°6 TT

LEerlz
Leese9EE°ST 9EE°6Z
99Z°91
ZS°ST
1Z8°OT

Aewaq
926°6

7T Ory"

nen"! o00°¢
£1ZL°0 £98°0Z Ly9°<Z
£6°61 LLS°0Z
nc9°Z1
Tesrel
IZL°11

871°O S61°Z 8ze"e LeS°S


19°) £66°9
LICL 871°8 C8I°LI1Z0°61 Z6L°1Z
TOI°s! 61L°2Z 80S°SZ
uz/-
- |
88S°1Z

n9E"¢T
Z00°Z1
L0<°01
est°ll

076°81
LS8°Z1

0Z8°6I

SL7°2Z
s77°st
PIS°9l
85°71

290°81
FSET
9£9°8

7XZ/\
uolssaidxe
$00"|
679°1

OLore

O8<°S
6L1°9

L907°6

¢<0L°02
L8°LI

686°9
2£08°L
[0I11SNDISspoyja=
4of
£9o°7

cZ8°¢
21790°0
9007°0

"69°07

$80°01
s98°01

e772
819°S

Z70°L

8s10°0
O19"
790°1 89I°y OVZ°<T878°7T
170°71 £Ly°9T
6S9°ST VIT°SI
Z6Z°L1 89L°6I
6£6°81
L0S°8
O6L°L
70s°9

ZIL"6

1s9°TT

06°0 78S°0
1120 £oB°e
702°Z 060"¢ S98"
66S°0Z
ueY] ‘O¢ ay3
8ee°Z1

826°91

67°81


£6¢00°0
tI 119°

ZS¢°0 LIZ GCSE


se9"l 922°S T9Z°L LITOl
878°<1

6LE°ST

QLS°ST 802°LT

“IaYSi
ISst°9T
16S°TT

160°<T

Wopaady Jaqea1H

S6°0 Sv"
<01°O 1120 eel°e Ov6"¢ SLO? Z68°S cL9°8
1L49°9 C96°L 06£°6 1S8°OT
009°01

90¢°9T
266°1T

ScI°vl
69°21
£62°1T

600°<
S16°6

cSL°0 799°
MET"T BLT" 892°S SSo-L
906°L
Wor “WY
98°

86°0 829000°0
¢8t°0
7070°0620°0 CES°Z
c£0°7 6S0°¢ 609°¢ S9L°Y S86°S
v19°9 L9S°8
LE2°6
1

poyuudey

£s1000°0
saaibap
092°8
Z18°S
622°S
6£2°1

yo

SII°O 979°T
880°2 961°OI
20S°6 029° 6L8°Z1
9S8°01 952°
S9S°¢T
807°9

CBOE:
<s0°¢

STOL
LOI”
099°
IZs°¢
wopaady

TT
7SS°0
10Z0°0462°0 2L8°0 8S9°2 L68°8 1
861°ZI £S6°71
saaibaq
4 JO

ANN
TN OM ONO
So 92

-X
uoHNgLysip
qe
¢-g jo Iz cz SZ 92 Le 82 62 O¢

a gd
148,p-¢7 uONNgLysIp

550
¢ yUI0I0d URWIOY)
(adA} pue
| yusoIadore) (ad) syutod
Joy ay} UoRNGMIST
JO 4 p

saaibaq
jo
wopaaiy
JO} saaibag
jo wopaaiy
10} s0ye1awnu
(17)
JoJeulwousp ae
Fe ae +
(am) I é € 0 S 9 L 8 6 ol Il | PA val 91 0z 02 o¢ Ov os SL oor 002 00s oo

I 191 002 912 S2Z Oe nEZ Lez 6272 19Z CZ £92 902 SZ 90Z 802 6nZ= OSZ 1SZ 752 £SZ £62 SZ
ZS0P 666% E€0PS SZ9S 0SZ 967
POLS 6585 Bz6S T86S zZz09 9509 72809 90T9 ZbI9 6919 80€9 >PEZI BSz9 9879 ZOE9 EZE9 HEED ZSE9 I9E9 99€9
z 15°81 OO'6I YI°6I SZ°6I OF6I <°6I 95°61 LEI BEI 6S°6I Ov6l I7'6I ZH'6Il £761 7°61 sv6l 9n6I Ly6l LH’6I BEI
6P°86 TO°66 LT°66 SZ°66 6H°6I 6Y°EI OS°6I OS*61
OF66 CECE PEEE 9E°66 BE°66 OF°66 Tb*66 7Zh°66 Eb°66 bb°66 SV°66 96°66 LP*6E6 8F°66 8h°66 6F°66 6h°66 6°66 OS°66 OS*66
€ <I°OI SS 826 ZI°6 [0° 76°8 88°8 48° I8°8 8L°8 9/L°8 %2°8 IZL°8 698 99°8 99°8 z9°g 09°g 86°8 1S6°8
ZI*PE TB*OE 9F°6Z TL°8Z 95°8 S°8 75°8 £S°8
Prez TELE LO*LZ 6P°LE HELE ES*LZ ET°LE SOLE 26°9Z EB°9% 69°9c O9°9c 05°92 TH*9t OE9Z Lz*9% EL°9Z BI-9Z PI-9Z cT°9f
Y TY SE) SRE) (SEE) CYA SII) YSU VED aR) SYS SARIS YSIS TEKS) KS (UES (HESS isc Tig.
Oz*Iz OO°8T SOLS) 99° 599°C S975 THIS, 9S
69°9T 86°ST ZS°ST T2°ST 86°PT O8*PT 99°PT SPT SH°PT LET H2*PT ST°PT ZOOL E6°ET E€8°ET PLTET 69°ET I9°ET LS*ET ZS*ET BP°ET OP°ET
S 19°9 6/6 Ih¢ 6IS GOIS G6h 88°) Zev 8Li7 Wis OLth 89°h 49° 09°F 95°7 £67 OS°7 97°%
9z°9T LZ*ET 90°ZT GE*IT thr Zhh On BSD LED 92°77
LOOT L9°OT SP°OT Lz*OT SOTOT:«<ST*
= §=—96*E
OT «6686 §=LL°E 6896 =6SS°6 47°6 «=CLT°G6
«CETTE)—Z*EC*E
=6—LO"E sEG
«=bOE cO°6
9 §=66°S DI°S OLY Sth 6ER 82h 127 Sih Ol” 90°7 <0°7 HOt
| 968G) Z6rGe BIG. VOI LIEGE
«LET ZE*OT
= =—BL*6 KG PAE SU RS GES MSS
ST*E) =6SL°B Lee 92°38 «COTS «6B86L LBL 6L°L ZL*L O9L CSL 6&L TEL SPT*LSCE*L
GOL ZO*L 669. 6°9 06°9 88°9
L §65°S Dl’? Geer HnZL oC CumO°mnoC OGL
Sie Com
9: ESOS O91 EL Re GIS Tae 1RS MUEIS WSIS
Sz7T $5°6 SPB TOG BGT BCs Gren CAS CISION
SB*L 9F°L 61°L O00°L B8°9 «z9r9ss—TZ"
«PS*9
9 LBRO SE°9: LZ*9 ST*9 L0°9 86°S OBS SEs 82:5) SLES O£5 £95 SHS
8 ZE°S
= «98° LO°7? HBS 69°C BSS OSE HE BES MEE ISS B8ce repeats
| (aPAYe AGI ASS RS US
92°IT $9°8 65°L TOL <Ore OOS 862 96°% 62 £6°C
£999 LE*9 619 €0°9 T6e°S zB =—L9°S)—BL*
«69G*S
S «BRS OES =82°S Of°S IT'S 90°S 00°S 96% T6% 88> 98>
6 =ZI°S 92°7 9RtG -<o"G ahie. EG cccs
6 ECC less
G See Shi OLS LOS ZO°S 862 S6°% O62 «(98°
95°0T 20°8 CBZ Q*Z,
en 127 maf OSC CAROL. CLe MOS
66°99 Zr9 90°9 08S Z9°S LS SES 92°5 Bs ITS O00°S ZE*P O08 ELD vS> 96°p Ise Shr Tre OF €€b T&>
or 96°47 Oh =f Bye 6S 2226) ISG. LOLS COLwe G:CNL WG:C CeNG DErCM ZOtz. LEZ) aUlsc. NOLS
pOrOT 95°2 LDC, WIC 1927 EGG Ce MIGCemGGeC Ro
S5°9 66°S P9°S 6ES TZS 90°S S6*b S8b BLE TLR OOF ZS Tre E€&b Sob 4Tb ZIP Sor TOP 96°F EGE T6°E
09° os*Z 12°2 L8°7 26°LS°Z 88°T 78°1 18°TIES Toye: IL IZT
0n°Z 9E°E oie £12oo°e L0°Z 10°ZSLE 96°1S97 67°e cre 8L°I Clare StayTez £re
ERE

BEE

c2°Z

COPE
1¢°Z

80°Z

LLS
17°Z

ere

912

68°72

Z0°7
7L°7 92°7 I~ oz 66°1 Ly°T 78°T 6L°TcEZ 9L"T oL*T
18°T
c7°7
E99° THe T&E€ 90°E 26°7 70°Z0a°z Olt S6°1coz 16°oS°7 L8°1 cre LEZ Lez EEE

LEZ

08*T
EES
LOE EERE, LO°Z 06°1

Ly°z

98°T

z8°T
cre
S77OLE So°z90°F 90°72 612 VeLE°7 9L°7 86°1897 76°09°z
98°7 ZO°Z ESS 23° |(YES62°Z
9¢°7 82°2Of°E SsV°Z 68°7 TL°Z OSC TS*Z THz o8*T
LZOLE 6r°E 122bre oo°e 60°Z 90°Z6L°2 00°Z 96x£9°~ z6"1 68°1 "19n°C 78°1 z8*19E°% CES
0n7°Z 22°Z 92°Z 81°2 <1°Z 80°Z 70°Z8L°7 o0°zOL°e 9671 £6°18S°~ 16°I€S°e bre
os°2ose aoe. Lone T2°€ LO°E 96°7 98°7 ESS 88°TBre 98°1 78°1one
n2°7CHE LOS 12°Zere 91°72
TO°E 11°Z76°7 LO0°Z Z0°Z 96°1 161€S°Z 68°1
£S°798°E 7°72T9°E 92°E E€8°Z 9L°7 667169°C €9°% £6°18S°e 6r°z L8°1Sz
OL°E 1<°Z sv% 12T6°Z o8°~ 70°Z L9H 76°1
L9°2v6°E 97°7 8<°ZTS°E bee S220z°E 02°2ore oore L0°2 LL°Z 00°ZcL°7 86° 9671c9°z 8S°e z6"1SZ
cO°o

BL°E

62°7

92°7

80°€

0o°€

26°S
19°%

S0°Z
80°Z

coe
SZ
os°2

6S°E
Z7°7

6E°S

8T°E
SoZ
ERE

98°C

08°z

99°72
612

£0°Z

8671
00°2
OL°e

96°1
SLE
IZ

98°E SOE 00°e 88°~ €a°z 70°ZBL°Z bL°e


A
ore 95° 97°7HOES, 6£°ZTSE £0°79E°E 82° £2°2
ETE
CH 61°ZLO°E SZ ZZ76°7 60°Z L0°Z Z0°Z 00°zOL°z
86°F

Bbr°e
77°Z

LOE
6£°2Z

ere
19°2

SO°E

66°C
S12

£12

Ora
125%

90°2
Ta°z
Te*p
OL°Z

09°Z

BL°E

cOrE

£o°7
LEE

62°7

6re
S27

8I°Z

b6°~

68°7

60°2
SB°~
HL°Z OL°E LETS SOC 0z°2 ZO°E L6°7 11°Z
62°D 99°ZSO°D S9°ZSBE 87°72 £7°Z9S°E SHE £o°7SEE 62°ZLEE 9C°%6T°E ETE LO°E 81°Z VIZ £V°Z£6°% 68°7
6L°Z 09°2 oa°e 87°Z 7E°7 OE cre 02°72 IVS
Opp 69°7
booT 96°F £S°Z LYE cv°zSS°E 82°ZSh°E LEE Lez 82°7EZ°E S22LTE £2°7 LOPE 8I°ZEO°E 66°S
ELSE
cao

s7°Z

60°E
O&E

SO°E
2°72

02°2
IEC
172
9b°>D
Z8°Z

cO°n
ZL°Z

£9°7

99°7
98°E

cS°e

7¢°7
1o°Z

oo°E
82°7
THE

ore
LES

92°7
GEE

T<°z

Ble

ore
Of>

orp
96°7

9L°%

09°2

69°E
S92

67°7

S77

TS°€

8¢°~

92°E
SEZ

ones
os°z

92°C

72°CETE
oS°d

D6°E
L9°%

oa°e

6S°E

17°Z

ERE

LEE

TEE
ofS

Toe
82°Z

6&°D 60D 68°E Oore cS°€ ope OEE Os*zSOE 82°C


06°Z£9 08*Z cL*Z S92£0°D 69°Z 99°C8L°E 0s°Z89°E 97°Z £77 On°ZSRE Le“? SezGEE CES T2°e
Os*b
bL’D

S8°Z
S62

83°D SoD 82°D OL°z


oT’o £0°D SB°E LLE S77
1o°¢ 26°C 78°Zbb’o EIS 99°T c9°ZE6°E 89°Z SS°Z cS°ZTLE 67°7CES LY°C
6S°E bS°E £7°Zos°é 17°Z9n°E

CES
O2z"¢ 69°D 06°2
96°C 18°ZDED 99S66°E 99°T
bere 29°Z06°E 09°Z
I<90°S c0"¢98°D 9S°b S8°Zbro LECSZ°D VL°Z40°? WEOr’e 89°Cb0°D 98°E
9°¢ NERS 8I°e 8S°D Os*> 78°C Z8°Z 9L°7
LES Tres 0z°s Ile£0°S 90°¢68°R To"eLLP 96°7L9°b £6°S 06°2 2°72ern LED Tee 08°Z
90° 8L°Zcoe 8T°o
6S°¢f2°9 67°¢S6°S bL°S 62°¢crs 62°S O2*¢ 9I°<60°S c8°p To"<cL°p
Ie nee9S°S 92° 8r°s cI¢TO°S ors6°? LO*¢L8°R sO°¢ <0°¢9L°9 66°289°F
Of*Z 08*< Ci9, To°9 Sas 8L°S CLS

86°¢ €6°9
88°< 0L°9 yL'¢qTs°9 89°¢9E°9 £9°€EZ59) 69°¢ SS°E cs"¢E6°S 60°¢ Lv¢ ne cre99-9 onsoS BereLS°S:
08°17 €E°6 40°6 98°32 0S" €s°8 ov? 17782°38 88°Z 92°17

$9°6 sly L9°7 09°) 89°8 67°17 sty 8T°8 sonors conzo°s os"bErL 82°07
8<°7 L£2°2
ca°L 92°07
él
II

<1 aI SI 91 LI 6l 02
81

IZ ce £2 ne ‘Go

551
148.p-@ (panunuoD)

LT
2a)w
¢ JUDdI0d UeWIOY)
(adAj pue
| yusdIadoye!) (adAy syutod
Joy 94} UORNQUIsIp
JO 4

saaibaq
jo
wopaad}
JO} saasb6aq
yo Wopaady
10} s0je1ausnNU
(Ia)
Jojeulwousp
(22) I Z £ ” S 9 i ie 6 I él val gT Oz 92 0c On Os SL O01 002

8
00s oo

Or
2@ mmecosy eco CueL
GR aCe 6G 2e CML
CSi LZOZ
0 <ZZs20 BI5Zs Iscn eG sce OGulCSGu
mm SOs]limeO6ul
en COnt aun
ulOL
RGO;CMOl Nanos
cL Olsen CPs
ELIZ «ESS OH PTR BE =6S°E ChE LTE 60°F COE 96°C «98+ «772% 997% 8S*S. OST THT. GEST. Boe,

c£°7
Soe, GIS, SIC) oie,

62°E
LE Cat,
CGOG SSC LSC NEC DSC Ose ES cdOG Sher Cama ipa kd TU Tesi) GERI ah TIRKIC Cyl Mie WEN GERI
B9°Z 6RS O9°F TI*R 6L°E 9S°E 6E°E PTE 90°F 86°F £6°C EGC PL*Z E997 GG8Ca LACE BESTE EELS SCHC Caceat DEC CTsCeMN OTRO

O¢°292°E
- 8c OGSY: ComSS SOcCun LC IGS NUNC FELCH 72°E 612 Gime vAGe chips
= aia SYSit Rit [GP TGR GYR) Gif iil GEA IE SEA
8S9*Z SPSS LS*P LO°P) 9ZIE) VESTER DEE ITE EO°9E S6S 06° BT Lr O9% EE HHS SEZ OFS fe BT

62°72
ETS 60°F 90°C

ECE
6z OT" OEE CEBTZ «(OLS 69S £72 Sot Cee Ce
lex
8 CeWN O1eZ <O%z. 0072) svGel [Link] SOrI OSs at
/ eh Eg alee eeena eae
09°L 25°S PS*R POR ELE OS*E EEE BOE S—C«OO"EE-s« ZO"ESC LL°ELB*E
89° LS: «6FT THT TET LEZ 61% SIZ, ONC. IO:Ce- ECS

82°7OE
o< TAY) RASS (ASA CRG SESS LAI UESCE crc cc,we Chad MOO DOC G6| -<6al) 6Bcl) Vealh
ee nGlall Ola Coen Gone eee ale Cee
9672= (6ESS TS¢h. CORE OL{Em LELEN ORES 90°E 86°F OB PE? PLT 99° SEZ LHT BET OC HES OTF ETS LOS EOS TOS

LCseLEE:
ce ST OSS «(COGS CLO IG (ORS CSC lec econame OleC Oram COpCm
ee EGal NSAI SiR rah CNR TH (ERE TER UERIE LER TUN
OS*Z PESGS OPP ZELEN (IFES CHER, SCIEN TOE FEZ 98°F 08° OL zZerZ Ist Ze PET Sez OFZ ZT BOS ZO"? 86°T 96°T

LALAere
nE Iv BE BB S97 G6H% 88 Os Lee Clete ZeeOO: OOsCGOrCe
Gu Gaal Gslen st 0Gst leery
Salma Oem Sela GSalime ES
PPTL 6F°S CH*H EGE TOE BEE To*E LENT 6B CBS QL 99°C BS% LHe BEC

£2°7
OF ez STZ BO°X OT 86°T POT T6T
9¢ Tce 9ZIE 9852 $952 BN SESS 82o ec cme :Cum90;C
oO OUilame
aeOl 6a Gel Oct G/alan Clale 69u SIs CIN OS Sale Seale
GEL StS BER 68°F BSE SEE 8T°E ENT OBE BLT «CLT CHS OSS ER SET

80°e 12°72
9% LIZ 2I% POTS O07 ET O6T 48°T
8¢ Oley GCse SBC CIC CONC EGS DCC HI'Z 60°% <O°% 20% 961 C61 S81 GBI CYA (VR Cie WS SURE VASA SRLS
SEZ FS PER 98°F; PSE CEE STE
(SEN
652 CBS SEZ 69°C) Ce NGS ESSER

bO°E 61%
OF zET eee PIe BOS 00% LET O6T 98T Pe°T

core
Ov BD CEZ"E HBTZ
«OO OI9Z «SHS SS S22 ZI*z LOZ ¥0°2Z 00:2: S61 0671) VOT GET “HLL. 69°T (99 19%L) 65:1. SS eGo
“TESL STS) TESDY PESSE TSIE. 6CLES CIFES
Rast
(8822 OBS ELST) 995ZE 9S:Z) 6RiC) LEZ bez ore IT? Sot LOT PET B88T PST T8°T

81°66°C
cy CLO OCZZE BZ 6SZ (MS 28S "oo II°Z 90% ZO°Z 66° él 68°l ZBI BLT Sil 89° 7921) O9Te LST Poel
ZEAL STSS] 6CSP OBER (6526= SC°ES ODSES* TER Cita
Wiha MACE EE ATE ME SEZ 927 LIZ Bor Zor POT Té6T SBT O8T SLT
ny «(90H «CIZE «7B Bc°Z EHS 4I1€% <o2 ~=—CoT°*z «<0 «1O"%. 86" Z6zl 885i 18 ILA /et
zc 9977 S977 OG 9Sct leecoo beerOSt
PEL CET*S OZ*H
= «BLE =—ORE =—PE°E LOE Syn
«GL*ZOvB*T
BOT EIT CSS bee

LUZ96°C Dee
CES peer SI*z 90% 00% 726I 88T Z8°T 8£°T SLT

b6°S
90 SsOr OE 18% LOZ cyt Of coe OZ «=00TZ_—SsswOTZ_~=s
«=L6I:«G «6=T6"I =L8°T (O8l SZ°T) TILT Corl COTE Ge Ssheee
TzZ OTS PER OLE PRE 22°E SO°E a Siete
Pee
COS
| VELS. IPT ORC OSC PZ
912
OFZ eer ETZ PO'Z B86°T O6T 98T O8°T 9LT CLT

f6°7
87 “HORY, 6IeS. (ORFs ‘9SiZe ined)a OSC Se VIC 8024 SOi%s 665) PSIG GRO6 FBT GL VEST OLS VOSS SEIT SEISI SSN OSes WEVA GOA
ETL 8802S ©2S°R FLOE) ERE 0CSE. PO°E ((06°S; OBS, TLE «PSE A8S*S BRE OWE BE OC Tics. “COC 96T A8EE PET S8L5T* TELE OLT-

Os Ony, Olson eG/aCer OSC© CamOlcG mcr OCC. elie LO eCOrce —“86ply meSGrll meOGnl| SOrl aie
8 Lala
7 Osan Ose OOetm MSGR SEN MOC
Oy OVA im
ZTSZ -90%S OSB. UEZLSE. TEE’ BITE COPE BEE)
6 «BLE «COL «68COE SE OG OPS GES OES 8T5o) (OTSe -OO2S PET 9ST (CET OLE TET OST.

SS vcOny, /lecee Bloce SSIZe OlcCeecCumascc,


Nee “SOLer (OD:2y F6rl= Reon RBS Tee
sesc8 al emo Ceol L9sT Sele sel meme ColBa aSeo Dla
Ted, mc TOCG GISa) (BS2E ELESE nSICe CaBb S8iSm SLC "99°C" 6S:05 C) SMES REEDS GEC ECS VSI 90°C '9650) LOOT» amcor OL OFLA
TIM

09 ‘SUGEOOat,
“Isc VESrCw GOCE, LC Oi VOR 66ylee S6ule ec6rl)”
= 9Beee Bal Lelie:mS OLae GI GSolen OGateee OSme Oi7 alieny yale 6Fnv
BOLLar (86:7en CISD =SOtE PEE SacIne SCS COC CLS ESSN (9GTE GOG:S ONC CEC. [Link] CILES EO;CN 86.0) OE LE CLT TOLER OIA NESE OFT,

s9 uen6Gre a sy -Glacem Stee CmmvCrcmmoecc


Gli COrGOncue
me (Son avGmle GOnemObul!
se
OG ee
« merc Oa em COnPe eae Sr)
mn Yee al ec Sa LAG
POTZ SE5P) OTP E9°E.
© TEE| 6O°E EGE OL. OL° T9°S, PSS%, LPS NOES. LES BIEN (60°C 00%. -O6;E GPET S9LST, eTL2D PIT) O90 OST)

OL OG.S Clee aULic OSG. MESCee ConG Com iy LOccemme (Orc <L6ulen -SOnle mvanlarGGnl.
laGL Ziel Sule Cosme OG eGo meU eeYl eG maSoa
LO-La o6<ha 80th ~O9Sfs WGCee, LOEa 160 LLG LOC 65:0 TS: aese.e. Secwe eee SEC LOC lanS86 SE
| =eCOue Lal mee TODD OSESESSE

08 9656 Uileeae CEC Bye SSe crc Ce cI SOle) 66ull S6ule N6aly eSSel COnlae VL Oa Soot JOSTes 47Set Gal) Svolaeem Sole Sa CONS
—~96°9) BaF - Oth 95°F §SELE (POLE LBZ PLS, HOF SSS- BPS THCA CES vere? TEE E€O°% FET, F8°T 82° VOLT SST LSE) CSE 6U-T

ool vbr GU; OLce Ship OlicerGliGanOSc,


es SOc L6n Gn Bale SBallee GL GLo 898Ts SOC SHI Su Oval macula 6S Sete OS Scale
0629) (SE°r 8656. TSG OCcE, 66S. (CBS 691% OSS, TSE) Ete NOESS 9C:S (OTS: 90°C. 8671 Tes (682 6L5T)© ELIE VOT 6ST ISS OVE ET

scl Uicuicone
ame O9:cwo oie mGGeCu Ll mO0\Cmm Occ Sonlamel OGrlen Sue cOells
em ol S9dleurcaleane
aanOS GSelen Mi OV MGV Ore) sl Cal Cale GCL
88s9) BLP B68 (LESEP LICE SG6e, GL SIC OSC) “(L92Cm (OPIS SEEESC: ECLS: c AST COC P6rTs Set Ta (SL OS OSit SEB be OV OVE LET

OsI eGse IOS /9°Cmme Suscu icc’ lca O0LCaLO;CamO


V6ulen Goran GO7le
= almecOnl
eee Ll Onl GSidlume Salen EVA eel) cal Cole Cala Colo
S Coe
Lol) Glebeon LGsG (Pi-Ga| EGIL CmECOLC
OL CM CO ESrCue CmMTP LEC mmOGC. mm0CsC Cc CaacCl OO;an LGul eCGrhmn
enal NCL OSs GOS otS sme eee LE CET

002 Gum60sc VO; GOecm “lizen megcrcn. O6ulmGOccmmilcc


LOnenConlen
a SOulem eeOSEl sew OmteGo
wc Coo Coulee <Gvile- lmeSeolemecwel
oo Calee eo CCal Glisten
9229 TLZeR BBSEL «COTE TEE 06S ELS 09°F 0S°S, THE FEST 8ese LTT, (60°S L650 8858 GLUT. 69°T) COLE ESE SVT 688 CEL SCT

00” IESG SCOree 829ICR \GSncu Send: nClecme Cu 96ulabsSO;


OGrlee SOaleee Bhl e Ola lemC/oan ON 7SplenOIs
Nalin -Zral=s OS mac olalecccaenOCol
Sale
«OLD §=099°R «=CEBTE «=COETE GOCE GB*Z «= OZ S«GG*Z HZ OE*ZCLECT—_
SC ETZCCEZ*Z
«OOP «66°T:) «6Pet 3«(LE OT LS°T LHT CHT ZET per 6TT

ooo! SO OU; 95cm OSiCue


se OlacueCore,
CanCO; GOtleSosle
(7O2le TOSdia mmole emOso Sol Salae mera cGo( Crime aiae etySal ano eOSr Cl9 Glee Aisle SOT
9929 T9°F O8%E. PELE POLE. CSS 99S, ESS ERE FES, OFT 0S, 60°C OS \G82be)
| T82h TLES “TST PST, errr, Sota 8C-h Es GT CEE

NOLS OGcCumpm O9cC Lorem menliCaC rc O;CsenOU


al Gulan iueet/
CBr lee
Oz eas GLa /al mS 9s mmG 7971en CSalemeGrl
dlen
D7 Olsens
Ww cule mSaleC Calo ley
el iisieee COTE
-99°9 (09°F) BZ°E CELE CO°E 0BSC| 89S TSS. TPS CES Fe*Z_ 8 20°C {6651 2870 6Z°T
— 69°T| 6S°T CSs0 eet SET SCT SST 00-E
©

“H

Aq
Aq
AQ

“MA
YT

pue
0861

wWoIy
PMO]
a1P1S

We

IBIOaH
WUIAag

"WMO]

553
spoyjay
‘UONIPA

‘Ssoig
‘SoUTY
JOIIpsug
‘ueIyIOD

poiuUdsay
jpI11S1NHIg

UOISstUIad

AVISIOATUA)
JGeT.C-g A-UIGING
UOS}E INSHL}IS M-UIABS)
9914 (so]qu} -UTGING
A uOs}e :onsneis
| yuso10d souvoyruais
sjutod
jo Tp pue 2fP

1=4 c=

554
oA =A S=At 91 l= 8=1 6=4
u Tp Np Tp p Tp Np Tp Mp Tp Np Tp Np Tp Op Tp p Tp Op

O1=.4
9 O6¢"0 ZvVI°l AS See re Oe eae es ——— a oe aaa aon =e ad ---- “see
L S£v"0 9<0°I 762°0 9L9°T eee ee ——— ee ee a ere oe —— —_—_
8 L6t°0 <00°I Sve"O 687°l 622°0 ZOI°Z ere a oe oe anne —<<= ee —_-
6 SS°0 866°0 807°O 68¢°1 622°0 SL8°l ¢BI°O se7°Z ee = cso cacao ee anos a oor
ol 09°0 100°! -—-
990°0 ece"l OvE'O <cL°l O€Z°0 €61°Z2 OSI°O 069°2 aea oe te,
i Abou© cana s =k)
Il ¢£s9°0 OIo’l 61S5°0 L6Z°1_ 96¢°0 079°T 98Z2°0 O<0°Z <61°O £s7°z ~Z1°0 268°2 n= Ps, Sree ee. aie
Zl 1L69°0 £20"! 695°0 HLZ°1 6707°0 SLS°Il 6<<°O €16°l 7Z°0 082°Z 79I°0 S99°2 SOI°O €S0°€ ees! pn ee ee
<I 8£Z°0 8<0°l 919°0 19z°T 667°0 92S°1 16¢°O 978°I 762°0 OSI°Z 11Z°0 067°2 OVID 8<8°Z 060°0 Z8I°< oeoe eae
val 9LL°O 7sO0°l 099°0 #S2°1 LS°0 064° I%77°0 LSL°1 £€H<°0 60”0°2 LSZ°0 9So°% ¢8I°0 L99°% ZZ1°0 186°Z 820°0 L82°¢
SI =118°0 oLo’t O0ZL°O
+ ZS2°T 165°0 797°T 88°0 7OL°l 16¢°0 L£96°1T ~<0E"0 97Z°% 92Z2°0 0<s°Z I91°O L18°Z LOO Tore 890°0 ples
91 78°0 980°! CSCHLCLSO£€£9°0 9n7°T Z€S°0 £991 L¢é"0 006° 675°O £SI*Z 692°0 917°Z 002°0 189°Z ZHI"O 9H6°Z %60°0 10z*€
<I ~=%728°0 coll ZLL°0 SSz°l Z29°0 zev"l 7LS°0 O<9°T O8t°0 Lv8°I <£6£°0 820°2 <€1¢°0 61<°2 12°0 99S°Z 6Z1°O T18°Z LZ1°0 <s0°e
81 206°0 8II°l $08°0 6SZ2°1 ~=80L°0 ZZ7°I ¢€19°0 709° Z2Zs°0 <08°T S¢v°O S10°2 ~SS¢°0 8E72°Z 78Z°0 L9”°Z 91Z°0 L69°Z O91°0 $26°Z
61 826°0 ZEIT S¢8°0 s9z°1l Z72°0 SIv*l 0S9°0 78S°T 1995°0 L9L°T 9Ln°0 £96°1 96¢°0 691°Z 72Z€°0 1g<°Z $SZ°0 L6S°Z 961°0 18°C
02 2S66°0 LYI°l £98°0 TZZ2*1 ¢Z2°0 Tiv*t $89°0 L9S°1 86S°0 LEL*l SIs°O 816°l 9¢°0 OII°Z 79¢°0 g0¢°Z %62°0 OIS°2 2€Z°0 HIL°Z
1Z SZ6°0 T9T*T 068°0 LLZ°T <£08°0 807°T 8IZL°O 7SS°I €¢9°O
+ ZILT ZSS°0 188° 7L7°0 6S0°2 O00°0 90Z°Z 1€€°O ven"z 89Z°0 SZ9°Z
2 L66°0 HLT 16°0 "8Z°1 1£€8°0 LOv°l 8Z°0 £9S°T 299°0 169°I L8S°0 678° OIS°0 S10°2 LEv"O 88I°Z 89¢°0 L9€°Z 70E"0 87S°2
£2 8Iovl Z8I°l 8¢6°0 162°T 8sS8°0 LOv°T LLL°O nes°t 869°0 £L9°T 0z9°O 1Z8°l $S°0 LL6"1 ¢Lv°O OVI°Z 07°D g0¢°Z OVE'O 6L7°Z
02 LEO 66I1°T 096°0 862°I Z88°0 LOv*l S08°0 8ZS°I 8Z2°0 8S9°1T 7S9°0 L6L°T 8LS°0 976°1 L0S°0 460°Z 6£7°0 SS2°Z SL¢°0 LI”°Z
SZ S$S0°T T1121 186°0 sO¢e*T 906°0 607°T 1¢€8°0 €ZS°I_ 9SL°0 S79°T 789°0 99L°T O19°0 S16°I 0”S°0 6S0°2 ¢Lv°0 602°Z 60°0
+ Z9€°Z
92 COCMNCLOM
100° 21g" 826°0 TI?"l SS8°0 8IS°T €8L°0 scl TIL°O 6SL°1 079°0 688°I ZLS°0 970°Z S0S°0 89T°Z I%7"0 e1¢°Z
Le 680°1 €e2°1 610°I 61g*l 676°0 <17°l 828°0 SIS*I 808°0 979°1 BELO €vL°l 699°0 L98°T Z09°0 L£66°1 9€S5°0 TE1°Z €Ln°O 69772
82 vOr'l
= nZ°1 e207 S207 696°0 SI?°T 006°0 €1S*t z<8°O 819°T 79L°0 6ZL°1 969°0 L”8°T 0£9°0 OL6"1 995°0 860°Z +0S°0 622°Z
62 6II°T "S21 S01 cee"l 886°0 817°l 126°0 ZIS°I $S8°0 TI9°T 88Z°0 8IZ°T ¢ZL°0 O<e"l 8S59°0 L”6°1 S6S°0 890°Z £€S°0
O¢ tial Culses
SO OLO'T 6€<°1 900°
£61°2
TZ? 176°0 IIS*T LL8°0 909°T ZI8°0 LOL*I 87Z°0 VI8°T 789°0 Sz6°I_ 72Z9°0 170°Z 29S°0 O91°Z
l¢ Lytl €LZ2°1 S80 Stel ¢Z0°1 Scv"l 096°0 OIS*T 268°0 109°T 7¢8°0 869°Il ZZL°0 aos"! OIZL°O 906°1 69°0 L10°Z
ce ~O9I"T c8Z°1 OOI'T 2Se°T
685°0 T<1°Z
Ov0°T 8Z7°T 626°0 OIS*I LI6°0 L6S°T 958°0 069°I 76Z°0 88Z°I 7EL°0 688°I %719°0 S66°I S19°0 HOI°Z
ce ZLI°I 162° PITT 8ST SSO°T ZE°T 966°0 OIS"T 9£€6°0 76S°1 928°0 £89°l LetO PELE LSL°O HL8°1 869°0 SL6"Il
ne PBI" 662° BZI'T 9S" OLO'I
1179°0 080°2
sev" ZIO°I TIS*] 756°0 16S°I 968°0 LL9°l LEB°O 99L*l 6ZL°0
Se 098°1 ZZL°0 LS6°I S99°0 Ls0°Z
lal LOSSES OVI'T OLE*T S80°l 6e7°1l 8ZO'l ZIS*T 126°0 68S°T 716°0 1Z9°Il LS8°0 LSL°T 008°0 Z£?8°I 77L°O Ov6°T 689°0
9¢ 90271 STE*T €ST°l 9LE°T 86071 cyl £70°T
2¢0°Z
€1S*l 886°0 88S"I 7Z£6°0 999°1 LL8°0 60L°1 128°0 9¢8°1
le SLA SCOT 99L°0 SZ6°I TIL°0 810°2
~SIT*T CBE" ZIT 9071 8S0°I vIS*T 700°T 98S°I 0S6°0 Z99°I S68°0 ZvL*l 178°0 sz8°I 28Z°0 T16*l ¢¢L°0 100°Z
8¢ L221 O<e"l IZLII 88e"l PZI°l 6071 ZLO°I SIS°I 610°T S8S°I 996°0 8S9°l ¢€16°0 S¢L*l 098°0 918°T Z08°0
6£ Lez Leet LBITT 668°1 S2°0 S86°I
£6E°1 LETT ESV S80°l LIS°T EOI 78S°T Z86°0 Ss9°I 0£€6°0 6ZL°T 8Z8°0 208°! 928°0 L88°I PLL°O OL6°T
Ov 92°71 Nye] Bé6I"T 86o°I 8vI°l Losv°l 860°I 8I1S°T 80°I 78S°T L66°0 ZS9°I 976°0 HZL°l S68°0 66L°1 778°0
7] 8821 9LE°T SHz°l 9L8°l 68Z°0 9S6°1
€Zvl 102°1 glyl 9ST" 8ZS°I =TIT'l "8ST S90°T £79°T 6I10°T POL*I 26°0 892°I LZ6°0 7E8°l 188°0 Z06"I
Os 7Z<C°l <Onl S8Z°1 9071 S77! T6n"l SO0Z°T 8ES°I P9T"l
= L8S°T €ZI°l 629° 180°1 Z69°1 6£<0°1 872° 266°0
SS 9S¢°T Levl OZE"l 997°T s08°! $S6°0 798°1
8Z°T 90S°I LaZ°I BYS°T 602° Z6S°1 ZLIT 8e9°T PET°l $89°l S60°T MEL°l LSO°T S8Z°T BI0°I L¢8°l
o9 <BErl
l 607 OSE"l perl LIET Ozs*l €82°1 8SS°T 61271 86S°I 1 VIZ 6E9°T 6LI°T 289° DPT] 97L°1 S8OI°l TZZ°*T
S9 LOnT 8907°T LL¢°T oos*T 9NE°T NEST
Z20°1 Z18°1
SIE] 89S°T <8Z°1 709° =1S2°1 f09°T 8Iz°l 089°T
OL 98I°T OZL°I €ST°T T9L°T =OZI*I ZOB*T
627°1 SBn"l T OOF SIS*T ZLET 97S°T €HE"I BLS°Il Sele LISS €82°1 S79°l <¢S7°1 O89°T | ERLE CHEN Z6I°T 9SL°l Z9T°T
SL 8t7h"T
= 10s*T 220° 62S°T S6E°T LSS°T 89" L8S°T Z6L°1
OVE"T LI9*T CIE" 979°T =782°1 789°l 962°1 9JIL*T Leeal SUE
08 990° SIS*T 17" T9s°t) 9It°l 89ST 06¢°I
66I°T SBL°T
S6S°1 79E°1 HZ9°T BEE" £s9°T ZIT €89°T S8Z°1 MILT 6S2°1
S8 Z8t°T 8ZS°1 8S9°T £6S°T SEnT SPL°l CCEA PEPE
BLS°T TIT <O9°T 98E°T Os9°T 29st LOOT Leet SB9°l
06 96n°T O7S°1 ZI< MILT L8Z°T Sol" Ck SEERCI
PLP°l £9S°T 2S" L8S°I 624°1 T1191 900°T 9E9°T ¢€8E"T 199°T
S6 OIS*T O9¢"T L89°T 9¢E°T HILT Z1S*T Tel? 88z°! 69L°1
2SS°T 689°1 €LS°l 89n°T 96S°T 991°T BINT SZ CONT ¢€0°T 999°T 18¢°T 069°T 8S¢°I SIL°T 9E¢"T Tell
Oot 22S°T Z9S°T <0S°T £8S°T Z8t°T 709°I 72991
Eel LOLT
SZ9°1 TT Leo") 12771 OL9T O01 £69°T SZ¢rl LILI
osI ATS LEIA 86S°T 1S9°T 78S"T LS¢"I Tell SEErl S9L*Il
S99°l 1LS°1 649° LSS*T £69°T €9S°T 80L°T O€G"l ZL"
002 799° 789°l SIS*T LEL’l 10S°T ZSL°T 98h°T L9L°T
£€S9°T £691 €79°T HOLT ¢€¢9°T SILT €Z9°T SZL°l €19°T SEL] €09°T 9PL°T Z6S°1 LSL°T Z8S°1 892°1 ILS°1 6LL°1
21981,S- (panuyuoy)

T= Z1= c1=1 bl=1 S1=1 1 91= LI=1 8I= 61=4 0z=41


Mpa

pu
Cleat 7 tis Tp Mm 27, 15
| 090°0 907°< ket ae haan y E+. Opes i ene dea eas a Wake} mK: ea aa oe “eee ang oe aa
LI 180°0 982°¢ ¢sS0°0 90s°¢ eas ow eae ca) ee Medes Tae ae ee re ce ee eer sar eee)
81 “SUTCO SPIE $/0°0 see L0°0 £ss°¢€ Cher Re @==s> See ae ey betes See SSeS ste sane seem oe
61 ~SHI"O <20°e =ZOl°O Lé2°< 190°0 Ozv"¢ ¢t70°0 Ta9°< Soe eetSa Toe. ae ee es sa sern
a sre A
02 B8ZT°O v16°2 T<1°O 60I°< +260°0 L6Z°€ 190°0 pLy"< 8<0°0 6£9°€ prs SSe50 ests S254 weses are ses Ss ae
1Z RCO ATSIC Z9IT°O 700"< 611°0 s8i°< 1780°0 gc<e"¢ $S0°0 1Zs*€ S¢0°O 1L9°€ tare Sy emer ara Aas Se ewan care
cc 992°0 62L°% 761°0 606°2 87I°O ”80°¢ ~601°O zsz*< LL0°0 ZIv7°< 0S0°0 z9S°¢€ Z<£0°0 OOL*< Seas Sees seer, 2 ea eS
<Z 182°0 169°Z ~LZ2°0 228° 8ZI°0 166°2 9¢T°O SST°€ OOO Tl¢"¢ OL0°0 6S7°¢ 970°0 L46S°€ 6Z20°0 ScL°€ ease Sas Sar Soe
92 SI¢°O Ogs*z 09Z°0 WHL°Z 602°0 906°% S9T°O s90°< SZI°O 81Z°€ 760°0 £9¢°€ $90°0 10s°¢ £€70°0 629°€ LZ0°0 Lyl°< eo sooo
SZ 8v<°0 LIs°2 762°0 HL9°Z OVZ°0 628°2 761°0 786°Z 72S1°O Tete 9IT°O Hle< sS80°0 OIv7°’< 090°0 8<s°¢ 6£€0°0 Ls9°€ $Z0°0 99L°€
92 18¢°0 o9n°z ~HZ<E"0 OI9°Z ZLZ°0 8S/°2 2Z2°0 906°2 O8t*O aso’s I71°0 Té6te LOT°O sze"e 6L0°0 Zs" $S0°0 ZLS°€ 9€0°0 Z89°¢
LZ <17°0 607°2 96¢°0 CSS°Z ¢€0E°0 769°% €SZ°0 9€8°Z 802°0 9L6°Z LITO elle TET" snee OOO IZe"¢ ¢L0°0 067°< 160°0 zo9°¢
82 HHO £9¢°% L8C°0 667°2 EEE°0 Se9°% €82°0 CLL°Z L¢z°0 L06°2 761°0 O70°< 9ST°O 691°< ~ZZ1°0 "62°E €60°0 ZIv°< +890°0 9ZS°€
62 L070 12e°z LIv°O Isv°Z £9¢°0 c8S°Z <I¢eO ¢IL°% 992°0 £VR°Z 722°0 ZL6°Z Z8I°O 860°¢ 97T°O Oz2°< VITO sees L80°0 Ose
o¢ £0S°0 £82°2 Lth°0 Lov°2 £6¢°0 ££S°% ZHE°O 6S9°2 62°0 S8L°2 6%72°0 606°2 80Z°0 z<0°< {LTO cst°e LET°O L9z°< LOTTO 6Le"¢
I¢ 1£S°0 872°2 SLv°O L£9¢°% ZZ07°0 L87°Z IL¢°O 609°2 Z22¢°0 O£L°2 LL2°0 168°2 EZ°0 OL6°2 961°O L80°¢ O9T°O 10z*< 8ZI°O Ile"
ce 8SS°0 9122 €0S°0 Oge"Z OS7°0 9072 66¢°0 £9S°2 OS<"0 089°Z 70¢°0 L6L°% 192°0 Z16°Z 122°0 920°€ 81°0 Lets STO ICE
£¢ S8S°0 L81°Z 0£S°0 962°2 LLY°O 80”°zZ 927°0 O2S°Z SELEO ESIC EEO 9PL°% L8Z°0 8S8°2 97Z°0 696°% 602°0 8Z0°< L1°O mete
¢ O19°0 o9t°z 965°0 99Z°2 £0S°0 eL¢°% ZS7°0 18#°2 0%°0 06S5°2 LS¢°0 669°Z <I<°0 s08°z ZLZ°0 S16°2 ~£EZ°0 ZZ0°s L6T°O 9z1°<
Se 729°0 9<1°% =18S°0 L¢z°z 62S5°0 Ove"Z 8L7°0 vonZ O€7°O 0Sss°2 ¢8E"O SS9°Z 6£€£°0 19L°% 1L62°0 $98°% LSZ°0 696° 122°0 TLO°€
9¢ 8S9°0 ¢1I°Z S09°0 O1Z°Z 769°0 O1<*2 0S°0 O12 SSO ZIS°Z 601°0 V19°Z 9<°0 LIL°2 ZZ<°0 818°Z 7282°0 616°Z 772°0 610°
Le 089°0 Z60°Z 82z9°0 981°Z ~8LS°0 Z82°2 82S5°0 6L¢°Z O8%°0 LLv°Z ~HE7°0 §=9LS*% 682°0) SL9°Z Leo PLL°Z 90¢°0
+ ZL8°Z 89Z2°0 696°Z
8¢ ZOL°O ¢L0°Z 159°0 991°Z 10950 ISC2 7Z6S5°0 Os¢°Z 70S°0 sz 8S7°0 Ovs*z I1v"0 L¢9°2 LEO eel°Z O<E"0 8zZ8°Z 16Z2°0 £26°Z
6¢ €ZL°0 Ss0°Z €19°0 spr°c 7£z°Z€79°0
SLS°O £éo°% 8ZS°0 VI17°Z 28”°0 L0s°Z. 8¢°0 009°2 S6¢°O 769°% S¢°0 L8L°2 SIE°O 6L8°2
oO” 77L°0 6£0°2 69°0 £212 S9°0 OIZ°2 L6S°0 L6Z°Z 16S5°0 98¢°Z S0S°0 9L7°% 19%°0 995°2 8I7°0 £69°2 LLE"O 80L°2 BEe"0 8e8°Z
7] S¢8°0 cL6"1 O6L°0 ”70°Z 7L°0 8I1°Z OOL"O €61°Z_:
= S690 692° 719°0 9967 ~OLS°0 727°Z 82S°0 £0S°Z 88°0
+ Z8S°Z 8t707°0 199°Z
Os £160 SZ6"I IZ28°0 £86°1 628°0 1s0°Z L8L°0 911°Z ~97L°0 Z8I°Z SOL°0 0S2°Z $99°0 8I1g°Z $29°0 L8¢°2 985°0 9S7°Z 81S°0 92S°7
SS 626°0 168°T 07670 Sv6°l Z06°0 cO0°Z £98°0 §=©6S0°2 S28°O LIT ~98L°0 =ILI°Z BLO LEz°Z IIL°0 862°2 7L9°0 6S€°Z L¢9°O 127°Z
09 L¢eOrl s98°T 100°T v16"I S96°0 996°! 626°0 S10°Z £68°0 L£90°2 LS8°0 OZ1°Z 7278°0 £L1°Z -98L°0 L222 I1SL°0 €82°Z 9IL°O 8ee°2
S9 80°! S78" €S0°I 688° OZO"T 7€6°1 986°0 086°I £€56°0 L£ZO°Z 616°0 SLO°Z 988°0 €Z1°Z 72S8°0 ZLI°Z 618°0 1ZZ°Z 98L°0 CLEC
OL TEI°l T<8°T 660°1 OL8°T 890°T TT6°l LEO"l £S6°1 SO0°T S66°T 17L6°0 8<£0°2 £76°0
+ Z80°2 116°0 L212 088°0 ZL1°Z 6178°0 L12°Z
SL OLI"I 618°1 ItvI'l 9S68°T TIT'T 68° Z80°1 1g6°T ZSO°T OL6°T €ZO°l 600°2 £66°0 670°2 96°0 060°Z 7€6°0 Tet°% $06°0 ZLI°Z
08 s0Z°T O18T LLL*T p7s8"t OSI"l 828°T cele SlGylie 760°T 676° 990°T 786°1 ~6€0°1 220° I10°l £g0°2 £€86°0 L60°Z $S66°0 SE1°Z
S8 9€Z°I £08°l O12 HBT vBI°l 998°T 8cI°l 868°T cole RSG 9OI°T S96°1 O80"! 666°1 €S0°T €<0°2 LZ0°I 890°Z O00"! YOI°Z
06 79271 86L°1 O21 Lé8°1 SIZ7T 9S8°T I6I°l 988°T 99T"T LI6°T IIT 876° 9II"I 6L6°T =T60°I c10°% 990°1 770°Z 1470°1 LLO°%
S6 06271 £6L°1 LOceT ZEA 7Z°T 878°l 12271 9LB°T L6I'l S06°T LIT HE6°l OST°l £961 9ZI°T £66°1 ZOI"l ¢Z0°2 6L0°1 9S0°2
Ol VIET O6L°I 26271 918°T OLZ°1 Tv8°l 81Z°I 898° SZZ°T S68°T £0Z°1 f26°1 I8I"l 676° 8SI°l LL6°T 9<I°I 900°Z ENT DEO°S
OST <lV°l £8L°1 8Sh"I 66L°1 7t7h"T HIBT 620°1 O<el PIn'T LO8°l O0"T £98°1 s8¢"l ose’l OLE"! L68°T Scerl €16°l OvE"l Te6°l
002 196°T T6Z°T OSs°T TO8*T 6£S°1 £18°l 8ZS°T H28°T BIS"T 928° LOS*I Z£98°t sSén"T O98"! 78t7°T TZ8°T L7°T £88°l Z9%°1 968°1

,% si ay} Jaquinu
jo siossaibai burpntaoxe
ay] *ydaasaqur

555
71421S- (panuuod)

l= c= f= 9=1 S=1 1 9= L=A 8=.4 6=1 OI=.4


i 1p Np Ip Np Ip M 1p Np Tp Np 1p Mp Tp Mp 1p Np Ip NM Vp [
«9 01990 OOv| SESSne a === ona-
OL DOLO 95¢°1 ae wan- =e wnn- a ---- ---- a
L970 9681 aes ema = ——- aee wero re aoe
E9L70—ZEE*18 §=s 655°0 LLL*I =e sn-- ao ---- ---- === ———
8950 L82°2 a Se oom wane nee weet nen wars
C6 dOOZE*I_
= HZBO 6290 -SSH'O)BZI°Z
GE9"I
= 962°0 B8S°2 anne wn-- a ---= ---
ES paca aeas oa nee crr- ---- ----0 ----
OO <6LB°O.— OZS*1= «L690 «SZS°0—Ss9IO"Z
=
sIH9TT 9LE°0 PIMZ 8 <H2'0) eee caelA cake =o wa--
OMT «ZED HZEE ss oa Sine ceeek aaah aa ae Seay See
-BSL°O PONT= S650 -BZEI © VHD <BZ"Z = 9IED SH9'Z = £020 s00"€ Sy
«1Z6°0—OSE"ZI -ZIB'O «SLT Sie ss a arias pail
8590 P9BTT= «ZIS°O LLI'Z GLE 90S"Z = B92°0 28°= 1210 6vl'e ae ace Sree
EI = OIOI OvE"T = 198°0 z9¢71 SILO. -HLSO.O per
NEOTZ
9IBI
= SHY'O 6E"Z BZED —269"Z O70 $86°Z LYIO 99e°<
= SOT OSE*l $0670 Iss*l «L9L"O LL" -z69°0 SoZ -96Z"Z_-<
«= SOS*O GBSO ZLS"Z == ---- == w---
ST = LLOTI §=«19K*E 96D <HS"l = 9820 BHBZ= OOZO IIT" L2T0 O9ere So So
8 VIB'O OsL"| S890 LLU“ 2950 zzz © Loy ZLHZ © LZL'Z_<HEO“16270 LOZ
=O «9OTTI 2867O)~=—6ES"*I
=
sTLE*T LSB°O 82Z*I SLID 9IZ"E = TINO Bere
HEL"O SEB «S190 Z0S°0LBBE*Z «=
ST'z BESO HZIZ= HOS'O 09BZ= ZZ2°0 O60"E
CLI SEIT =sTBE*E STO EST LOBOS OIL"T 64L°0 ~—s006 SSI'O vOS"e
BT = ~—sHOT"Z_-999°0
55°70 «BIK*Z ISO -9SE0LES*Z ZLZOLSLO°Z SL'Z = BETO VIE
BSITI 161 OVO SES*I €£6°0 969° 0z8°0 ZBI OIL°0 ~—«090'Z =—-LSZ"Z_-<09°0
6 «OBIT §=sTOH*T «PLOT 9ES*T _7ZOS*O §=«19H'Z LOW'O L99'Z 1260 <LB°Z = He £L0"€
= L967) SB9"T = 6S8°0) <ZO'Z_-
= ZSL°0—a69ND ahnaT -90Z"Z
© -96E"Z_6
© SYD
oz = «Oz THT= OOII =LES*T 8660 ISH BBS'Z 69ED €BL°Z = O6Z°0 7L6°
1 IL -468"0 _Z6L°0—sBZAYI. «-Z69°0~—sZ9T"Z «dL66TI.
= S6S°D GEEZ ZOSO IZ"=
sdZ «ZZ OZHTSs SZITI. BESTT= 9ZO"I 6991= L260 ZIB8'l 91ND VOL'Z = 960 S88"
ZB'O
6 9G HZI'Z_-Z
= ELO LENO O6Z"Z © LHS O9HZ 19H'0 =£E9Z OBS"0
ZZ GZ 6ZHT «= LYITl IHS*T= ESOT 9991= 85670 L"I
LE ~—sONO-£98°0 908°2
gz LEZ LET= «BOTTI HST= 060"Z_69
= L°0 9HZ"Z_L
= L9DBBS°O LOWZ= 70S°0 ILS*Z= 7270 mele
BLOT 0991 986°0 SB8L°I S68°0 OZEI © OBO ©=—«190°Z SILO BUZZ=
nz SLZI OHH«= BBITT 9HS*T= BZ9'D ONEZ SSO) HIS'Z = SHO OL9°Z
“TOIT 9S9°T= €10°T SLL°1 $z6°0 «206° -LEB‘O SEO'Z © ISL°0 HLI"Z 9990 BIEZ
sz BBZIT HST= 90Z*T OSS*T= EZT*T HS9T= BEOl LOL"T BSD 99"Z 9050 £19°%
<<6"0 «OBB 98D ZI0'Z = «BLO MHI"Z = ZOL°O —O8Z"Z 1290 §=6IMZ
9 «= ZOS*I 19HT= HZZ1 ESS*1 «= HITT ZS9°T= 2901 6SL"l 6LE"0 LBL PSO O9S"2
hz STaey) 69HT© LOBOS 2661= 9IB"O CLII'Z SEL°*O HZZ= LSND 6LEZ 18S°0 €1S°Z
ObzeT “9S6Gute Oleh 1S vee “SL TOONS 98° <SZ60~—«TL DLE
=z Bzel 9ZnI SSz*l O96°1 —SHB'O <60'Z LILO 9IZ"Z = 1690 ZHE"Z= INO OLHZ
I8I"l OS9T VOI LUL"| —0S8°T_-8Z0°T 4LB°O
§=—-BSBTT_-156°0
éz ~~ «IHEtt BHT= OLZ*T £9571= BEIT «OSI -B6L°D——-BBIZ
sILOZ EZL°0 60E"Z= OS9'0 lepr2
ZIT =<HL°T = OsOT HBT «dT SLED. MHI= OOO ZS0°Z= 9ZBD HITZ
~—sOK ZSE"1 BHTT «= «HBZ*I LOST«= VIZ OSHT= EPIL 6EL"T SLO BLZ"Z = 7890 966°%
T2401 S58 8660 -~—sC«dTS6"1 «9760 HEOTZ= HSB°O HI'Z ZBL 16°
KOKO «SHH «LOZ OLSTT= 6ZZ*T OS9T= -O9T*T SEL"1= O60°T ScB"l ZILO £902
=o <<LEtl ZOSI= <6OS*1 HLSTT= OS6°0~—sBIO"Z
OZE"T.~O
= ZOTI -OIB"D=—9ZZ"Z MLO
OOZT"Z_-6LB"D. €eo°%
HHZT OSI= LALIT ZEk"t= 6OIT I8°T
6 «IDOL -Z60~SPOO'ZGO6T
§=— 70670 ZOI"Z
gg CBSE OST~S «ZEIT LLG] «BSZT ISHT= LOTT OLL"T = = 9680 €0Z"Z = 69L°0 90E"2
oe ~~ LeU €18T «190°T ODT. Y66"D 166" -LZ6°D SB0°Z= 198°O 181%
cect HIS«= ESI OBS IZZ71 299° BOZT 8ze1 = VII BOS S6L"0 182"Z
Sg ZOMI «BIS CHS] PAST O80°T ~—s168"I_ SIO"! 6LEI_= 05670 90°Z
6 S880 Z91°Z= 128°0 Lsc°%
EBT ES 2221 92L"1 OFT £08" LEOTI HEBTT «= HSOE L9G"=
9g «TIMI SZS*i «= «SEI LOSE= S6Z*T SOT IL6D HSOZ= BOD MHI"Z SUBD 9<e°%
EZ V2L"1 = SLIT 6641 HITT LZB"T SC <SOTT. LSOTE
LE «GID «=OOSSTL HOE OBST= LOSI SST «© GHZT EZLI= 166°) 1HO'Z= OLED LZT°Z| B9B°O 91e°%
~~ «Lz «SES O6I"T SOL" CISTI ~—OLB"L -ILOI SHOT«= 6ZOZ_“T
= IOT
156° ZITZ
SLE HOST «= BIST SHI «= 1921 Z2L"1 POC 26L"T TT = 168°O B62
6 Scr «ONS ON H9BTT
= BOI 66"«= 6ZO"l §=LIO"Z OL6°D —860°Z ZI6°O
ZBS*E LOST «= BZE1 BS9TT= ELZ*1 ZZLT= BIZ*T 6BL"I CISTI SBT«= O8I"Z
Ob ZT «HST «6S ONT= BEST BST POI] ZE6"1 «= LHOl LOO"Z= 066°0 SB0°Z = £60 "9V"%
SBZ*T IZL"T= OEe"T 9BL"T— S4V"T OZII~—y bZ6T
SaTT
«= «P90
Sh Cee 99671 «= OST «SITE SBE «99971 «= «9EE*I OZL* L66"I
«= BOO ZLO°Z = S60 6PI°Z
Os SUSI -SBSTT LBZ" SLL"T Beet «—<BTT «GRIT. SBT«= SII BS6"I = 680 2002
«= «-Z9HT BZ= «ZT «LOTT BLE IZLE See" TLL"T= B<O"l 880°
ZEIS~=—«LOST S “OGL THOTT Veer ONz"l~—SLOTzzert
= «1OZ*T OLET= 9ST*T =—986"T= OIL
25H BOT«= VIM «PALL Mel BOLT "ECT TTI m70"2
9 BHStT «=—o9IT HIST 7ST«= OBMT 6BITT«= ~HEZTT ~=T9BTT <<SZ*I BOBTT= ZIZ*1 §=6S6I OLTI O1O'Z
PPT LZL"I= BONT LOL*T e4e"1 SES" =—OSBTI
«9 LETT ZT§=— -9EG1 299= COST 9691 «= LVL ISLE BOZI EBT= O9Z*1 6671 Z2Z*1 MB6"l
OL BCT THOTT. BET =—LOLT rOvt -SoB'T ~~ OLET CHBTT«= OEE -ZBBT «= TOS" €Z6T
SST ZL9TT se <SZSI. SOLT «= VENT SELTI «= MIMI BOLT = 9921 "961
SL 86S°1 cS9°1 IZS°I
<£7"T ZB" «TOT LET~—s 69ST ELT= Lett Ol6T SOS*T
Os9"T <7S"T 60L°T SIS"I 6EL°1l L8yT OLL*T 8M6"l
08 119° Z99°T_=o 8¢"T 108°T 8ZHT MEB°l 66€°I L498°1 69¢°T 106°I
98ST B89"= <O9S*I «SILT HESTT EHLTL LOST ZLL"T= OB 6EE"T SE6l
S8 9Z9°1 IL9°T 009°1 969°1
VOS"T cen] ~—oSB*T SzHT IB«= LEST <68T 69271
SLS°T TZL°T OSS*T LoL°T SZS"T PLL'I 00S°T T08°T
SCO
06 Seo" 6L9°T ZI9"I <OL°T 7Lh°T 6281 8hh"T L4S8°l 7ZH°T 988° 96E°T
1 FIG
68S°T 97L°l 999°] ISsZ°T 29S"1 QLLT 8Is'I 108°T
S6 St79°T L£89°I €29°T 60L°T 76n°1 L78°l 69n°T 0S8°l Sih*T 188°l

556
Z09°T Zell 6LS°I SSL°T LSS°T BLL°T OZ7"T 606"1
O0T 7S9°IT 769° SEs" cO8*l ZIS*T LeB°l 68h°T ZS8°l S9n"T
pE9"T SILT E€I19°T 9EL°T Z6S°I 8SL°I IZS°T LLB" ZHI £06°T
OST “OZL*T 9HL*T O8Z°T 0SS°I £08°T 8ZS°T 9Z8°l 90S°T OS8°T
= 9OLT ~—O9L"T C691 MLL= 6L9I BBLTT= 18h"T PLB] 971 868"T
«BSL"I_~BLLTT_- S991 ZOB*T= S91 LIST= Le"1 zeg*T 2291 §=LHBT
—o0z= -BHL*T EBL*T
= BELT 66L"T= B2L"l 18s] SLE OZER= BOT Z9BI= "6ST CLB°l
LO/e1 L691 vOut
") eH ec ee DT
7198],S-q (panuuoy)

11=41 Z1=41 €1=41 I= S1=4 91=1 LI=1 81=41 61=1 Oz=


ny 1p Np Ip Np Dp M 1p Mp 1p Np Ip Mp Tp Mp Ip Np Tp Mp Ip A

91 86070 €0s°€ See| Sars ee at == me pace eos<a cotter ee cole Tae Sg


LT BETO. BLE“E = L80°O Lss¢ SSF Sem gm Saree mae Se Po gee== Ets az apa pag Tor Sen ese Soa
BI ~LLIO )=—-S9Z"E SZI7O §=Ivhs BLO €09"s re* Sse ss sas ES Go eo cae ane a == a2
6 OZZ°0 6SI"E OIIO SEE IIT 96H OLO'O Z79°€ gar Seoae Soe -2as
FP SoM ee a = caw
=O <9Z70 90° BOZO §=HEz"E SHIDO Sé6Es § ODIO §=2HSE <90°0 9L9°< SoS ee ea =-=8 === ----| wo--
9LU"Z_-LOS*O—CsdIZ
=: «OHZO =dIWI"E ZBI“O OOSS § ZEI°O Brhs 160°0 <8S°E BS0°O SOL See ose Ao ee =m aoe
ZZ 6HEO LBZ = -1BZ°O LSO"E= ~OZZO =sCATIZ*E -99T°O BGE*E§=—- OZIO S6HE= FBO'D G9"= 7S0°O IeL°"€ eee Baa oon wee-
EZ «16S70—Ss -ZZEO—Ss«O9ZB"Z
GLGU*ZSs -GSZO BZ= ZOZO ZLZ*E ESI°O 60S 8 DITO §=SEs*s 9LO'O O59 870°D esl'e ESS Sas eae eee)
OZ «SHO ~=Ss«d9L*Z 2950 BUS'Z = ~L6Z°0 EGOTE Ss <GEZ7D KHIE= 9BITO LZEE 8 IHI"O HSHE 8 IOTO §=2LS°E OLO'O BL9°<
8 OO ELLE ES Sp
SZ 3 OLY'O ZOLZ OOWO =OoHHBYZ SEED) EBETZ Ss <SLZ7O BIE 386 12270 §=1SZ°E ZLI"O 9LE"E OLIO ENS H60'OD ONE S9ID;O ZOLE 1700 O6L"<
<9Z = BOS'O -6HNZ . ~BSHO WBLTZ = LEO GHZ}§=— ZIS*O «ISOS «<9GZ7OD—S GLT"E = GOZO OEE= ONTO §=—OZHE OZIO §=6IESE LBO'O §8—zE9E 0900 els
LZ -0009°Z_HHG"D
= ~—SLHOLZ §=—-OOE 6OHD GSBZ «= -LBE"Z_BHE"D
= 16770 §=—ZIT"E BEZO EAE= 1610 «BREE EID O9HE ZITO 9S"E§ 180°O Bsoe
saz §=SS6°Z_-BLS°0.—
DISD OBZ ShH'O)§©= §=s-SEOB*Z -EBC°O BZU*Z = SZE"O OSOS ILZ°70 §=B9I*E ZZZ°0 <BZE BLITO ZEEE BETO Sév's
8 VOID Z6S°€
Z6 -ZI9°0.~—s STS°Z«= HED S9°Z = §8=SGL°Z_-LYTD BIND PLZ = -6SE°0) §=-Z66°Z SOLD LOT"E = HSZ"O GIZ*E86 BOZO 8=LZEE 99TO §=eWE 6210 Bzs"s
SOO «CH9TD YZ CLL LLS*O ZHS°Z = BOL"Z_—-ZISTO.
= ISO EZB*Z = Z6E0 LEE~Z LEEO ASO'E 9BZO OII"E BEZO 99Z°E S6I°O BIE ISTO S9VE
HLID~—OdI¢ <hh'z
= 80990 GHStO)€SG"z
S99°Z HBY'O ILLZ SZHO LB6°% OLED 966% LIE OIE 69270 BZ 72270 60E°E <81°0 90"<
ze «SOLO. ~=CsT'TW°Z 8E9°0 -9LG°D)~—sCLIS*Z
GZ9°Z
«= SISO -LSHOEELZ §=OWB'Z ~10V'D 8=9HE"Z HED §=—OSOE 66270 ESTE £S7°0 Z627e 1120 Byes
«ISLOCSg zBE*Z= 89970 HAZ= 9090 BBS*Z = 9HS7D ZH9"Z_ §©=- BBD Y9L*Z 86 ZENO 66B"Z GLE"OD :ODO"E
= «6ZE"0 ODE 8=— £8770 BBE= 6€2°0 £6c°<
CW ~BGL°D. SSE°Z = S69°0 «GHZ E9°D ~SLS*OGG°Z
7S9°Z
86 BISO HSL*Z 29ND) SB"Z86 6OVO §=—7S6°Z BSED §=6ISOE Zig 8=LYIE L970 Ones
OCS BLO [Link]«COKE"Z
SZHZ= «ZNO -dZS*Z = HOD 19°Z = LHS'O 9TL"Z = ZENO 8<IB'Z 6E70 O62 BBE°D —SOO’E
8 OWED 660" S62°O O6I"€
9g BORO -90S*z = BHL°O. BETZ=: -6B9TD ZZ§= «190 089"Z_-SLS*OS«9BS*Z
= OSD HLLZ L9H°0 8982 LIVO §=«196" 69670 =sS0"S
8 €2e°0 eul"e
CLE ISBO SBZZ 80 ZLLIO. LE"Z 86 HILO HOZ= L590 SSS°Z= §8=6—9H9'Z_-ZOIN'D
BYSO §8BEL'Z S6V0 628°C SHH'D OZE'% LEED 600° Ise"0 C60°s
HSB‘. ~—«S9Z"Z 96L°0. «OCdISS*Z GELID BENZ= -€B9N'N ~BZ9°D=-«9ZS*Z
WI9°Z
Ss SSO) KOL'Z «2750S Z6L"Z= ZL7VO 86088 H270 88962 BLE°0 9sO"s
6S -SLB‘O 9HZ"Z © GIBCO ZETZ 6 COLO I7Z«EE LOLTO GHZ = £590 §8=S8SZ DONO sTL9°Z = «6HS°0 LSL*Z
= 6670 8=EVBT 15770 6262 VOV'O <10°<
BZZ"Z_-«96BN.—Ss«OO
=: OBO OE*Z «4G= ©=«d16E°Z_-SBL°O.
IEL°O ELHZ = BLINN LGS*Z = «97ND «SLS"OsIH9°Z
PZL°Z
= SZS°0) 86—B0B°Z LLH'O ZEB? Ofv'0 L6°%
OSH BBE s«ST*Z_ «= §=—-SZZ"Z_-BKS'OD Z"Z_-LBBO 649K «BEBO «LEZ BBL°O EZ§= OLD ZIG*Z= -Z69'N 9BS*Z= HID 6S9°Z= BESO €El"Z= £SS°0 LOB"2
~Y9O"I~=—SsiS
~=s<OT'Z BIO" OTZ= LO SZZ*Z= LZ )8=EsLBZ"Z }§=OSE"Z_-ZBB'O
«980 IZ Z6L70 «LZ LHLO —S*Z EOL0 8=OI9Z 099°0 s49°%
SSS ~<6ZITI Z90°Z «= «Ss(I"Z_«LBO"ISHO, TZ «=sCOOL KOOL SZZ*Z_—- )§= 19670 G6I6°D)§=sdT'8Z*Z
|=BEEZ LLB'O «9EE"Z= 9EB°D SZ SEL'O ZIS*Z SLO ILS°2
~©=—-09 ~HBITI ~=sdTS0°Z SHIT] LOZ«= «OTT LZI°Z = «B9O"T LLI'Z Ss ZO" LZZ"Z 66D BLZ"Z = «1S6°0) OOEE"Z
= KIBO ZBE"Z860 LBD vee = 968°0) Lev"
«EZ ~S(«9Z SOI. HZ «=H OI] S6O'Z «= ZIT BETZ«= BBO BIZ= ZSO0°l 6222 9IOI 9LZ°% 0B6°D E22 HED ILE°% 8060 617%
COL -ZLZI «=—s9BETN EZ 9ZO"Z’_—- §«=: 9OZ*I 990°Z «= ZLT*I 9OTZ «=: «GEITI SHIZ «= SOIT GBI'Z ZLOTI«S ZEZ*Z
«= BENT SLZ*Z«= SOO"I =8Ie°2 1160 e9e"c
SL BOSTI =O ~LLZ*1 900°Z«© Lyz1 <HOTZ= SIZ OBO"Z «= «“HBITI STI°Z «= EST*T STZ= «ZTE S6T°Z «= OBO SEZ*Z«= BSO'T GLZ°Z= L2O°l Stg%
08 ~=— ~“ObEI OLS"«= «CTIS*I =SCdLOOTT 1 «CBZ PZO'Z = GTI GSOVZ «= ZZ €60°Z= SETI OZTZ «= GOTT S9T"Z «= «=—=s-*TOZ*ZET"T
BOI"T BETZ = 9LO"T SLe°%
SB «= 69ST OHETT~~ OZHETT LLO"I «= GIST GOO'Z«= -LBZI OOHOZ «= ~O9Z*T ELOTZ Ss -ZEZ*T SOT"Z= SOZ*T =GETZ LLATT ZAI = ERI 902= W2TT Were
6 = S6S*I LEST«= 69STI «=(99GT WHS] SETI«= BIST SZO'Z= -ZHZ*I SGO"Z «= 997" SBO"Z_- «= —OWZT 9TT*Z«= E121 BIS= LETT 6LT% O9T Ue
S6 «= «BIN «=C6ZOTI «HOST «=«9SG*T “OLE BOTT «=— «SHETI ZIO'Z = “ZEIT OHO'Z«= 9671 B9OZ = «LZ LOO'Z = Let 92S 22271 =9ST%
= LET IBIS
~—CODT HSHT EZOTI«= «SII HEI«= «S651 LOTT= §«=—OOO"Z_—sILEI
Lvsl 9ZO'Z «= HZE"I EGO"= «TOSI OBO"Z= LLZ°1 BOT"Z = EST SETS= 62271 IVS
OSI 6LS°I Z68°l 19ST 8061 OSS*I ZO SES*I Ov6l GIST 95671 POST ZL6T 68Hl 686°T HLT 9OOZ BSI 20% <vHT OVO"e
~—cgoZ ~HS9"T GBBT«= SHOT 96BT«=: ZE9"l BOG= «1291 TO"T
6 §«= OINT §=—sdTS6"T <66S°T EHG"T= 88ST S671= 9LS*I L9G*T= S9S*T 6L6"l SST 166°T

,% si ayy Jaquinu
yo siossaibas Bulpnjaxa
ay} *ydaosajur

payudsy
Kq UoIsstULIAd
Wo ‘vaLJaWoU0rg
"JOA ‘Sp ‘OU ‘8 ‘1161“dd '$661-2661

557
558 ECONOMETRIC METHODS

Table B-6 Wallis statistic for fourth-order autocorrelation

5 percent significance points of d, , and d 4, u for regressions


without quarterly dummy variables (k = k’ + 1)

ktz] k'=2 kla3 kag k'=5


Ba Sa Se a) Saneane, Moai area ete enty
16 0.774 0.982 0.662 1.109 0.549 1.275 0.435 1.381 0.350 1.532
20. 0.924 1.102 0.827 1.203 0.728 1.327 0.626 1.428 0.544 1.556
24 1.036 1.189 0.953 1.273 0.867 1.371 0.779 10459 0.702 1.565
28 1.123 1.257 1.050 1.328 0.975 1.410 0.898 1.487 0.828 1.576
32 14192 1.311 1.127 1.373 1.061 1.443 0.993 1.511 0.929 1.587
36 1.248 1.355 L191 1.410 1.131 1.471 1.070 1.532 1.013 1.598
40 1.295 1.392 1.243 1.442 1.190 1.496 1135 1.550 1.082 1.609
44 1.335 1.423 1.288 1.469 1.239 1.518 1.189 1.567 1.141 1.620
48 1.369 1.451 1.326 1.493 1.281 1.537 1.236 1.582 1.191 1.630
52. 1.399 1.475 1.359 1.513 1.318 1.554 1.276 1.595 1.235. 1.639
56 1.426 1.496 1.389 1.532 1.351 1.569 1.312 1.608 1.273 1.648
60 1.449 1.515 1.415 1.548 1.379 1.583 1.343 1.619 1.307 1.656
64 1.470 1.532 1.438 1.563 1.405 1.596 1.371 1.629 1.337 1664
68 1.489 1.548 1.459 1.577 1.427 1.608 1.396 1.639 1.364 1.671
72 1,507 1.562 1.478 1.589 1.448 1.618 1.418 1.648 1.388 1.678
76 1.522 1.574 1.495 1.601. 1.467 1.628 1.439 1.656 Lall 1.685
60 1.537 1.586 1.511 1.611 1,484 1.637 1.457 1.663 1.431 1:69]
84 1.550 1.597 1.525 1.621 1.500 1.646 1.475 1.671 1.449 1.696
88 1.562 1.607 1.539 1.630 1.515 1.654 1.490 1.677 1.466 1.702
92 1.574 1.617 1.55) 1.639 1.528 1.661 1.505 1.684 1.682 1-907
96 1.584 1.626 1.563 1.647 1.541 1.668 1.519 1.690 1.496 1.712
100 1.594 1.634. 1.573 1.654 1.552 1.674 1.531 1.695 1.510 L717

5 percent significance points of d, ; and d4 y for regressions including


a constant term and quarterly dummy variables (k = k” + 4)

kM kM=2 kMa3 kMa4 kMa5


a SO ea TS praia ay Sat Ga ime aneds cy Sy ey
16 1.156 1.381 1.031 1.532 0.902 1.776 0.777 2.191 0.693 2.238
20 1.228 1.428 1.123 1.556 1.013 1.726 0.899 1.954 0.806 2.042
24 1.287 1.459 1.199 1,565 1.107 1.694 1,011 1.856 0.928 1.949
28 1.337 1.487 1.261 1.576 1.181 1.679 1.099 1.803 1.025 1.889
5214379 1511 1.312 1.587. 1,263 1.673 1171 1773 16108 1.850
36 1014 1.532 1.355 1.598 1.293 1.672 1.230 1.755 1.170 1.824
4O 1.445 1.550 1.391 1.609 1.336 1.674 1.279 1.745 1.225 1.807
44 1471-1567 1.422 1.620 1.373 1.677 1.321 1.739 1.972 1.795
48 1.494 1.582 1.450 1.630 1.404 1.681 1.357 1.737 1312 1788
521.514 1.595 1.674 1.639 1.432 1.686 1.389 1.736 1.347 1.792
56 1.533 1.608 1.495 1.648 1.456 1.691 1.416 1.736 1.377 1.979
60 1.549 1.619 1.514 1.656 1.478 1.696 1.441 1.737 1.404 1.997
64 1.564 1.629 1.531 1.664 1.497 1.700 1.463 1.739 1.429 1.776
68 1.577 1.639 1.546 1.671 1.515 1.705 1.482 1.741 1.450 1.775
72 1,590 1.648 1.560 1.678 1.531 1.710 1.500 1.743 1.470 1.776
76 1.601 1.656 1.573 1.685 1.545 1.714 1517 1.746 1.488. 1.776
BO 1-611 1.663 1.585 1.691 1.559 1.719 S31 1.748 1.504 1-977
84 1.621 1.671 1.596 1.696 1.571 1.723 1.545 1.751 1.519 1.778
88 = 1.630 1.677 1.607 1.702 1.582 1.727 1.558 1.753 1.533 1.979
92. 1.639 1.684 | 1.616 1.707 1.593 1.731 "1.570. $1756" 1/546. 1761
96 1.647 1.690 1.625 1.712 1.603 1.735 1.580 1.759 1.558 1-782
100 1.654 1.695 1.633 1.717 1.612 1.739 1.59] 1.76) 1.569 L786

Reprinted by permission from Econometrica, vol. 40, no. 0, 1972,


pp.
623-625.
STATISTICAL TABLES 559

Table B-7 The modified Von Neumann ratio


5 percent, | percent, and .1 percent points of the modified Von Neumann ratio

Degrees Degrees
of 5% 1% 1% 5% 1% 1% of 5% 1% 1% 5% 1% 1%
Freedom Freedom

One-tailed test One-tailed test One-tailed test One-tailed test


against positive against negative against positive against negative
autocorrelation autocorrelation autocorrelation autocorrelation

31 1.410 1.186 .955 2.595 2.826 3.066


2 025 .001 .000 3.975 3.999 4.000 32 1.419 1.198 .970 2.585 2.813 3.051
3 252 052 005 4.142 4.427 4.493 33 1.428 1.209 .984 2.576 2.801 3.036
4 474 .170 = .037 3.827 4.295 4.496 34 1.437 1.221 .997 2.567 2.789 3.021
5 598 .292 .095 3.571 4.076 4.378 35 1.445 1.231 1.010 2.559 2.778 3.007

6 -J01 .386 = .163 3.413 3.881 4.233 36 1.452 1.241 1.022 2.551 2.767 2.994
a -790 .464 .228 3.299 3.731 4.095 37 1.460 1.251 1.034 2.544 2.757 2.982
8 861 537 .285 3.206 3.618 3.973 38 1.467 1.261 1.045 2.536 2.747 2.969
a 22 601 2539 3.131 3.524 3.871 39 1.474 1.270 1.057 2.529 2.738 2.957
10 iD: %657, 590 3.069 3.445 3.784 40 1.480 1.279 1.067 2.522 2.729 2.946

11 1.020 .708 .438 3.016 3.378 3.710 41 1.487 1.287 1.078 2.516 2.720 2.935
12 1.060 .753 .482 2.970 3.319 3.645 42 1.493 1.295 1.088 2.510 2.711 2.925
13 V096 .795 .523 2.930 3.268 3.587 43 1.499 1.303 1.097 2.504 2.703 2.914
14 1.128 .832 .561 2.8959 5.222) 3.955 44 1.504 1.311 1.107 2.498 2.695 2.904
15 1157, 6866-597 2.863 3.181 3.488 45 1.510 1.318 1.116 2.492 2.687 2.895

16 1.183 .898 .630 2.835 3.144 3.445 46 1.515 1.325 1.125 2.487 2.680 2.885
17 1.207 .927 .661 2.809 3.110 3.406 47 1.520 1.332 1.133 2.482 2.673 2.876
18 12228°) 954.6911 2.785 3.079 3.370 48 1.525 1.339 1.142 2.477 2.666 2.868
19 15249 979 718 2.764 3.051 3.337 49 1.530 1.346 1.150 2.472 2.659 2.859
20 1.267 1.003 .744 2.744 3.025 3.306 50 1.535 1.352 1.158 2.467 2.653 2.851

21 1.285 1.024 .769 2.725 3.000 3.277 51 1.540 1.358 1.165 2.462 2.646 2.843
22 1.301 1.045 .792 2.708 2.978 3.250 52 1.544 1.364 1.173 2.458 2.640 2.835
23 1.316 1.064 .814 2.692 2.957 3.225 53 1.548 1.370 1.180 2.453 2.634 2.828
24 1.330 1.082 .834 2.677 2.937 3.201 54 1.552 1.376 1.187 2.449 2.628 2.820
25 1.344 1.100 .854 2.663 2.918 3.179 55 1.557 1.381 1.194 2.445 2.623 2.813

26 1.356 1.116 .873 2.650 2.901 3.157 56 1.561 1.387 1.201 2.441 2.617 2.806
27 1.368 1.131 .891 2.638 2.884 3.137 oy) 1.564 1.392 1.207 2.437 2.612 2.799
28 1.380 1.146 .908 2.626 2.868 3.118 58 1.568 1.397 1.214 2.433 2.606 2.793
29 155909 PV6O) 9.925 2.615 2.854 3.100 59 1.572 1.402 1.220 2.429 2.601 2.786
30 1.400 1.173 .940 2.605 2.839 3.083 60 1.575 1.407 1.226 2.426 2.596 2.780
ee Sa

Reprinted by permission of S. J. Press and R. B. Brooks from Report No. 6911, Center for
Mathematical Studies in Business and Economics, University of Chicago, Chicago, 1969.
560 ECONOMETRIC METHODS

Table B-8 Significance values for cy in the cusum of squares test


ao
a a
ee
m 0-10 0-05 0-025 0-01 0-005 m
II EN 0-10 0-05 0-025
ae 0-01
ge 0-005
1 0.40000 0.45000 0.47500 0.49000 0.49500 41 0.14916 0.17215 0.19254 0.21667. .233310
2 35044 44306 = .50855 456667 59596 42 14761 417034 = 1905021436 =~ .23081
3 35477 41811 46702-53456 ~—.57900 43 14611 16858 = 61885221212 .22839
4 33435 3907544641 = 50495 -.54210 44 14466 = 16688618661 «20995 .22605
5 31556 63735942174 4769251576 45 14325 16524 618475 20785 = .22377
6 30244 3552240045 645440 .48988 46 14188 = 16364 = «18295. 20581 .22157
7 -28991 33905 38294 = 443337 .46761 47 14055 16208 = .18120 20383 .21943
8 -27828 = 32538 = .36697 41522 .44819 48 13926 «1605817950 .20190 «21735
9 -26794 31325 35277439922 .43071 49 13800 «15911 «17785 20003-21534
1025884 630221 = 3402238481 41517 50 13678 = «15769417624 = 19822 .21337
IL 25071629227 43289437187 .40122 51 13559415630 «17468 = «19645 .21146
12 24325 .28330 431869 36019 .38856 52 13443 1549517316 419473 .20961
13 423639427515 3093534954 37703 53 1333015363 .17168 = «19305 .20780
14 23010 «26767 = .30081 = 33980-36649 54 1322115235 .17024 = 419142 .20604
15 22430» .26077 «29296 ~— 33083 = .35679 55 .ISM13 J ISI10— 16884 £18983 © .20432
16 21895... .25439 .28570 32256 ~—.34784 56 13009414989 .16746 «18828 = .20265
1721397 .24847-— 427897431489 33953 57 12907 14870 .16613 «18677. 20101
18 .20933-.24296 = 2727030775 ~—.33181 58 1280714754 16482418529 .19942
19 20498-23781 = «2668530108 =~ .32459 59 12710414641 .16355 £18385 19786
20 .20089 23298 )=—.26137 .29484 31784 60 12615 1453016230 41824519635
21 19705 22844 «25622 «28898 ~—.31149 62 12431 14316 = 1599017973 .19341
22 1934322416 25136 © .28346 ~—.30552 64 12255 14112415760 17713— 19061
23 19001-22012, .24679-—-.27825 29989 66 12087 13916 15540 £17464 = .18792
24 1867721630 «24245 427333 29456 68 11926 13728 41532917226 18535
25 1837021268 = .23835 626866 =~ .28951 70 1177113548 415127 16997 .18288
26 = «18077 20924 = .23445 26423 .28472 72 41162213375 414932416777. 18051
27° «17799-20596 23074 = .26001—-. 28016 74 41147913208 £14745 16566 17823
28 = «17533 20283 .22721-~S .25600~—-«.27582 76 11341 13048 = 1456516363 17604
29 41728019985 .22383 25217127168 78 11208 12894414392 16167417392
30. «17037419700 22061 = «24851 .26772 80 11079412745. 414224 15978 17188
31 16805 1942721752 .24501~—.26393 82 10955 41260114063 15795 16992
32 616582419166 = 2145724165 .26030 84 10835 12462413907 1561916802
33 1636818915 2117323843 125683 86 10719 412327413756 £15449 16618
34 61616218674 = .20901 23534 = 125348 88 10607 12197413610 .15284 =, 16440
35 615964 18442 .20639-— 23237125027 90 .10499 12071 £13468 = 15124 16268
36 «1577418218 .20387.22951-—.24718 92 10393 11949413331 = 414970. 16101
3741559018003 20144 .22676 =~ .24421 94 1029111831 13198 = 14820-15940
38 = «15413417796 19910 622410124134 96 1019211716 = 413070 14674415783
39415242, .17595 19684-22154 123857 98 10096 11604 = 12944 14533
401507617402 19465421906 =23589 «= 100 «10002 115631
«(11496 512823 «514396. 15485
Values for odd n greater than 650 are available from the author on request.

The values of cy are used to determine the pair of lines, 5, = +c¢9


+ (r — k)/(n — k). For n
observations, k explanatory variables (including the intercept, if there
is one) and a given significance
level a, co is found by entering the table at m =1(n — k) — 1 and 7a. For
a one-sided test, enter at
m =3(n—k)—1 and a. When (n — k) is odd, the procedure suggested is to interpol
ate linearly
between m = 3(n — k)—3 and m =4(n — k) —4.

Reprinted by permission of the Biometrika Trustees from Biometr


ika, vol. 56, 1969, p. 4.
INDEX

Adaptive expectations (see Lagged Autocorrelation coefficient, 304


variables) Autocovariance, 304
Adelman, I., 541 Autoregressive, moving average
Almon, S., 352 (ARMA) processes, 306, 308,
Amemiya, T., 409, 428 375-381
Analysis of variance (ANOVA): Autoregressive (AR) processes, 306
in general linear model, 186-192
in two-variable linear model,
39-42 Balestra, P., 405, 407
Anderson, O. D., 378 Baltagi, B. H., 407
Anderson, T. W., 484, 485 Bartlett, M. S., 431
Arrow, K. J., 62 slepvelny (Co IM, SWS), ST)
Asymptotic properties of estimators, Beguin, J. M., 378
268 -274 Belsley, D. A., 249
Autocorrelated disturbances, Berndt, E. R., 333, 340
304-309 Best linear unbiased estimator
consequences of, for ordinary (bs 1s ures):
least squares (OLS), definition of, 32
310-313 in general linear model, 171-174
estimation procedures with, Betancourt, R., 367
321-329 Box, [Link] 62, 306; 345, 372:
fourth-order, 317 374,376, 3773 379, 380
lagged dependent variable and, Box-Cox transformation, 62-72
362-371 Breusch, T. S., 300, 319
prediction with, 329-330 Bronowski, J., 516
reasons for, 309-310 Brooks, R. B., 389
spatial, 305 Brown, R. L., 387, 390, 409 .
tests for, 313-321 Brundy, J. M., 482

561
562 INDEX

Burman, P., 542 Determinant (Cont.):


Buse, A., 263, 304, 395 minor, 125
of partitioned matrix, 137-138
Cairncross, Sir Alec, 510 properties of, 127-133
Dhrymes, P. J., 153, 358, 365, 509
Cairncross test, 509-510
Charatsis, E. G., 327 Disturbances:
Chatfield, C., 377 sources of nonspherical, 287-290
Chebysheff’s theorem, 270 spherical, 287
Chenery, H. B., 62 (See also Autocorrelated
Chiang."A..C410; 517 disturbances)
Chow, G. C., 508 Disturbances term:
Christensen, L. R., 335, 340 properties of, 15-16
reasons for, 14-15
Cochrane, D., 323, 366, 367
Duesenberry, J., 483
Collier, P., 390
Collinearity (see Multicollinearity) Dummy variables
Confidence intervals: use of as regressors, 225-233
in general linear model, 181-198 use of, in seasonal adjustment
in two-variable linear model, procedures, 234-239
34-39 Durbin, J. M., 314, 318, 324, 387,
Convergence: 390, 409, 431
in distribution, 272-274 Durbin test with lagged dependent
in probability, 269-272 variable, 318
Cooley, 1) Reals; 3:15 Durbin-Watson test, 314-317
Correlation coefficient:
multiple: adjustment of, for
degrees of freedom, 177-178
in general linear model, Efficiency of an estimator, 31
176-177 Endogenous variables, 7
in three-variable case, 78, 84 Errors in variables, 428 --435
partial, 82-84 Evans, J. M., 387, 390, 409
in two-variable linear model, Exogenous variables, 7
23-25
Correlation ratio, 53-56
Correlogram, 306
Fair, R. C., 409
Cost share equations, 336-337
Farebrother, R. W., 316
Courant, R., 353
Feldstein, M. S., 255
Cox, DRe 6214204 425..408
Finney, D. J., 420, 428
Cramer-Rao theorem, 276-277
Fisher, F. M., 220, 455, 467, 483
Cramer’s rule, 138-140
Fisher, R. A., 422
Fisher, W. D., 480
Data mining, 501-504 Friedman, M., 351
Degenerate distribution, 272 Frisch, R., 237
Determinant: Fromm, G., 483
cofactor, 125 Full-information maximum
definition of, 123-127 likelihood (FIML), 490-492
INDEX 563

Garbade, K., 392 Hood, W. C., 484, 485


Garber, S. G., 392 Houck, JP. 41!
Gauss-Markov theorem, 173 Hudson, E. A., 333
Gaver, K. M., 511 Hunter, J., 508
Geiseletvins. 11 Hussain, A., 407
Generalized least-squares (GLS) Huxley, Sir Julian, 61
estimator, 291-293 Hyman, H. H., 541
Georgopoulou, A., 239
Gilbert, R. F., 338
Giles D. BSA.; 317
Glesjer, H., 301 Identification:
Godfrey, L. G., 319 examples of, 456-460
Goldberger, A. S., 330 general statement on, 450-456
Goldfeld, S. M., 301, 302, 389, 407, inhomogeneous linear restrictions,
409 461
Goodnight, J., 257 necessary condition, 454-455
Gourieroux, C., 378 restrictions across equations,
Granger, C. W. J., 377 462-463
Graybill, F. A., 285 restrictions on structural
Gregory, P. R., 333 coefficients, 452-460
Griffin, J. M., 333, 406 restrictions on variance matrix,
Griffiths, W. E., 279, 303, 341 463-467
Griliches, Z., 324 simple two-equation illustration,
444-450
treatment of identities, 460-461
Hadley, G., 144 Indirect least squares (ILS), 442
Hannan, E. J., 316 in recursive systems, 469-472
Hanssens, D. M., 378 Information matrix, 276-277
Hart, B. I., 389 Instrumental variable (IV)
Harvey,.A..C. +279; 301; 326,327; estimation, 363-366
388, 389 with errors in variables, 430-432
Haugh, L. D., 379, 380 with lagged dependent variable,
Hausman, J. A., 402 364-365
Hendry, D. F., 506-508 (See also Simultaneous equation
Heteroscedasticity: systems, instrumental
definition of, 169 variable estimation)
estimation under, 302-304
in grouped data, 293-296
in replicated data, 296-298
tests for, 298-302 Jaffee, D. M., 409
Hildreth, C., 411 Jenkins, G. M., 306, 345, 372, 376,
Hill, R. C., 279, 303, 341 S77,
Hoel, P. G., 274, 277, 409, 524 Johnston, J., 239, 500
Hoerl, A. E., 252 Jorgenson, D..W 5 235, 333; 339,
Homoscedasticity, definition of, 482, 508
169 Judge, G. G., 279, 303, 341, 428
564 INDEX

Kelejian, H., 367, 409 LeRoy, iS. .o2o


Kendall, M. G., 274, 277, 278, 298, Lim less
435, 540-542 Limited-information maximum
Kennard, R. W., 252 likelihood (LIML) estimators,
King, M. L., 317 483-486
Klein, L. R., 482, 483 Limiting distribution (see
Kloek, T., 323, 481 Convergence, in distribution)
Kmenta, J., 338, 398 Lindgren, B. W., 531
Koopmans, T. C., 484, 485 Linear restrictions:
Koyck, L. M., 347 estimation subject to, 204-207
Koyck scheme (see Lagged in general linear model, 182-184
variables, Koyck scheme) inference procedures, 184-198
Kuh, E., 249, 483 tests of, in sets of equations,
338-341
Linearity:
Lag operator, 289-290, 308 test of, in two-variable linear
Lagged variables, 343-381 model, 48-60
adaptive expectations, 348-349, transformations inducing, 61-74
351-352 iu, L2 M:, 378
Almon lags, 352-368 Lovell, M. C., 237
Koyck scheme, 346-348
direct estimation of, 358-360 McAvinchey, I. D., 326, 327
lagged dependent variables: and McCallum, B. T., 544
autocorrelated disturbances, McFadden, D., 425
362-371 MacKinnon, J. G., 325, 327
and well-behaved Maddala, G. S., 405, 407, 409, 462
disturbances, 360-362 Malinvaud, E., 279, 361, 469
lagged explanatory variable, Mann, H. B., 362, 365
343-346 Matrices:
mean lag, 344-346 addition of, 94-95
partial adjustment, 349-352 differentiation of, 102-104
Laus a Ji 335 equality of, 95
Reamer, EE. S01, 505,510 55.13- Kronecker (direct) product of,
515 136
Least squares: inverse of, 137
geometric treatment of, 104-113 multiplication of, 92-94
properties of estimators, 25-34 partitioned, 99-100
(See also Indirect least squares; addition and multiplication of,
Three-stage least squares; 100
Two-stage least squares) determinant of, 137-138
Least-squares principle, 17 inverse of, 135
Least variance ratio (LVR) transposition of, 92-93, 96-97
estimators, 483-486 Matrix:
Eecemin ©, 2792 303.2341 adjoint (adjugate), 125
Lempets, F2B=323 characteristic equation for, 141
INDEX 565

Matrix (Cont.): Maximum likelihood (ML) (Cont.):


characteristic roots and vectors with lagged dependent variable,
(see eigenvalues and 366-371
eigenvectors, below) limited-information (LIML),
definition of, 90 483-486
determinant of (see Determinant) Mean-squared error, 27-28
diagonal, 98 Mennes, L. B. M., 481
eigenvalues and eigenvectors, Mincer, J., 66
141-143 Minhas, B. S., 62
properties of, 143-150 Minimum variance bound (MVB)
idempotent, 99 (see Cramer-Rao theorem)
information, 276-277 Mitchell, B. M., 326
inverse, 113-114, 122-127 Mizon, G. E., 506
properties of, 133-138 Model:
latent roots and vectors (see reduced form, 7-8, 440,
eigenvalues and 443-450
eigenvectors, above) structural form, 6-7, 443-450
null, 99 Model selection:
nullspace (kernel) of, 117 Bayesian approach, 510-516
dimension (nullity) of, 117-118 criteria for, 504-510
orthogonal, 145 Monfort, A., 378
partitioned, 99-100 Morris, C. T., 541
addition and multiplication of, Moving average (MA) processes,
100 306
determinant of, 137-138 Multicollinearity, 239-241
inverse of, 135 definition of, 240
positive definite, 151-153 detection of, 249-250
variance matrix, 162-163 effects of, 240-241, 245-249
rank of, 114-116 and estimable functions, 241-245
summary on, 122 remedies in, 250-259
scalar, 98 Multipliers:
symmetric, 96 impact, 9
trace of, 147-148 interim, 10
unit (identity), 98 total, 10
Matrix operations, summary on, Mundlak, Y., 407
100
Maximum likelihood (ML) Nadiri, M. I., 508
estimator, 274-275 Nagar, A. L., 316
with autocorrelated disturbances, Nelson, F. D., 409
325-329 Nerlove, M., 316, 405, 407, 428
with errors in variables, 432-435 Newbold, P., 377
full-information (FIML), Normal equations:
490-492 in general case, 104, 111-113
in general linear model, 275 in three-variable case, 76
general properties of, 276-279 in two-variable case, 18
566 INDEX

Oliver. Fo R373 Quadratic forms:


Orcutt GH 323536083567 definition of, 150-151
Ordinary least squares (OLS): distribution of, 165-167
consequences of autocorrelated independence of, 167
disturbances for, 310-313 Qualitative dependent variables,
in general linear model, 171-174 419-428
inconsistency of, 440-441 Quandt, R. E., 301, 302, 389, 407,
409

Pagan, A. R., 300 Rao, C. R., 280


Parka heb 326 Rao, P., 324
Parks, R. W., 333, 341 Recursive residuals, 384-392
Partial adjustment (see Lagged Recursive systems (see
variables) Simultaneous equation
Pearsonian correlation coefficient, systems, recursive systems)
23-25 Reduced form equations, 7-8, 440,
Reckee Ja Kees 443-450
Perry, G.S0285 Riddell, W. C., 263, 387, 408
Phillips, G. D. A., 301, 388, 389 Ridge regression, 252
Phlips, L., 330 Rubin, H., 484, 485
RidoteG aber inawo42
Pierce, DivA™ 379
Plackett, R. L., 436 Sargan, J. D., 506, 507
Plim (see Probability limit) Savin, N. E., 316
Poirier, D. J., 392, 395, 396 Schmidt, P.; 252; 279, 28.1, 303, 338
Pooling of time-series and Scot J. Le) p..o4u
cross-section data, 396-407 Seasonal adjustment, 234-239
Praiscess dees 24 Seemingly unrelated regression
Predetermined variables, 7 equations (SURE), 337-341
Prediction: Sets of equations, 330-341
with autocorrelated disturbances, tests of linear restrictions in,
329 - 330 338-341
in general linear model, 193-198 Sewell, W. P., 270
with stochastic X variables, SHAZAM program, 326
198 -200 Shephard’s lemma, 334
in two-variable linear model, Significance tests:
42-45 in general linear model, 181-198
Prescott ECG 415 in two-variable linear model,
Press, S..J., 389, 428 34-39
Price, =D) Dao42 Simultaneous equation systems,
Principal components, use of, in 439-492
two-stage least-squares (2SLS) examples of, 439-441
estimation, 481-482 full-information maximum
Probability limit (plim), 269-272 likelihood (FIML),
Probits, 426 490-492
INDEX 567

Simultaneous equation systems Switching regressions (see


(Cont.): Variable-parameter models)
identification problem (see Sylwestrowicz, J. D., 507
Identification)
inconsistency of ordinary least
squares (OLS), 440-441 Taylor, L.D., 274) 277,524,027,
indirect least-squares (ILS) 335
estimation, 442, 467-472 Terrell, sR. D393 16
instrumental variable (IV) Theil, H.5 25952 7 1P 27312745316
estimation, 441-442, 385, 479, 486, 488, 490, 492,
477-479, 482-483 504
least variance ratio (LVR) Three-stage least squares (3SLS),
estimators, 483-486 486-490
limited-information maximum Time-series methods, 371-381
likelihood (LIML) stationarity, 372-375
estimators, 483-486 transfer function, 372
recursive systems, 466 Toro-Vizcarrondo, C., 257
estimation of, 467-469 Transcendental logarithmic
three-stage least-squares (3SLS) (translog) functions, 335-337
estimation, 486-490 Two-stage least squares (2SLS),
two-stage least-squares (2SLS) 442-443, 472-477
estimation, 442-443, as an instrumental variable (IV)
472-483 estimator, 477-483
Solow, R. M., 62
Specification error, 259-264 Variable-parameter models,
Spline functions, 392-396 407-419
Standard error of an estimator, 27 Vector:
Structural change, tests of, 207-225 definition of, 90
Structural form equations, 6-7, length (norm) of, 109
443-450 null, 99
Structure, 450 Vector geometry, summary on, 113
Stuart, A., 274, 277, 278, 298, 435, Vector space, 107
540, 541 basis, 107
Sum of squares: spanning set, 109
decomposition of: in general Vectors:
linear model, 176-177 geometric representation of,
104-113
in three-variable model, 76-78
linear dependence and
in two-variable regression,
independence, 107-111
21-22
orthogonal, 110
explained (ESS), 21-22 orthogonal set, 145
residual (unexplained) (RSS), parallelogram law for addition of,
21-22 105
total (TSS), 21-22 scalar, dot, or inner product,
Swamy, P. A. V. B., 412, 415 93-94
568 INDEX

von Neumann, J., 389 White, J. S., 534


von Neumann ratio, 389 White, K. J., 316, 326
Wilks, S. S., 528
Winsten, C. B., 324
Wadycki, W. J., 480 Wood, D. O., 333
Waelbroeck, J., 492
Wald, A., 362, 365, 431
Wallace, T. D.; 257; 407 Yates, F., 422
WallisssK. Fs, 3:16, 317
Watson, G. S., 314
Waud, R. N., 358 Zarembka, P., 62, 415, 425, 511
Waugh, F. V., 237 Zellner, A., 337, 486, 488, 490, 510,
Welsch, R. E., 249 SA2e513
EEE

Ee
ee
eae
De
EIS EL
EASA ea
z 7 Pere
Eee
aE
ie:

LIE,
ee

aa
ee
EE
Soa LE

es
LY
Zep LL
Ze
Ee
eeeEEO EE
SEES
Les ie

Common questions

Powered by AI

Heteroskedasticity, which refers to the presence of non-constant variance in errors across observations, significantly affects econometric models as it violates the assumptions of homoscedasticity necessary for best linear unbiased estimators (BLUE). It complicates the prediction and estimation outputs, leading to inefficient estimators and invalid hypothesis tests if not corrected. Often seen in regression analyses involving cross-sectional data, addressing heteroskedasticity involves using generalized least squares (GLS) or robust standard errors to ensure reliable interpretations .

Transposition in matrix operations alternates the arrangement of elements from rows to columns (and vice versa) without changing the actual values. This process allows for multiplication operations that require compatible dimensions, such as turning a column vector into a row vector for the inner product calculation. Transposition thus facilitates various mathematical manipulations necessary in linear algebra and ensures operations adhere to required structural formats .

Matrix algebra simplifies the solution of least squares problems by facilitating the manipulation of entire datasets in a structured way. Using matrix operations, problems are reduced to solving normal equations like (X'X)b = X'y, where the vector b is computed using inverses of matrix products. This algebraic approach enables efficient calculations and generalization across different datasets and models within econometrics .

In matrix algebra, vectors and matrices are integral in defining the inner product. The inner product of two vectors is calculated by transposing one vector and multiplying it by another vector, which results in a scalar value. This involves row vectors from one matrix and column vectors from another, highlighting the role of vector alignment and dimensions in determining the matrix product .

Using restrictions in simultaneous equation systems aids in the model identification process. Restrictions, such as exclusion restrictions or linear homogeneous restrictions, guide the assignation of variables to specific equations, simplifying the complexity of multi-equation models. They ensure unique and meaningful solutions by allowing the determination of unknowns in structurally defined relationships, addressing identification challenges like multicollinearity and simultaneity biases .

The F-statistic in one-way ANOVA assesses the variation among class means by comparing it to the variation within classes. It effectively tests the homogeneity of class means by evaluating the sum of squares due to the difference between the class means and the overall mean, relative to the sum of squares due to variation within the classes .

Logistic regression plays a crucial role in addressing the inherent limitations of linear regression types, notably when the dependent variable is categorical. By employing the logistic function, it maps predicted probabilities to a range between 0 and 1, mitigating issues like heteroscedasticity and ensuring meaningful, bounded predictions. It successfully models binary, and sometimes multinomial outcomes, providing more robust, interpretable models under conditions where traditional linear regression assumptions do not hold .

The identifiability condition ensures that econometric models yield unique solutions to their parameters by setting necessary constraints, often through rank conditions and restrictions on variables. It prevents overspecification and ensures that each equation in a model can be isolated and estimated accurately. This condition is vital in complex multi-equation or structural models, where failure to meet identifiability can lead to indefinite parameter estimation, introducing ambiguity and bias .

Sums of squares are utilized in regression models to quantify different sources of variability within the data. Total sum of squares (TSS) is partitioned into explained sum of squares (ESS) and residual sum of squares (RSS). ESS measures the explained variability due to the regression model, while RSS measures the unexplained variability. This partitioning enables tests of significance, such as F-tests, to determine the contribution of certain variables or the overall fit of the model .

The concept of linear independence is crucial for the existence of an inverse matrix because only if the columns (and equivalently, rows) of a matrix are linearly independent can it span the full space necessary to form a basis. This ensures the matrix is non-singular and invertible, allowing for the unique solution of equations, as seen in the case of the matrix X'X in linear regression .

PANN 
Rg 
Ae 
a 
OT 
Loe 
i ile 
alt 
abit 
AU Rei 
Vea, 
faite 
o, 
ONIN deter 
LPP RANE 
danni
Johnston, J. (John), 
i 
mui 
2931 0135033
Te DUE 
by 
Mm 
| 
2 
as 
a 
: Ve 
Pp | 
7 
> | co 
mt 
jo 
CN) 
< 
eel = 
at 
[AO O
ECONOMETRIC 
METHODS 
Third Edition 
J. J ohnston 
University of California, Irvine 
McGraw-Hill Book Company 
New York 
St.L
IN MEMORY OF 
B. and J. 
This book was set in Times Roman by Science Typographers, Inc. 
The editors were Patricia A. Mitchel
Chapter 1 
il 
1-2 
1-3 
1-4 
1-5 
1-6 
Chapter 2 
ek 
39 
2-3 
2-4 
Ons 
2-6 
7 
Chapter 3 
3-1 
a) 
3-3 
3-4 
Chapter 4 
4-
iv CONTENTS 
4-4 
4-5 
4-6 
4-7 
Chapter 5 
5-1 
5 
53 
5-4 
Chapter 6 
6-1 
6-2 
6-3 
6-4 
6-5 
6-6 
Chapter 7 
Tel 
1p 
73
Chapter 11 
11-1 
11-2 
11-3 
Chapter 12 
A-1 
A-2 
A-3 
A-5 
A-6 
A-7 
A-8 
A-9 
A-10 
B-1 
B-3 
B-4 
B-5 
B-6 
B-7 
B-8 
Si
ae 
vache 
fate 
ee 
ie: Mt end ig Mi 
ie “whe 
Oe 
‘WAySee iér ne sil 
mi 
ole 
«MC 
NAA 
Mien 
ines 
SS 
’ 
To 
; ais OThii

You might also like