All Chap 1
All Chap 1
Marsh
PowerPoint Slides
for
Undergraduate Econometrics
by
Lawrence C. Marsh
To accompany: Undergraduate Econometrics
by R. Carter Hill, William E. Griffiths and George G. Judge
Publisher: John Wiley & Sons, 1997
Copyright 1996 Lawrence C. Marsh
The Role of
Econometrics
in Economic Analysis
Chapter 1
Copyright 1997 John Wiley & Sons, Inc. All rights reserved. Reproduction or translation of this work beyond
that permitted in Section 117 of the 1976 United States Copyright Act without the express written permission of the
copyright owner is unlawful. Request for further information should be addressed to the Permissions Department,
John Wiley & Sons, Inc. The purchaser may make back-up copies for his/her own use only and not for distribution
or resale. The Publisher assumes no responsibility for errors, omissions, or damages, caused by the use of these
programs or from the use of the information contained herein.
1.1
Copyright 1996 Lawrence C. Marsh
Using Information:
1. Information from economic theory.
2. Information from economic data.
The Role of Econometrics
1.2
Copyright 1996 Lawrence C. Marsh
Understanding Economic Relationships:
federal
budget
Dow-Jones
Stock Index
trade
deficit
Federal Reserve
Discount Rate
capital gains tax
rent
control
laws
short term
treasury bills
power of
labor unions
crime rate
inflation
unemployment
money supply
1.3
Copyright 1996 Lawrence C. Marsh
economic theory
economic data
}
economic
decisions
To use information effectively:
*Econometrics* helps us combine
economic theory and economic data .
Economic Decisions
1.4
Copyright 1996 Lawrence C. Marsh
Consumption, c, is some function of income, i :
c = f(i)
For applied econometric analysis
this consumption function must be
specified more precisely.
The Consumption Function
1.5
Copyright 1996 Lawrence C. Marsh
demand, q
d
, for an individual commodity:
q
d
= f( p, p
c
, p
s
, i )
supply, q
s
, of an individual commodity:
q
s
= f( p, p
c
, p
f
)
p = own price; p
c
= price of complements;
p
s
= price of substitutes; i = income
p = own price; p
c
= price of competitive products;
p
s
= price of substitutes; p
f
= price of factor inputs
demand
supply
1.6
Copyright 1996 Lawrence C. Marsh
Listing the variables in an economic relationship is not enough.
For effective policy we must know the amount of change
needed for a policy instrument to bring about the desired
effect:
How much ?
By how much should the Federal Reserve
raise interest rates to prevent inflation?
By how much can the price of football tickets
be increased and still fill the stadium?
1.7
Copyright 1996 Lawrence C. Marsh
Answering the How Much? question
Need to estimate parameters
that are both:
1. unknown
and
2. unobservable
1.8
Copyright 1996 Lawrence C. Marsh
Average or systematic behavior
over many individuals or many firms.
Not a single individual or single firm.
Economists are concerned with the
unemployment rate and not whether
a particular individual gets a job.
The Statistical Model
1.9
Copyright 1996 Lawrence C. Marsh
The Statistical Model
Actual vs. Predicted Consumption:
Actual = systematic part + random error
Systematic part provides prediction, f(i),
but actual will miss by random error, e.
Consumption, c, is function, f, of income, i, with error, e:
c = f(i) + e
1.10
Copyright 1996 Lawrence C. Marsh
c = f(i) + e
Need to define f(i) in some way.
To make consumption, c,
a linear function of income, i :
f(i) =
1
+
2
i
The statistical model then becomes:
c =
1
+
2
i + e
The Consumption Function
1.11
Copyright 1996 Lawrence C. Marsh
Dependent variable, y, is focus of study
(predict or explain changes in dependent variable).
Explanatory variables, X
2
and X
3
, help us explain
observed changes in the dependent variable.
y =
1
+
2
X
2
+
3
X
3
+ e
The Econometric Model
1.12
Copyright 1996 Lawrence C. Marsh
Statistical Models
Controlled (experimental)
vs.
Uncontrolled (observational)
Uncontrolled experiment (econometrics) explaining consump-
tion, y : price, X
2
, and income, X
3
, vary at the same time.
Controlled experiment (pure science) explaining mass, y :
pressure, X
2
, held constant when varying temperature, X
3
,
and vice versa.
1.13
Copyright 1996 Lawrence C. Marsh
Econometric model
economic model
economic variables and parameters.
statistical model
sampling process with its parameters.
data
observed values of the variables.
1.14
Copyright 1996 Lawrence C. Marsh
Uncertainty regarding an outcome.
Relationships suggested by economic theory.
Assumptions and hypotheses to be specified.
Sampling process including functional form.
Obtaining data for the analysis.
Estimation rule with good statistical properties.
Fit and test model using software package.
Analyze and evaluate implications of the results.
Problems suggest approaches for further research.
The Practice of Econometrics
1.15
Copyright 1996 Lawrence C. Marsh
Note: the textbook uses the following symbol
to mark sections with advanced material:
Skippy
1.16
Copyright 1996 Lawrence C. Marsh
Some Basic
Probability
Concepts
Chapter 2
Copyright 1997 John Wiley & Sons, Inc. All rights reserved. Reproduction or translation of this work beyond
that permitted in Section 117 of the 1976 United States Copyright Act without the express written permission of the
copyright owner is unlawful. Request for further information should be addressed to the Permissions Department,
John Wiley & Sons, Inc. The purchaser may make back-up copies for his/her own use only and not for distribution
or resale. The Publisher assumes no responsibility for errors, omissions, or damages, caused by the use of these
programs or from the use of the information contained herein.
2.1
Copyright 1996 Lawrence C. Marsh
random variable:
A variable whose value is unknown until it is observed.
The value of a random variable results from an experiment.
The term random variable implies the existence of some
known or unknown probability distribution defined over
the set of all possible values of that variable.
In contrast, an arbitrary variable does not have a
probability distribution associated with its values.
Random Variable
2.2
Copyright 1996 Lawrence C. Marsh
Controlled experiment values
of explanatory variables are chosen
with great care in accordance with
an appropriate experimental design.
Uncontrolled experiment values
of explanatory variables consist of
nonexperimental observations over
which the analyst has no control.
2.3
Copyright 1996 Lawrence C. Marsh
discrete random variable:
A discrete random variable can take only a finite
number of values, that can be counted by using
the positive integers.
Example: Prize money from the following
lottery is a discrete random variable:
first prize: $1,000
second prize: $50
third prize: $5.75
since it has only four (a finite number)
(count: 1,2,3,4) of possible outcomes:
$0.00; $5.75; $50.00; $1,000.00
Discrete Random Variable
2.4
Copyright 1996 Lawrence C. Marsh
continuous random variable:
A continuous random variable can take
any real value (not just whole numbers)
in at least one interval on the real line.
Examples:
Gross national product (GNP)
money supply
interest rates
price of eggs
household income
expenditure on clothing
Continuous Random Variable
2.5
Copyright 1996 Lawrence C. Marsh
A discrete random variable that is restricted
to two possible values (usually 0 and 1) is
called a dummy variable (also, binary or
indicator variable).
Dummy variables account for qualitative differences:
gender (0=male, 1=female),
race (0=white, 1=nonwhite),
citizenship (0=U.S., 1=not U.S.),
income class (0=poor, 1=rich).
Dummy Variable
2.6
Copyright 1996 Lawrence C. Marsh
A list of all of the possible values taken
by a discrete random variable along with
their chances of occurring is called a probability
function or probability density function (pdf).
die x f(x)
one dot 1 1/6
two dots 2 1/6
three dots 3 1/6
four dots 4 1/6
five dots 5 1/6
six dots 6 1/6
2.7
Copyright 1996 Lawrence C. Marsh
A discrete random variable X
has pdf, f(x), which is the probability
that X takes on the value x.
f(x) = P(X=x)
0 < f(x) < 1
If X takes on the n values: x
1
, x
2
, . . . , x
n
,
then f(x
1
) + f(x
2
)+. . .+f(x
n
) = 1.
Therefore,
2.8
Copyright 1996 Lawrence C. Marsh
Probability, f(x), for a discrete random
variable, X, can be represented by height:
0
1 2 3 X
number, X, on Deans List of three roommates
f(x)
0.2
0.4
0.1
0.3
2.9
Copyright 1996 Lawrence C. Marsh
A continuous random variable uses
area under a curve rather than the
height, f(x), to represent probability:
f(x)
X $34,000
$55,000
. .
per capita income, X, in the United States
0.1324
0.8676
red area
green area
2.10
Copyright 1996 Lawrence C. Marsh
Since a continuous random variable has an
uncountably infinite number of values,
the probability of one occurring is zero.
P [ X = a ] = P [ a < X < a ] = 0
Probability is represented by area.
Height alone has no area.
An interval for X is needed to get
an area under the curve.
2.11
Copyright 1996 Lawrence C. Marsh
P [ a < X < b ] = f(x) dx
b
a
The area under a curve is the integral of
the equation that generates the curve:
For continuous random variables it is the
integral of f(x), and not f(x) itself, which
defines the area and, therefore, the probability.
2.12
Copyright 1996 Lawrence C. Marsh
n
Rule 2:
ax
i
= a
x
i
i = 1 i = 1
n
Rule 1:
x
i
= x
1
+ x
2
+ . . . + x
n
i = 1
n
Rule 3:
(x
i
+ y
i
) =
x
i
+
y
i
i = 1
i = 1 i = 1
n n n
Note that summation is a linear operator
which means it operates term by term.
Rules of Summation
2.13
Copyright 1996 Lawrence C. Marsh
Rule 4:
(ax
i
+ by
i
) = a
x
i
+ b
y
i
i = 1
i = 1 i = 1
n n n
Rules of Summation (continued)
Rule 5: x =
x
i
=
i = 1
n
n
1
x
1
+ x
2
+ . . . + x
n
n
The definition of x as given in Rule 5 implies
the following important fact:
(x
i
x) = 0
i = 1
n
2.14
Copyright 1996 Lawrence C. Marsh
Rule 6:
f(x
i
) = f(x
1
) + f(x
2
) + . . . + f(x
n
)
i = 1
n
Notation:
f(x
i
) =
f(x
i
)
=
f(x
i
)
n
x i i = 1
n
Rule 7:
f(x
i
,y
j
) =
[ f(x
i
,y
1
) + f(x
i
,y
2
)+. . .+ f(x
i
,y
m
)]
i = 1
i = 1
n m
j = 1
The order of summation does not matter :
f(x
i
,y
j
) =
f(x
i
,y
j
)
i = 1
n m
j = 1 j = 1
m n
i = 1
Rules of Summation (continued)
2.15
Copyright 1996 Lawrence C. Marsh
The mean or arithmetic average of a
random variable is its mathematical
expectation or expected value, EX.
The Mean of a Random Variable
2.16
Copyright 1996 Lawrence C. Marsh
Expected Value
There are two entirely different, but mathematically
equivalent, ways of determining the expected value:
1. Empirically:
The expected value of a random variable, X,
is the average value of the random variable in an
infinite number of repetitions of the experiment.
In other words, draw an infinite number of samples,
and average the values of X that you get.
2.17
Copyright 1996 Lawrence C. Marsh
Expected Value
2. Analytically:
The expected value of a discrete random
variable, X, is determined by weighting all
the possible values of X by the corresponding
probability density function values, f(x), and
summing them up.
E[X] = x
1
f(x
1
) + x
2
f(x
2
) + . . . + x
n
f(x
n
)
In other words:
2.18
Copyright 1996 Lawrence C. Marsh
In the empirical case when the
sample goes to infinity the values
of X occur with a frequency
equal to the corresponding f(x)
in the analytical expression.
As sample size goes to infinity, the
empirical and analytical methods
will produce the same value.
Empirical vs. Analytical
2.19
Copyright 1996 Lawrence C. Marsh
x =
x
i
n
i = 1
where n is the number of sample observations.
Empirical (sample) mean:
E[X] =
x
i
f(x
i
)
i = 1
n
where n is the number of possible values of x
i
.
Analytical mean:
Notice how the meaning of n changes.
2.20
Copyright 1996 Lawrence C. Marsh
E X = x
i
f(x
i
)
i=1
n
The expected value of X-squared:
E X = x
i
f(x
i
)
i=1
n
2
2
It is important to notice that f(x
i
) does not change!
The expected value of X-cubed:
E X = x
i
f(x
i
)
i=1
n
3
3
The expected value of X:
2.21
Copyright 1996 Lawrence C. Marsh
EX = 0 (.1) + 1 (.3) + 2 (.3) + 3 (.2) + 4 (.1)
2
EX
= 0 (.1) + 1 (.3) + 2 (.3) + 3 (.2) + 4 (.1)
2 2 2 2 2
= 1.9
= 0 + .3 + 1.2 + 1.8 + 1.6
= 4.9
3
EX = 0 (.1) + 1 (.3) + 2 (.3) + 3 (.2) +4 (.1)
3 3 3 3 3
= 0 + .3 + 2.4 + 5.4 + 6.4
= 14.5
2.22
Copyright 1996 Lawrence C. Marsh
E [g(X)] = g(x
i
)
f(x
i
)
n
i = 1
g(X) = g
1
(X) + g
2
(X)
E [g(X)] = [g
1
(x
i
) + g
2
(x
i
)]
f(x
i
)
n
i = 1
E [g(X)] = g
1
(x
i
)
f(x
i
) + g
2
(x
i
)
f(x
i
)
n
i = 1
n
i = 1
E [g(X)] = E [g
1
(X)]
+ E [g
2
(X)]
2.23
Copyright 1996 Lawrence C. Marsh
Adding and Subtracting
Random Variables
E(X-Y) = E(X) - E(Y)
E(X+Y) = E(X) + E(Y)
2.24
Copyright 1996 Lawrence C. Marsh
E(X+a) = E(X) + a
Adding a constant to a variable will
add a constant to its expected value:
Multiplying by constant will multiply
its expected value by that constant:
E(bX) = b E(X)
2.25
Copyright 1996 Lawrence C. Marsh
var(X) = average squared deviations
around the mean of X.
var(X) = expected value of the squared deviations
around the expected value of X.
var(X) = E [(X - EX) ]
2
Variance
2.26
Copyright 1996 Lawrence C. Marsh
var(X) = E [(X - EX) ]
= E [X - 2XEX + (EX) ]
2
2
2
= E(X ) - 2 EX EX + E (EX)
2
2
= E(X ) - 2 (EX) + (EX)
2 2 2
= E(X ) - (EX)
2
2
var(X) = E [(X - EX) ]
2
var(X) = E(X ) - (EX)
2
2
2.27
Copyright 1996 Lawrence C. Marsh
variance of a discrete
random variable, X:
standard deviation is square root of variance
var ( X) =
(x
i
- EX)
2
f(x
i
)
i = 1
n
2.28
Copyright 1996 Lawrence C. Marsh
x
i
f(x
i
) (x
i
- EX) (x
i
- EX) f(x
i
)
2 .1 2 - 4.3 = -2.3 5.29 (.1) = .529
3 .3 3 - 4.3 = -1.3 1.69 (.3) = .507
4 .1 4 - 4.3 = - .3 .09 (.1) = .009
5 .2 5 - 4.3 = .7 .49 (.2) = .098
6 .3 6 - 4.3 = 1.7 2.89 (.3) = .867
x
i
f(x
i
) = .2 + .9 + .4 + 1.0 + 1.8 = 4.3
(x
i
- EX) f(x
i
) = .529 + .507 + .009 + .098 + .867
= 2.01
2
2
calculate the variance for a
discrete random variable, X:
i = 1
n
n
i = 1
2.29
Copyright 1996 Lawrence C. Marsh
Z = a + cX
var(Z) = var(a + cX)
= E [(a+cX) - E(a+cX)]
= c var(X)
2
2
var(a + cX) = c var(X)
2
2.30
Copyright 1996 Lawrence C. Marsh
A joint probability density function,
f(x,y), provides the probabilities
associated with the joint occurrence
of all of the possible pairs of X and Y.
Joint pdf
2.31
Copyright 1996 Lawrence C. Marsh
college grads
in household
.15
.05
.45
.35
joint pdf
f(x,y)
Y = 1 Y = 2
vacation
homes
owned
X = 0
X = 1
Survey of College City, NY
f(0,1) f(0,2)
f(1,1)
f(1,2)
2.32
Copyright 1996 Lawrence C. Marsh
E[g(X,Y)] = g(x
i
,y
j
) f(x
i
,y
j
)
i
j
E(XY) = (0)(1)(.45)+(0)(2)(.15)+(1)(1)(.05)+(1)(2)(.35)=.75
E(XY) = x
i
y
j
f(x
i
,y
j
)
i
j
Calculating the expected value of
functions of two random variables.
2.33
Copyright 1996 Lawrence C. Marsh
The marginal probability density functions,
f(x) and f(y), for discrete random variables,
can be obtained by summing over the f(x,y)
with respect to the values of Y to obtain f(x)
with respect to the values of X to obtain f(y).
f(x
i
) = f(x
i
,y
j
) f(y
j
) = f(x
i
,y
j
)
i j
Marginal pdf
2.34
Copyright 1996 Lawrence C. Marsh
.15
.05
.45
.35
marginal
Y = 1 Y = 2
X = 0
X = 1
.60
.40
.50 .50
f(X = 1)
f(X = 0)
f(Y = 1)
f(Y = 2)
marginal
pdf for Y:
marginal
pdf for X:
2.35
Copyright 1996 Lawrence C. Marsh
The conditional probability density
functions of X given Y=y , f(x
|
y),
and of Y given X=x , f(y
|
x),
are obtained by dividing f(x,y) by f(y)
to get f(x
|
y) and by f(x) to get f(y
|
x).
f(x
|
y) =
f(y
|
x) =
f(x,y)
f(x,y)
f(y)
f(x)
Conditional pdf
2.36
Copyright 1996 Lawrence C. Marsh
.15
.05
.45
.35
conditonal
Y = 1
Y = 2
X = 0
X = 1
.60
.40
.50 .50
.25 .75
.875 .125
.90
.10 .70
.30
f(Y=2
|
X= 0)=.25
f(Y=1
|
X = 0)=.75
f(Y=2
|
X = 1)=.875
f(X=0
|
Y=2)=.30
f(X=1
|
Y=2)=.70
f(X=0
|
Y=1)=.90
f(X=1
|
Y=1)=.10
f(Y=1
|
X = 1)=.125
2.37
Copyright 1996 Lawrence C. Marsh
X and Y are independent random
variables if their joint pdf, f(x,y),
is the product of their respective
marginal pdfs, f(x) and f(y) .
f(x
i
,y
j
) = f(x
i
) f(y
j
)
for independence this must hold for all pairs of i and j
Independence
2.38
Copyright 1996 Lawrence C. Marsh
.15
.05
.45
.35
not independent
Y = 1 Y = 2
X = 0
X = 1
.60
.40
.50 .50
f(X = 1)
f(X = 0)
f(Y = 1)
f(Y = 2)
marginal
pdf for Y:
marginal
pdf for X:
.50x.60=.30
.50x.60=.30
.50x.40=.20 .50x.40=.20
The calculations
in the boxes show
the numbers
required to have
independence.
2.39
Copyright 1996 Lawrence C. Marsh
The covariance between two random
variables, X and Y, measures the
linear association between them.
cov(X,Y) = E[(X - EX)(Y-EY)]
Note that variance is a special case of covariance.
cov(X,X) = var(X) = E[(X - EX) ]
2
Covariance
2.40
Copyright 1996 Lawrence C. Marsh
cov(X,Y) = E [(X - EX)(Y-EY)]
= E [XY - X EY - Y EX + EX EY]
= E(XY) - 2 EX EY + EX EY
= E(XY) - EX EY
cov(X,Y) = E [(X - EX)(Y-EY)]
cov(X,Y) = E(XY) - EX EY
= E(XY) - EX EY - EY EX + EX EY
2.41
Copyright 1996 Lawrence C. Marsh
.15
.05
.45
.35
Y = 1 Y = 2
X = 0
X = 1
.60
.40
.50 .50
EX=0(.60)+1(.40)=.40
EY=1(.50)+2(.50)=1.50
E(XY) = (0)(1)(.45)+(0)(2)(.15)+(1)(1)(.05)+(1)(2)(.35)=.75
EX EY = (.40)(1.50) = .60
cov(X,Y) = E(XY) - EX EY
= .75 - (.40)(1.50)
= .75 - .60
= .15
covariance
2.42
Copyright 1996 Lawrence C. Marsh
The correlation between two random
variables X and Y is their covariance
divided by the square roots of their
respective variances.
Correlation is a pure number falling between -1 and 1.
cov(X,Y)
(X,Y) =
var(X) var(Y)
Correlation
2.43
Copyright 1996 Lawrence C. Marsh
.15
.05
.45
.35
Y = 1 Y = 2
X = 0
X = 1
.60
.40
.50 .50
EX=.40
EY=1.50
cov(X,Y) = .15
correlation
EX=0(.60)+1(.40)=.40
2 2 2
var(X) = E(X ) - (EX)
= .40 - (.40)
= .24
2 2
2
EY=1(.50)+2(.50)
= .50 + 2.0
= 2.50
2
2 2
var(Y) = E(Y ) - (EY)
= 2.50 - (1.50)
= .25
2 2
2
(X,Y) =
cov(X,Y)
var(X) var(Y)
(X,Y) = .61
2.44
Copyright 1996 Lawrence C. Marsh
Independent random variables
have zero covariance and,
therefore, zero correlation.
The converse is not true.
Zero Covariance & Correlation
2.45
Copyright 1996 Lawrence C. Marsh
The expected value of the weighted sum
of random variables is the sum of the
expectations of the individual terms.
Since expectation is a linear operator,
it can be applied term by term.
E[c
1
X + c
2
Y] = c
1
EX + c
2
EY
E[c
1
X
1
+...+ c
n
X
n
] = c
1
EX
1
+...+ c
n
EX
n
In general, for random variables X
1
, . . . , X
n
:
2.46
Copyright 1996 Lawrence C. Marsh
The variance of a weighted sum of random
variables is the sum of the variances, each times
the square of the weight, plus twice the covariances
of all the random variables times the products of
their weights.
var(c
1
X + c
2
Y)=c
1
var(X)+c
2
var(Y) + 2c
1
c
2
cov(X,Y)
2 2
var(c
1
X c
2
Y) = c
1
var(X)+c
2
var(Y) 2c
1
c
2
cov(X,Y)
2 2
Weighted sum of random variables:
Weighted difference of random variables:
2.47
Copyright 1996 Lawrence C. Marsh
The Normal Distribution
Y ~ N(,
2
)
f(y) =
2
2
1
exp
y
f(y)
2
2
(y - )
2
-
2.48
Copyright 1996 Lawrence C. Marsh
The Standardized Normal
Z ~ N(0,1)
f(z) =
2
1
exp
2
z
2
-
Z = (y - )/
2.49
Copyright 1996 Lawrence C. Marsh
P [ Y > a ] = P > = P Z >
a -
a -
Y -
y
f(y)
a
Y ~ N(,
2
)
2.50
Copyright 1996 Lawrence C. Marsh
P [ a < Y < b ] = P < <
= P < Z <
a - Y -
b -
a -
b -
y
f(y)
a
Y ~ N(,
2
)
b
2.51
Copyright 1996 Lawrence C. Marsh
Y
1
~ N(
1
,
1
2
), Y
2
~ N(
2
,
2
2
), . . . , Y
n
~ N(
n
,
n
2
)
W = c
1
Y
1
+ c
2
Y
2
+ . . . + c
n
Y
n
Linear combinations of jointly
normally distributed random variables
are themselves normally distributed.
W ~ N[ E(W), var(W) ]
2.52
Copyright 1996 Lawrence C. Marsh
mean: E[V] = E[
(m)
] = m
If Z
1
, Z
2
, . . . , Z
m
denote m independent
N(0,1) random variables, and
V = Z
1
+ Z
2
+ . . . + Z
m
, then V ~
(m)
2 2
2 2
V is chi-square with m degrees of freedom.
Chi-Square
variance: var[V] = var[
(m)
] = 2m
If Z
1
, Z
2
, . . . , Z
m
denote m independent
N(0,1) random variables, and
V = Z
1
+ Z
2
+ . . . + Z
m
, then V ~
(m)
2 2
2
2
V is chi-square with m degrees of freedom.
2
2
2.53
Copyright 1996 Lawrence C. Marsh
mean: E[t] = E[t
(m)
] = 0 symmetric about zero
variance: var[t] = var[t
(m)
] = m / (m2)
If Z ~ N(0,1) and V ~
(m)
and if Z and V
are independent then,
~ t
(m)
t is student-t with m degrees of freedom.
2
t =
Z
V
m
Student - t
2.54
Copyright 1996 Lawrence C. Marsh
If V
1
~
(m
1
)
and V
2
~
(m
2
)
and if V
1
and V
2
are independent, then
~ F
(m
1
,m
2
)
F is an F statistic with m
1
numerator
degrees of freedom and m
2
denominator
degrees of freedom.
2
F =
V
1
m
1
V
2
m
2
2
F Statistic
2.55
Copyright 1996 Lawrence C. Marsh
The Simple Linear
Regression
Model
Chapter 3
Copyright 1997 John Wiley & Sons, Inc. All rights reserved. Reproduction or translation of this work beyond
that permitted in Section 117 of the 1976 United States Copyright Act without the express written permission of the
copyright owner is unlawful. Request for further information should be addressed to the Permissions Department,
John Wiley & Sons, Inc. The purchaser may make back-up copies for his/her own use only and not for distribution
or resale. The Publisher assumes no responsibility for errors, omissions, or damages, caused by the use of these
programs or from the use of the information contained herein.
3.1
Copyright 1996 Lawrence C. Marsh
1. Estimate a relationship among economic
variables, such as y = f(x).
2. Forecast or predict the value of one
variable, y, based on the value of
another variable, x.
Purpose of Regression Analysis
3.2
Copyright 1996 Lawrence C. Marsh
Weekly Food Expenditures
y = dollars spent each week on food items.
x = consumers weekly income.
The relationship between x and the expected
value of y , given x, might be linear:
E(y|x) =
1
+
2
x
3.3
Copyright 1996 Lawrence C. Marsh
f(y|x=480)
f(y|x=480)
y
y|x=480
Figure 3.1a Probability Distribution f(y|x=480)
of Food Expenditures if given income x=$480.
3.4
Copyright 1996 Lawrence C. Marsh
f(y|x)
f(y|x=480) f(y|x=800)
y
y|x=480
y|x=800
Figure 3.1b Probability Distribution of Food
Expenditures if given income x=$480 and x=$800.
3.5
Copyright 1996 Lawrence C. Marsh
{
1
x
E(y|x)
E(y|x)
Average
Expenditure
x (income)
E(y|x)=
1
+
2
x
2
=
E(y|x)
x
Figure 3.2 The Economic Model: a linear relationship
between avearage expenditure on food and income.
3.6
Copyright 1996 Lawrence C. Marsh
.
.
x
t x
1
=480 x
2
=800
y
t
f(y
t
)
Figure 3.3. The probability density function
for y
t
at two levels of household income, x
t
e
x
p
e
n
d
i
t
u
r
e
Homoskedastic Case
income
3.7
Copyright 1996 Lawrence C. Marsh
.
x
t
x
1
x
2
y
t f(y
t
)
Figure 3.3+. The variance of y
t
increases
as household income, x
t
, increases.
e
x
p
e
n
d
i
t
u
r
e
Heteroskedastic Case
x
3
.
.
income
3.8
Copyright 1996 Lawrence C. Marsh
Assumptions of the Simple Linear
Regression Model - I
1. The average value of y, given x, is given by
the linear regression:
E(y) =
1
+
2
x
2. For each value of x, the values of y are
distributed around their mean with variance:
var(y) =
2
3. The values of y are uncorrelated, having zero
covariance and thus no linear relationship:
cov(y
i
,y
j
) = 0
4. The variable x must take at least two different
values, so that x c, where c is a constant.
3.9
Copyright 1996 Lawrence C. Marsh
5. (optional) The values of y are normally
distributed about their mean for each
value of x:
y ~ N [(
1
+
2
x),
2
]
One more assumption that is often used in
practice but is not required for least squares:
3.10
Copyright 1996 Lawrence C. Marsh
The Error Term
y is a random variable composed of two parts:
I. Systematic component: E(y) =
1
+
2
x
This is the mean of y.
II. Random component: e = y - E(y)
= y -
1
-
2
x
This is called the random error.
Together E(y) and e form the model:
y =
1
+
2
x + e
3.11
Copyright 1996 Lawrence C. Marsh
Figure 3.5 The relationship among y, e and
the true regression line.
.
.
.
.
y
4
y
1
y
2
y
3
x
1
x
2
x
3
x
4
}
}
{
{
e
1
e
2
e
3
e
4
E(y) =
1
+
2
x
x
y 3.12
Copyright 1996 Lawrence C. Marsh
}
.
}
.
.
.
y
4
y
1
y
2
y
3
x
1
x
2
x
3
x
4
{
{
e
1
e
2
e
3
e
4
x
y
Figure 3.7a The relationship among y, e and
the fitted regression line.
^
y = b
1
+ b
2
x
^
.
.
.
.
y
1
y
2
y
3
y
4
^
^
^
^
^
^
^
^
3.13
Copyright 1996 Lawrence C. Marsh
{
{
.
.
.
.
.
y
4
y
1
y
2 y
3
x
1
x
2
x
3
x
4
x
y
Figure 3.7b The sum of squared residuals
from any other line will be larger.
y = b
1
+ b
2
x
^
.
.
.
y
1
^
y
3
^
y
4
^
y = b
1
+ b
2
x
^* * *
*
e
1
^
*
e
2
^*
y
2
^*
e
3
^*
*
e
4
^
*
*
{
{
3.14
Copyright 1996 Lawrence C. Marsh
f(.)
f(e) f(y)
Figure 3.4 Probability density function for e and y
0
1
+
2
x
3.15
Copyright 1996 Lawrence C. Marsh
The Error Term Assumptions
1. The value of y, for each value of x, is
y =
1
+
2
x + e
2. The average value of the random error e is:
E(e) = 0
3. The variance of the random error e is:
var(e) =
2
= var(y)
4. The covariance between any pair of es is:
cov(e
i
,e
j
) = cov(y
i
,y
j
) = 0
5. x must take at least two different values so that
x c, where c is a constant.
6. e is normally distributed with mean 0, var(e)=
2
(optional) e ~ N(0,
2
)
3.16
Copyright 1996 Lawrence C. Marsh
Unobservable Nature
of the Error Term
1. Unspecified factors / explanatory variables,
not in the model, may be in the error term.
2. Approximation error is in the error term if
relationship between y and x is not exactly
a perfectly linear relationship.
3. Strictly unpredictable random behavior that
may be unique to that observation is in error.
3.17
Copyright 1996 Lawrence C. Marsh
Population regression values:
y
t
=
1
+
2
x
t
+ e
t
Population regression line:
E(y
t
|x
t
) =
1
+
2
x
t
Sample regression values:
y
t
= b
1
+ b
2
x
t
+ e
t
Sample regression line:
y
t
= b
1
+ b
2
x
t
^
^
3.18
Copyright 1996 Lawrence C. Marsh
y
t
=
1
+
2
x
t
+ e
t
Minimize error sum of squared deviations:
S(
1
,
2
) =
(y
t
-
1
-
2
x
t
)
2
(3.3.4)
t=1
T
e
t
= y
t
-
1
-
2
x
t
3.19
Copyright 1996 Lawrence C. Marsh
Minimize w. r. t.
1
and
2
:
S(
1
,
2
) =
(y
t
-
1
-
2
x
t
)
2
(3.3.4)
t =1
T
= - 2
(y
t
-
1
-
2
x
t
)
= - 2
x
t
(y
t
-
1
-
2
x
t
)
S(.)
1
S(.)
2
Set each of these two derivatives equal to zero and
solve these two equations for the two unknowns:
1
2
3.20
Copyright 1996 Lawrence C. Marsh
S(.)
S(.)
i
b
i
.
.
.
Minimize w. r. t.
1
and
2
:
S(.) =
(y
t
-
1
-
2
x
t
)
2
t =1
T
S(.)
i
<
0
S(.)
i
>
0
S(.)
i
=0
3.21
Copyright 1996 Lawrence C. Marsh
To minimize S(.), you set the two
derivatives equal to zero to get:
= - 2
(y
t
- b
1
- b
2
x
t
) = 0
= - 2
x
t
(y
t
- b
1
- b
2
x
t
) = 0
S(.)
1
S(.)
2
When these two terms are set to zero,
1
and
2
become b
1
and b
2
because they no longer
represent just any value of
1
and
2
but the special
values that correspond to the minimum of S(.) .
3.22
Copyright 1996 Lawrence C. Marsh
- 2
(y
t
- b
1
- b
2
x
t
) = 0
- 2
x
t
(y
t
- b
1
- b
2
x
t
) = 0
y
t
- Tb
1
- b
2
x
t
= 0
x
t
y
t
- b
1
x
t
- b
2
x
t
= 0
2
Tb
1
+ b
2
x
t
=
y
t
b
1
x
t
+ b
2
x
t
=
x
t
y
t
2
3.23
Copyright 1996 Lawrence C. Marsh
Solve for b
1
and b
2
using definitions of x and y
Tb
1
+ b
2
x
t
=
y
t
b
1
x
t
+ b
2
x
t
=
x
t
y
t
2
T
x
t
y
t
-
x
t
y
t
T
x
t
- (
x
t
)
2 2
b
2
=
b
1
= y - b
2
x
3.24
Copyright 1996 Lawrence C. Marsh
elasticities
percentage change in y
percentage change in x
=
=
x/x
y/y
=
y x
x y
Using calculus, we can get the elasticity at a point:
= lim
=
y x
x y
y x
x y x 0
3.25
Copyright 1996 Lawrence C. Marsh
E(y) =
1
+
2
x
E(y)
x
=
2
applying elasticities
E(y)
x
=
2
=
E(y)
x
E(y)
x
3.26
Copyright 1996 Lawrence C. Marsh
estimating elasticities
y
x
= b
2
=
y
x
y
x
^
y
t
= b
1
+ b
2
x
t
= 4 + 1.5 x
t
^
x
= 8 = average number of years of experience
y
= $10 = average wage rate
= 1.5 = 1.2
8
10
= b
2
y
x ^
3.27
Copyright 1996 Lawrence C. Marsh
Prediction
y
t
= 4 + 1.5 x
t
^
Estimated regression equation:
x
t
= years of experience
y
t
= predicted wage rate
^
If x
t
= 2 years, then y
t
= $7.00 per hour.
^
If x
t
= 3 years, then y
t
= $8.50 per hour.
^
3.28
Copyright 1996 Lawrence C. Marsh
log-log models
ln(y) =
1
+
2
ln(x)
ln(y)
x
ln(x)
x
=
2
y
x
=
2
1
y
x
x
1
x
3.29
Copyright 1996 Lawrence C. Marsh
y
x
=
2
1
y
x
x
1
x
=
2
y
x
x
y
elasticity of y with respect to x:
=
2
y
x
x
y
=
3.30
Copyright 1996 Lawrence C. Marsh
Properties of
Least Squares
Estimators
Chapter 4
Copyright 1997 John Wiley & Sons, Inc. All rights reserved. Reproduction or translation of this work beyond
that permitted in Section 117 of the 1976 United States Copyright Act without the express written permission of the
copyright owner is unlawful. Request for further information should be addressed to the Permissions Department,
John Wiley & Sons, Inc. The purchaser may make back-up copies for his/her own use only and not for distribution
or resale. The Publisher assumes no responsibility for errors, omissions, or damages, caused by the use of these
programs or from the use of the information contained herein.
4.1
Copyright 1996 Lawrence C. Marsh
y
t
= household weekly food expenditures
Simple Linear Regression Model
y
t
=
1
+
2
x
t
+
t
x
t
= household weekly income
For a given level of x
t
, the expected
level of food expenditures will be:
E(y
t
|x
t
) =
1
+
2
x
t
4.2
Copyright 1996 Lawrence C. Marsh
1. y
t
=
1
+
2
x
t
+
t
2. E(
t
) = 0 <=> E(y
t
) =
1
+
2
x
t
3. var(
t
) =
2
= var(y
t
)
4. cov(
i
,
j
) = cov(y
i
,y
j
) = 0
5. x
t
c for every observation
6.
t
~N(0,
2
) <=> y
t
~N(
1
+
2
x
t
,
2
)
Assumptions of the Simple
Linear Regression Model
4.3
Copyright 1996 Lawrence C. Marsh
The population parameters
1
and
2
are unknown population constants.
The formulas that produce the
sample estimates b
1
and b
2
are
called the estimators of
1
and
2
.
When b
0
and b
1
are used to represent
the formulas rather than specific values,
they are called estimators of
1
and
2
which are random variables because
they are different from sample to sample.
4.4
Copyright 1996 Lawrence C. Marsh
If the least squares estimators b
0
and b
1
are random variables, then what are their
their means, variances, covariances and
probability distributions?
Compare the properties of alternative
estimators to the properties of the
least squares estimators.
Estimators are Random Variables
( estimates are not )
4.5
Copyright 1996 Lawrence C. Marsh
The Expected Values of b1 and b2
The least squares formulas (estimators)
in the simple regression case:
b
2
=
Tx
t
y
t
- x
t
y
t
Tx
t
-(x
t
)
2 2
b
1
= y - b
2
x
where y = y
t
/ T and x = x
t
/ T
(3.3.8a)
(3.3.8b)
4.6
Copyright 1996 Lawrence C. Marsh
Substitute in y
t
=
1
+
2
x
t
+
t
to get:
b
2
=
2
+
Tx
t
t
- x
t
t
Tx
t
-(x
t
)
2 2
The mean of b
2
is:
Eb
2
=
2
+
Tx
t
E
t
- x
t
E
t
Tx
t
-(x
t
)
2 2
Since E
t
= 0, then Eb
2
=
2
.
4.7
Copyright 1996 Lawrence C. Marsh
The result Eb
2
=
2
means that
the distribution of b
2
is centered at
2
.
Since the distribution of b
2
is centered at
2
,we say that
b
2
is an unbiased estimator of
2
.
An Unbiased Estimator
4.8
Copyright 1996 Lawrence C. Marsh
The unbiasedness result on the
previous slide assumes that we
are using the correct model.
If the model is of the wrong form
or is missing important variables,
then E
t
0, then Eb
2
2
.
Wrong Model Specification
4.9
Copyright 1996 Lawrence C. Marsh
Unbiased Estimator of the Intercept
In a similar manner, the estimator b
1
of the intercept or constant term can be
shown to be an unbiased estimator of
1
when the model is correctly specified.
Eb
1
=
1
4.10
Copyright 1996 Lawrence C. Marsh
b
2
=
Tx
t
y
t
x
t
y
t
Tx
t
(x
t
)
2
2
(3.3.8a)
(4.2.6)
Equivalent expressions for b
2
:
Expand and multiply top and bottom by T:
b
2
=
(x
t
x )(y
t
y )
(x
t
x
)
2
4.11
Copyright 1996 Lawrence C. Marsh
Variance of b
2
Given that both y
t
and
t
have variance
2
,
the variance of the estimator b
2
is:
b
2
is a function of the y
t
values but
var(b
2
) does not involve y
t
directly.
(x
t
x)
2
2
var(b
2
) =
4.12
Copyright 1996 Lawrence C. Marsh
Variance of b
1
(x
t
x)
2
var(b
1
) =
2
x
t
2
the variance of the estimator b
1
is:
b
1
= y b
2
x Given
4.13
Copyright 1996 Lawrence C. Marsh
Covariance of b
1
and b
2
(x
t
x)
2
cov(b
1
,b
2
) =
2
x
If x = 0, slope can change without affecting
the variance.
4.14
Copyright 1996 Lawrence C. Marsh
What factors determine
variance and covariance ?
1.
2
: uncertainty about y
t
values uncertainty about
b
1
, b
2
and their relationship.
2. The more spread out the x
t
values are then the more
confidence we have in b
1
, b
2
, etc.
3. The larger the sample size, T, the smaller the
variances and covariances.
4. The variance b
1
is large when the (squared) x
t
values
are far from zero (in either direction).
5. Changing the slope, b
2
, has no effect on the intercept,
b
1
, when the sample mean is zero. But if sample
mean is positive, the covariance between b
1
and
b
2
will be negative, and vice versa.
4.15
Copyright 1996 Lawrence C. Marsh
Gauss-Markov Theorm
Under the first five assumptions of the
simple, linear regression model, the
ordinary least squares estimators b
1
and b
2
have the smallest variance of
all linear and unbiased estimators of
1
and
2
. This means that b
1
and b
2
are the Best Linear Unbiased Estimators
(BLUE) of
1
and
2
.
4.16
Copyright 1996 Lawrence C. Marsh
implications of Gauss-Markov
1. b
1
and b
2
are best within the class
of linear and unbiased estimators.
2. Best means smallest variance
within the class of linear/unbiased.
3. All of the first five assumptions must
hold to satisfy Gauss-Markov.
4. Gauss-Markov does not require
assumption six: normality.
5. G-Markov is not based on the least
squares principle but on b
1
and b
2
.
4.17
Copyright 1996 Lawrence C. Marsh
G-Markov implications (continued)
6. If we are not satisfied with restricting
our estimation to the class of linear and
unbiased estimators, we should ignore
the Gauss-Markov Theorem and use
some nonlinear and/or biased estimator
instead. (Note: a biased or nonlinear
estimator could have smaller variance
than those satisfying Gauss-Markov.)
7. Gauss-Markov applies to the b
1
and b
2
estimators and not to particular sample
values (estimates) of b
1
and b
2
.
4.18
Copyright 1996 Lawrence C. Marsh
Probability Distribution
of Least Squares Estimators
b
2
~ N
2
,
(x
t
x)
2
2
b
1
~ N
1
,
(x
t
x)
2
2
x
t
2
4.19
Copyright 1996 Lawrence C. Marsh
y
t
and
t
normally distributed
The least squares estimator of
2
can be
expressed as a linear combination of y
t
s:
b
2
= w
t
y
t
b
1
= y b
2
x
(x
t
x)
2
where w
t
=
(x
t
x)
This means that b
1
and b
2
are normal since
linear combinations of normals are normal.
4.20
Copyright 1996 Lawrence C. Marsh
normally distributed under
The Central Limit Theorem
If the first five Gauss-Markov assumptions
hold, and sample size, T, is sufficiently large,
then the least squares estimators, b
1
and b
2
,
have a distribution that approximates the
normal distribution with greater accuracy
the larger the value of sample size, T.
4.21
Copyright 1996 Lawrence C. Marsh
Consistency
We would like our estimators, b
1
and b
2
, to collapse
onto the true population values,
1
and
2
, as
sample size, T, goes to infinity.
One way to achieve this consistency property is
for the variances of b
1
and b
2
to go to zero as T
goes to infinity.
Since the formulas for the variances of the least
squares estimators b
1
and b
2
show that their
variances do, in fact, go to zero, then b
1
and b
2
,
are consistent estimators of
1
and
2
.
4.22
Copyright 1996 Lawrence C. Marsh
Estimating the variance
of the error term,
2
e
t
= y
t
b
1
b
2
x
t
^
e
t
^
t =1
T
2
T 2
2
=
2
is an unbiased estimator of
2
^
^
4.23
Copyright 1996 Lawrence C. Marsh
The Least Squares
Predictor, y
o
^
Given a value of the explanatory
variable, X
o
, we would like to predict
a value of the dependent variable, y
o
.
The least squares predictor is:
y
o
= b
1
+ b
2
x
o
(4.7.2)
^
4.24
Copyright 1996 Lawrence C. Marsh
Inference
in the Simple
Regression Model
Chapter 5
Copyright 1997 John Wiley & Sons, Inc. All rights reserved. Reproduction or translation of this work beyond
that permitted in Section 117 of the 1976 United States Copyright Act without the express written permission of the
copyright owner is unlawful. Request for further information should be addressed to the Permissions Department,
John Wiley & Sons, Inc. The purchaser may make back-up copies for his/her own use only and not for distribution
or resale. The Publisher assumes no responsibility for errors, omissions, or damages, caused by the use of these
programs or from the use of the information contained herein.
5.1
Copyright 1996 Lawrence C. Marsh
1. y
t
=
1
+
2
x
t
+
t
2. E(
t
) = 0 <=> E(y
t
) =
1
+
2
x
t
3. var(
t
) =
2
= var(y
t
)
4. cov(
i
,
j
) = cov(y
i
,y
j
) = 0
5. x
t
c for every observation
6.
t
~N(0,
2
) <=> y
t
~N(
1
+
2
x
t
,
2
)
Assumptions of the Simple
Linear Regression Model
5.2
Copyright 1996 Lawrence C. Marsh
Probability Distribution
of Least Squares Estimators
b
1
~ N
1
,
(x
t
x)
2
2
x
t
2
b
2
~ N
2
,
(x
t
x)
2
2
5.3
Copyright 1996 Lawrence C. Marsh
2 ^
=
2
e
t
^
2
x)
2
2
Create a standardized normal random variable, Z,
by subtracting the mean of b
2
and dividing by its
standard deviation:
b
2
2
var(b
2
)
=
(0,1)
5.6
Copyright 1996 Lawrence C. Marsh
Simple Linear Regression
y
t
=
1
+
2
x
t
+
t
where E
t
= 0
y
t
~ N(
1
+
2
x
t
,
2
)
since Ey
t
=
1
+
2
x
t
t
= y
t
1
2
x
t
Therefore,
t
~ N(0,
2
) .
5.7
Copyright 1996 Lawrence C. Marsh
Create a Chi-Square
t
~ N(0,
2
) but want N(0,1) .
(
t
/) ~ N(0,1) Standard Normal .
(
t
/)
2
~
2
(1)
Chi-Square .
5.8
Copyright 1996 Lawrence C. Marsh
Sum of Chi-Squares
t
=1
(
t
/)
2
=
(
1
/)
2
+ (
2
/)
2
+. . .+ (
T
/)
2
2
(1)
+
2
(1)
+. . .+
2
(1)
=
2
()
Therefore,
t
=1
(
t
/)
2
2
()
5.9
Copyright 1996 Lawrence C. Marsh
Since the errors
t
= y
t
1
2
x
t
are not observable, we estimate them with
the sample residuals e
t
= y
t
b
1
b
2
x
t
.
Unlike the errors, the sample residuals are
not independent since they use up two degrees
of freedom by using b
1
and b
2
to estimate
1
and
2
.
We get only T2 degrees of freedom instead of T.
Chi-Square degrees of freedom
5.10
Copyright 1996 Lawrence C. Marsh
Student-t Distribution
t = ~ t
(m)
Z
V / m
where Z ~ N(0,1)
and V ~
(m)
2
5.11
Copyright 1996 Lawrence C. Marsh
t = ~ t
(m)
Z
V / ( T 2)
where Z =
(b
2
2
)
var(b
2
)
and var(b
2
) =
2
( x
i
x )
2
5.12
Copyright 1996 Lawrence C. Marsh
t =
Z
V / (T-2)
(b
2
2
)
var(b
2
)
t =
(T2)
2
2
^
( T 2)
V =
(T2)
2
2
^
5.13
Copyright 1996 Lawrence C. Marsh
var(b
2
) =
2
( x
i
x )
2
(b
2
2
)
2
( x
i
x )
2
t = =
(T2)
2
2
^
( T 2)
(b
2
2
)
2
( x
i
x )
2
^
notice the
cancellations
5.14
Copyright 1996 Lawrence C. Marsh
(b
2
2
)
2
( x
i
x )
2
^
t = =
(b
2
2
)
var(b
2
)
^
t =
(b
2
2
)
se(b
2
)
5.15
Copyright 1996 Lawrence C. Marsh
Students t - statistic
t = ~ t
(T2)
(b
2
2
)
se(b
2
)
t has a Student-t Distribution
with T 2 degrees of freedom.
5.16
Copyright 1996 Lawrence C. Marsh
Figure 5.1 Student-t Distribution
(1)
t
0
f(t)
-t
c
t
c
/2
/2
red area = rejection region for 2-sided test
5.17
Copyright 1996 Lawrence C. Marsh
probability statements
P(-t
c
t t
c
) = 1
P( t < -t
c
) = P( t > t
c
) = /2
P(-t
c
t
c
) = 1
(b
2
2
)
se(b
2
)
5.18
Copyright 1996 Lawrence C. Marsh
Confidence Intervals
Two-sided (1)x100% C.I. for
1
:
b
1
t
/2
[se(b
1
)], b
1
+ t
/2
[se(b
1
)]
b
2
t
/2
[se(b
2
)], b
2
+ t
/2
[se(b
2
)]
Two-sided (1)x100% C.I. for
2
:
5.19
Copyright 1996 Lawrence C. Marsh
Student-t vs. Normal Distribution
1. Both are symmetric bell-shaped distributions.
2. Student-t distribution has fatter tails than the normal.
3. Student-t converges to the normal for infinite sample.
4. Student-t conditional on degrees of freedom (df).
5. Normal is a good approximation of Student-t for the first few
decimal places when df > 30 or so.
5.20
Copyright 1996 Lawrence C. Marsh
Hypothesis Tests
1. A null hypothesis, H
0
.
2. An alternative hypothesis, H
1
.
3. A test statistic.
4. A rejection region.
5.21
Copyright 1996 Lawrence C. Marsh
Rejection Rules
1. Two-Sided Test:
If the value of the test statistic falls in the critical region in either
tail of the t-distribution, then we reject the null hypothesis in favor
of the alternative.
2. Left-Tail Test:
If the value of the test statistic falls in the critical region which lies
in the left tail of the t-distribution, then we reject the null
hypothesis in favor of the alternative.
2. Right-Tail Test:
If the value of the test statistic falls in the critical region which lies
in the right tail of the t-distribution, then we reject the null
hypothesis in favor of the alternative.
5.22
Copyright 1996 Lawrence C. Marsh
Format for Hypothesis Testing
1. Determine null and alternative hypotheses.
2. Specify the test statistic and its distribution
as if the null hypothesis were true.
3. Select and determine the rejection region.
4. Calculate the sample value of test statistic.
5. State your conclusion.
5.23
Copyright 1996 Lawrence C. Marsh
practical vs. statistical
significance in economics
Practically but not statistically significant:
When sample size is very small, a large average gap between
the salaries of men and women might not be statistically
significant.
Statistically but not practically significant:
When sample size is very large, a small correlation (say, =
0.00000001) between the winning numbers in the PowerBall
Lottery and the Dow-Jones Stock Market Index might be
statistically significant.
5.24
Copyright 1996 Lawrence C. Marsh
Type I and Type II errors
Type I error:
We make the mistake of rejecting the null
hypothesis when it is true.
= P(rejecting H
0
when it is true).
Type II error:
We make the mistake of failing to reject the null
hypothesis when it is false.
= P(failing to reject H
0
when it is false).
5.25
Copyright 1996 Lawrence C. Marsh
Prediction Intervals
A (1)x100% prediction interval for y
o
is:
y
o
t
c
se( f )
^
se( f ) = var( f )
^
f = y
o
y
o
^
(x
t
x)
2
var( f ) =
2
1 + +
^
1
(x
o
x)
2
^
5.26
Copyright 1996 Lawrence C. Marsh
The Simple Linear
Regression Model
Chapter 6
Copyright 1997 John Wiley & Sons, Inc. All rights reserved. Reproduction or translation of this work beyond
that permitted in Section 117 of the 1976 United States Copyright Act without the express written permission of the
copyright owner is unlawful. Request for further information should be addressed to the Permissions Department,
John Wiley & Sons, Inc. The purchaser may make back-up copies for his/her own use only and not for distribution
or resale. The Publisher assumes no responsibility for errors, omissions, or damages, caused by the use of these
programs or from the use of the information contained herein.
6.1
Copyright 1996 Lawrence C. Marsh
Explaining Variation in y
t
Predicting y
t
without any explanatory variables:
y
t
=
1
+ e
t
e
t
=
(y
t
1
)
2 2
t = 1 t = 1
T
T
= 2
(y
t
b
1
) = 0
e
t
2
t = 1
t = 1
T
T
(y
t
b
1
) = 0
t = 1
T
y
t
Tb
1
= 0
t = 1
T
b
1
= y
Why not y?
6.2
Copyright 1996 Lawrence C. Marsh
Explaining Variation in y
t
y
t
= b
1
+ b
2
x
t
+ e
t
^
Unexplained variation:
y
t
= b
1
+ b
2
x
t
^
Explained variation:
e
t
= y
t
y
t
= y
t
b
1
b
2
x
t
^ ^
6.3
Copyright 1996 Lawrence C. Marsh
Explaining Variation in y
t
y
t
= y
t
+ e
t
^
Why not y?
^
y
t
y = y
t
y + e
t
^
^
using y as baseline
SST = SSR + SSE
(y
t
y)
2
=
(y
t
y)
2
+
e
t
t = 1
T
^
^
T T
t = 1
t = 1
2
cross
product
term
drops
out
6.4
Copyright 1996 Lawrence C. Marsh
Total Variation in y
t
SST = total sum of squares
SST measures variation of y
t
around y
(y
t
y)
2
t = 1
T
SST =
6.5
Copyright 1996 Lawrence C. Marsh
Explained Variation in y
t
SSR = regression sum of squares
y
t
= b
1
+ b
2
x
t
^
Fitted y
t
values:
^
SSR measures variation of y
t
around y
^
(y
t
y)
2
t = 1
T
SSR =
^
6.6
Copyright 1996 Lawrence C. Marsh
Unexplained Variation in y
t
SSE = error sum of squares
SSE measures variation of y
t
around y
t
^
e
t
= y
t
y
t
= y
t
b
1
b
2
x
t
^ ^
(y
t
y
t
)
2
=
e
t
2
t = 1
T
SSE =
^
t = 1
T
^
6.7
Copyright 1996 Lawrence C. Marsh
Analysis of Variance Table
^
Table 6.1 Analysis of Variance Table
Source of Sum of Mean
Variation DF Squares Square
Explained 1 SSR SSR/1
Unexplained T-2 SSE SSE/(T-2)
[=
2
]
Total T-1 SST
6.8
Copyright 1996 Lawrence C. Marsh
Coefficient of Determination
0 R
2
1
What proportion of the variation
in y
t
is explained?
SSR
SST
R
2
=
6.9
Copyright 1996 Lawrence C. Marsh
Coefficient of Determination
SST = SSR + SSE
SST SSR SSE
SST SST SST
= +
SSR SSE
SST SST
1 = +
Dividing
by SST
SSR
SST
R
2
= = 1
SSE
SST
6.10
Copyright 1996 Lawrence C. Marsh
R
2
is only a descriptive measure.
R
2
does not measure the quality
of the regression model.
Focusing solely on maximizing
R
2
is not a good idea.
Coefficient of Determination
6.11
Copyright 1996 Lawrence C. Marsh
cov(X,Y)
=
var(X) var(Y)
Correlation Analysis
cov(X,Y)
r =
var(X) var(Y)
Population:
^
^ ^
Sample:
6.12
Copyright 1996 Lawrence C. Marsh
Correlation Analysis
var(X) =
^
(x
t
x)
2
/
(T1)
t = 1
T
var(Y) =
^
(y
t
y)
2
/
(T1)
t = 1
T
cov(X,Y) =
^
(x
t
x)(y
t
y)
/
(T1)
t = 1
T
6.13
Copyright 1996 Lawrence C. Marsh
Correlation Analysis
(x
t
x)
2
(y
t
y)
2
t = 1
T
(x
t
x)(y
t
y)
t = 1
T
r =
t = 1
T
Sample Correlation Coefficient
6.14
Copyright 1996 Lawrence C. Marsh
Correlation Analysis and R
2
For simple linear regression analysis:
r
2
= R
2
R
2
is also the correlation
between y
t
and y
t
measuring goodness of fit.
^
6.15
Copyright 1996 Lawrence C. Marsh
Regression Computer Output
Table 6.2 Computer Generated Least Squares Results
(1) (2) (3) (4) (5)
Parameter Standard T for H0:
Variable Estimate Error Parameter=0 Prob>|T|
INTERCEPT 40.7676 22.1387 1.841 0.0734
X 0.1283 0.0305 4.201 0.0002
Typical computer output of regression estimates:
6.16
Copyright 1996 Lawrence C. Marsh
Regression Computer Output
se(b
1
) = var(b
1
) = 490.12 = 22.1287
^
se(b
2
) = var(b
2
) = 0.0009326 = 0.0305
^
b
1
= 40.7676
b
2
= 0.1283
se(b
1
)
t = = = 1.84
b
1
40.7676
22.1287
se(b
2
)
b
2
t = = = 4.20
0.1283
0.0305
6.17
Copyright 1996 Lawrence C. Marsh
Regression Computer Output
Table 6.3 Analysis of Variance Table
Sum of Mean
Source DF Squares Square
Explained 1 25221.2229 25221.2229
Unexplained 38 54311.3314 1429.2455
Total 39 79532.5544
R-square: 0.3171
Sources of variation in the dependent variable:
6.18
Copyright 1996 Lawrence C. Marsh
Regression Computer Output
SSR
SST
R
2
= = 1 = 0.317
SSE
SST
SSE /(T-2) =
2
= 1429.2455
^
SSE =
e
t
2
= 54311
^
SST =
(y
t
y)
2
= 79532
SSR =
(y
t
y)
2
= 25221
^
6.19
Copyright 1996 Lawrence C. Marsh
y
t
= 40.7676 + 0.1283x
t
(s.e.) (22.1387) (0.0305)
y
t
= 40.7676 + 0.1283x
t
(t) (1.84) (4.20)
Reporting Regression Results
6.20
Copyright 1996 Lawrence C. Marsh
R
2
= 0.317
Reporting Regression Results
This R
2
value may seem low but it is
typical in studies involving cross-sectional
data analyzed at the individual or micro level.
A considerably higher R
2
value would be
expected in studies involving time-series data
analyzed at an aggregate or macro level.
6.21
Copyright 1996 Lawrence C. Marsh
Effects of Scaling the Data
Changing the scale of x
y
t
=
1
+ (c
2
)(x
t
/c)
+ e
t
y
t
=
1
+
2
x
t
+ e
t
y
t
=
1
+
2
x
t
+ e
t
*
*
2
= c
2
*
x
t
= x
t
/c
*
where
and
The estimated
coefficient and
standard error
change but the
other statistics
are unchanged.
6.22
Copyright 1996 Lawrence C. Marsh
Effects of Scaling the Data
Changing the scale of y
y
t
/c = (
1
/c)
+ (
2
/c)x
t
+ e
t
/c
y
t
=
1
+
2
x
t
+ e
t
1
=
1
/c
*
and
All statistics
are changed
except for
the t-statistics
and R
2
value.
y
t
=
1
+
2
x
t
+ e
t
*
*
*
*
2
=
2
/c
*
*
e
t
= e
t
/c
y
t
= y
t
/c
where
*
6.23
Copyright 1996 Lawrence C. Marsh
Effects of Scaling the Data
Changing the scale of x and y
y
t
/c = (
1
/c)
+ (c
2
/c)x
t
/c
+ e
t
/c
y
t
=
1
+
2
x
t
+ e
t
1
=
1
/c
*
and
No change in
the R
2
or the
t-statistics or
in regression
results for
2
but all other
stats change.
y
t
=
1
+
2
x
t
+ e
t
*
*
*
*
x
t
= x
t
/c
*
*
e
t
= e
t
/c
y
t
= y
t
/c
where
*
6.24
Copyright 1996 Lawrence C. Marsh
Functional Forms
The term linear in a simple
regression model does not mean
a linear relationship between
variables, but a model in which
the parameters enter the model
in a linear way.
6.25
Copyright 1996 Lawrence C. Marsh
y
t
=
1
+
2
x
t
+ e
t
Linear Statistical Models:
Nonlinear Statistical Models:
ln(y
t
) =
1
+
2
x
t
+ e
t
y
t
=
1
+
2
ln(x
t
)
+ e
t
y
t
=
1
+
2
x
t
+ e
t
2
y
t
=
1
+
2
x
t
+ e
t
3
y
t
=
1
+
2
x
t
+ exp(
3
x
t
)
+ e
t
y
t
=
1
+
2
x
t
+ e
t
3
Linear vs. Nonlinear
6.27
Copyright 1996 Lawrence C. Marsh
y
x
nonlinear
relationship
between food
expenditure and
income
Linear vs. Nonlinear
food
expenditure
income
0
6.27
Copyright 1996 Lawrence C. Marsh
Useful Functional Forms
1. Linear
2. Reciprocal
3. Log-Log
4. Log-Linear
5. Linear-Log
6. Log-Inverse
Look at
each form
and its
slope and
elasticity
6.28
Copyright 1996 Lawrence C. Marsh
Linear
y
t
=
1
+
2
x
t
+ e
t
slope:
2
elasticity:
2
y
t
Useful Functional Forms
x
t
6.29
Copyright 1996 Lawrence C. Marsh
Reciprocal
y
t
=
1
+
2
+ e
t
Useful Functional Forms
1
x
t
slope:
elasticity:
1
x
t
2
2
1
x
t
y
t
2
6.30
Copyright 1996 Lawrence C. Marsh
x
t
y
t
Log-Log
ln(y
t
)=
1
+
2
ln(x
t
) + e
t
slope:
2
elasticity:
2
Useful Functional Forms
6.31
Copyright 1996 Lawrence C. Marsh
Log-Linear
ln(y
t
)=
1
+
2
x
t
+ e
t
slope:
2
y
t
elasticity:
2
x
t
Useful Functional Forms
6.32
Copyright 1996 Lawrence C. Marsh
Linear-Log
y
t
=
1
+
2
ln(x
t
)
+ e
t
_
slope:
2
elasticity:
2
1
x
t
y
t
1
_
Useful Functional Forms
6.33
Copyright 1996 Lawrence C. Marsh
Useful Functional Forms
ln(y
t
) =
1
-
2
+ e
t
1
x
t
Log-Inverse
slope:
2
elasticity:
2
x
2
t
y
t
1
x
t
6.34
Copyright 1996 Lawrence C. Marsh
1. E (e
t
) = 0
2. var (e
t
) =
2
3. cov(e
i
, e
j
) = 0
4. e
t
~ N(0,
2
)
Error Term Properties
6.35
Copyright 1996 Lawrence C. Marsh
Economic Models
1. Demand Models
2. Supply Models
3. Production Functions
4. Cost Functions
5. Phillips Curve
6.36
Copyright 1996 Lawrence C. Marsh
1. Demand Models
* quality demanded (y
d
) and price (x)
* constant elasticity
Economic Models
ln(y
t
)=
1
+
2
ln(x)
t
+ e
t
d
6.37
Copyright 1996 Lawrence C. Marsh
2. Supply Models
* quality supplied (y
s
) and price (x)
* constant elasticity
Economic Models
ln(y
t
)=
1
+
2
ln(x
t
) + e
t
s
6.38
Copyright 1996 Lawrence C. Marsh
3. Production Functions
* output (y) and input (x)
* constant elasticity
Economic Models
ln(y
t
)=
1
+
2
ln(x
t
) + e
t
Cobb-Douglas Production Function:
6.39
Copyright 1996 Lawrence C. Marsh
4a. Cost Functions
* total cost (y) and output (x)
Economic Models
y
t
=
1
+
2
x
2
t
+ e
t
6.40
Copyright 1996 Lawrence C. Marsh
4b. Cost Functions
* average cost (x/y) and output (x)
Economic Models
(y
t
/x
t
) =
1
/x
t
+
2
x
t
+ e
t
/x
t
6.41
Copyright 1996 Lawrence C. Marsh
5. Phillips Curve
* wage rate (w
t
) and time (t)
Economic Models
unemployment rate, u
t
w
t-1
% w
t
=
w
t
w
t-1
= +
u
t
1
nonlinear in both variables and parameters
6.42
Copyright 1996 Lawrence C. Marsh
The Multiple
Regression Model
Chapter 7
Copyright 1997 John Wiley & Sons, Inc. All rights reserved. Reproduction or translation of this work beyond
that permitted in Section 117 of the 1976 United States Copyright Act without the express written permission of the
copyright owner is unlawful. Request for further information should be addressed to the Permissions Department,
John Wiley & Sons, Inc. The purchaser may make back-up copies for his/her own use only and not for distribution
or resale. The Publisher assumes no responsibility for errors, omissions, or damages, caused by the use of these
programs or from the use of the information contained herein.
7.1
Copyright 1996 Lawrence C. Marsh
Two Explanatory Variables
y
t
=
1
+
2
x
t2
+
3
x
t3
+ e
t
y
t
x
t2
=
2
x
t3
y
t
=
3
x
t
s affect y
t
separately
But least squares estimation of
2
now depends upon both x
t2
and x
t3
.
7.2
Copyright 1996 Lawrence C. Marsh
Correlated Variables
y
t
= output
x
t2
= capital x
t3
= labor
Always 5 workers per machine.
If number of workers per machine
is never varied, it becomes impossible
to tell if the machines or the workers
are responsible for changes in output.
y
t
=
1
+
2
x
t2
+
3
x
t3
+ e
t
7.3
Copyright 1996 Lawrence C. Marsh
The General Model
y
t
=
1
+
2
x
t2
+
3
x
t3
+. . .+
K
x
tK
+ e
t
The parameter
1
is the intercept (constant) term.
The variable attached to
1
is x
t1
= 1.
Usually, the number of explanatory variables
is said to be K1 (ignoring x
t1
= 1), while the
number of parameters is K. (Namely:
1
. . .
K
).
7.4
Copyright 1996 Lawrence C. Marsh
1. E(e
t
) = 0
2. var(e
t
) =
2
3. cov(e
t
, e
s
) = 0 for t s
4. e
t
~ N(0,
2
)
Statistical Properties of e
t
7.5
Copyright 1996 Lawrence C. Marsh
1. E (y
t
) =
1
+
2
x
t2
+. . .+
K
x
tK
2. var(y
t
) = var(e
t
) =
2
3. cov(y
t
,y
s
) = cov(e
t
, e
s
) = 0 ts
4. y
t
~ N(
1
+
2
x
t2
+. . .+
K
x
tK
,
2
)
Statistical Properties of y
t
7.6
Copyright 1996 Lawrence C. Marsh
Assumptions
1. y
t
=
1
+
2
x
t2
+. . .+
K
x
tK
+ e
t
2. E (y
t
) =
1
+
2
x
t2
+. . .+
K
x
tK
3. var(y
t
) = var(e
t
) =
2
4. cov(y
t
,y
s
) = cov(e
t
,e
s
) = 0 t s
5. The values of x
tk
are not random
6. y
t
~ N(
1
+
2
x
t2
+. . .+
K
x
tK
,
2
)
7.7
Copyright 1996 Lawrence C. Marsh
Least Squares Estimation
y
t
=
1
+
2
x
t2
+
3
x
t3
+ e
t
S S(
1
,
2
,
3
) =
(
y
t
1
2
x
t2
3
x
t3
)
2
t = 1
T
Define: y
t
= y
t
y
*
x
t2
= x
t2
x
2
*
x
t3
= x
t3
x
3
*
7.8
Copyright 1996 Lawrence C. Marsh
b
1
= y b
1
b
2
x
2
b
3
x
3
b
3
=
(
y
t
x
t3
)(
x
t2
)
(
y
t
x
t2
)(
x
t3
x
t2
)
*
* *
* * * *
2
(
x
t2
)(
x
t3
) (
x
t2
x
t3
)
* *
* *
2
2
2
b
2
=
(
y
t
x
t2
)(
x
t3
)
(
y
t
x
t3
)(
x
t2
x
t3
)
*
* *
* * * *
2
(
x
t2
)(
x
t3
) (
x
t2
x
t3
)
* *
* *
2
2
2
Least Squares Estimators
7.9
Copyright 1996 Lawrence C. Marsh
Dangers of Extrapolation
Statistical models generally are good only
within the relevant range. This means
that extending them to extreme data values
outside the range of the original data often
leads to poor and sometimes ridiculous results.
If height is normally distributed and the
normal ranges from minus infinity to plus
infinity, pity the man minus three feet tall.
7.10
Copyright 1996 Lawrence C. Marsh
Error Variance Estimation
2 ^
=
e
t
^
2
(x
t3
x
3
)
2
2
2
var(b
2
) =
(1
r
23
)
(x
t2
x
2
)
2
2
(x
t2
x
2
)
2
(x
t3
x
3
)
2
where r
23
=
(x
t2
x
2
)(x
t3
x
3
)
When r
23
= 0
these reduce
to the simple
regression
formulas.
7.13
Copyright 1996 Lawrence C. Marsh
Variance Decomposition
The variance of an estimator is smaller when:
1. The error variance,
2
, is smaller:
2
0 .
2. The sample size, T, is larger:
(x
t2
x
2
)
2
.
3. The variables values are more spread out:
(x
t2
x
2
)
2
.
4. The correlation is close to zero: r
23
0 .
2
t = 1
T
7.14
Copyright 1996 Lawrence C. Marsh
Covariances
y
t
=
1
+
2
x
t2
+
3
x
t3
+ e
t
where r
23
=
(x
t2
x
2
)
2
(x
t3
x
3
)
2
(x
t2
x
2
)(x
t3
x
3
)
(1
r
23
)
(x
t2
x
2
)
2
(x
t3
x
3
)
2
cov(b
2
,b
3
) =
2
r
23
2
7.15
Copyright 1996 Lawrence C. Marsh
Covariance Decomposition
1. The error variance,
2
, is larger.
2. The sample size, T, is smaller.
3. The values of the variables are less spread out.
4. The correlation, r
23
, is high.
The covariance between any two estimators
is larger in absolute value when:
7.16
Copyright 1996 Lawrence C. Marsh
Var-Cov Matrix
y
t
=
1
+
2
x
t2
+
3
x
t3
+ e
t
var(b
1
) cov(b
1
,b
2
) cov(b
1
,b
3
)
cov(b
1
,b
2
,b
3
) = cov(b
1
,b
2
) var(b
2
) cov(b
2
,b
3
)
cov(b
1
,b
3
) cov(b
2
,b
3
) var(b
3
)
The least squares estimators b
1
, b
2
, and b
3
have covariance matrix:
7.17
Copyright 1996 Lawrence C. Marsh
Normal
y
t
=
1
+
2
x
2t
+
3
x
3t
+. . .+
K
x
Kt
+ e
t
y
t
~N (
1
+
2
x
2t
+
3
x
3t
+. . .+
K
x
Kt
),
2
e
t
~ N(0,
2
)
This implies and is implied by:
b
k
~ N
k
, var(b
k
)
z = ~ N(0,1) for k = 1,2,...,K
b
k
k
var(b
k
)
Since b
k
is a linear
function of the y
t
s:
7.18
Copyright 1996 Lawrence C. Marsh
Student-t
b
k
k
var(b
k
)
^
t = =
b
k
k
se(b
k
)
Since generally the population variance
of b
k
, var(b
k
) , is unknown, we estimate
it with which uses
2
instead of
2
. var(b
k
)
^ ^
t has a Student-t distribution with df=(TK).
7.19
Copyright 1996 Lawrence C. Marsh
Interval Estimation
b
k
k
se(b
k
)
P t
c
t
c
= 1
t
c
is critical value for (T-K) degrees of freedom
such that P(t
t
c
) = /2.
P b
k
t
c
se(b
k
)
k
b
k
+ t
c
se(b
k
)
= 1
Interval endpoints:
b
k
t
c
se(b
k
) , b
k
+ t
c
se(b
k
)
7.20
Copyright 1996 Lawrence C. Marsh
Hypothesis Testing
and
Nonsample Information
Chapter 8
Copyright 1997 John Wiley & Sons, Inc. All rights reserved. Reproduction or translation of this work beyond
that permitted in Section 117 of the 1976 United States Copyright Act without the express written permission of the
copyright owner is unlawful. Request for further information should be addressed to the Permissions Department,
John Wiley & Sons, Inc. The purchaser may make back-up copies for his/her own use only and not for distribution
or resale. The Publisher assumes no responsibility for errors, omissions, or damages, caused by the use of these
programs or from the use of the information contained herein.
8.1
Copyright 1996 Lawrence C. Marsh
1. Student-t Tests
2. Goodness-of-Fit
3. F-Tests
4. ANOVA Table
5. Nonsample Information
6. Collinearity
7. Prediction
Chapter 8: Overview
8.2
Copyright 1996 Lawrence C. Marsh
Student - t Test
y
t
=
1
+
2
X
t2
+
3
X
t3
+
4
X
t4
+ e
t
Student-t tests can be used to test any linear
combination of the regression coefficients:
H
0
:
2
+
3
+
4
= 1 H
0
:
1
= 0
H
0
: 3
2
7
3
= 21 H
0
:
2
3
5
Every such t-test has exactly TK degrees of freedom
where K=#coefficients estimated(including the intercept).
8.3
Copyright 1996 Lawrence C. Marsh
One Tail Test
y
t
=
1
+
2
X
t2
+
3
X
t3
+
4
X
t4
+ e
t
H
0
:
3
0
H
1
:
3
> 0
b
3
se(b
3
)
t = ~ t
(TK)
t
c
0
df = T K
= T 4
(1 )
8.4
Copyright 1996 Lawrence C. Marsh
Two Tail Test
y
t
=
1
+
2
X
t2
+
3
X
t3
+
4
X
t4
+ e
t
H
0
:
2
= 0
H
1
:
2
0
b
2
se(b
2
)
t = ~ t
(TK)
t
c
0
df = T K
= T 4
/2
(1 )
-t
c
/2
8.5
Copyright 1996 Lawrence C. Marsh
Goodness - of - Fit
0 R
2
1
Coefficient of Determination
SST
R
2
= =
(y
t
y)
2
t = 1
T
^
SSR
(y
t
y)
2
t = 1
T
8.6
Copyright 1996 Lawrence C. Marsh
Adjusted R-Squared
Adjusted Coefficient of Determination
Original:
Adjusted:
SST/(T1)
R
2
= 1
SSE/(TK)
SST
= 1
SSE
R
2
=
SST
SSR
8.7
Copyright 1996 Lawrence C. Marsh
Computer Output
Table 8.2 Summary of Least Squares Results
Variable Coefficient Std Error t-value p-value
constant 104.79 6.48 16.17 0.000
price 6.642 3.191 2.081 0.042
advertising 2.984 0.167 17.868 0.000
b
2
se(b
2
)
t = =
6.642
3.191
2.081 =
8.8
Copyright 1996 Lawrence C. Marsh
Reporting Your Results
y
t
= 104.79 6.642 X
t2
+ 2.984 X
t3
^
(6.48) (3.191) (0.167) (s.e.)
y
t
= 104.79 6.642 X
t2
+ 2.984 X
t3
^
(16.17) (-2.081) (17.868) (t)
Reporting t-statistics:
Reporting standard errors:
8.9
Copyright 1996 Lawrence C. Marsh
Single Restriction F-Test
y
t
=
1
+
2
X
t2
+
3
X
t3
+
4
X
t4
+ e
t
H
0
:
2
= 0
H
1
:
2
0
df
d
= T K = 49
df
n
= J = 1
(SSE
R
SSE
U
)/J
SSE
U
/(TK)
F =
(1964.758 1805.168)/1
1805.168/(52 3)
=
= 4.33
By definition this is the t-statistic squared:
t = 2.081 F = t
2
= 4.33
8.10
Copyright 1996 Lawrence C. Marsh
Multiple Restriction F-Test
y
t
=
1
+
2
X
t2
+
3
X
t3
+
4
X
t4
+ e
t
H
0
:
2
= 0,
4
= 0
H
1
: H
0
not true
df
d
= T K = 49
df
n
= J = 2
(SSE
R
SSE
U
)/J
SSE
U
/(TK)
F =
First run the restricted
regression by dropping
X
t2
and X
t4
to get SSE
R
.
Next run unrestricted regression to get SSE
U
.
8.11
Copyright 1996 Lawrence C. Marsh
F-Tests
(SSE
R
SSE
U
)/J
SSE
U
/(TK)
F =
F-Tests of this type are always right-tailed,
even for left-sided or two-sided hypotheses,
because any deviation from the null will
make the F value bigger (move rightward).
0 F
c
(1 )
f(F)
F
8.12
Copyright 1996 Lawrence C. Marsh
F-Test of Entire Equation
y
t
=
1
+
2
X
t2
+
3
X
t3
+ e
t
H
0
:
2
=
3
= 0
H
1
: H
0
not true
df
d
= T K = 49
df
n
= J = 2
(SSE
R
SSE
U
)/J
SSE
U
/(TK)
F =
(13581.35 1805.168)/2
1805.168/(52 3)
=
= 159.828
We ignore
1
.
Why?
F
c
= 3.187
= 0.05
Reject H
0
!
8.13
Copyright 1996 Lawrence C. Marsh
ANOVA Table
Table 8.3 Analysis of Variance Table
Sum of Mean
Source DF Squares Square F-Value
Explained 2 11776.18 5888.09 158.828
Unexplained 49 1805.168 36.84
Total 51 13581.35 p-value: 0.0001
SST
R
2
=
=
SSR
=
0.867
11776.18
13581.35
8.14
Copyright 1996 Lawrence C. Marsh
Nonsample Information
ln(y
t
) =
1
+
2
ln(X
t2
)
+
3
ln(X
t3
)
+
4
ln(X
t4
)
+ e
t
A certain production process is known to be
Cobb-Douglas with constant returns to scale.
2
+
3
+
4
= 1 where
4
= (1
2
3
)
ln(y
t
/X
t4
) =
1
+
2
ln(X
t2
/X
t4
)
+
3
ln(X
t3
/X
t4
)
+ e
t
y
t
=
1
+
2
X
t2
+
3
X
t3
+
4
X
t4
+ e
t
*
* * *
Run least squares on the transformed model.
Interpret coefficients same as in original model.
8.15
Copyright 1996 Lawrence C. Marsh
Collinear Variables
The term independent variable means
an explanatory variable is independent of
of the error term, but not necessarily
independent of other explanatory variables.
Since economists typically have no control
over the implicit experimental design,
explanatory variables tend to move
together which often makes sorting out
their separate influences rather problematic.
8.16
Copyright 1996 Lawrence C. Marsh
Effects of Collinearity
1. no least squares output when collinearity is exact.
2. large standard errors and wide confidence intervals.
3. insignificant t-values even with high R
2
and a
significant F-value.
4. estimates sensitive to deletion or addition of a few
observations or insignificant variables.
5. good within-sample(same proportions) but poor
out-of-sample(different proportions) prediction.
A high degree of collinearity will produce:
8.17
Copyright 1996 Lawrence C. Marsh
Identifying Collinearity
Evidence of high collinearity include:
1. a high pairwise correlation between two
explanatory variables.
2. a high R-squared when regressing one
explanatory variable at a time on each of the
remaining explanatory variables.
3. a statistically significant F-value when the
t-values are statistically insignificant.
4. an R-squared that doesnt fall by much when
dropping any of the explanatory variables.
8.18
Copyright 1996 Lawrence C. Marsh
Mitigating Collinearity
Since high collinearity is not a violation of
any least squares assumption, but rather a
lack of adequate information in the sample:
1. collect more data with better information.
2. impose economic restrictions as appropriate.
3. impose statistical restrictions when justified.
4. if all else fails at least point out that the poor
model performance might be due to the
collinearity problem (or it might not).
8.19
Copyright 1996 Lawrence C. Marsh
Prediction
Given a set of values for the explanatory
variables, (1 X
02
X
03
), the best linear
unbiased predictor of y is given by:
y
t
=
1
+
2
X
t2
+
3
X
t3
+ e
t
This predictor is unbiased in the sense
that the average value of the forecast
error is zero.
y
0
= b
1
+ b
2
X
02
+ b
3
X
03
^
8.20
Copyright 1996 Lawrence C. Marsh
Extensions
of the Multiple
Regression Model
Chapter 9
Copyright 1997 John Wiley & Sons, Inc. All rights reserved. Reproduction or translation of this work beyond
that permitted in Section 117 of the 1976 United States Copyright Act without the express written permission of the
copyright owner is unlawful. Request for further information should be addressed to the Permissions Department,
John Wiley & Sons, Inc. The purchaser may make back-up copies for his/her own use only and not for distribution
or resale. The Publisher assumes no responsibility for errors, omissions, or damages, caused by the use of these
programs or from the use of the information contained herein.
9.1
Copyright 1996 Lawrence C. Marsh
Topics for This Chapter
1. Intercept Dummy Variables
2. Slope Dummy Variables
3. Different Intercepts & Slopes
4. Testing Qualitative Effects
5. Are Two Regressions Equal?
6. Interaction Effects
7. Dummy Dependent Variables
9.2
Copyright 1996 Lawrence C. Marsh
Intercept Dummy Variables
Dummy variables are binary (0,1)
D
t
= 1 if red car, D
t
= 0 otherwise.
y
t
=
1
+
2
X
t
+
3
D
t
+ e
t
y
t
= speed of car in miles per hour
X
t
= age of car in years
Police: red cars travel faster.
H
0
:
3
= 0
H
1
:
3
> 0
9.3
Copyright 1996 Lawrence C. Marsh
y
t
=
1
+
2
X
t
+
3
D
t
+ e
t
red cars: y
t
= (
1
+
3
) +
2
x
t
+ e
t
other cars: y
t
=
1
+
2
X
t
+ e
t
y
t
X
t
miles
per
hour
age in years
0
1
+
3
2
r
e
d
c
a
r
s
o
t
h
e
r
c
a
r
s
9.4
Copyright 1996 Lawrence C. Marsh
Slope Dummy Variables
y
t
=
1
+
2
X
t
+
3
D
t
X
t
+ e
t
y
t
=
1
+ (
2
+
3
)X
t
+ e
t
y
t
=
1
+
2
X
t
+ e
t
y
t
X
t
value
of
porfolio
years
0
2
+
2
stocks
bonds
Stock portfolio: D
t
= 1 Bond portfolio: D
t
= 0
1
= initial
investment
9.5
Copyright 1996 Lawrence C. Marsh
Different Intercepts & Slopes
y
t
=
1
+
2
X
t
+
3
D
t
+
4
D
t
X
t
+ e
t
y
t
= (
1
+
3
) + (
2
+
4
)X
t
+ e
t
y
t
=
1
+
2
X
t
+ e
t
y
t
X
t
harvest
weight
of corn
rainfall
2
+
2
miracle
regular
miracle seed: D
t
= 1 regular seed: D
t
= 0
1
+
3
9.6
Copyright 1996 Lawrence C. Marsh
y
t
=
1
+
2
X
t
+
3
D
t
+ e
t
2
1
+
3
1
y
t
X
t
Men
Women
0
y
t
=
1
+
2
X
t
+ e
t
For men: D
t
= 1.
For women: D
t
= 0.
years of experience
y
t
= (
1
+
3
) +
2
X
t
+ e
t
wage
rate
H
0
:
3
= 0
H
1
:
3
> 0 .
.
Testing for
discrimination
in starting wage
9.7
Copyright 1996 Lawrence C. Marsh
y
t
=
1
+
5
X
t
+
6
D
t
X
t
+ e
t
5
+
6
1
y
t
X
t
Men
Women
0
y
t
=
1
+ (
5
+
6
)X
t
+ e
t
y
t
=
1
+
5
X
t
+ e
t
For men D
t
= 1.
For women D
t
= 0.
Men and women have the same
starting wage,
1
, but their wage rates
increase at different rates (diff.=
6
).
6
> 0 means that mens wage rates are
increasing faster than women's wage rates.
years of experience
wage
rate
9.8
Copyright 1996 Lawrence C. Marsh
y
t
=
1
+
2
X
t
+
3
D
t
+
4
D
t
X
t
+ e
t
1
+
3
2
2
+
4
y
t
X
t
Men
Women
0
y
t
= (
1
+
3
) + (
2
+
4
) X
t
+ e
t
y
t
=
1
+
2
X
t
+ e
t
Women are given a higher starting wage,
1
,
while men get the lower starting wage,
1
+
3
,
(
3
< 0 ). But, men get a faster rate of increase
in their wages,
2
+
4
, which is higher than the
rate of increase for women,
2
, (since
4
> 0 ).
years of experience
An Ineffective Affirmative Action Plan
women are started
at a higher wage.
Note:
(
3
< 0 )
wage
rate
9.9
Copyright 1996 Lawrence C. Marsh
Testing Qualitative Effects
1. Test for differences in intercept.
2. Test for differences in slope.
3. Test for differences in both
intercept and slope.
9.10
Copyright 1996 Lawrence C. Marsh
H
0
:
3
0 vs.
1
:
3
> 0
H
0
:
4
0 vs.
1
:
4
> 0
Y
t
=
1
+
2
X
t
+
3
D
t
+
4
D
t
X
t
b 0
3
Est. Var b
3
t
n 4
b 0
4
Est. Var b
4
t
n 4
men: D
t
= 1 ; women: D
t
= 0
Testing for
discrimination in
starting wage.
Testing for
discrimination in
wage increases.
intercept
slope
+ e
t
9.11
Copyright 1996 Lawrence C. Marsh
Testing: H
o
: 3 = 4 = 0
H
1
: otherwise
and
SSE
R
= (
y
t
b
1
b
2
X
t
)
2
t = 1
T
SSE
U
= (y
t
b
1
b
2
X
t
b
3
D
t
b
4
D
t
X
t
)
2
t=1
T
( SSE
R
SSE
U
) / 2
SSE
U
/ ( T 4 )
F
T 4
2
intercept and slope
9.12
Copyright 1996 Lawrence C. Marsh
Are Two Regressions Equal?
y
t
=
1
+
2
X
t
+
3
D
t
+
4
D
t
X
t
+ e
t
variations of The Chow Test
I. Assuming equal variances (pooling):
men: D
t
= 1 ; women: D
t
= 0
H
o
:
3
=
4
= 0 vs. H
1
: otherwise
y
t
= wage rate
This model assumes equal wage rate variance.
X
t
= years of experience
9.13
Copyright 1996 Lawrence C. Marsh
y
t
=
1
+
2
X
t
+ e
t
II. Allowing for unequal variances:
y
tm
=
1
+
2
X
tm
+ e
tm
y
tw
=
1
+
2
X
tw
+ e
tw
Everyone:
Men only:
Women only:
SSE
R
Forcing men and women to have same
1
,
2
.
Allowing men and women to be different.
SSE
m
SSE
w
where SSE
U
= SSE
m
+
SSE
w
F =
(SSE
R
SSE
U
)/J
SSE
U
/(TK)
J = # restrictions
K=unrestricted coefs.
(running three regressions)
J = 2 K = 4
9.14
Copyright 1996 Lawrence C. Marsh
Interaction Variables
1. Interaction Dummies
2. Polynomial Terms
(special case of continuous interaction)
3. Interaction Among Continuous Variables
9.15
Copyright 1996 Lawrence C. Marsh
1. Interaction Dummies
y
t
=
1
+
2
X
t
+
3
M
t
+
4
B
t
+ e
t
For men: M
t
= 1. For women: M
t
= 0.
For black: B
t
= 1. For nonblack: B
t
= 0.
No Interaction: wage gap assumed the same:
y
t
=
1
+
2
X
t
+
3
M
t
+
4
B
t
+
5
M
t
B
t
+ e
t
Interaction: wage gap depends on race:
Wage Gap between Men and Women
y
t
= wage rate; X
t
= experience
9.16
Copyright 1996 Lawrence C. Marsh
2. Polynomial Terms
y
t
=
1
+
2
X
t
+
3
X
2
t
+
4
X
3
t
+ e
t
Linear in parameters but nonlinear in variables:
y
t
= income; X
t
= age
Polynomial Regression
y
t
X
t
People retire at different ages or not at all.
90
20 30 40 50 60 80 70
9.17
Copyright 1996 Lawrence C. Marsh
y
t
=
1
+
2
X
t
+
3
X
2
t
+
4
X
3
t
+ e
t
y
t
= income; X
t
= age
Polynomial Regression
Rate income is changing as we age:
y
t
X
t
=
2
+ 2
3
X
t
+ 3
4
X
2
t
Slope changes as X
t
changes.
9.18
Copyright 1996 Lawrence C. Marsh
3. Continuous Interaction
y
t
=
1
+
2
Z
t
+
3
B
t
+
4
Z
t
B
t
+ e
t
Exam grade = f(sleep:Z
t
, study time:B
t
)
Sleep and study time do not act independently.
More study time will be more effective
when combined with more sleep and less
effective when combined with less sleep.
9.19
Copyright 1996 Lawrence C. Marsh
Your mind sorts
things out while
you sleep (when you have things to sort out.)
y
t
=
1
+
2
Z
t
+
3
B
t
+
4
Z
t
B
t
+ e
t
Exam grade = f(sleep:Z
t
, study time:B
t
)
y
t
B
t
=
2
+
4
Z
t
Your studying is
more effective
with more sleep.
y
t
Z
t
=
2
+
4
B
t
continuous interaction
9.20
Copyright 1996 Lawrence C. Marsh
y
t
=
1
+
2
Z
t
+
3
B
t
+
4
Z
t
B
t
+ e
t
Exam grade = f(sleep:Z
t
, study time:B
t
)
If Z
t
+ B
t
= 24 hours, then B
t
= (24 Z
t
)
y
t
=
1
+
2
Z
t
+
3
(24 Z
t
)
+
4
Z
t
(24 Z
t
) + e
t
y
t
= (
1
+24
3
) + (
2
3
+24
4
)Z
t
4
Z
2
t
+ e
t
y
t
=
1
+
2
Z
t
+
3
Z
2
t
+ e
t
Sleep needed to maximize your exam grade:
y
t
Z
t
=
2
+ 2
3
Z
t
= 0
where
2
> 0 and
3
< 0
2
2
3
Z
t
=
9.21
Copyright 1996 Lawrence C. Marsh
1. Linear Probability Model
2. Probit Model
3. Logit Model
Dummy Dependent Variables
9.22
Copyright 1996 Lawrence C. Marsh
Linear Probability Model
y
i
=
1
+
2
X
i2
+
3
X
i3
+
4
X
i4
+ e
i
X
i2
= total hours of work each week
1 quits job
0 does not quit
y
i
=
X
i3
= weekly paycheck
X
i4
= hourly pay (X
i3
divided by X
i2
)
9.23
Copyright 1996 Lawrence C. Marsh
X
i2
y
i
=
1
+
2
X
i2
+
3
X
i3
+
4
X
i4
+ e
i
y
t
= 1
0
y
t
=
total hours of work each week
y
i
= b
1
+ b
2
X
i2
+ b
3
X
i3
+ b
4
X
i4
^
y
i
^
Read predicted values of y
i
off the regression line:
Linear Probability Model
9.24
Copyright 1996 Lawrence C. Marsh
1. Probability estimates are sometimes
less than zero or greater than one.
2. Heteroskedasticity is present in that
the model generates a nonconstant
error variance.
Linear Probability Model
Problems with Linear Probability Model:
9.25
Copyright 1996 Lawrence C. Marsh
Probit Model
z
i
=
1
+
2
X
i2
+ . . .
2
f(z
i
) = e
0.5z
i
2
1
F(z
i
) = P[
Z
z
i
] = e
0.5u
2
du
2
1
Normal probability density function:
Normal cumulative probability function:
z
i
latent variable, z
i
:
9.26
Copyright 1996 Lawrence C. Marsh
p
i
= P[
Z
1
+
2
X
i2
] = F(
1
+
2
X
i2
)
Since z
i
=
1
+
2
X
i2
+ . . .
, we can
substitute in to get:
Probit Model
X
i2
total hours of work each week
y
t
= 1
0
y
t
=
9.27
Copyright 1996 Lawrence C. Marsh
Logit Model
p
i
=
1
1 + e
(
1
+
2
X
i2
+ . . .)
Define p
i
:
For
2
> 0, p
i
will approach 1 as X
i2
+
p
i
is the probability of quitting the job.
For
2
> 0, p
i
will approach 0 as X
i2
9.28
Copyright 1996 Lawrence C. Marsh
Logit Model
X
i2
total hours of work each week
y
t
= 1
0
y
t
=
p
i
=
1
1 + e
(
1
+
2
X
i2
+ . . .)
p
i
is the probability of quitting the job.
9.29
Copyright 1996 Lawrence C. Marsh
Maximum Likelihood
Maximum likelihood estimation (MLE)
is used to estimate Probit and Logit functions.
The small sample properties of MLE
are not known, but in large samples
MLE is normally distributed, and it is
consistent and asymptotically efficient.
9.30
Copyright 1996 Lawrence C. Marsh
Heteroskedasticity
Chapter 10
Copyright 1997 John Wiley & Sons, Inc. All rights reserved. Reproduction or translation of this work beyond
that permitted in Section 117 of the 1976 United States Copyright Act without the express written permission of the
copyright owner is unlawful. Request for further information should be addressed to the Permissions Department,
John Wiley & Sons, Inc. The purchaser may make back-up copies for his/her own use only and not for distribution
or resale. The Publisher assumes no responsibility for errors, omissions, or damages, caused by the use of these
programs or from the use of the information contained herein.
10.1
Copyright 1996 Lawrence C. Marsh
The Nature of Heteroskedasticity
Heteroskedasticity is a systematic pattern in
the errors where the variances of the errors
are not constant.
Ordinary least squares assumes that all
observations are equally reliable.
For efficiency (accurate estimation/prediction)
reweight observations to ensure equal error
variance.
10.2
Copyright 1996 Lawrence C. Marsh
y
t
=
1
+
2
x
t
+ e
t
Regression Model
E(e
t
) = 0
var(e
t
) =
2
zero mean:
homoskedasticity:
nonautocorrelation: cov(e
t
, e
s
) = 0 t s
heteroskedasticity: var(e
t
) =
t
2
10.3
Copyright 1996 Lawrence C. Marsh
Homoskedastic pattern of errors
x
t
y
t
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
. .
.
.
.
.
.
.
.
.
income
consumption
10.4
Copyright 1996 Lawrence C. Marsh
.
.
x
t
x
1
x
2
y
t
f(y
t
)
The Homoskedastic Case
.
.
x
3
x
4
income
c
o
n
s
u
m
p
t
i
o
n
10.5
Copyright 1996 Lawrence C. Marsh
Heteroskedastic pattern of errors
x
t
y
t
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
. .
.
.
.
.
.
.
income
consumption
10.6
Copyright 1996 Lawrence C. Marsh
.
x
t
x
1
x
2
y
t
f(y
t
)
c
o
n
s
u
m
p
t
i
o
n
x
3
.
.
The Heteroskedastic Case
income
rich people
poor people
10.7
Copyright 1996 Lawrence C. Marsh
Properties of Least Squares
1. Least squares still linear and unbiased.
2. Least squares not efficient.
3. Usual formulas give incorrect standard
errors for least squares.
4. Confidence intervals and hypothesis tests
based on usual standard errors are wrong.
10.8
Copyright 1996 Lawrence C. Marsh
y
t
=
1
+
2
x
t
+ e
t
heteroskedasticity: var(e
t
) =
t
2
incorrect formula for least squares variance:
var(b
2
) =
2
(
x
t
x
)
2
correct formula for least squares variance:
var(b
2
) =
t
2
(
x
t
x
)
2
[ (
x
t
x
)
2
]
2
10.9
Copyright 1996 Lawrence C. Marsh
Hal Whites Standard Errors
Whites estimator of the least squares variance:
[Link](b
2
) =
e
t
2
(
x
t
x
)
2
[ (
x
t
x
)
2
]
2
^
In large samples Whites standard error
(square root of estimated variance) is a
correct / accurate / consistent measure.
10.10
Copyright 1996 Lawrence C. Marsh
Two Types of Heteroskedasticity
1. Proportional Heteroskedasticity.
(continuous function(of x
t
, for example))
2. Partitioned Heteroskedasticity.
(discrete categories/groups)
10.11
Copyright 1996 Lawrence C. Marsh
Proportional Heteroskedasticity
y
t
=
1
+
2
x
t
+ e
t
where
var(e
t
) =
t
2
E(e
t
) = 0 cov(e
t
, e
s
) = 0 t s
t
2
=
2
x
t
The variance is
assumed to be
proportional to
the value of x
t
10.12
Copyright 1996 Lawrence C. Marsh
t
2
=
2
x
t
y
t
=
1
+
2
x
t
+ e
t
[Link]. proportional to x
t
variance:
standard deviation:
t
=
x
t
y
t
1 x
t
e
t
=
1
+
2
+
x
t
x
t
x
t
x
t
To correct for heteroskedasticity divide the model by
x
t
var(e
t
) =
t
2
10.13
Copyright 1996 Lawrence C. Marsh
y
t
1 x
t
e
t
=
1
+
2
+
x
t
x
t
x
t
x
t
y
t
=
1
x
t1
+
2
x
t2
+ e
t
* * * *
var(e
t
) = var( ) = var(e
t
) =
2
x
t
*
e
t
x
t
1
x
t
1
x
t
var(e
t
) =
2 *
e
t
is heteroskedastic, but e
t
is homoskedastic.
*
10.14
Copyright 1996 Lawrence C. Marsh
1. Decide which variable is proportional to the
heteroskedasticity (x
t
in previous example).
2. Divide all terms in the original model by the
square root of that variable (divide by x
t
).
3. Run least squares on the transformed model
which has new y
t
, x
t1
and x
t2
variables
but no intercept.
Generalized Least Squares
These steps describe weighted least squares:
* * *
10.15
Copyright 1996 Lawrence C. Marsh
Partitioned Heteroskedasticity
y
t
=
1
+
2
x
t
+ e
t
var(e
t
) =
1
2
var(e
t
) =
2
2
error variance of field corn:
error variance of sweet corn:
y
t
= bushels per acre of corn
x
t
= gallons of water per acre (rain or other)
t = 1, . . . ,100
t = 1, . . . ,80
t = 81, . . . ,100
10.16
Copyright 1996 Lawrence C. Marsh
y
t
=
1
+
2
x
t
+ e
t
var(e
t
) =
1
2
field corn:
y
t
=
1
+
2
x
t
+ e
t
var(e
t
) =
2
2
sweet corn:
y
t
1 x
t
e
t
=
1
+
2
+
1
1
1
y
t
1 x
t
e
t
=
1
+
2
+
2
2
2
Reweighting Each Groups Observations
t = 1, . . . ,80
t = 81, . . . ,100
10.17
Copyright 1996 Lawrence C. Marsh
Apply Generalized Least Squares
Run least squares separately on data for each group.
1
2
provides estimator of
1
2
using
the 80 observations on field corn.
^
2
2
provides estimator of
2
2
using
the 20 observations on sweet corn.
^
10.18
Copyright 1996 Lawrence C. Marsh
1. Residual Plots provide information on the
exact nature of heteroskedasticity (partitioned
or proportional) to aid in correcting for it.
2. Goldfeld-Quandt Test checks for presence
of heteroskedasticity.
Detecting Heteroskedasticity
Determine existence and nature of heteroskedasticity:
10.19
Copyright 1996 Lawrence C. Marsh
Residual Plots
e
t
0
x
t
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
Plot residuals against one variable at a time
after sorting the data by that variable to try
to find a heteroskedastic pattern in the data.
10.20
Copyright 1996 Lawrence C. Marsh
Goldfeld-Quandt Test
The Goldfeld-Quandt test can be used to detect
heteroskedasticity in either the proportional case
or for comparing two groups in the discrete case.
For proportional heteroskedasticity, it is first necessary
to determine which variable, such as x
t
, is proportional
to the error variance. Then sort the data from the
largest to smallest values of that variable.
10.21
Copyright 1996 Lawrence C. Marsh
H
o
:
1
2
=
2
2
H
1
:
1
2
>
2
2
GQ = ~ F
[T
1
-K
1
, T
2
-K
2
]
1
2
2
2
^
^
In the proportional case, drop the middle
r observations where r T/6, then run
separate least squares regressions on the first
T
1
observations and the last T
2
observations.
Small values of GQ support H
o
while large values support H
1
.
Goldfeld-Quandt
Test Statistic
Use F
Table
10.22
Copyright 1996 Lawrence C. Marsh
t
2
=
2
exp{
1
z
t1
+
2
z
t2
}
More General Model
Structure of heteroskedasticity could be more complicated:
z
t1
and z
t2
are any observable variables upon
which we believe the variance could depend.
Note: The function exp{
.
} ensures that
t
2
is positive.
10.23
Copyright 1996 Lawrence C. Marsh
t
2
=
2
exp{
1
z
t1
+
2
z
t2
}
More General Model
ln(
t
2
)
= ln(
2
) +
1
z
t1
+
2
z
t2
ln(
t
2
)
=
0
+
1
z
t1
+
2
z
t2
where
0
= ln(
2
)
H
o
:
1
= 0,
2
= 0
H
1
:
1
0,
2
0
and/or
Least squares residuals, e
t
^
ln(e
t
2
)
=
0
+
1
z
t1
+
2
z
t2
+
t
^
the usual F test
10.24
Copyright 1996 Lawrence C. Marsh
Autocorrelation
Chapter 11
Copyright 1997 John Wiley & Sons, Inc. All rights reserved. Reproduction or translation of this work beyond
that permitted in Section 117 of the 1976 United States Copyright Act without the express written permission of the
copyright owner is unlawful. Request for further information should be addressed to the Permissions Department,
John Wiley & Sons, Inc. The purchaser may make back-up copies for his/her own use only and not for distribution
or resale. The Publisher assumes no responsibility for errors, omissions, or damages, caused by the use of these
programs or from the use of the information contained herein.
11.1
Copyright 1996 Lawrence C. Marsh
The Nature of Autocorrelation
For efficiency (accurate estimation/prediction)
all systematic information needs to be incor-
porated into the regression model.
Autocorrelation is a systematic pattern in the
errors that can be either attracting (positive)
or repelling (negative) autocorrelation.
11.2
Copyright 1996 Lawrence C. Marsh
Postive
Auto.
No
Auto.
Negative
Auto.
e
t
.
0
e
t
0
e
t
0
t
t
t
.
.
. .
.
.
.
.
.
.
. .
.
. .
.
.
. .
.
.
.
.
.
.
.
.
.
.
.
.
. .
.
.
.
.
.
.
.
.
.
..
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
crosses line not enough (attracting)
crosses line randomly
crosses line too much (repelling)
11.3
Copyright 1996 Lawrence C. Marsh
y
t
=
1
+
2
x
t
+ e
t
Regression Model
E(e
t
) = 0
var(e
t
) =
2
zero mean:
homoskedasticity:
nonautocorrelation: cov(e
t
, e
s
) = 0 t s
autocorrelation: cov(e
t
, e
s
) 0 t s
11.4
Copyright 1996 Lawrence C. Marsh
Order of Autocorrelation
y
t
=
1
+
2
x
t
+ e
t
e
t
= e
t1
+
t
e
t
=
1
e
t1
+
2
e
t2
+
t
e
t
=
1
e
t1
+
2
e
t2
+
3
e
t3
+
t
1st Order:
2nd Order:
3rd Order:
We will assume First Order Autocorrelation:
e
t
= e
t1
+
t
AR(1) :
11.5
Copyright 1996 Lawrence C. Marsh
First Order Autocorrelation
y
t
=
1
+
2
x
t
+ e
t
e
t
= e
t1
+
t
where 1 < < 1
E(
t
) = 0 var(
t
) =
2
cov(
t
,
s
) = 0 t s
These assumptions about
t
imply the following about e
t
:
E(e
t
) = 0
var(e
t
) =
e
2
=
cov(e
t
, e
tk
) =
e
2
k
for k > 0
corr(e
t
, e
tk
) =
k
for k > 0
2
1
2
11.6
Copyright 1996 Lawrence C. Marsh
Autocorrelation creates some
Problems for Least Squares:
1. The least squares estimator is still linear
and unbiased but it is not efficient.
2. The formulas normally used to compute
the least squares standard errors are no
longer correct and confidence intervals and
hypothesis tests using them will be wrong.
11.7
Copyright 1996 Lawrence C. Marsh
Generalized Least Squares
y
t
=
1
+
2
x
t
+ e
t
e
t
= e
t1
+
t
y
t
=
1
+
2
x
t
+ e
t1
+
t
substitute
in for e
t
Now we need to get rid of e
t1
(continued)
AR(1) :
11.8
Copyright 1996 Lawrence C. Marsh
y
t
=
1
+
2
x
t
+ e
t
y
t
=
1
+
2
x
t
+ e
t1
+
t
e
t
= y
t
1
2
x
t
e
t1
= y
t1
1
2
x
t1
y
t
=
1
+
2
x
t
+ (y
t1
1
2
x
t1
)
+
t
lag the
errors
once
(continued)
11.9
Copyright 1996 Lawrence C. Marsh
y
t
=
1
+
2
x
t
+ (y
t1
1
2
x
t1
)
+
t
y
t
=
1
+
2
x
t
+ y
t1
1
2
x
t1
+
t
y
t
y
t1
=
1
(1) +
2
(x
t
x
t1
)
+
t
y
t
=
1
+
2
x
t2
+
t
* * *
y
t
= y
t
y
t1
*
1
=
1
(1)
*
x
t2
= (x
t
x
t1
)
*
11.10
Copyright 1996 Lawrence C. Marsh
y
t
=
1
+
2
x
t2
+
t
* * *
y
t
= y
t
y
t1
*
1
=
1
(1)
*
x
t2
= x
t
x
t1
*
Problems estimating this model with least squares:
1. One observation is used up in creating the
transformed (lagged) variables leaving only
(T1) observations for estimating the model.
2. The value of is not known. We must find
some way to estimate it.
11.11
Copyright 1996 Lawrence C. Marsh
Recovering the 1st Observation
Dropping the 1st observation and applying least squares
is not the best linear unbiased estimation method.
Efficiency is lost because the variance
of the error associated with the 1st observation
is not equal to that of the other errors.
This is a special case of the heteroskedasticity
problem except that here all errors are assumed
to have equal variance except the 1st error.
11.12
Copyright 1996 Lawrence C. Marsh
Recovering the 1st Observation
y
1
=
1
+
2
x
1
+ e
1
The 1st observation should fit the original model as:
We could include this as the 1st observation for our
estimation procedure but we must first transform it so
that it has the same error variance as the other observations.
with error variance: var(e
1
) =
e
2
=
2
/(1-
2
).
Note: The other observations all have error variance
2
.
11.13
Copyright 1996 Lawrence C. Marsh
y
1
=
1
+
2
x
1
+ e
1
with error variance: var(e
1
) =
e
2
=
2
/(1-
2
).
The other observations all have error variance
2
.
Given any constant c : var(ce
1
) = c
2
var(e
1
).
If c = 1-
2
, then var( 1-
2
e
1
) = (1-
2
) var(e
1
).
= (1-
2
)
e
2
= (1-
2
)
2
/(1-
2
)
=
2
The transformation
1
= 1-
2
e
1
has variance
2
.
11.14
Copyright 1996 Lawrence C. Marsh
y
1
=
1
+
2
x
1
+ e
1
The transformed error
1
= 1-
2
e
1
has variance
2
.
Multiply through by 1-
2
to get:
1-
2
y
1
= 1-
2
1
+ 1-
2
2
x
1
+ 1-
2
e
1
This transformed first observation may now be
added to the other (T-1) observations to obtain
the fully restored set of T observations.
11.15
Copyright 1996 Lawrence C. Marsh
Estimating Unknown Value
e
t
= e
t1
+
t
First, use least squares to estimate the model:
If we had values for the e
t
s, we could estimate:
y
t
=
1
+
2
x
t
+ e
t
The residuals from this estimation are:
e
t
= y
t
- b
1
- b
2
x
t
^
11.16
Copyright 1996 Lawrence C. Marsh
e
t
= y
t
- b
1
- b
2
x
t
^
e
t
= e
t1
+
t
^ ^ ^
Next, estimate the following by least squares:
The least squares solution is:
e
t
e
t-1
e
t-1
T
T
t = 2
t = 2
2
^ ^
^
=
^
11.17
Copyright 1996 Lawrence C. Marsh
Durbin-Watson Test
H
o
: = 0 vs. H
1
: 0 , > 0, or < 0
(
e
t
e
t-1
)
e
t
T
T
t = 2
t = 1
2
^ ^
^
d =
2
The Durbin-Watson Test statistic, d, is :
11.18
Copyright 1996 Lawrence C. Marsh
Testing for Autocorrelation
The test statistic, d, is approximately related to as:
^
d 2(1)
^
When = 0 , the Durbin-Watson statistic is d 2.
^
When = 1 , the Durbin-Watson statistic is d 0.
^
Tables for critical values for d are not always
readily available so it is easier to use the p-value
that most computer programs provide for d.
Reject H
o
if p-value < , the significance level.
11.19
Copyright 1996 Lawrence C. Marsh
Prediction with AR(1) Errors
When errors are autocorrelated, the previous periods
error may help us predict next periods error.
The best predictor, y
T+1
, for next period is:
y
T+1
=
1
+
2
x
T+1
+ e
T
^
^ ^
^
~
where
1
and
2
are generalized least squares
estimates and e
T
is given by:
~
^ ^
e
T
= y
T
1
2
x
T
^ ^ ~
11.20
Copyright 1996 Lawrence C. Marsh
y
T+h
=
1
+
2
x
T+h
+
h
e
T
^
^ ^
^
~
For h periods ahead, the best predictor is:
Assuming | | < 1, the influence of
h
e
T
diminishes the further we go into the future
(the larger h becomes).
^ ^
~
11.21
Copyright 1996 Lawrence C. Marsh
Pooling
Time-Series and
Cross-Sectional Data
Chapter 12
Copyright 1997 John Wiley & Sons, Inc. All rights reserved. Reproduction or translation of this work beyond
that permitted in Section 117 of the 1976 United States Copyright Act without the express written permission of the
copyright owner is unlawful. Request for further information should be addressed to the Permissions Department,
John Wiley & Sons, Inc. The purchaser may make back-up copies for his/her own use only and not for distribution
or resale. The Publisher assumes no responsibility for errors, omissions, or damages, caused by the use of these
programs or from the use of the information contained herein.
12.1
Copyright 1996 Lawrence C. Marsh
Pooling Time and Cross Sections
y
it
=
1it
+
2it
x
2it
+
3it
x
3it
+ e
it
If left unrestricted,
this model requires different equations
for each firm in each time period.
for the i
th
firm in the t
th
time period
12.2
Copyright 1996 Lawrence C. Marsh
Seemingly Unrelated Regressions
y
it
=
1i
+
2i
x
2it
+
3i
x
3it
+ e
it
SUR models impose the restrictions:
1it
=
1i
2it
=
2i
3it
=
3i
Each firm gets its own coefficients:
1i
,
2i
and
3i
but those coefficients are constant over time.
12.3
Copyright 1996 Lawrence C. Marsh
The investment expenditures (INV) of General Electric (G)
and Westinghouse(W) may be related to their stock market
value (V) and actual capital stock (K) as follows:
INV
Gt
=
1G
+
2G
V
Gt
+
3G
K
Gt
+ e
Gt
INV
Wt
=
1W
+
2W
V
Wt
+
3W
K
Wt
+ e
Wt
i = G, W t = 1, . . . , 20
Two-Equation SUR Model
12.4
Copyright 1996 Lawrence C. Marsh
Estimating Separate Equations
For now make the assumption of no correlation
between the error terms across equations:
We make the usual error term assumptions:
cov(e
Gt
, e
Gs
) = 0 cov(e
Wt
, e
Ws
) = 0
var(e
Gt
) =
G
2
var(e
Wt
) =
W
2
E(e
Gt
) = 0 E(e
Wt
) = 0
cov(e
Gt
, e
Wt
) = 0 cov(e
Gt
, e
Ws
) = 0
12.5
Copyright 1996 Lawrence C. Marsh
homoskedasticity assumption:
G
=
W
2 2
INV
t
=
1G
+
1
D
t
+
2G
V
t
+
2
D
t
V
t
+
3G
K
t
+
3
D
t
K
t
+ e
t
Dummy variable model assumes that :
G
=
W
2 2
For Westinghouse observations D
t
= 1; otherwise D
t
= 0.
1W
=
1G
+
1
2W
=
2G
+
2
3W
=
3G
+
3
12.6
Copyright 1996 Lawrence C. Marsh
Problem with OLS on Each Equation
The first assumption of the Gauss-Markov
Theorem concerns the model specification.
If the model is not fully and correctly specified
the Gauss-Markov properties might not hold.
Any correlation of error terms across equations
must be part of model specification.
12.7
Copyright 1996 Lawrence C. Marsh
Any correlation between the
dependent variables of two or
more equations that is not due
to their explanatory variables
is by default due to correlated
error terms.
Correlated Error Terms
12.8
Copyright 1996 Lawrence C. Marsh
1. Sales of Pepsi vs. sales of Coke.
(uncontrolled factor: outdoor temperature)
2. Investments in bonds vs. investments in stocks.
(uncontrolled factor: computer/appliance sales)
3. Movie admissions vs. Golf Course admissions.
(uncontrolled factor: weather conditions)
4. Sales of butter vs. sales of bread.
(uncontrolled factor: bagels and cream cheese)
Which of the following models would
be likely to produce positively correlated
errors and which would produce
negatively correlations errors?
12.9
Copyright 1996 Lawrence C. Marsh
Joint Estimation of the Equations
INV
Gt
=
1G
+
2G
V
Gt
+
3G
K
Gt
+ e
Gt
INV
Wt
=
1W
+
2W
V
Wt
+
3W
K
Wt
+ e
Wt
cov(e
Gt
, e
Wt
) =
GW
12.10
Copyright 1996 Lawrence C. Marsh
Seemingly Unrelated Regressions
When the error terms of two or more equations
are correlated, efficient estimation requires the use
of a Seemingly Unrelated Regressions (SUR)
type estimator to take the correlation into account.
Be sure to use the Seemingly Unrelated Regressions (SUR)
procedure in your regression software program to estimate
any equations that you believe might have correlated errors.
12.11
Copyright 1996 Lawrence C. Marsh
Separate vs. Joint Estimation
SUR will give exactly the same results as estimating
each equation separately with OLS if either or both
of the following two conditions are true:
1. Every equation has exactly the same set of
explanatory variables with exactly the same
values.
2. There is no correlation between the error
terms of any of the equations.
12.12
Copyright 1996 Lawrence C. Marsh
Test for Correlation
:
GW
= 0
Test the null hypothesis of zero correlation
GW
W
^
^ ^
r
GW
=
2
2
2
2 = T r
GW
2
2
(1)
asy.
12.13
Copyright 1996 Lawrence C. Marsh
Start with
the residuals
e
Gt
and e
Wt
from each
equation
estimated
separately.
^
^
GW
W
^
^ ^
r
GW
=
2
2
2
2
= T r
GW
2
2
(1)
asy.
GW
= e
Gt
e
Wt
1
T
^ ^
^
G
= e
Gt
1
T
^ ^
2
2
W
= e
Wt
1
T
^ ^
2 2
12.14
Copyright 1996 Lawrence C. Marsh
Fixed Effects Model
y
it
=
1it
+
2it
x
2it
+
3it
x
3it
+ e
it
y
it
=
1i
+
2
x
2it
+
3
x
3it
+ e
it
Fixed effects models impose the restrictions:
1it
=
1i
2it
=
2
3it
=
3
For each i
th
cross section in the t
th
time period:
Each i
th
cross-section has its own constant
1i
intercept.
12.15
Copyright 1996 Lawrence C. Marsh
The Fixed Effects Model is conveniently
represented using dummy variables:
y
it
=
11
D
1i
+
12
D
2i
+
13
D
3i
+
14
D
4 i
+
2
x
2it
+
3
x
3it
+ e
it
D
1i
=1 if North
D
1i
=0 if not N
D
2i
=1 if East
D
2i
=0 if not E
D
3i
=1 if South
D
3i
=0 if not S
D
4i
=1 if West
D
4i
=0 if not W
y
it
= millions of bushels of corn produced
x
2it
= price of corn in dollars per bushel
x
3it
= price of soybeans in dollars per bushel
Each cross-sectional unit gets its own intercept,
but each cross-sectional intercept is constant over time.
12.16
Copyright 1996 Lawrence C. Marsh
H
o
:
11
=
12
=
13
=
14
Test for Equality of Fixed Effects
H
1
: H
o
not true
The H
o
joint null hypothesis may be tested with F-statistic:
(SSE
R
SSE
U
) / J
SSE
U
/ (NT K)
F = ~ F
(NT K)
J
SSE
R
is the restricted error sum of squares (one intercept)
SSE
U
is the unrestricted error sum of squares (four intercepts)
N is the number of cross-sectional units (N = 4)
K is the number of parameters in the model (K = 6)
J is the number of restrictions being tested (J = N1 = 3)
T is the number of time periods
12.17
Copyright 1996 Lawrence C. Marsh
Random Effects Model
y
it
=
1i
+
2
x
2it
+
3
x
3it
+ e
it
1i
=
1
+
i
1
is the population mean intercept.
i
is an unobservable random error that
accounts for the cross-sectional differences.
12.18
Copyright 1996 Lawrence C. Marsh
1i
=
1
+
i
i
are independent of one another and of e
it
E(
i
) = 0
var(
i
) =
2
where i = 1, ... ,N
Consequently,
E(
1i
) =
1
var(
1i
) =
2
Random Intercept Term
12.19
Copyright 1996 Lawrence C. Marsh
y
it
=
1i
+
2
x
2it
+
3
x
3it
+ e
it
y
it
= (
1
+
i
) +
2
x
2it
+
3
x
3it
+ e
it
y
it
=
1
+
2
x
2it
+
3
x
3it
+ (
i
+e
it
)
y
it
=
1
+
2
x
2it
+
3
x
3it
+
it
Random Effects Model
12.20
Copyright 1996 Lawrence C. Marsh
it
= (
i
+e
it
)
y
it
=
1
+
2
x
2it
+
3
x
3it
+
it
it
has zero mean: E(
it
) = 0
it
is homoskedastic: var(
it
) =
+
e
2 2
The errors from the same firm in different time periods
are correlated:
The errors from different firms are always uncorrelated:
cov(
it
,
is
) =
2
cov(
it
,
js
) = 0
t s
i j
12.21
Copyright 1996 Lawrence C. Marsh
Simultaneous
Equations
Models
Chapter 13
Copyright 1997 John Wiley & Sons, Inc. All rights reserved. Reproduction or translation of this work beyond
that permitted in Section 117 of the 1976 United States Copyright Act without the express written permission of the
copyright owner is unlawful. Request for further information should be addressed to the Permissions Department,
John Wiley & Sons, Inc. The purchaser may make back-up copies for his/her own use only and not for distribution
or resale. The Publisher assumes no responsibility for errors, omissions, or damages, caused by the use of these
programs or from the use of the information contained herein.
13.1
Copyright 1996 Lawrence C. Marsh
Keynesian Macro Model
Assumptions of Simple Keynesian Model
1. Consumption, c, is function of income, y.
2. Total expenditures = consumption + investment.
3. Investment assumed independent of income.
13.2
Copyright 1996 Lawrence C. Marsh
consumption is a function of income:
income is either consumed or invested:
c =
1
+
2
y
y = c + i
The Structural Equations
13.3
Copyright 1996 Lawrence C. Marsh
The Statistical Model
c
t
=
1
+
2
y
t
+ e
t
y
t
= c
t
+ i
t
The consumption equation:
The income identity:
13.4
Copyright 1996 Lawrence C. Marsh
The Simultaneous Nature
of Simultaneous Equations
c
t
=
1
+
2
y
t
+ e
t
y
t
= c
t
+ i
t
Since y
t
contains e
t
they are
correlated
2. 1.
3.
4.
5.
13.5
Copyright 1996 Lawrence C. Marsh
The Failure of Least Squares
The least squares estimators of
parameters in a structural simul-
taneous equation is biased and
inconsistent because of the cor-
relation between the random error
and the endogenous variables on
the right-hand side of the equation.
13.6
Copyright 1996 Lawrence C. Marsh
Single Equation:
Simultaneous Equations:
Single vs. Simultaneous Equations
c
t
y
t
e
t
c
t
y
t
i
t
e
t
13.7
Copyright 1996 Lawrence C. Marsh
Deriving the Reduced Form
c
t
=
1
+
2
y
t
+ e
t
y
t
= c
t
+ i
t
c
t
=
1
+
2
(c
t
+ i
t
) + e
t
(1
2
)c
t
=
1
+
2
i
t
+ e
t
13.8
Copyright 1996 Lawrence C. Marsh
Deriving the Reduced Form
(1
2
)c
t
=
1
+
2
i
t
+ e
t
c
t
= + i
t
+ e
t
(1
2
)
(1
2
) (1
2
)
1
2
c
t
=
11
+
21
i
t
+
t
The Reduced Form Equation
13.9
Copyright 1996 Lawrence C. Marsh
Reduced Form Equation
c
t
=
11
+
21
i
t
+
t
(1
2
)
11
=
(1
2
)
21
=
(1
2
)
1
t
= + e
t
and
13.10
Copyright 1996 Lawrence C. Marsh
y
t
= c
t
+ i
t
where c
t
=
11
+
21
i
t
+
t
y
t
=
12
+
22
i
t
+
t
It is sometimes useful to give this equation
its own reduced form parameters as follows:
y
t
=
11
+ (1+
21
) i
t
+
t
13.11
Copyright 1996 Lawrence C. Marsh
y
t
=
12
+
22
i
t
+
t
c
t
=
11
+
21
i
t
+
t
Since c
t
and y
t
are related through the identity:
y
t
= c
t
+ i
t
, the error term,
t
, of these two
equations is the same, and it is easy to
show that:
(1
2
)
11
=
12
=
(1
2
)
22
= (1
21
)
=
1
13.12
Copyright 1996 Lawrence C. Marsh
Identification
The structural parameters are
1
and
2
.
The reduced form parameters are
11
and
21
.
Once the reduced form parameters are estimated,
the identification problem is to determine if the
orginal structural parameters can be expressed
uniquely in terms of the reduced form parameters.
(1+
21
)
2
=
21
^
^
^
(1+
21
)
1
=
11
^
^
^
13.13
Copyright 1996 Lawrence C. Marsh
Identification
An equation is exactly identified if its structural
(behavorial) parameters can be uniquely expres-
sed in terms of the reduced form parameters.
An equation is over-identified if there is more
than one solution for expressing its structural
(behavorial) parameters in terms of the reduced
form parameters.
An equation is under-identified if its structural
(behavorial) parameters cannot be expressed
in terms of the reduced form parameters.
13.14
Copyright 1996 Lawrence C. Marsh
The Identification Problem
A system of M equations
containing M endogenous
variables must exclude at least
M1 variables from a given
equation in order for the
parameters of that equation to
be identified and to be able to
be consistently estimated.
13.15
Copyright 1996 Lawrence C. Marsh
Two Stage Least Squares
Problem: right-hand endogenous variables
y
t2
and y
t1
are correlated with the error terms.
y
t1
=
1
+
2
y
t2
+
3
x
t1
+ e
t1
y
t2
=
1
+
2
y
t1
+
3
x
t2
+ e
t2
13.16
Copyright 1996 Lawrence C. Marsh
Problem: right-hand endogenous variables
y
t2
and y
t1
are correlated with the error terms.
Solution: First, derive the reduced form equations.
y
t1
=
1
+
2
y
t2
+
3
x
t1
+ e
t1
y
t2
=
1
+
2
y
t1
+
3
x
t2
+ e
t2
y
t1
=
11
+
21
x
t1
+
31
x
t2
+
t1
y
t2
=
12
+
22
x
t1
+
32
x
t2
+
t2
Solve two equations for two unknowns, y
t1
, y
t2
:
13.17
Copyright 1996 Lawrence C. Marsh
y
t1
=
11
+
21
x
t1
+
31
x
t2
+
t1
y
t2
=
12
+
22
x
t1
+
32
x
t2
+
t2
Use least squares to get fitted values:
2SLS: Stage I
y
t1
=
11
+
21
x
t1
+
31
x
t2
^
^ ^ ^
y
t2
=
12
+
22
x
t1
+
32
x
t2
^
^ ^ ^ y
t2
= y
t2
+
t2
^
^
y
t1
= y
t1
+
t1
^
^
13.18
Copyright 1996 Lawrence C. Marsh
2SLS: Stage II
y
t2
= y
t2
+
t2
^
^
y
t1
= y
t1
+
t1
^
^
and
y
t1
=
1
+
2
y
t2
+
3
x
t1
+ e
t1
y
t2
=
1
+
2
y
t1
+
3
x
t2
+ e
t2
Substitue in
for y
t1
, y
t2
y
t1
=
1
+
2
(y
t2
+
t2
) +
3
x
t1
+ e
t1
^
^
y
t2
=
1
+
2
(y
t1
+
t1
) +
3
x
t2
+ e
t2
^ ^
13.19
Copyright 1996 Lawrence C. Marsh
2SLS: Stage II (continued)
y
t1
=
1
+
2
y
t2
+
3
x
t1
+ u
t1
^
y
t2
=
1
+
2
y
t1
+
3
x
t2
+ u
t2
^
^
u
t1
=
2
t2
+ e
t1
u
t2
=
2
t1
+ e
t2
^
where and
Run least squares on each of the above equations
to get 2SLS estimates:
1
,
2
,
3
,
1
,
2
and
3
~ ~ ~
~ ~ ~
13.20
Copyright 1996 Lawrence C. Marsh
Nonlinear
Least
Squares
Chapter 14
Copyright 1997 John Wiley & Sons, Inc. All rights reserved. Reproduction or translation of this work beyond
that permitted in Section 117 of the 1976 United States Copyright Act without the express written permission of the
copyright owner is unlawful. Request for further information should be addressed to the Permissions Department,
John Wiley & Sons, Inc. The purchaser may make back-up copies for his/her own use only and not for distribution
or resale. The Publisher assumes no responsibility for errors, omissions, or damages, caused by the use of these
programs or from the use of the information contained herein.
14.1
Copyright 1996 Lawrence C. Marsh
(A.) Regression model with only an intercept term:
Review of Least Squares Principle
y
t
= + e
t
e
t
= y
t
e
t
= (y
t
)
2
2
SSE = (y
t
)
2
SSE
= 2 (y
t
) = 0
^
y
t
= 0
^
y
t
= 0
^
= y
t
= y
^
1
T
(minimize the sum of squared errors)
Yields an exact analytical solution:
14.2
Copyright 1996 Lawrence C. Marsh
Review of Least Squares
(B.) Regression model without an intercept term:
y
t
= x
t
+ e
t
e
t
= y
t
x
t
e
t
= (y
t
x
t
)
2
2
SSE = (y
t
x
t
)
2
SSE
= 2 x
t
(y
t
x
t
)= 0
^
x
t
y
t
x
t
= 0
^ 2
x
t
y
t
x
t
= 0
^ 2
x
t
= x
t
y
t
^
2
=
^
x
t
y
t
2
x
t
This yields an exact
analytical solution:
14.3
Copyright 1996 Lawrence C. Marsh
Review of Least Squares
(C.) Regression model with both an intercept and a slope:
y
t
= + x
t
+ e
t
SSE = (y
t
x
t
)
2
SSE
= 2 (y
t
x
t
) = 0
^
^
SSE
= 2 x
t
(y
t
x
t
) = 0
^
^
y
x = 0
^
^
=
^
(x
t
x)(y
t
y)
(x
t
x)
2
x
t
y
t
x
t
x
t
= 0
^
^
2
This yields an exact
analytical solution:
= y
x
^
^
14.4
Copyright 1996 Lawrence C. Marsh
Nonlinear Least Squares
(D.) Nonlinear Regression model:
y
t
= x
t
+ e
t
SSE = (y
t
x
t
)
2
PROBLEM: An exact
analytical solution to
this does not exist.
SSE
= 2 x
t
ln(x
t
)(y
t
x
t
) = 0
^ ^
[x
t
ln(x
t
)y
t
]
[x
t
2
ln(x
t
)] = 0
^ ^
Must use numerical
search algorithm to
try to find value of
to satisfy this.
14.5
Copyright 1996 Lawrence C. Marsh
Find Minimum of Nonlinear SSE
^
SSE
SSE = (y
t
x
t
)
2
14.6
Copyright 1996 Lawrence C. Marsh
The least squares principle
is still appropriate when the
model is nonlinear, but it is
harder to find the solution.
Conclusion
14.7
Copyright 1996 Lawrence C. Marsh
Nonlinear least squares
optimization methods:
The Gauss-Newton Method
Optional Appendix
14.8
Copyright 1996 Lawrence C. Marsh
The Gauss-Newton Algorithm
1. Apply the Taylor Series Expansion to the
nonlinear model around some initial b
(o)
.
2. Run Ordinary Least Squares (OLS) on the
linear part of the Taylor Series to get b
(m)
.
3. Perform a Taylor Series around the new b
(m)
to get b
(m+1)
.
4. Relabel b
(m+1)
as b
(m)
and rerun steps 2.-4.
5. Stop when (b
(m+1)
b
(m)
) becomes very small.
14.9
Copyright 1996 Lawrence C. Marsh
The Gauss-Newton Method
y
t
= f(X
t
,b) +
t
for t = 1, . . . , n.
Do a Taylor Series Expansion around the vector b = b
(o)
as follows:
y
t
= f(X
t
,b
()
) + f(X
t
,b
()
)(b - b
()
) +
t
where
t
(b - b
(o)
)
T
f(X
t
,b
()
)(b - b
()
) + R
t
+
t
f(X
t
,b) = f(X
t
,b
()
) + f(X
t
,b
()
)(b - b
()
)
+ (b - b
()
)
T
f(X
t
,b
()
)(b - b
()
) + R
t
14.10
Copyright 1996 Lawrence C. Marsh
y
t
= f(X
t
,b
()
) + f(X
t
,b
()
)(b - b
()
) +
t
y
t
- f(X
t
,b
()
) = f(X
t
,b
()
)b - f(X
t
,b
()
) b
()
+
t
y
t
- f(X
t
,b
()
) + f(X
t
,b
()
) b
()
= f(X
t
,b
()
)b +
t
y
t
()
= f(X
t
,b
()
)b +
t
where y
t
()
y
t
- f(X
t
,b
()
) + f(X
t
,b
()
) b
()
This is linear in b .
Gauss-Newton just runs OLS on this
transformed truncated Taylor series.
14.11
Copyright 1996 Lawrence C. Marsh
y
t
()
= f(X
t
,b
()
)b +
t
()
= f(X,b
()
)b +
()
^
This is analogous to linear OLS where
y = Xb + led to the solution: b = (X
T
X)
1
X
T
y
^
except that X is replaced with the matrix of first
partial derivatives: f(X
t
,b
()
) and y is replaced by y
()
(i.e. y = y*
()
and X = f(X,b
()
) )
14.12
Copyright 1996 Lawrence C. Marsh
Recall that: y
*
(o)
y f(X,b
(o)
) + f(X,b
()
) b
()
Now define: y
()
y f(X,b
(o)
)
Therefore: y
()
= y
()
+ f(X,b
()
) b
()
b = [ f(X,b
()
)
T
f(X,b
()
)]
-1
f(X,b
()
)
T
y
()
^
Now substitute in for y
in Gauss-Newton solution:
to get:
b = b
(o)
+
[ f(X,b
()
)
T
f(X,b
()
)]
-1
f(X,b
()
)
T
y
()
^
14.13
Copyright 1996 Lawrence C. Marsh
b = b
(o)
+
[ f(X,b
()
)
T
f(X,b
()
)]
-1
f(X,b
()
)
T
y
()
^
b
(1)
= b
()
+
[ f(X,b
()
)
T
f(X,b
()
)]
-1
f(X,b
()
)
T
y
()
Now call this b value b
(1)
as follows:
^
More generally, in going from interation m to
iteration (m+1) we obtain the general expression:
b
(m+1)
= b
(m)
+
[ f(X,b
(m)
)
T
f(X,b
(m)
)]
-1
f(X,b
(m)
)
T
y
(m)
14.14
Copyright 1996 Lawrence C. Marsh
b
(m+1)
= [ f(X,b
(m)
)
T
f(X,b
(m)
)]
-1
f(X,b
(m)
)
T
y
*
(m)
b
(m+1)
= b
(m)
+
[ f(X,b
(m)
)
T
f(X,b
(m)
)]
-1
f(X,b
(m)
)
T
y
(m)
Thus, the Gauss-Newton (nonlinear OLS) solution
can be expressed in two alternative, but equivalent,
forms:
1. replacement form:
2. updating form:
14.15
Copyright 1996 Lawrence C. Marsh
For example, consider Durbins Method of estimating
the autocorrelation coefficient under a first-order
autoregression regime:
y
t
= b
1
+ b
2
X
t 2
+ . . . + b
K
X
t K
+
t
for t = 1, . . . , n.
t
=
t - 1
+ u
t
where u
t
satisfies the conditions
E u
t
= 0 , E u
2
t
= s
u
2
, E u
t
u
s
= 0 for s t.
Therefore, u
t
is nonautocorrelated and homoskedastic.
Durbins Method is to set aside a copy of the equation,
lag it once, multiply by and subtract the new equation
from the original equation, then move the y
t-1
term to
the right side and estimate along with the b
s
by OLS.
14.16
Copyright 1996 Lawrence C. Marsh
Durbins Method is to set aside a copy of the equation,
lag it once, multiply by and subtract the new equation
from the original equation, then move the y
t-1
term to
the right side and estimate along with the bs by OLS.
y
t
= b
1
+ b
2
X
t 2
+ b
3
X
t 3
+
t
for t = 1, . . . , n.
where
t
=
t - 1
+ u
t
y
t-1
= b
1
+ b
2
X
t -1, 2
+ b
3
X
t -1, 3
+
t -1
Lag once and multiply by :
Subtract from the original and move y
t-1
to right side:
y
t
= b
1
(1-) + b
2
(X
t 2
- X
t-1, 2
)
+ b
3
(X
t 3
X
t-1, 3
)+ y
t-1
+ u
t
14.17
Copyright 1996 Lawrence C. Marsh
y
t
= b
1
(1-) + b
2
(X
t 2
- X
t-1, 2
)
+ b
3
(X
t 3
- X
t-1, 3
) + y
t-1
+ u
t
Now Durbin separates out the terms as follows:
y
t
= b
1
(1-) + b
2
X
t 2
- b
2
X
t-1 2
+ b
3
X
t 3
- b
3
X
t-1 3
+ y
t-1
+ u
t
The structural (restricted,behavorial) equation is:
The corresponding reduced form (unrestricted) equation is:
y
t
=
1
+
2
X
t, 2
+
3
X
t-1, 2
+
4
X
t, 3
+
5
X
t-1, 3
+
6
y
t-1
+ u
t
1
= b
1
(1-)
2
= b
2
3
= - b
2
4
= b
3
5
= - b
3
6
=
14.18
Copyright 1996 Lawrence C. Marsh
Given OLS estimates:
1
6
^ ^ ^ ^ ^ ^
we can get three separate and distinct estimates for :
=
3
2
^
^
^
=
5
4
^
^
^
=
6
^ ^
These three separate estimates of are in conflict !!!
It is difficult to know which one to use as the
legitimate estimate of . Durbin used the last one.
1
= b
1
(1-)
2
= b
2
3
= - b
2
4
= b
3
5
= - b
3
6
=
14.19
Copyright 1996 Lawrence C. Marsh
The problem with Durbins Method is that it ignores
the inherent nonlinear restrictions implied by this
structural model. To get a single (i.e. unique) estimate
for the implied nonlinear restrictions must be
incorporated directly into the estimation process.
Consequently, the above structural equation should be
estimated using a nonlinear method such as the
Gauss-Newton algorithm for nonlinear least squares.
y
t
= b
1
(1-) + b
2
X
t 2
- b
2
X
t -1, 2
+ b
3
X
t 3
- b
3
X
t -1, 3
+ y
t-1
+ u
t
14.20
Copyright 1996 Lawrence C. Marsh
y
t
= b
1
(1-) + b
2
X
t 2
- b
2
X
t-1, 2
+ b
3
X
t 3
- b
3
X
t-1, 3
+ y
t-1
+ u
t
f(X
t
,b) = [ ]
y
t
= (1 )
= (X
t, 2
X
t-1,2
)
= (X
t, 3
X
t-1,3
)
y
t
= ( - b
1
- b
2
X
t-1,2
- b
3
X
t-1,3
+ y
t-1
)
y
t
b
1
y
t
b
2
y
t
b
3
y
t
b
1
y
t
b
2
y
t
b
3
14.21
Copyright 1996 Lawrence C. Marsh
where y
t
(m)
y
t
- f(X
t
,b
(m)
) + f(X
t
,b
(m)
) b
(m)
(m+1)
= [ f(X,b
(m)
)
T
f(X,b
(m)
)]
-1
f(X,b
(m)
)
T
y
(m)
^
f(X
t
,b) = b
1
(1-) + b
2
X
t 2
- b
2
X
t-1 2
+ b
3
X
t 3
- b
3
X
t-1 3
+ y
t-1
b
(m)
=
b
1(m)
(m)
b
2(m)
b
3(m)
Iterate until convergence.
f(X
t
,b
(m)
) = [ ]
y
t
(m)
y
t
b
1(m)
y
t
b
2(m)
y
t
b
3(m)
14.22
Copyright 1996 Lawrence C. Marsh
Distributed
Lag Models
Chapter 15
Copyright 1997 John Wiley & Sons, Inc. All rights reserved. Reproduction or translation of this work beyond
that permitted in Section 117 of the 1976 United States Copyright Act without the express written permission of the
copyright owner is unlawful. Request for further information should be addressed to the Permissions Department,
John Wiley & Sons, Inc. The purchaser may make back-up copies for his/her own use only and not for distribution
or resale. The Publisher assumes no responsibility for errors, omissions, or damages, caused by the use of these
programs or from the use of the information contained herein.
15.1
Copyright 1996 Lawrence C. Marsh
The Distributed Lag Effect
Economic action
at time t
Effect
at time t
Effect
at time t+1
Effect
at time t+2
15.2
Copyright 1996 Lawrence C. Marsh
Unstructured Lags
y
t
= +
0
x
t
+
1
x
t-1
+
2
x
t-2
+ . . . +
n
x
t-n
+ e
t
n unstructured lags
no systematic structure imposed on the s
the s are unrestricted
15.3
Copyright 1996 Lawrence C. Marsh
Problems with Unstructured Lags
1. n observations are lost with n-lag setup.
2. high degree of multicollinearity among x
t-j
s.
3. many degrees of freedom used for large n.
4. could get greater precision using structure.
15.4
Copyright 1996 Lawrence C. Marsh
The Arithmetic Lag Structure
proposed by Irving Fisher (1937)
the lag weights decline linearly
Imposing the relationship:
#
= (n - # + 1)
0
= (n+1)
1
= n
2
= (n-1)
3
= (n-2)
.
.
n-2
= 3
n-1
= 2
n
=
only need to estimate one coefficient, ,
instead of n+1 coefficients,
0
, ... ,
n
.
15.5
Copyright 1996 Lawrence C. Marsh
Arithmetic Lag Structure
y
t
= +
0
x
t
+
1
x
t-1
+
2
x
t-2
+ . . . +
n
x
t-n
+ e
t
y
t
= + (n+1) x
t
+ n x
t-1
+ (n-1) x
t-2
+ . . . + x
t-n
+ e
t
Step 1: impose the restriction:
#
= (n - # + 1)
Step 2: factor out the unknown coefficient, .
y
t
= + [(n+1)x
t
+ nx
t-1
+ (n-1)x
t-2
+ . . . + x
t-n
] + e
t
15.6
Copyright 1996 Lawrence C. Marsh
Arithmetic Lag Structure
Step 3: Define z
t
.
y
t
= + [(n+1)x
t
+ nx
t-1
+ (n-1)x
t-2
+ . . . + x
t-n
] + e
t
z
t
= [(n+1)x
t
+ nx
t-1
+ (n-1)x
t-2
+ . . . + x
t-n
]
Step 5: Run least squares regression on:
y
t
= + z
t
+ e
t
Step 4: Decide number of lags, n.
For n = 4: z
t
= [ 5x
t
+ 4x
t-1
+ 3x
t-2
+ 2x
t-3
+ x
t-4
]
15.7
Copyright 1996 Lawrence C. Marsh
Arithmetic Lag Structure
i
0
= (n+1)
1
= n
2
= (n-1)
n
=
.
.
.
0 1 2 . . . . . n n+1
.
.
.
.
linear
lag
structure
15.8
Copyright 1996 Lawrence C. Marsh
Polynomial Lag Structure
proposed by Shirley Almon (1965)
the lag weights fit a polynomial
where i = 1, . . . , n
p = 2 and n = 4
For example, a quadratic polynomial:
0
=
0
1
=
0
+
1
+
2
2
=
0
+ 2
1
+ 4
2
3
=
0
+ 3
1
+ 9
2
4
=
0
+ 4
1
+ 16
2
n = the length of the lag
p = degree of polynomial
where i = 1, . . . , n
i
=
0
+
1
i +
2
i
+...+
p
i
2 p
i
=
0
+
1
i +
2
i
2
15.9
Copyright 1996 Lawrence C. Marsh
Polynomial Lag Structure
y
t
= +
0
x
t
+
1
x
t-1
+
2
x
t-2
+
3
x
t-3
+
4
x
t-4
+ e
t
y
t
= +
0
x
t
+ (
0
+
1
+
2
)x
t-1
+ (
0
+ 2
1
+ 4
2
)x
t-2
+ (
0
+ 3
1
+ 9
2
)x
t-3
+ (
0
+ 4
1
+ 16
2
)x
t-4
+ e
t
Step 2: factor out the unknown coefficients:
0
,
1
,
2
.
y
t
= +
0
[x
t
+ x
t-1
+ x
t-2
+ x
t-3
+ x
t-4
]
+
1
[x
t
+ x
t-1
+ 2x
t-2
+ 3x
t-3
+ 4x
t-4
]
+
2
[x
t
+ x
t-1
+ 4x
t-2
+ 9x
t-3
+ 16x
t-4
] + e
t
Step 1: impose the restriction:
i
=
0
+
1
i +
2
i
2
15.10
Copyright 1996 Lawrence C. Marsh
Polynomial Lag Structure
Step 3: Define z
t0
, z
t1
and z
t2
for
0
,
1
, and
2
.
y
t
= +
0
[x
t
+ x
t-1
+ x
t-2
+ x
t-3
+ x
t-4
]
+
1
[x
t
+ x
t-1
+ 2x
t-2
+ 3x
t-3
+ 4x
t-4
]
+
2
[x
t
+ x
t-1
+ 4x
t-2
+ 9x
t-3
+ 16x
t-4
] + e
t
z
t0
= [x
t
+ x
t-1
+ x
t-2
+ x
t-3
+ x
t-4
]
z
t1
= [x
t
+ x
t-1
+ 2x
t-2
+ 3x
t-3
+ 4x
t- 4
]
z
t2
= [x
t
+ x
t-1
+ 4x
t-2
+ 9x
t-3
+ 16x
t- 4
]
15.11
Copyright 1996 Lawrence C. Marsh
Polynomial Lag Structure
Step 4: Regress y
t
on z
t0
, z
t1
and z
t2
.
y
t
= +
0
z
t0
+
1
z
t1
+
2
z
t2
+ e
t
Step 5: Express
i
s in terms of
0
,
1
, and
2
.
^
^ ^ ^
0
=
0
1
=
0
+
1
+
2
2
=
0
+ 2
1
+ 4
2
3
=
0
+ 3
1
+ 9
2
4
=
0
+ 4
1
+ 16
2
^
^
^
^
^
^
^ ^ ^
^ ^ ^
^ ^ ^
^ ^ ^
15.12
Copyright 1996 Lawrence C. Marsh
Polynomial Lag Structure
.
.
.
.
.
0 1 2 3 4 i
i
Figure 15.3
4
15.13
Copyright 1996 Lawrence C. Marsh
Geometric Lag Structure
y
t
= +
0
x
t
+
1
x
t-1
+
2
x
t-2
+ . . . + e
t
infinite distributed lag model:
y
t
= +
i
x
t-i
+ e
t
i=0
(15.3.1)
geometric lag structure:
i
=
i
where || < 1 and
i
> 0 .
15.14
Copyright 1996 Lawrence C. Marsh
Geometric Lag Structure
y
t
= +
0
x
t
+
1
x
t-1
+
2
x
t-2
+
3
x
t-3
+ . . . + e
t
y
t
= + (x
t
+ x
t-1
+
2
x
t-2
+
3
x
t-3
+ . . .) + e
t
infinite unstructured lag:
infinite geometric lag:
Substitute
i
=
i
0
=
1
=
2
=
2
3
=
3
.
.
.
15.15
Copyright 1996 Lawrence C. Marsh
Geometric Lag Structure
interim multiplier (3-period) :
impact multiplier :
long-run multiplier :
y
t
= + (x
t
+ x
t-1
+
2
x
t-2
+
3
x
t-3
+ . . .) + e
t
+ +
2
(1 + +
2
+
3
+ . . . ) =
1
15.16
Copyright 1996 Lawrence C. Marsh
Geometric Lag Structure
i
Figure 15.5
.
.
.
.
.
0 1 2 3 4 i
1
=
2
=
2
3
=
3
4
=
4
0
=
geometrically
declining
weights
15.17
Copyright 1996 Lawrence C. Marsh
Geometric Lag Structure
Problem:
How to estimate the infinite number
of geometric lag coefficients ???
y
t
= + (x
t
+ x
t-1
+
2
x
t-2
+
3
x
t-3
+ . . .) + e
t
Answer:
Use the Koyck transformation.
15.18
Copyright 1996 Lawrence C. Marsh
The Koyck Transformation
y
t
= + (x
t
+ x
t-1
+
2
x
t-2
+
3
x
t-3
+ . . .) + e
t
y
t
y
t-1
= (1 ) + x
t
+ (e
t
e
t-1
)
Lag everything once, multiply by and subtract from original:
y
t-1
= + ( x
t-1
+
2
x
t-2
+
3
x
t-3
+ . . .) + e
t-1
15.19
Copyright 1996 Lawrence C. Marsh
The Koyck Transformation
y
t
y
t-1
= (1 ) + x
t
+ (e
t
e
t-1
)
y
t
= (1 ) + y
t-1
+ x
t
+ (e
t
e
t-1
)
Solve for y
t
by adding y
t-1
to both sides:
y
t
=
1
+
2
y
t-1
+
3
x
t
+
t
15.20
Copyright 1996 Lawrence C. Marsh
The Koyck Transformation
y
t
= (1 ) + y
t-1
+ x
t
+ (e
t
e
t-1
)
y
t
=
1
+
2
y
t-1
+
3
x
t
+
t
Defining
1
= (1 ) ,
2
= , and
3
= ,
use ordinary least squares:
=
3
^ ^
=
2
^
^
=
1
/ (1
2
)
^
^ ^
The original structural
parameters can now be
estimated in terms of
these reduced form
parameter estimates.
15.21
Copyright 1996 Lawrence C. Marsh
Geometric Lag Structure
0
=
1
=
2
=
2
3
=
3
.
.
.
^
^
^ ^
^
^ ^
^
^ ^
^
y
t
= + (x
t
+ x
t-1
+
2
x
t-2
+
3
x
t-3
+ . . .) + e
t
^
^
^
^ ^ ^
y
t
= +
0
x
t
+
1
x
t-1
+
2
x
t-2
+
3
x
t-3
+ . . . + e
t
^ ^
^ ^ ^ ^
15.22
Copyright 1996 Lawrence C. Marsh
Durbins h-test
for autocorrelation
T 1
1 ( T 1)[se(b
2
)]
2
h = 1
d
2
h = Durbins h-test statistic
d = Durbin-Watson test statistic
se(b
2
) = standard error of the estimate b
2
T = sample size
Estimates inconsistent if geometric lag model is autocorrelated,
but Durbin-Watson test is biased in favor of no autocorrelation.
15.23
Copyright 1996 Lawrence C. Marsh
y
t
= + x*
t
+ e
t
Adaptive Expectations
y
t
= credit card debt
x*
t
= expected (anticipated) income
(x*
t
is not observable)
15.24
Copyright 1996 Lawrence C. Marsh
Adaptive Expectations
x*
t
- x*
t-1
= (x
t-1
- x*
t-1
)
adjust expectations
based on past realization:
15.25
Copyright 1996 Lawrence C. Marsh
Adaptive Expectations
x*
t
- x*
t-1
= (x
t-1
- x*
t-1
)
x*
t
= x
t-1
+ (1- ) x*
t-1
rearrange to get:
x
t-1
= [x*
t
- (1- ) x*
t-1
]
or
15.26
Copyright 1996 Lawrence C. Marsh
Adaptive Expectations
y
t
= + x*
t
+ e
t
Lag this model once and multiply by (1 ):
y
t
= - (1 )y
t-1
+ [x*
t
- (1 )x*
t-1
]
+ e
t
- (1 )e
t-1
subtract this from the original to get:
(1 )y
t-1
= (1 ) + (1 ) x*
t-1
+ (1 )e
t-1
15.27
Copyright 1996 Lawrence C. Marsh
Adaptive Expectations
y
t
= - (1 )y
t-1
+ [x*
t
- (1 )x*
t-1
]
+ e
t
- (1 )e
t-1
Since x
t-1
= [x*
t
- (1- ) x*
t-1
]
we get:
y
t
= - (1 )y
t-1
+ x
t-1
+ u
t
where u
t
= e
t
- (1 )e
t-1
15.28
Copyright 1996 Lawrence C. Marsh
Adaptive Expectations
y
t
= - (1 )y
t-1
+ x
t-1
+ u
t
y
t
=
1
+
2
y
t-1
+
3
x
t-1
+ u
t
Use ordinary least squares regression on:
and we get:
=
(1
2
)
3
^
^
^
= (1
2
)
^ ^
=
(1
2
)
1
^
^
^
15.29
Copyright 1996 Lawrence C. Marsh
Partial Adjustment
y
t
- y
t-1
= (y*
t
- y
t-1
)
inventories partially adjust , 0 < < 1,
towards optimal or desired level, y*
t
:
y*
t
= + x
t
+ e
t
15.30
Copyright 1996 Lawrence C. Marsh
Partial Adjustment
y
t
- y
t-1
= (y*
t
- y
t-1
)
= ( + x
t
+ e
t
- y
t-1
)
= + x
t
- y
t-1
+ e
t
y
t
= + (1
- )y
t-1
+ x
t
+ e
t
Solving for y
t
:
15.31
Copyright 1996 Lawrence C. Marsh
Partial Adjustment
y
t
= + (1
- )y
t-1
+ x
t
+ e
t
y
t
=
1
+
2
y
t-1
+
3
x
t
+
t
=
(1
2
)
3
^
^
^
= (1
2
)
^
^
=
(1
2
)
1
^
^
^
Use ordinary least squares regression to get:
15.32
Copyright 1996 Lawrence C. Marsh
Time
Series
Analysis
Chapter 16
Copyright 1997 John Wiley & Sons, Inc. All rights reserved. Reproduction or translation of this work beyond
that permitted in Section 117 of the 1976 United States Copyright Act without the express written permission of the
copyright owner is unlawful. Request for further information should be addressed to the Permissions Department,
John Wiley & Sons, Inc. The purchaser may make back-up copies for his/her own use only and not for distribution
or resale. The Publisher assumes no responsibility for errors, omissions, or damages, caused by the use of these
programs or from the use of the information contained herein.
16.1
Copyright 1996 Lawrence C. Marsh
Previous Chapters used Economic Models
1. economic model for dependent variable of interest.
2. statistical model consistent with the data.
3. estimation procedure for parameters using the data.
4. forecast variable of interest using estimated model.
Times Series Analysis does not use this approach.
16.2
Copyright 1996 Lawrence C. Marsh
Time Series Analysis is useful for short term forecasting only.
Time Series Analysis does not generally
incorporate all of the economic relationships
found in economic models.
Times Series Analysis uses
more statistics and less economics.
Long term forecasting requires incorporating more involved
behavioral economic relationships into the analysis.
16.3
Copyright 1996 Lawrence C. Marsh
Univariate Time Series Analysis can be used
to relate the current values of a single economic
variable to:
1. its past values
2. the values of current and past random errors
Other variables are not used
in univariate time series analysis.
16.4
Copyright 1996 Lawrence C. Marsh
1. autoregressive (AR)
2. moving average (MA)
3. autoregressive moving average (ARMA)
Three types of Univariate Time Series Analysis
processes will be discussed in this chapter:
16.5
Copyright 1996 Lawrence C. Marsh
1. its past values.
2. the past values of the other forecasted variables.
3. the values of current and past random errors.
Multivariate Time Series Analysis can be
used to relate the current value of each of
several economic variables to:
Vector autoregressive models discussed later in
this chapter are multivariate time series models.
16.6
Copyright 1996 Lawrence C. Marsh
First-Order Autoregressive Processes, AR(1):
y
t
= +
1
y
t-1
+ e
t
, t = 1, 2,...,T. (16.1.1)
is the intercept.
1
is parameter generally between -1 and +1.
e
t
is an uncorrelated random error with
mean zero and variance
e
2
.
16.7
Copyright 1996 Lawrence C. Marsh
Autoregressive Process of order p, AR(p) :
y
t
= +
1
y
t-1
+
2
y
t-2
+...+
p
y
t-p
+ e
t
(16.1.2)
is the intercept.
i
s are parameters generally between -1 and +1.
e
t
is an uncorrelated random error with
mean zero and variance
e
2
.
16.8
Copyright 1996 Lawrence C. Marsh
AR models always have one or more lagged
dependent variables on the right hand side.
Consequently, least squares is no longer a
best linear unbiased estimator (BLUE),
but it does have some good asymptotic
properties including consistency.
Properties of least squares estimator:
16.9
Copyright 1996 Lawrence C. Marsh
AR(2) model of U.S. unemployment rates
y
t
= 0.5051 + 1.5537 y
t-1
- 0.6515 y
t-2
(0.1267) (0.0707) (0.0708)
Note: Q1-1948 through Q1-1978 from [Link] (1986) see [Link]
positive
negative
16.10
Copyright 1996 Lawrence C. Marsh
Choosing the lag length, p, for AR(p):
The Partial Autocorrelation Function (PAF)
The PAF is the sequence of correlations between
(y
t
and y
t-1
), (y
t
and y
t-2
), (y
t
and y
t-3
), and so on,
given that the effects of earlier lags on y
t
are
held constant.
16.11
Copyright 1996 Lawrence C. Marsh
Partial Autocorrelation Function
y
t
= 0.5 y
t-1
+ 0.3 y
t-2
+ e
t
0
2 / T
2 / T
1
1
k
kk
is the last (k
th
) coefficient
in a k
th
order AR process.
This sample PAF suggests a second
order process AR(2) which is correct.
Data simulated
from this model:
kk
^
16.12
Copyright 1996 Lawrence C. Marsh
Using AR Model for Forecasting:
unemployment rate: y
T-1
= 6.63 and y
T
= 6.20
y
T+1
= +
1
y
T
+
2
y
T-1
= 0.5051 + (1.5537)(6.2) - (0.6515)(6.63)
= 5.8186
^
^ ^ ^
y
T+2
= +
1
y
T+1
+
2
y
T
= 0.5051 + (1.5537)(5.8186) - (0.6515)(6.2)
= 5.5062
^
^ ^ ^
y
T+1
= +
1
y
T
+
2
y
T-1
= 0.5051 + (1.5537)(5.5062) - (0.6515)(5.8186)
= 5.2693
^
^ ^ ^
16.13
Copyright 1996 Lawrence C. Marsh
Moving Average Process of order q, MA(q):
y
t
= + e
t
+
1
e
t-1
+
2
e
t-2
+...+
q
e
t-q
+ e
t
(16.2.1)
is the intercept.
i
s are unknown parameters.
e
t
is an uncorrelated random error with
mean zero and variance
e
2
.
16.14
Copyright 1996 Lawrence C. Marsh
An MA(1) process:
y
t
= + e
t
+
1
e
t-1
(16.2.2)
Minimize sum of least squares deviations:
S(,
1
) = e
t
= (y
t
-
-
1
e
t-1
) (16.2.3)
2
t=1
T
t=1
T
2
16.15
Copyright 1996 Lawrence C. Marsh
stationary:
A stationary time series is one whose mean, variance,
and autocorrelation function do not change over time.
nonstationary:
A nonstationary time series is one whose mean,
variance or autocorrelation function change over time.
Stationary vs. Nonstationary
16.16
Copyright 1996 Lawrence C. Marsh
y
t
= z
t
- z
t-1
First Differencing is often used to transform
a nonstationary series into a stationary series:
where z
t
is the original nonstationary series
and y
t
is the new stationary series.
16.17
Copyright 1996 Lawrence C. Marsh
Choosing the lag length, q, for MA(q):
The Autocorrelation Function (AF)
The AF is the sequence of correlations between
(y
t
and y
t-1
), (y
t
and y
t-2
), (y
t
and y
t-3
), and so on,
without holding the effects of earlier lags
on y
t
constant.
The PAF controlled for the effects of previous lags
but the AF does not control for such effects.
16.18
Copyright 1996 Lawrence C. Marsh
Autocorrelation Function
y
t
= e
t
0.9 e
t-1
0
2 / T
2 / T
1
1
k
r
kk
r
kk
is the last (k
th
) coefficient
in a k
th
order MA process.
This sample AF suggests a first order
process MA(1) which is correct.
Data simulated
from this model:
16.19
Copyright 1996 Lawrence C. Marsh
Autoregressive Moving Average
ARMA(p,q)
An ARMA(1,2) has one autoregressive lag
and two moving average lags:
y
t
= +
1
y
t-1
+ e
t
+
1
e
t-1
+
2
e
t-2
16.20
Copyright 1996 Lawrence C. Marsh
Integrated Processes
A time series with an upward or downward
trend over time is nonstationary.
Many nonstationary time series can be made
stationary by differencing them one or more times.
Such time series are called integrated processes.
16.21
Copyright 1996 Lawrence C. Marsh
The number of times a series must be
differenced to make it stationary is the
order of the integrated process, d.
An autocorrelation function, AF,
with large, significant autocorrelations
for many lags may require more than
one differencing to become stationary.
Check the new AF after each differencing
to determine if further differencing is needed.
16.22
Copyright 1996 Lawrence C. Marsh
Unit Root
z
t
=
1
z
t-1
+ + e
t
+
1
e
t-1
(16.3.2)
-1 <
1
< 1 stationary ARMA(1,1)
1
= 1 nonstationary process
1
= 1 is called a unit root
16.23
Copyright 1996 Lawrence C. Marsh
Unit Root Tests
z
t
=
1
z
t-1
+ + e
t
+
1
e
t-1
(16.3.3)
Testing
1
= 0 is equivalent to testing
1
= 1
z
t
-
z
t-1
= (
1
-
1)z
t-1
+ + e
t
+
1
e
t-1
*
where z
t
= z
t
-
z
t-1
and
1
=
1
-
1
*
*
16.24
Copyright 1996 Lawrence C. Marsh
Unit Root Tests
H
0
:
1
= 0 vs. H
1
:
1
< 0 (16.3.4)
* *
Computer programs typically use one of
the following tests for unit roots:
Dickey-Fuller Test
Phillips-Perron Test
16.25
Copyright 1996 Lawrence C. Marsh
Autoregressive Integrated Moving Average
ARIMA(p,d,q)
An ARIMA(p,d,q) model represents an
AR(p) - MA(q) process that has been
differenced (integrated, I(d)) d times.
y
t
= +
1
y
t-1
+...+
p
y
t-p
+ e
t
+
1
e
t-1
+... +
q
e
t-q
16.26
Copyright 1996 Lawrence C. Marsh
The Box-Jenkins approach:
1. Identification
determining the values of p, d, and q.
2. Estimation
linear or nonlinear least squares.
3. Diagnostic Checking
model fits well with no autocorrelation?
4. Forecasting
short-term forecasts of future y
t
values.
16.27
Copyright 1996 Lawrence C. Marsh
Vector Autoregressive (VAR) Models
y
t
=
0
+
1
y
t-1
+...+
p
y
t-p
+
1
x
t-1
+... +
p
x
t-p
+ e
t
x
t
=
0
+
1
y
t-1
+...+
p
y
t-p
+
1
x
t-1
+... +
p
x
t-p
+ u
t
Use VAR for two or more interrelated time series:
16.28
Copyright 1996 Lawrence C. Marsh
1. extension of AR model.
2. all variables endogenous.
3. no structural (behavioral) economic model.
4. all variables jointly determined (over time).
5. no simultaneous equations (same time).
Vector Autoregressive (VAR) Models
16.29
Copyright 1996 Lawrence C. Marsh
The random error terms in a VAR model
may be correlated if they are affected by
relevant factors that are not in the model
such as government actions or
national/international events, etc.
Since VAR equations all have exactly the
same set of explanatory variables, the usual
seemingly unrelation regression estimation
produces exactly the same estimates as
least squares on each equation separately.
16.30
Copyright 1996 Lawrence C. Marsh
Consequently, regardless of whether
the VAR random error terms are
correlated or not, least squares estimation
of each equation separately will provide
consistent regression coefficient estimates.
Least Squares is Consistent
16.31
Copyright 1996 Lawrence C. Marsh
VAR Model Specification
To determine length of the lag, p, use:
2. Schwarzs SIC criterion
1. Akaikes AIC criterion
These methods were discussed in Chapter 15.
16.32
Copyright 1996 Lawrence C. Marsh
Spurious Regressions
y
t
=
1
+
2
x
t
+
t
where
t
=
1
t-1
+
t
-1 <
1
< 1 I(0) (i.e. d=0)
1
= 1 I(1) (i.e. d=1)
If
1
=1 least squares estimates of
2
may
appear highly significant even when true
2
= 0 .
16.33
Copyright 1996 Lawrence C. Marsh
Cointegration
y
t
=
1
+
2
x
t
+
t
If x
t
and y
t
are nonstationary I(1)
we might expect that
t
is also I(1).
However, if x
t
and y
t
are nonstationary I(1)
but
t
is stationary I(0), then x
t
and y
t
are
said to be cointegrated.
16.34
Copyright 1996 Lawrence C. Marsh
Cointegrated VAR(1) Model
y
t
=
0
+
1
y
t-1
+
1
x
t-1
+ e
t
x
t
=
0
+
1
y
t-1
+
1
x
t-1
+ u
t
VAR(1) model:
If x
t
and y
t
are both I(1) and are cointegrated,
use an Error Correction Model, instead of VAR(1).
16.35
Copyright 1996 Lawrence C. Marsh
Error Correction Model
y
t
=
0
+ (
1
-1)y
t-1
+
1
x
t-1
+ e
t
x
t
=
0
+
1
y
t-1
+ (
1
-1)x
t-1
+ u
t
y
t
= y
t
- y
t-1
and x
t
= x
t
- x
t-1
(continued)
16.36
Copyright 1996 Lawrence C. Marsh
Error Correction Model
y
t
=
0
+
1
(y
t-1
-
1
-
2
x
t-1
)
+ e
t
*
x
t
=
0
+
2
(y
t-1
-
1
-
2
x
t-1
)
+ u
t
*
0
=
0
+
1
1
*
0
=
0
+
2
1
*
2
=
1
1
=
1
- 1
2
=
1
1 -
1
16.37
Copyright 1996 Lawrence C. Marsh
y
t-1
=
1
+
2
x
t-1
+
t-1
Estimate by least squares:
to get the residuals:
t-1
= y
t-1
-
1
-
2
x
t-1
^
^ ^
Estimating an Error Correction Model
Step 1:
16.38
Copyright 1996 Lawrence C. Marsh
Estimate by least squares:
Estimating an Error Correction Model
Step 2:
y
t
=
0
+
1
t-1
+ e
t
*
x
t
=
0
+
2
t-1
+ u
t
*
^
^
16.39
Copyright 1996 Lawrence C. Marsh
Using cointegrated I(1) variables in a
VAR model expressed solely in terms
of first differences and lags of first
differences is a misspecification.
The correct specification is to use an
Error Correction Model
16.40
Copyright 1996 Lawrence C. Marsh
Chapter 17
Copyright 1997 John Wiley & Sons, Inc. All rights reserved. Reproduction or translation of this work beyond
that permitted in Section 117 of the 1976 United States Copyright Act without the express written permission of the
copyright owner is unlawful. Request for further information should be addressed to the Permissions Department,
John Wiley & Sons, Inc. The purchaser may make back-up copies for his/her own use only and not for distribution
or resale. The Publisher assumes no responsibility for errors, omissions, or damages, caused by the use of these
programs or from the use of the information contained herein.
Guidelines for
Research Project
17.1
Copyright 1996 Lawrence C. Marsh
Formulation
economic ====> econometric.
Estimation
selecting appropriate method.
Interpretation
how the x
t
s impact on the y
t
.
Inference
testing, intervals, prediction.
What Book Has Covered
17.2
Copyright 1996 Lawrence C. Marsh
Topics for This Chapter
1. Types of Data by Source
2. Nonexperimental Data
3. Text Data vs. Electronic Data
4. Selecting a Topic
5. Writing an Abstract
6. Research Report Format
17.3
Copyright 1996 Lawrence C. Marsh
Types of Data by Source
i) Experimental Data
from controlled experiments.
ii) Observational Data
passively generated by society.
iii) Survey Data
data collected through interviews.
17.4
Copyright 1996 Lawrence C. Marsh
Time vs. Cross-Section
Time Series Data
data collected at distinct points in time
(e.g. weekly sales, daily stock price, annual
budget deficit, monthly unemployment.)
Cross Section Data
data collected over samples of units, individuals,
households, firms at a particular point in time.
(e.g. salary, race, gender, unemployment by state.)
17.5
Copyright 1996 Lawrence C. Marsh
Micro vs. Macro
Micro Data:
data collected on individual economic
decision making units such as individuals,
households or firms.
Macro Data:
data resulting from a pooling or aggregating
over individuals, households or firms at the
local, state or national levels.
17.6
Copyright 1996 Lawrence C. Marsh
Flow vs. Stock
Flow Data:
outcome measured over a period of time,
such as the consumption of gasoline during
the last quarter of 1997.
Stock Data:
outcome measured at a particular point in
time, such as crude oil held by Chevron in
US storage tanks on April 1, 1997.
17.7
Copyright 1996 Lawrence C. Marsh
Quantitative vs. Qualitative
Quantitative Data:
outcomes such as prices or income that may
be expressed as numbers or some transfor-
mation of them (e.g. wages, trade deficit).
Qualitative Data:
outcomes that are of an either-or nature
(e.g. male, home owner, Methodist, bought
car last year, voted in last election).
17.8
Copyright 1996 Lawrence C. Marsh
International Data
International Financial Statistics (IMF monthly).
Basic Statistics of the Community (OECD annual).
Consumer Price Indices in the European
Community (OECD annual).
World Statistics (UN annual).
Yearbook of National Accounts Statistics (UN).
FAO Trade Yearbook (annual).
17.9
Copyright 1996 Lawrence C. Marsh
United States Data
Survey of Current Business (BEA monthly).
Handbook of Basic Economic Statistics (BES).
Monthly Labor Review (BLS monthly).
Federal Researve Bulletin (FRB monthly).
Statistical Abstract of the US (BC annual).
Economic Report of the President (CEA annual).
Economic Indicators (CEA monthly).
Agricultural Statistics (USDA annual).
Agricultural Situation Reports (USDA monthly).
17.10
Copyright 1996 Lawrence C. Marsh
State and Local Data
State and Metropolitan Area Data Book
(Commerce and BC, annual).
CPI Detailed Report (BLS, annual).
Census of Population and Housing
(Commerce, BC, annual).
County and City Data Book
(Commerce, BC, annual).
17.11
Copyright 1996 Lawrence C. Marsh
Citibase on CD-ROM
Financial series: interest rates, stock market, etc.
Business formation, investment and consumers.
Construction of housing.
Manufacturing, business cycles, foreign trade.
Prices: producer and consumer price indexes.
Industrial production.
Capacity and productivity.
Population.
17.12
Copyright 1996 Lawrence C. Marsh
Citibase on CD-ROM
(continued)
Labor statistics: unemployment, households.
National income and product accounts in detail.
Forecasts and projections.
Business cycle indicators.
Energy consumption, petroleum production, etc.
International data series including trade
statistics.
17.13
Copyright 1996 Lawrence C. Marsh
Resources for Economists
Resources for Economists by Bill Goffe
[Link]
Bill Goffe provides a vast database of information
about the economics profession including economic
organizations, working papers and reports,
and economic data series.
17.14
Copyright 1996 Lawrence C. Marsh
Internet Data Sources
Shortcut to All Resources.
Macro and Regional Data.
Other U.S. Data.
World and Non-U.S. Data.
Finance and Financial Markets.
Data Archives.
Journal Data and Program Archives.
A few of the items on Bill Goffes Table of Contents:
17.15
Copyright 1996 Lawrence C. Marsh
Useful Internet Addresses
[Link]
[Link]
[Link] FED RESERVE BK - ST. LOUIS
[Link] BUREAU OF LABOR STATISTICS
[Link] NATL BUR. ECON. RESEARCH
[Link]
.www/[Link] UNIVERSITY OF MARYLAND
[Link] FEB BOARD OF GOVERNORS
[Link]
17.16
Copyright 1996 Lawrence C. Marsh
Data from Surveys
i) identify the population of interest.
ii) designing and selecting the sample.
iii) collecting the information.
iv) data reduction, estimation and inference.
The survey process has four distinct aspects:
17.17
Copyright 1996 Lawrence C. Marsh
Controlled Experiments
1. Labor force participation: negative income tax:
guaranteed minimum income experiment.
2. National cash housing allowance experiment:
impact on demand and supply of housing.
3. Health insurance: medical cost reduction:
sensitivity of income groups to price change.
4. Peak-load pricing and electricity use:
daily use pattern of residential customers.
Controlled experiments were done on these topics:
17.18
Copyright 1996 Lawrence C. Marsh
Economic Data Problems
I. poor implicit experimental design
(i) collinear explanatory variables.
(ii) measurement errors.
II. inconsistent with theory specification
(i) wrong level of aggregation.
(ii) missing observations or variables.
(iii) unobserved heterogeneity.
17.19
Copyright 1996 Lawrence C. Marsh
Selecting a Topic
What am I interested in?
Well-defined, relatively simple topic.
Ask prof for ideas and references.
Journal of Economic Literature (ECONLIT)
Make sure appropriate data are available.
Avoid extremely difficult econometrics.
Plan your work and work your plan.
General tips for selecting a research topic:
17.20
Copyright 1996 Lawrence C. Marsh
Writing an Abstract
(i) concise statement of the problem.
(ii) key references to available information.
(iii) description of research design including:
(a) economic model
(b) statistical model
(c) data sources
(d) estimation, testing and prediction
(iv) contribution of the work
Abstract of less than 500 words should include:
17.21
Copyright 1996 Lawrence C. Marsh
Research Report Format
1. Statement of the Problem.
2. Review of the Literature.
3. The Economic Model.
4. The Statistical Model.
5. The Data.
6. Estimation and Inferences Procedures.
7. Empirical Results and Conclusions.
8. Possible Extensions and Limitations.
9. Acknowledgments.
10. References.
17.22