0% found this document useful (0 votes)
8 views15 pages

Ch17 Dynamic Models Notes

Chapter 17 of 'Basic Econometrics' discusses dynamic econometric models, focusing on autoregressive and distributed-lag models that incorporate time lags in economic responses. It covers key concepts such as the Koyck model, adaptive expectations, and the partial adjustment model, emphasizing the reasons for lags and the implications for estimation methods. The chapter also provides examples and statistical measures related to lag distributions, along with the challenges of ad hoc estimation methods.

Uploaded by

nimish.eco2527
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
8 views15 pages

Ch17 Dynamic Models Notes

Chapter 17 of 'Basic Econometrics' discusses dynamic econometric models, focusing on autoregressive and distributed-lag models that incorporate time lags in economic responses. It covers key concepts such as the Koyck model, adaptive expectations, and the partial adjustment model, emphasizing the reasons for lags and the implications for estimation methods. The chapter also provides examples and statistical measures related to lag distributions, along with the challenges of ad hoc estimation methods.

Uploaded by

nimish.eco2527
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

GUJARATI & PORTER · BASIC ECONOMETRICS · 5TH ED.

Dynamic Econometric Models:


Autoregressive and
Distributed-Lag Models
Lags in economics, Koyck transformation, adaptive expectations, partial adjustment, Almon PDL,
and Granger causality

Chapter 17 Pages 617–662 High Priority Part III — Topics in Econometrics

CONTENTS

1. Role of Lags & Multipliers


2. Reasons for Lags
3. Ad Hoc Estimation Problems
4. The Koyck Model
5. Adaptive Expectations Model
6. Partial Adjustment Model (PAM)
7. Estimation of Autoregressive Models
8. Instrumental Variables (IV) Method
9. Durbin h Test
10. Almon Polynomial Distributed Lag
11. Granger Causality Test
12. Revision Summary

The Role of Lags & Multipliers


17.1

KEY CONCEPT

In economics, Y rarely responds to X instantaneously. The time it takes for Y to respond is


called a lag. Models that incorporate current and past values of regressors are called
distributed-lag models; models where lagged Y appears as a regressor are autoregressive
models. Together these are dynamic models.

The general finite distributed-lag model with k lags is:

FINITE DISTRIBUTED-LAG MODEL

Yt = α + β₀Xt + β₁Xt₋₁ + β₂Xt₋₂ + ··· + βkXt₋k + ut ...(17.1.2)

Multiplier Terminology:
β₀ = short-run (impact) multiplier: change in mean Y per unit ΔX this period
β₀ + β₁ = interim multiplier after 1 period
β₀ + β₁ + β₂ = interim multiplier after 2 periods
β₀+β₁+···+βk = β = long-run (total) distributed-lag multiplier

Standardised coefficients: βᵢ* = βᵢ/β


→ partial sums of βᵢ* give the proportion of the long-run impact felt by period i

Illustrative Examples of Lags in Economics

▸ Consumption function (Example 17.1): A $2,000 permanent income rise spreads over 3 years: $800
spent in year 1, $600 in year 2, $400 in year 3. Short-run MPC = 0.4; long-run MPC = 0.4 + 0.3 + 0.2 =
0.9.

▸ Bank money multiplier (Example 17.2): A $1,000 injection with 20% reserve requirement generates
$5,000 in deposits — but the process unfolds across multiple stages over time.

▸ Money and prices (Example 17.3): A 1% change in M1 affects inflation over 20 quarters. Short-run
elasticity ≈ 0.04 (not significant); long-run elasticity ≈ 1.03 (monetarist prediction). Mean lag ≈ 11
quarters.

▸ R&D and productivity (Example 17.4): There are multiple lags: invention lag, development lag, and
diffusion lag.

▸ J curve (Example 17.5): Currency depreciation worsens the trade balance initially (import costs rise
immediately) before improving it (exports eventually rise).

▸ Accelerator model (Example 17.6): Investment is proportional to the change in output: It = β(Xt −
Xt−1), a model with a built-in one-period lag.

Mean and Median Lags

SUMMARY MEASURES OF THE LAG DISTRIBUTION

Mean lag = Σ(k·βk) / Σβk (weighted average of lag lengths; βk as weights)


→ measures average time for full impact to be felt

Median lag = time for first 50% of total long-run impact to occur
→ indicates speed of adjustment

For the Koyck model (see Section 17.4):


Mean lag = λ/(1−λ)
Median lag = −log(2)/log(λ)

17.2Why Do Lags Exist?


KEY CONCEPT

Three fundamental reasons explain why Y responds to X with a time lag rather than
instantaneously.

▸ Psychological reasons (habit/inertia): People do not immediately revise consumption or investment


behaviour after income or price changes. Lottery winners, for instance, may not change lifestyle
immediately. Agents also need time to distinguish between permanent and transitory changes — a one-
off income rise may be saved rather than spent.

▸ Technological reasons (gestation period): Adding capital takes time to install and commission. If a
price change is expected to be temporary, firms may not bother to adjust. Imperfect information also
contributes — consumers researching competing products before buying.

▸ Institutional reasons (contractual lock-in): Contractual obligations prevent immediate switching of


labour or raw material sources. Long-term fixed deposits lock funds in for years. Health insurance
choices may be locked in for a year. These rigidities create institutional lags.

These reasons explain why short-run elasticities are generally smaller (in absolute value) than long-
run elasticities — a fundamental short-run/long-run distinction in all applied economics.

17.3Ad Hoc Estimation & Its Problems

The straightforward approach to estimating the distributed-lag model is to regress Y on Xt, then Xt
and Xt−1, then add Xt−2, stopping when lagged coefficients become insignificant or change sign.
This is the Alt–Tinbergen ad hoc method.

Four major problems make this approach unreliable:

▸ No guide to maximum lag length: Choosing too few lags causes omitted variable bias; choosing too
many wastes degrees of freedom and inflates standard errors.

▸ Loss of degrees of freedom: Each additional lag costs one observation. With limited data, estimation
becomes unreliable.

▸ Multicollinearity: Successive lags of X (e.g., Xt, Xt−1, Xt−2) are highly correlated in time series. This
inflates standard errors, causing individually insignificant coefficients even when the group is jointly
significant.

▸ Data mining: Sequential searching for the best lag structure increases the chance of finding
spuriously significant results. The nominal significance level overstates true significance.

Conclusion: Ad hoc estimation has little to recommend it in practice. Some structure must be
imposed on the β coefficients via economic theory or a parametric assumption. The Koyck and Almon
methods do precisely this.

17.4The Koyck Model


KEY CONCEPT

Koyck proposes that in an infinite distributed-lag model, the lag coefficients decline
geometrically: βk = β0λk, where 0 < λ < 1. This ensures all β’s are positive, declining
weights, with a finite long-run sum. The infinite model is then transformed into a tractable
autoregressive model by a clever algebraic trick.

KOYCK GEOMETRIC DECAY ASSUMPTION


βk = β₀λk k = 0, 1, 2, ... where 0 < λ < 1

Properties:
• All βk ≥ 0 (same sign as β₀) — no sign changes allowed
• β declines toward zero as k increases (distant past has less weight)
• Long-run multiplier: Σβk = β₀/(1−λ) (finite, since λ < 1)

λ is the rate of decay (or rate of decline):


λ near 1 → slow decay → distant past matters a lot → high mean lag
λ near 0 → fast decay → only recent past matters → low mean lag

The Koyck Transformation

KOYCK TRANSFORMATION: FROM INFINITE DISTRIBUTED LAG TO AUTOREGRESSIVE MODEL

Step 1 — Original infinite model:


Yt = α + β₀Xt + β₀λXt₋₁ + β₀λ²Xt₋₂ + ··· + ut ...(17.4.3)

Step 2 — Lag by one period and multiply by λ:


λYt₋₁ = λα + β₀λXt₋₁ + β₀λ²Xt₋₂ + ··· + λut₋₁ ...(17.4.5)

Step 3 — Subtract (step 2) from (step 1):


Yt − λYt₋₁ = α(1−λ) + β₀Xt + (ut − λut₋₁)

Step 4 — Rearrange:
Yt = α(1−λ) + β₀Xt + λYt₋₁ + vt where vt = ut − λut₋₁ ...(17.4.7)

RESULT: Instead of estimating infinitely many β's, now estimate just 3:


α(1−λ), β₀, and λ.
No multicollinearity from the lagged X's since they are all collapsed into Yt₋₁.

Four Important Features of the Koyck Model

▸ Distributed lag → autoregressive: The transformation converts an infinite distributed-lag model into
an autoregressive model with Yt−1 as a regressor. This shows the equivalence between the two model
types.

▸ Stochastic regressor problem: Yt−1 is stochastic. The classical OLS assumption requires regressors
to be uncorrelated with the error term. This may be violated here.

▸ Serial correlation in vt: The transformed error vt = ut − λut−1 is a moving average of ut. If ut is white

noise, vt is serially correlated: E(vtvt−1) = −λσ2 ≠ 0. So OLS faces both a stochastic regressor and
autocorrelated errors.

▸ DW test invalid: Presence of lagged Y means DW d tends toward 2, masking genuine autocorrelation.
The Durbin h test (Section 17.10) must be used instead.

Mean and Median Lag for Koyck

KOYCK LAG SUMMARY STATISTICS

Mean lag = λ/(1−λ) [λ = 0.5 → mean lag = 1; λ = 0.8 → mean lag = 4]


Median lag = −log(2)/log(λ) [λ = 0.2 → 0.43 periods; λ = 0.8 → 3.11 periods]

Long-run multiplier = β₀/(1−λ)


Short-run multiplier = β₀

EXAMPLE 17.7 — PPCE on PPDI, U.S. 1959–2006 (Koyck Model)

Short-run regression (Koyck):


PPCEt = −252.92 + 0.2139·PPDIt + 0.7971·PPCEt₋₁
se = (157.35) (0.0706) (0.0733)
t = (−1.61) (3.03) (10.87)
R² = 0.998 d = 0.962 Durbin h = 3.83 (significant → autocorrelation present)

Estimated λ = 0.7971 (coefficient on lagged PPCE)


Estimated β₀ = 0.2139 (short-run MPC)

Distributed lag coefficients:


β₀ = 0.2139, β₁ = 0.2139×0.7971 = 0.1705, β₂ ≈ 0.1359, ...

Long-run MPC = β₀/(1−λ) = 0.2139/(1−0.7971) ≈ 1.054


→ A sustained $1 rise in PPDI raises PPCE by $1.05 in the long run,
but only $0.21 immediately.

Mean lag = 0.7971/(1−0.7971) = 3.93 years


Median lag = −log(2)/log(0.7971) = 3.06 years
→ Takes about 3–4 years to feel half (or the average) of the full impact.

Recovering the long-run consumption function: set PPCEt = PPCEt−1 (equilibrium), divide the
short-run equation by (1−λ) = 0.2029, and drop the lagged term:

Long-run: PPCEt = −1247.14 + 1.054·PPDIt

The Adaptive Expectations Model (AEM)


17.5

KEY CONCEPT

The AEM provides an economic rationale for the Koyck structure. Y depends on the
expected (unobservable) value X* of a regressor. Expectations are formed by adjusting last
period’s expectation by a fraction γ of the forecasting error. The resulting estimating equation
is mathematically identical to the Koyck model.

ADAPTIVE EXPECTATIONS HYPOTHESIS (CAGAN & FRIEDMAN)

Underlying model: Yt = β₀ + β₁Xt* + ut (Y depends on expected X*)


Expectation formation:
Xt* − Xt₋₁* = γ(Xt − Xt₋₁*) 0 < γ ≤ 1

Equivalently: Xt* = γXt + (1−γ)Xt₋₁*

γ = coefficient of expectation
γ = 1: current expectations = current actuals (instantaneous learning)
γ = 0: static expectations (never revise — "what is true today will always be")

After substituting and simplifying, the estimating equation becomes:


Yt = γβ₀ + γβ₁Xt + (1−γ)Yt₋₁ + [ut − (1−γ)ut₋₁] ...(17.5.5)

Comparing with Koyck (17.4.7): structurally identical with λ ↔ (1−γ)

Interpretation:
Coefficient on Xt = γβ₁ → short-run response to OBSERVED X
Coefficient on Yt₋₁ = (1−γ) → from this, recover γ = 1 − coefficient
Long-run response = γβ₁/γ = β₁ (via Xt* mechanism)

EXAMPLE 17.8 — AEM Interpretation of PPCE Regression

From Table 17.3: coefficient on PPCEt₋₁ = 0.7971 = (1 − γ̂)


→ γ̂ ≈ 0.2029 (coefficient of expectation)

Meaning: Each year, consumers close about 20% of the gap between
actual and expected disposable income.
80% of the previous expectation is retained → very slow learning.

Short-run MPC (wrt observed PPDI) = γ̂β̂₁ = 0.2139


Long-run MPC (wrt permanent PPDI) = β̂₁ = 0.2139/0.2029 ≈ 1.054

Adaptive vs Rational Expectations

The AEM was dominant until Muth, Lucas, and Sargent proposed the Rational Expectations (RE)
hypothesis: agents use all currently available relevant information to form expectations, not just
the past history of one variable. Under RE, agents do not make systematic forecast errors. The RE
hypothesis is more theoretically demanding and has itself attracted significant criticism. The AEM
remains defensible as a practical working hypothesis.

The Partial Adjustment Model (PAM)


17.6

KEY CONCEPT

PAM, developed by Nerlove, provides another economic rationale yielding the same
autoregressive estimating equation as Koyck/AEM. Y* is the desired (long-run equilibrium)
level; the partial adjustment hypothesis says actual Y only closes a fraction δ of the gap
between desired and actual Y each period.

PARTIAL ADJUSTMENT (STOCK ADJUSTMENT) MODEL

Long-run desired level: Yt* = β₀ + β₁Xt + ut ...(17.6.1)

Adjustment mechanism: Yt − Yt₋₁ = δ(Yt* − Yt₋₁) 0 < δ ≤ 1 ...(17.6.2)


where δ = coefficient of adjustment
(Yt − Yt₋₁) = actual change in Y
(Yt* − Yt₋₁) = desired change in Y

δ = 1: instant full adjustment (actual = desired each period)


δ = 0: no adjustment (Y never moves toward Y*)
Typically 0 < δ < 1 due to inertia, rigidity, contractual obligations.

Equivalently: Yt = δYt* + (1−δ)Yt₋₁ (actual Y = weighted avg of desired and lagged actual)

Substituting (17.6.1) into the adjustment equation:


Yt = δβ₀ + δβ₁Xt + (1−δ)Yt₋₁ + δut ...(17.6.5)

This is the SHORT-RUN demand/supply function for Y.


To recover LONG-RUN: divide δβ₀ and δβ₁ by δ, drop Yt₋₁.

Three Model Comparison

Koyck

Algebraic approach. Geometric decay of β's. Error: vt = ut−λut−1. Moving average error → OLS biased &
inconsistent.

Adaptive Expectations

Expectations-based. Error: ut−(1−γ)ut−1. Same structure as Koyck. OLS biased & inconsistent. Use IV.

Partial Adjustment

Adjustment-cost rationale. Error: δut (simple scalar). OLS is consistent (not biased in large samples).
Preferred estimation path.

Critical warning: All three models produce the same estimating equation: Yt = α0 + α1Xt + α2Yt−1
+ vt. You cannot tell from the regression output alone which model generated the data. The
theoretical justification must be stated in advance, not chosen for statistical convenience.

EXAMPLE — Demand for Money, Canada 1979–1988 (PAM Application)


Short-run: ln(Mt) = 0.856 − 0.0634 ln(Rt) − 0.0237 ln(GDPt) + 0.9607 ln(Mt₋₁)
t = (1.68) (−4.81) (−0.65) (23.20)
R² = 0.9482 d = 2.4582

δ̂ = 1 − 0.9607 = 0.0393
→ Only ~4% of the gap between desired and actual real cash balances is closed per
quarter.
Very slow adjustment — high δ = 0.9607 on lagged M.

Long-run interest elasticity = −0.0634/0.0393 = −1.613


Long-run: ln(M*) = 21.79 − 1.61 ln(R) − 0.60 ln(GDP)
→ Long-run elasticity is much larger (in absolute value) than short-run.

Estimation of Autoregressive Models


17.8

KEY CONCEPT

All three models share the form Yt = α0 + α1Xt + α2Yt−1 + vt. OLS faces two problems: (1)
Yt−1 is stochastic, and (2) vt may be serially correlated. Whether OLS is reliable depends
critically on which model is true.

OLS PROPERTIES BY MODEL

Model Error term vt OLS properties


─────────────────────────────────────────────────────────────────────
Koyck ut − λut₋₁ BIASED & INCONSISTENT
(Yt₋₁ correlated with vt via ut₋₁)
cov[Yt₋₁, vt] = cov[Yt₋₁, ut−λut₋₁] = −λσ² ≠ 0

Adaptive ut − (1−γ)ut₋₁ BIASED & INCONSISTENT


Expectations (same structure as Koyck)

Partial δut CONSISTENT (though still biased in small samples)


Adjustment Yt₋₁ depends on ut₋₁, ut₋₂, ... but NOT on current ut
→ as long as ut is serially independent, OLS is OK

The key distinction: in PAM, the error term vt = δut depends only on the current shock, while Yt−1
depends on past shocks. Since current and past shocks are uncorrelated (classical assumption), Yt−1
and vt are uncorrelated → OLS is consistent for PAM.

Do not choose PAM just because OLS is convenient — as Johnston warns, the PAM adjustment
pattern may sometimes be implausible. Choose the model on theoretical grounds, then solve the
estimation problem. Don’t let estimation convenience drive model choice.
The Instrumental Variables (IV) Method
17.9

KEY CONCEPT

For Koyck and AEM, OLS is inconsistent because Yt−1 is correlated with the error vt. The IV
method (Liviatan) replaces Yt−1 with a proxy (instrument) that is highly correlated with Yt−1
but uncorrelated with vt. OLS applied with this instrument yields consistent estimators.

LIVIATAN'S IV METHOD

Problem: in model Yt = α₀ + α₁Xt + α₂Yt₋₁ + vt


Yt₋₁ is correlated with vt = ut − λut₋₁ → OLS inconsistent

Solution: Use Xt₋₁ as the instrument for Yt₋₁


• Xt₋₁ is highly correlated with Yt₋₁ (both follow the series history)
• Xt₋₁ is uncorrelated with vt (X is nonstochastic)

IV Normal Equations (replace Yt₋₁ in the 3rd OLS equation with Xt₋₁):
ΣYt = nα₀ + α₁ΣXt + α₂ΣYt₋₁
ΣYtXt = α₀ΣXt + α₁ΣXt² + α₂ΣYt₋₁Xt
ΣYtXt₋₁ = α₀ΣXt₋₁ + α₁ΣXtXt₋₁ + α₂ΣYt₋₁Xt₋₁ ← Xt₋₁ replaces Yt₋₁ here

Liviatan shows: estimates from IV normal equations are consistent;


estimates from OLS normal equations are NOT consistent.

Limitation of IV: Multicollinearity

The instrument Xt−1 and the original regressor Xt both appear in the IV normal equations. Since
successive values of X in time series are typically highly correlated, multicollinearity is likely. This
makes IV estimates consistent but inefficient (large standard errors). Finding better instruments
that are more weakly correlated with each other is challenging in practice.

When a suitable instrument cannot be found, Maximum Likelihood Estimation (MLE) may be
required, though it is more computationally demanding and beyond this text’s scope. The Sargan
(SARG) test can be used to assess instrument validity.

17.10 Detecting Autocorrelation in Autoregressive Models:


Durbin h Test
KEY CONCEPT

In autoregressive models, the DW d statistic is biased toward 2, creating a built-in


tendency to miss autocorrelation. Durbin’s h statistic is designed specifically for detecting
first-order autocorrelation when the model contains Yt−1 as a regressor.

DURBIN H TEST
h = ρ̂ · √[n / (1 − n·var(α̂₂))] ...(17.10.1)

where:
n = sample size
ρ̂ ≈ 1 − d/2 (estimated from Durbin–Watson d statistic)
var(α̂₂) = variance of the lagged Y coefficient (Yt₋₁) in the regression

Under H₀: ρ = 0 (no first-order autocorrelation):


h ~asy N(0, 1) [asymptotically standard normal]

Decision rule: if |h| > 1.96, reject H₀ at 5% level (two-tailed).

Important notes:
1. Only need var(α̂₂), regardless of how many X's or lagged Y's are in the model.
2. Test is undefined if n·var(α̂₂) > 1 (rare in practice).
3. Large-sample test only — unreliable in small samples.
4. BG test (Chapter 12) is more powerful and preferred when feasible.

DURBIN h — PPCE Regression (Example 17.7, n=47)

From Table 17.3: d = 0.9619, var(α̂₂) = (0.0733)² = 0.005373

ρ̂ ≈ 1 − d/2 = 1 − 0.9619/2 = 0.5190

h = 0.5190 × √[47 / (1 − 47×0.005373)]


= 0.5190 × √[47 / 0.7475]
= 0.5190 × √62.87
= 0.5190 × 7.929
= 4.11

Since h = 4.11 >> 1.96 → REJECT H₀ → Strong positive autocorrelation confirmed.


(Note: even though d ≈ 1, which might wrongly suggest mild autocorrelation,
the h test correctly detects the problem)

BG test confirms: (n−p)R² = 15.39 with 7 df → p ≈ 3% → Autocorrelation present.


Remedy: Use Newey–West HAC standard errors (Table 17.4).

The Almon Polynomial Distributed Lag (PDL)


17.13

KEY CONCEPT

The Koyck approach assumes geometrically declining β’s — too restrictive when the lag
structure is hump-shaped, cyclical, or otherwise non-monotone. Almon’s PDL approach
approximates the β pattern using a polynomial in the lag length, allowing much more flexible
shapes. Crucially, the Yt−1 estimation problem disappears.
ALMON PDL: SETUP

Finite distributed-lag model (k lags):


Yt = α + Σᵢ₌₀ᵏ βᵢXt₋ᵢ + ut

Almon assumption: βᵢ can be approximated by an m-degree polynomial in i:


βᵢ = a₀ + a₁i + a₂i² + ··· + aₘiᵐ where m < k

For a 2nd-degree (quadratic) polynomial: βᵢ = a₀ + a₁i + a₂i²

Substituting into the distributed-lag model and constructing Z variables:


Z₀t = ΣXt₋ᵢ (sum of all k+1 lagged X values)
Z₁t = ΣiXt₋ᵢ (lag-weighted sum)
Z₂t = Σi²Xt₋ᵢ (squared-lag-weighted sum)

Estimating equation (OLS on constructed Z's):


Yt = α + a₀Z₀t + a₁Z₁t + a₂Z₂t + ut

Once â₀, â₁, â₂ are estimated, recover original β's:


β̂₀ = â₀
β̂₁ = â₀ + â₁ + â₂
β̂₂ = â₀ + 2â₁ + 4â₂
β̂₃ = â₀ + 3â₁ + 9â₂
etc. (βᵢ = â₀ + â₁i + â₂i²)

Practical Implementation Issues

▸ Choosing lag length k: Start with a large k and reduce. Use Akaike (AIC) or Schwarz (SIC)
information criteria. Too few lags → omitted variable bias. Too many → inclusion bias (less severe).

▸ Choosing polynomial degree m: m should be at least one more than the number of turning points in
the β pattern. One turning point (hump-shaped) → m = 2. Two turning points (S-shaped) → m = 3. Try
starting large and testing downward.

▸ Multicollinearity among Z’s: The Z variables are linear combinations of the same X lags → likely to
be highly correlated. Individual a’s may be insignificant even when the β’s are jointly significant. But
linear combinations of the â’s (the β̂’s) can still be estimated more precisely.

▸ Endpoint restrictions: Can constrain β0 = 0 (current X has no immediate effect) or βk = 0 (beyond lag
k, no further impact). These are “near-end” and “far-end” restrictions respectively.

EXAMPLE 17.11 — Almon PDL: U.S. Inventories on Sales (k=3, m=2)

Model: Yt = α + β₀Xt + β₁Xt₋₁ + β₂Xt₋₂ + β₃Xt₋₃ + ut


where Y = inventories, X = sales, 1954–1999 (n=43 after lagging)

Constructed Z variables (from Eq. 17.13.13):


Z₀t = Xt + Xt₋₁ + Xt₋₂ + Xt₋₃
Z₁t = Xt₋₁ + 2Xt₋₂ + 3Xt₋₃
Z₂t = Xt₋₁ + 4Xt₋₂ + 9Xt₋₃
Almon regression:
Ŷt = 25,845 + 1.1149Z₀t − 0.3713Z₁t − 0.0600Z₂t
t = (3.92) (2.07) (−0.27) (−0.13)
R² = 0.9755 (Z₁ and Z₂ individually insignificant but jointly significant via F
test)

Recovered β's:
β̂₀ = 1.1149 (current sales: significant; +ve as expected)
β̂₁ = â₀+â₁+â₂ = 1.1149−0.3713−0.0600 = 0.6836
β̂₂ = â₀+2â₁+4â₂ = 1.1149−0.7426−0.2400 = 0.1323
β̂₃ = â₀+3â₁+9â₂ = 1.1149−1.1139−0.5400 = −0.5390 ← negative; may impose endpoint
restriction

Note: Low d=0.164 likely signals model misspecification, not autocorrelation.

Advantages and Disadvantages of Almon PDL vs Koyck

FEATURE KOYCK ALMON PDL

Lag shape Must be monotone declining Any shape — hump, U, cyclical, inverted-U
(geometric)

Number of Just 3: α, β0, λ m+2: one for each polynomial degree +


params intercept

Lagged Y Yt−1 present → inconsistency risk No lagged Y → OLS is consistent and BLUE
problem

Multicollinearity Avoided (all lags collapsed to Yt−1) Z’s may be correlated; individual a’s imprecise

Choice required Just λ (estimated) Both k (lag length) and m (polynomial degree)

Economic basis Pure algebra; no theory Approximation; also largely data-driven

Granger Causality Test


17.14

KEY CONCEPT

In time series, the Granger causality test asks whether past values of X help predict Y
beyond what is already predicted by Y’s own past. It is a test of predictive precedence (does
X come before Y?), not philosophical causality. “X Granger-causes Y” = X contains useful
predictive information for Y, over and above Y’s own history.

GRANGER CAUSALITY FRAMEWORK: GDP AND MONEY SUPPLY

Two-variable system (bilateral):


GDPt = Σᵢαᵢ Mt₋ᵢ + Σⱼβⱼ GDPt₋ⱼ + u₁t ...(17.14.1)
Mt = Σᵢλᵢ Mt₋ᵢ + Σⱼδⱼ GDPt₋ⱼ + u₂t ...(17.14.2)

where u₁t and u₂t are assumed uncorrelated.

Four possible outcomes:


1. Unidirectional M → GDP: αᵢ ≠ 0 in (17.14.1); δⱼ = 0 in (17.14.2)
2. Unidirectional GDP → M: αᵢ = 0 in (17.14.1); δⱼ ≠ 0 in (17.14.2)
3. Bilateral (feedback): αᵢ ≠ 0 in (17.14.1) AND δⱼ ≠ 0 in (17.14.2)
4. Independence: αᵢ = 0 in (17.14.1) AND δⱼ = 0 in (17.14.2)

Test Procedure (F-test Approach)

GRANGER TEST STEPS (TESTING WHETHER M CAUSES GDP)

Step 1: Run restricted regression of GDPt on lagged GDP only → get RSSR
Step 2: Run unrestricted regression of GDPt on lagged GDP + lagged M → get RSSUR
Step 3: H₀: αᵢ = 0 ∀i (lagged M does not belong in the model)
Step 4: F = [(RSSR − RSSUR)/m] / [RSSUR/(n−k)] ~ F(m, n−k)
where m = number of lagged M terms, k = params in unrestricted model
Step 5: Reject H₀ if F > critical F → M Granger-causes GDP
Step 6: Repeat for M as dependent variable to test GDP → M direction

Important Caveats

▸ Stationarity required: Both variables must be stationary before running the test. First differencing
often achieves stationarity.

▸ Lag length sensitivity: The direction (and even existence) of Granger causality can depend critically
on the number of lags chosen. Always report results for multiple lag lengths. Use AIC/SIC to guide
choice.

▸ Spurious causality: If a third variable Z Granger-causes both X and Y, X may appear to Granger-cause
Y even though the true channel is X → Z → Y. Omitting Z creates false bilateral causality. Solution: use
vector autoregression (VAR) with multiple variables (Chapter 22).

▸ Not truly “causation”: Granger causality is really a test of predictive precedence — Leamer calls it
“precedence” and Diebold calls it “predictive causality.” It does not establish structural or policy
causation.

▸ Error terms must be uncorrelated: If u1t and u2t are correlated, appropriate transformation (as in
Chapter 12) may be required.

EXAMPLES OF GRANGER TESTS

Example 17.12 — Money and Income, U.S. 1960–1980 (4 lags):


Ṁ → GNP̈: F = 2.68 > critical F(4,71) = 2.50 → REJECT H₀
ĠNP → Ṁ: F = 0.56 → DO NOT REJECT H₀
Conclusion: Unidirectional causality — money growth Granger-causes GNP growth,
but not the reverse. Supports the monetarist view.

Example 17.13 — Money and Interest Rate, Canada 1979–1988:


At 2 lags: bilateral causality (R→M and M→R both significant)
At 4 lags: bilateral causality still holds
At 6 lags: bilateral causality still holds
At 8 lags: NEITHER direction is significant!
Conclusion: Results are highly sensitive to lag length — extreme caution warranted.

Granger Causality and Exogeneity

If X Granger-causes Y but Y does not Granger-cause X, does that make X exogenous? The answer
depends on which type of exogeneity is meant. Three types matter:

▸ Weak exogeneity (needed for efficient estimation): Y does not help explain X. Granger non-causality is
neither necessary nor sufficient for weak exogeneity.

▸ Strong exogeneity (needed for forecasting): no feedback from Y to X at any lag. Granger non-causality
is necessary (but not sufficient) for strong exogeneity.

▸ Super exogeneity (needed for policy analysis): parameters are invariant to changes in the X process
(counters the Lucas critique).

The practical takeaway: treat Granger causality as a useful descriptive and predictive tool for time
series, not as a definitive test of structural causation or exogeneity.

Chapter 17 — Revision Summary

▸ Lags in economics: Y responds to X with a delay due to psychological (habit, uncertainty),


technological (gestation periods), and institutional (contractual rigidities) reasons. This makes
short-run elasticities smaller than long-run elasticities.

▸ Distributed-lag model: Yt = α + ΣβiXt−i + ut. β0 = short-run (impact) multiplier; Σβi = β =


long-run multiplier. Standardised partial sums give cumulative impact by period.

▸ Ad hoc estimation fails due to: no guide on lag length, loss of degrees of freedom,
multicollinearity among lagged X values, and data mining.

▸ Koyck transformation: Assumes βk = β0λk (geometric decline, 0 < λ < 1). Converts infinite
distributed-lag to autoregressive: Yt = α(1−λ) + β0Xt + λYt−1 + vt. Mean lag = λ/(1−λ); median
lag = −log(2)/log(λ). Long-run multiplier = β0/(1−λ).

▸ Adaptive expectations (AEM): Economic rationale for Koyck structure. X*t = γXt +

(1−γ)X*t−1. Yields Yt = γβ0 + γβ1Xt + (1−γ)Yt−1 + vt. Identical estimating equation to Koyck;
coefficient on Yt−1 = (1−γ), so γ = 1 − that coefficient.

▸ Partial adjustment model (PAM): Yt − Yt−1 = δ(Y*t − Yt−1). Error term = δut (simple). OLS
is consistent (unlike Koyck/AEM). Short-run coefficient = δβ1; recover long-run by dividing by
δ.

▸ All three models yield the same estimating equation: Yt = α0 + α1Xt + α2Yt−1 + vt. Cannot
distinguish between them from output alone — theoretical justification is essential.

▸ OLS is biased and inconsistent for Koyck and AEM because Yt−1 is correlated with vt:
cov(Yt−1, vt) = −λσ2 ≠ 0. OLS is consistent for PAM because vt = δut and Yt−1 depends only on
past shocks, not the current ut.

▸ IV method (Liviatan): Use Xt−1 as an instrument for Yt−1. Yields consistent estimates for
Koyck/AEM. Drawback: multicollinearity between Xt and Xt−1 may reduce efficiency.

▸ Durbin h test: Use when model contains Yt−1 as a regressor (DW d is unreliable then). h =
ρ̂√[n/(1−n⋅var(α̂2))]. Under H0: ρ = 0, h ~ N(0,1). Large-sample test; BG test preferred for
small samples.

▸ Almon PDL: Approximates βi with a polynomial of degree m in i. Converts distributed-lag into

regression on constructed Z variables: Z0t=ΣXt−i, Z1t=ΣiXt−i, Z2t=Σi2Xt−i, etc. No Yt−1 problem


→ OLS is consistent. Requires choosing k (lag length) and m (polynomial degree) in advance.
Z’s may suffer from multicollinearity.

▸ Granger causality: “X Granger-causes Y” if past X significantly improves prediction of Y


beyond Y’s own history. Test: F-test comparing restricted (no lagged X) and unrestricted
(includes lagged X) regressions. Four outcomes: M→GDP, GDP→M, bilateral, or independence.
Requires stationarity; highly sensitive to lag length; does not establish structural causation or
exogeneity.

▸ Key numbers for the Koyck PPCE example: λ = 0.797; short-run MPC = 0.214; long-run
MPC = 1.054; mean lag ≈ 3.93 years; Durbin h = 4.11 (autocorrelation confirmed).

You might also like