GUJARATI & PORTER · BASIC ECONOMETRICS · 5TH ED.
Simultaneous-Equation
Models
Endogeneity, simultaneous-equation bias, inconsistency of OLS, structural vs reduced-form
equations, identification, and the Hausman test
Chapter 18 & 19 Pages 673–709 High Priority Part IV — Simultaneous-Equation Models
CONTENTS
1. Nature of Simultaneous-Equation Models
2. Five Classic Examples
3. Simultaneous-Equation Bias: The Proof
4. Monte Carlo Illustration
5. Structural vs Reduced-Form Equations
6. The Identification Problem
7. Order & Rank Conditions
8. Hausman Specification Test
9. Revision Summary
The Nature of Simultaneous-Equation Models
18.1
KEY CONCEPT
In single-equation models, causality runs one way: from X to Y. In simultaneous-equation
models, some Y variables are themselves determinants of other Y variables — there is two-
way (simultaneous) causation. This makes the distinction between dependent and
explanatory variables meaningless.
The defining characteristic is that the endogenous variable in one equation appears as an
explanatory variable in another equation. Because of this, that explanatory variable is stochastic
and correlated with the error term of the equation in which it appears — directly violating a core
OLS assumption.
GENERAL TWO-EQUATION SIMULTANEOUS SYSTEM
Y₁ᵢ = β₁₀ + β₁₂Y₂ᵢ + γ₁₁X₁ᵢ + u₁ᵢ ...(18.1.1)
Y₂ᵢ = β₂₀ + β₂₁Y₁ᵢ + γ₂₁X₁ᵢ + u₂ᵢ ...(18.1.2)
Y₁ and Y₂ are endogenous (jointly determined, stochastic)
X₁ is exogenous (nonstochastic, determined outside the system)
Problem: Y₂ in equation (18.1.1) is stochastic and correlated with u₁ᵢ.
Y₁ in equation (18.1.2) is stochastic and correlated with u₂ᵢ.
→ Applying OLS to each equation individually yields INCONSISTENT estimates.
They do not converge to true population values even as n → ∞.
Single-Equation Models (Parts 1–3)
One dependent Y, one or more X regressors. Causality runs X → Y only. OLS is BLUE if classical assumptions
hold. Estimation equation by equation is valid.
Simultaneous-Equation Models (Part 4)
Multiple jointly dependent Y variables. Causality is two-way. Endogenous Y appears as regressor. OLS is
inconsistent. Special methods required (IV, 2SLS).
The Two Types of Variables
▸ Endogenous variables: Jointly determined within the system. Always stochastic. In Ch. 18–19
notation: Y₁, Y₂, ..., YM. There must be exactly as many equations as endogenous variables.
▸ Predetermined variables: Values determined outside (or prior to) the current system. Treated as
nonstochastic. Include: current exogenous variables (e.g., government spending), lagged exogenous
variables, and lagged endogenous variables (whose values are known at time t). Notation: X₁, X₂, ..., XK.
18.2 Five Classic Examples
Example 18.1: Demand & Supply Model
DEMAND & SUPPLY (EXAMPLE 18.1)
Demand: Qᵈₜ = α₀ + α₁Pₜ + u₁ₜ (α₁ < 0, downward-sloping)
Supply: Qˢₜ = β₀ + β₁Pₜ + u₂ₜ (β₁ > 0, upward-sloping)
Equilibrium: Qᵈₜ = Qˢₜ
Endogenous: P and Q (jointly determined at the intersection of supply & demand)
Exogenous: none in this simple version
Problem: A shift in u₁ₜ (e.g., change in tastes) shifts the demand curve →
changes both P and Q simultaneously. Therefore corr(Pₜ, u₁ₜ) ≠ 0 and
corr(Pₜ, u₂ₜ) ≠ 0. OLS on either equation is inconsistent.
Example 18.2: Keynesian Model of Income Determination
SIMPLE KEYNESIAN MODEL (EXAMPLE 18.2)
Consumption function: Cₜ = β₀ + β₁Yₜ + uₜ (0 < β₁ < 1) ...(18.2.3)
Income identity: Yₜ = Cₜ + Iₜ ...(18.2.4)
Endogenous: C and Y (jointly determined)
Exogenous: I (investment, treated as given/fixed)
Problem: A shock to uₜ in the consumption function shifts C → changes Y (via
the income identity) → Yₜ and uₜ are correlated.
Proof: cov(Yₜ, uₜ) = σ²/(1−β₁) ≠ 0 (shown in Section 18.3)
→ OLS estimates of β₁ are inconsistent and upward biased.
Example 18.3: Wage–Price Phillips-Type Model
WAGE–PRICE MODEL (EXAMPLE 18.3)
Wage equation: Ẇₜ = α₀ + α₁UNₜ + α₂Ṗₜ + u₁ₜ
Price equation: Ṗₜ = β₀ + β₁Ẇₜ + β₂Ṙₜ + β₃Ṁₜ + u₂ₜ
Ẇ = wage change, Ṗ = price change (both endogenous: each appears in the other's equation)
UN = unemployment rate, Ṙ = cost of capital change, Ṁ = import price change (exogenous)
Wages affect prices AND prices affect wages → bilateral simultaneous causation.
Example 18.4: The IS Model
IS (GOODS MARKET) MODEL (EXAMPLE 18.4)
C = β₀ + β₁Yᵈ, T = α₀ + α₁Y, I = γ₀ + γ₁r
Yᵈ = Y − T, Y = C + I + Ḡ
Solving yields the IS curve: Y = π₀ + π₁r
Endogenous: C, T, I, Yᵈ, Y (all jointly determined within the system)
Exogenous: r (interest rate), Ḡ (government spending)
OLS on the consumption function alone ignores this interdependence → biased, inconsistent.
Example 18.5: The LM Model
LM (MONEY MARKET) MODEL (EXAMPLE 18.5)
Money demand: Mᵈₜ = a + bYₜ − crₜ
Money supply: Mˢₜ = M̄ (exogenously set by the Fed)
Equilibrium: Mᵈ = Mˢ → LM curve: Yₜ = λ₀ + λ₁M̄ + λ₂rₜ
The IS and LM curves together determine Y and r simultaneously.
Combining IS and LM: solve for both equilibrium Y and equilibrium r.
Example 18.6: Klein’s Model I
KLEIN'S MODEL I — CLASSIC MACRO ECONOMETRIC MODEL
Consumption: Cₜ = β₀ + β₁Pₜ + β₂(W+W')ₜ + β₃Pₜ₋₁ + u₁ₜ
Investment: Iₜ = β₄ + β₅Pₜ + β₆Pₜ₋₁ + β₇Kₜ₋₁ + u₂ₜ
Labor demand: Wₜ = β₈ + β₉(Y+T−W')ₜ + β₁₀(Y+T−W')ₜ₋₁ + β₁₁t + u₃ₜ
Identities: Y+T = C+I+G; Y = W'+W+P; K = Kₜ₋₁+I
Endogenous: C, I, W, Y, P, K (6 endogenous variables, 6 equations)
Predetermined: Pₜ₋₁, Kₜ₋₁, (Y+T−W')ₜ₋₁, W', G, t
These six jointly dependent variables cannot be estimated consistently
by OLS equation-by-equation.
Simultaneous-Equation Bias: The Proof
18.3
KEY CONCEPT
When an endogenous variable appears as a regressor, it is correlated with the error term
of that equation. This can be proven analytically. The result: OLS estimators are not only
biased but inconsistent — the bias does not vanish even as n → ∞.
Step 1: Prove cov(Yt, ut) ≠ 0
Using the simple Keynesian model Ct = β0 + β1Yt + ut with Yt = Ct + It, substitute to get the
reduced-form equation for Yt:
REDUCED FORM FOR INCOME
Yₜ = β₀/(1−β₁) + [1/(1−β₁)]Iₜ + [1/(1−β₁)]uₜ ...(18.3.1)
Taking deviations from expectations:
Yₜ − E(Yₜ) = uₜ/(1−β₁) ...(18.3.3)
uₜ − E(uₜ) = uₜ (since E(uₜ) = 0)
Therefore:
cov(Yₜ, uₜ) = E[Yₜ−E(Yₜ)][uₜ−E(uₜ)]
= E[uₜ/(1−β₁)] · uₜ
= σ²/(1−β₁) ...(18.3.5)
Since σ² > 0 and 0 < β₁ < 1:
cov(Yₜ, uₜ) = σ²/(1−β₁) > 0 ✗
Conclusion: Yₜ and uₜ are positively correlated.
OLS assumption of zero correlation between regressor and error term is VIOLATED.
Step 2: Prove OLS β̂1 is Inconsistent
OLS ESTIMATOR AND ITS PROBABILITY LIMIT
OLS estimator: β̂₁ = Σcₜyₜ / Σyₜ² = β₁ + Σyₜuₜ/Σyₜ² ...(18.3.7)
E(β̂₁) cannot be evaluated directly since E(A/B) ≠ E(A)/E(B).
But consistency requires: plim(β̂₁) = β₁.
Applying probability limit rules:
plim(β̂₁) = β₁ + plim(Σyₜuₜ/n) / plim(Σyₜ²/n)
As n → ∞: Σyₜuₜ/n → cov(Yₜ,uₜ) = σ²/(1−β₁)
Σyₜ²/n → var(Yₜ) = σ²Y
Therefore:
plim(β̂₁) = β₁ + σ²/[(1−β₁)σ²Y] ...(18.3.10)
Since σ² > 0, (1−β₁) > 0, σ²Y > 0:
plim(β̂₁) > β₁ ← OLS OVERESTIMATES the true β₁
This overestimation does NOT disappear as n → ∞.
β̂₁ is BIASED and INCONSISTENT.
The core message (White, Horsman & Wyatt): “In contrast to single-equation models, we can no
longer assume that variables on the right-hand side of the equation are uncorrelated with the error
term.” This is what simultaneous-equation bias is all about, and the bias persists even in very large
samples.
Monte Carlo Illustration of the Bias
18.4
KEY CONCEPT
A numerical Monte Carlo experiment confirms the theoretical result: OLS estimates of the
MPC from the consumption function are systematically upward biased by a precisely
predictable amount.
MONTE CARLO SETUP (TRUE VALUES: Β₀ = 2, Β₁ = 0.8)
Model: Cₜ = 2 + 0.8Yₜ + uₜ (true consumption function)
Yₜ = Cₜ + Iₜ (income identity)
uₜ ~ N(0, 0.04) (var σ² = 0.04)
Iₜ given (nonstochastic)
Generating Yₜ from reduced form (18.3.1) with the given Iₜ and uₜ values
→ obtain Cₜ from the consumption function
→ regress Cₜ on Yₜ using OLS
Theoretical bias:
β̂₁ = β₁ + Σyₜuₜ / Σyₜ²
= 0.8 + 3.8/184
= 0.8 + 0.02065
= 0.82065 ← upward biased as predicted
OLS regression results:
Ĉₜ = 1.4940 + 0.82065Yₜ
se = (0.354) (0.014) R² = 0.9945
Estimated β₁ = 0.82065 (TRUE = 0.80) → upward bias of 2.1%
The OLS estimate is EXACTLY what the theory predicted — confirming inconsistency.
Note that both β̂0 and β̂1 are biased. The magnitude of bias depends on β1, σ2, and var(Y). In
general, the direction of bias depends on the model structure and the sign of cov(Y, u).
Structural vs Reduced-Form Equations
19.1
KEY CONCEPT
Before estimation can proceed, the equations must be classified. Structural equations
describe economic behaviour directly and contain both endogenous and predetermined
variables. Reduced-form equations express each endogenous variable solely in terms of
predetermined variables — and can be estimated by OLS without bias.
Structural Equations
The equations as specified by economic theory (like the consumption function, the investment
function, the supply curve). They contain the structural (behavioural) parameters we care about
estimating. The coefficients are the structural parameters (β’s and γ’s).
Reduced-Form Equations
Obtained by solving the structural system to express each endogenous variable purely as a function
of all predetermined variables and stochastic disturbances. Since only predetermined variables
(assumed uncorrelated with the error) appear on the right-hand side, OLS can be applied directly.
KEYNESIAN MODEL — Reduced Forms
Structural equations:
Cₜ = β₀ + β₁Yₜ + uₜ (consumption function)
Yₜ = Cₜ + Iₜ (income identity)
Substituting (18.2.3) into (18.2.4) and solving:
Reduced form for Y:
Yₜ = Π₀ + Π₁Iₜ + wₜ ...(19.1.2)
where Π₀ = β₀/(1−β₁), Π₁ = 1/(1−β₁), wₜ = uₜ/(1−β₁)
Reduced form for C:
Cₜ = Π₂ + Π₃Iₜ + wₜ ...(19.1.4)
where Π₂ = β₀β₁/(1−β₁) [sic: β₀/(1−β₁) × β₁], Π₃ = β₁/(1−β₁)
Π₁ = 1/(1−β₁) = impact multiplier for Y with respect to I
→ If β₁ = 0.8: Π₁ = 5 (a $1 rise in investment raises income by $5)
Π₃ = β₁/(1−β₁) = impact multiplier for C with respect to I
→ If β₁ = 0.8: Π₃ = 4 (a $1 rise in investment raises consumption by $4)
KEY: Since only I (exogenous, uncorrelated with u) appears in the reduced forms,
OLS on the reduced-form equations gives CONSISTENT estimates of Π₀ and Π₁.
From these, we can recover β₀ and β₁ (the structural parameters):
β₁ = Π₃/Π₁ (= Π₃/(Π₃+1), since Π₁−Π₃ = 1).
This procedure is called Indirect Least Squares (ILS), studied in Chapter 20.
Reduced-form coefficients (the Π’s) are also called impact multipliers or short-run multipliers
— they measure the immediate effect of a one-unit change in a predetermined variable on an
endogenous variable. In the Keynesian model, Π1 = 1/(1−β1) is the famous income multiplier.
The Identification Problem
19.2
KEY CONCEPT
The identification problem asks: can we uniquely recover the structural parameters from
the estimated reduced-form coefficients? Without identification, we cannot tell which
equation we are actually estimating from the data, no matter how large the sample.
The problem arises because a given set of data (and hence the same reduced-form) may be
compatible with different structural models. In the demand–supply context: with data on P and Q
only and no additional information, we cannot distinguish whether our regression is estimating the
demand curve, the supply curve, or some linear combination (mongrel) of both.
THE "MONGREL" EQUATION PROBLEM
Demand: Qₜ = α₀ + α₁Pₜ + u₁ₜ
Supply: Qₜ = β₀ + β₁Pₜ + u₂ₜ
Multiply demand by λ and supply by (1−λ) and add:
Qₜ = γ₀ + γ₁Pₜ + wₜ ...(19.2.10)
where γ₀ = λα₀ + (1−λ)β₀, γ₁ = λα₁ + (1−λ)β₁
This "mongrel" equation is observationally INDISTINGUISHABLE from either the
demand or the supply function because all three involve only Q and P.
With data on P and Q only: the same data fits ANY linear combination of demand
and supply. There is no way to identify either curve.
→ Both equations are UNDERIDENTIFIED (unidentified).
The Three Cases of Identification
STATUS DEFINITION CONSEQUENCE
Underidentified Cannot obtain unique structural parameters Estimation impossible. No point
(unidentified) from reduced-form parameters proceeding.
Exactly (just) Unique values of all structural parameters One solution. ILS gives unique
identified can be obtained from reduced-form structural estimates.
Overidentified More than one value of one or more Multiple solutions. ILS not appropriate.
structural parameters can be obtained Need 2SLS or other methods.
How to Identify: The Key Insight
An equation can be identified if it excludes variables that appear in other equations. The
excluded variable provides additional information that allows us to distinguish this equation from a
mongrel combination.
IDENTIFICATION EXAMPLES
Case 1 — UNDERIDENTIFIED (both equations):
Demand: Qₜ = α₀ + α₁Pₜ + u₁ₜ
Supply: Qₜ = β₀ + β₁Pₜ + u₂ₜ
Same variables in both equations. No additional information. Neither identified.
Case 2 — Supply JUST IDENTIFIED, Demand UNDERIDENTIFIED:
Demand: Qₜ = α₀ + α₁Pₜ + α₂Iₜ + u₁ₜ (I = consumer income; exogenous)
Supply: Qₜ = β₀ + β₁Pₜ + u₂ₜ
Supply excludes I (present in demand) → Supply is just identified (K−k = m−1 = 1)
Demand excludes nothing from supply → Demand is underidentified.
Intuition: as income shifts the demand curve, we trace out the stable supply curve.
Case 3 — BOTH equations just identified:
Demand: Qₜ = α₀ + α₁Pₜ + α₂Iₜ + u₁ₜ (excludes Pₜ₋₁)
Supply: Qₜ = β₀ + β₁Pₜ + β₂Pₜ₋₁ + u₂ₜ (excludes Iₜ)
Each equation excludes exactly one predetermined variable present in the other.
Both are just identified.
Case 4 — Supply OVERIDENTIFIED:
Demand: Qₜ = α₀ + α₁Pₜ + α₂Iₜ + α₃Rₜ + u₁ₜ (I = income, R = wealth)
Supply: Qₜ = β₀ + β₁Pₜ + β₂Pₜ₋₁ + u₂ₜ
Supply excludes both Iₜ and Rₜ → two estimates of β₁ possible → overidentified.
Demand excludes only Pₜ₋₁ → just identified.
Order and Rank Conditions for Identification
19.3
Order and Rank Conditions for Identification
19.3
KEY CONCEPT
Rather than working through reduced-form algebra each time, two mechanical rules — the
order condition (necessary) and the rank condition (necessary and sufficient) — allow
quick identification checks.
Notation
IDENTIFICATION NOTATION
M = total number of endogenous variables in the full model
m = number of endogenous variables in the specific equation being checked
K = total number of predetermined variables in the model (including the intercept)
k = number of predetermined variables in the specific equation being checked
The Order Condition (Necessary but Not Sufficient)
ORDER CONDITION
Definition 1: The equation must EXCLUDE at least M−1 variables from the model.
Excludes exactly M−1 → just identified
Excludes more than M−1 → overidentified
Excludes fewer than M−1 → underidentified
Definition 2 (equivalent):
(K − k) ≥ (m − 1) ...(19.3.1)
K − k = number of predetermined variables EXCLUDED from the equation
m − 1 = number of endogenous variables INCLUDED in the equation, minus 1
K − k = m − 1 → just identified
K − k > m − 1 → overidentified
K − k < m − 1 → underidentified (not identified)
ORDER CONDITION — Applied to the Four Cases
Case 2: Demand Qₜ = α₀ + α₁Pₜ + α₂Iₜ + u₁ₜ vs Supply Qₜ = β₀ + β₁Pₜ + u₂ₜ
M=2 (Q,P endogenous), K=1 (I predetermined)
Demand: m=2, k=1 → K−k = 0; m−1 = 1 → 0 < 1 → NOT IDENTIFIED ✗
Supply: m=2, k=0 → K−k = 1; m−1 = 1 → 1 = 1 → JUST IDENTIFIED ✓
Case 3: Add Pₜ₋₁ to supply. K=2 (I, Pₜ₋₁ predetermined)
Demand: m=2, k=1 (has I) → K−k=1; m−1=1 → 1=1 → JUST IDENTIFIED ✓
Supply: m=2, k=1 (has Pₜ₋₁) → K−k=1; m−1=1 → 1=1 → JUST IDENTIFIED ✓
Case 4: Demand has I and R; Supply has Pₜ₋₁. K=3 (I, R, Pₜ₋₁)
Demand: m=2, k=2 (has I,R) → K−k=1; m−1=1 → 1=1 → JUST IDENTIFIED ✓
Supply: m=2, k=1 (has Pₜ₋₁) → K−k=2; m−1=1 → 2>1 → OVERIDENTIFIED ✓
The Rank Condition (Necessary and Sufficient)
The order condition is necessary but not sufficient. Even if K−k ≥ m−1, an equation may still be
unidentified if the excluded variables are linearly dependent. The rank condition provides the
definitive test.
RANK CONDITION
An equation is identified if and only if at least one NON-ZERO determinant of
order (M−1) × (M−1) can be constructed from the coefficients of the variables
EXCLUDED from that equation but INCLUDED in the other equations.
Procedure:
1. Write the full model in tabular form (Table 19.1 style): rows = equations,
columns = all variables; enter coefficient values.
2. For the equation being tested, strike out its entire row.
3. Also strike out columns for variables that DO appear in that equation
(i.e., keep only columns for variables EXCLUDED from it).
4. Form all possible (M−1)×(M−1) submatrices from remaining entries.
5. If at least one determinant ≠ 0 → IDENTIFIED.
If all determinants = 0 → NOT IDENTIFIED (despite passing order condition).
The rank of matrix A = M−1 → identified.
The rank of matrix A < M−1 → not identified even if order condition passes.
Practical note (Harvey): “The order condition is usually sufficient to ensure identifiability, and
although it is important to be aware of the rank condition, a failure to verify it will rarely result in
disaster.” For large models, applying the rank condition (which requires matrix determinants) is very
time-consuming. In practice, use the order condition routinely and apply the rank condition only
when suspicious.
General Principles of Identifiability
RANK (MATRIX
CONDITION ORDER (K−K VS M−1) STATUS
A)
Just identified K−k = m−1 rank = M−1 Unique structural estimates possible (ILS)
Overidentified K−k > m−1 rank = M−1 Multiple solutions for some parameters
(2SLS needed)
Not identified (by K−k ≥ m−1 (passes rank < M−1 Cannot recover structural parameters
rank) order)
Underidentified K−k < m−1 rank < M−1 Impossible to identify. No estimation
possible.
The Hausman Specification Test for Simultaneity
19.4
KEY CONCEPT
Before discarding OLS, we should test whether simultaneity actually exists. The
Hausman test detects whether a suspected endogenous regressor is indeed correlated with
the error term. If not, OLS remains valid and more efficient than alternatives.
The logic: if there is no simultaneity problem, the endogenous regressor Pt and the error u2t should
be uncorrelated. The Hausman test formalises this by checking whether the residuals from the
reduced-form regression of Pt on exogenous variables add any explanatory power when included in
the structural equation.
HAUSMAN TEST PROCEDURE
Model being tested (supply equation):
Qˢₜ = β₀ + β₁Pₜ + u₂ₜ ...(19.4.2)
Pₜ is suspected to be endogenous (correlated with u₂ₜ)
Step 1: Run the REDUCED-FORM regression of Pₜ on ALL exogenous variables
(here: Iₜ and Rₜ from the demand function):
Pₜ = Π₀ + Π₁Iₜ + Π₂Rₜ + vₜ ...(19.4.3)
Obtain the residuals v̂ₜ from this regression.
Step 2: Include v̂ₜ as an additional regressor in the structural supply equation:
Qₜ = β₀ + β₁P̂ₜ + β₁v̂ₜ + u₂ₜ ...(19.4.7)
(Note: P̂ₜ and v̂ₜ have the same coefficient β₁)
(Pindyck & Rubinfeld suggest using Pₜ instead of P̂ₜ)
Step 3: Test the hypothesis H₀: coefficient of v̂ₜ = 0
→ Use t test (one endogenous regressor) or F test (multiple)
If SIGNIFICANT: reject H₀ → Pₜ is endogenous → simultaneity problem exists
→ OLS inconsistent, use 2SLS or IV instead.
If NOT SIGNIFICANT: do not reject H₀ → no simultaneity → OLS is valid.
EXAMPLE 19.5 — Pindyck–Rubinfeld Model of Public Spending
Model:
EXP = β₁ + β₂AID + β₃INC + β₄POP + u (spending equation)
AID = δ₁ + δ₂EXP + δ₃PS + v (grants equation)
Endogenous: EXP, AID. Exogenous: INC, POP, PS.
Hausman test on the spending equation:
Step 1: Regress AID on INC, POP, PS → get residual ŵᵢ
Step 2: Regress EXP on AID, INC, POP, ŵᵢ:
EXP̂ = −89.41 + 4.50AID + 0.00013INC − 0.518POP − 1.39ŵᵢ
t = (−1.04) (5.89) (3.06) (−4.63) (−1.73)
R² = 0.99
Coefficient of ŵᵢ: t = −1.73
At 5% level: NOT significant → no simultaneity problem → OLS is valid at 5%.
At 10% level: SIGNIFICANT → possible simultaneity → use 2SLS as a precaution.
For comparison, plain OLS of spending equation:
EXP̂ = −46.81 + 3.24AID + 0.00019INC − 0.597POP
t = (−0.56) (13.64) (8.12) (−5.71) R² = 0.993
When simultaneity is explicitly allowed for, AID coefficient rises (4.50 vs 3.24)
but its t-statistic drops (5.89 vs 13.64) — OLS was overstating AID's significance.
Extension: Testing for Exogeneity of Multiple Variables
The Hausman test can be extended to test whether multiple Y variables in a structural equation are
truly endogenous (Section 19.5). Obtain predicted values Ŷ2i and Ŷ3i from their reduced-form
regressions, then include them as additional regressors alongside Y2i and Y3i in the structural
equation. Test H0: λ2 = λ3 = 0 using an F test. Rejection confirms endogeneity.
Bridge to Chapter 20: Once identification is confirmed (Chapter 19) and we know OLS is
inconsistent for identified equations (Chapter 18), Chapter 20 develops the estimation methods:
Indirect Least Squares (ILS) for just-identified equations, and Two-Stage Least Squares (2SLS) — the
workhorse method for overidentified equations. Both produce consistent estimates and properly
account for the simultaneity structure.
Chapters 18 & 19 — Revision Summary
▸ Simultaneous-equation models: Two or more endogenous variables jointly determined.
Endogenous Y in one equation appears as a regressor in another. This creates a two-way causal
flow that OLS cannot handle correctly.
▸ Endogenous vs Predetermined: Endogenous variables (Y's) are determined within the
system and are stochastic. Predetermined variables (X's: current exogenous, lagged exogenous,
lagged endogenous) are treated as nonstochastic and uncorrelated with current disturbances.
▸ OLS inconsistency — the proof: In the Keynesian model, cov(Yt, ut) = σ2/(1−β1) ≠ 0,
violating the OLS assumption of uncorrelated regressor and error. plim(β̂1) = β1 +
σ2/[(1−β1)σ2Y] > β1: OLS overestimates MPC. This bias does not vanish as n → ∞.
▸ Monte Carlo confirmation: With true β1 = 0.8, OLS yields β̂1 = 0.82065 — exactly the
upward-biased value predicted by the theory.
▸ Classic examples: Demand–supply (P, Q endogenous); Keynesian IS model (C, Y endogenous);
Phillips wage–price (Ẇ, Ṗ endogenous); IS–LM (Y, r, C, I, T endogenous); Klein’s Model I (C, I,
W, Y, P, K endogenous).
▸ Structural equations: Behavioural equations expressing economic theory directly; contain
both endogenous and predetermined variables. Parameters are the structural coefficients (β’s
and γ’s).
▸ Reduced-form equations: Obtained by solving the structural system so that each endogenous
variable is expressed only in terms of predetermined variables. Can be estimated consistently
by OLS. Coefficients are impact (short-run) multipliers (Π’s).
▸ Indirect Least Squares (ILS): Estimate reduced-form by OLS, then algebraically recover
structural coefficients. Valid only for just identified equations. Studied fully in Chapter 20.
▸ Identification problem: Can unique structural coefficients be recovered from the reduced-
form? Without identification, we cannot tell which equation the data is actually estimating —
we might be fitting a mongrel combination of demand and supply.
▸ Three identification states: (1) Underidentified: cannot estimate structural parameters. (2)
Just (exactly) identified: unique solution. (3) Overidentified: too many restrictions, multiple
solutions for some parameters.
▸ Order condition (necessary): K−k ≥ m−1. Variables excluded from the equation must be at
least as many as endogenous variables included (minus 1). Equivalently: the equation must
exclude at least M−1 variables from the model.
▸ Rank condition (necessary and sufficient): The matrix of coefficients of excluded variables
in other equations must have rank = M−1. In practice, the order condition is usually sufficient.
▸ Zero restrictions criterion: Identification is typically achieved by restricting certain variables
to zero in a given equation (excluding them). The excluded variable must genuinely belong to
other equations in the system.
▸ Hausman specification test: Tests whether a suspected endogenous regressor is actually
correlated with the error term. Step 1: regress the suspected endogenous variable on all
exogenous variables → get residuals v̂ . Step 2: include v̂ as an extra regressor in the structural
equation. Step 3: if the coefficient of v̂ is significant, simultaneity exists → use 2SLS; if not, OLS
is valid.
▸ Key identification numbers to check: For each equation, compute M (total endogenous), m
(endogenous in equation), K (total predetermined), k (predetermined in equation). If K−k =
m−1: just identified. If K−k > m−1: overidentified. If K−k < m−1: underidentified.
▸ Most macro models are overidentified rather than underidentified, so 2SLS is typically the
correct estimation method (Chapter 20).