Structural Analysis of VAR Models
Structural Analysis of VAR Models
net/publication/379248057
CITATIONS READS
0 418
1 author:
Christis Katsouris
University of Southampton
26 PUBLICATIONS 11 CITATIONS
SEE PROFILE
All content following this page was uploaded by Christis Katsouris on 25 March 2024.
Christis Katsouris†
University of Southampton
arXiv:2312.06402v9 [[Link]] 5 Feb 2024
(Work-in-progress)
February 6, 2024
Abstract
This set of lecture notes discuss key concepts for the Structural Analysis of Vector Autoregressive
models for the teaching of a course on Applied Macroeconometrics with Advanced Topics.
*I am grateful to Professor Markku Lanne and Professor Mika Meitz from the Faculty of Social Sciences, University
of Helsinki for helpful conversations. I also wish to thank Professor Jose Olmo and Professor Tassos Magdalinos from
the School of Economics as well as Professor Zudi Lu and Dr. Chao Zheng from the School of Mathematical Sciences,
University of Southampton for helpful discussions. Financial support from the Research Council of Finland (Grant 347986)
is also gratefully acknowledged. Address correspondence to Christis Katsouris, Faculty of Social Sciences, University of
Southampton, United Kingdom; Email Address: [Link]@[Link].
† Dr. Christis Katsouris, is currently a Postdoctoral Researcher at the Faculty of Social Sciences, University of Helsinki.
1
Contents
1. Introduction 5
1.1. The Identification Problem . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6
2
5. Dynamic Causal Effects 69
5.1. Identification and Estimation of Causal Effects . . . . . . . . . . . . . . . . . . . . . . . . . . . 69
5.1.1. Causal Effects and IV Regression . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 69
5.1.2. Causal Effects Identification with SVARs . . . . . . . . . . . . . . . . . . . . . . . . . . 70
5.2. Counterfactual Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 73
5.2.1. Causal Effect Inference with Time Series Data . . . . . . . . . . . . . . . . . . . . . . . 73
5.2.2. Potential Outcome Framework . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 73
5.3. Time Series Experiments and Causal Effects . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 75
5.3.1. Estimating Contemporaneous Effect of Treatment . . . . . . . . . . . . . . . . . . . . . . 75
5.3.2. Null Hypothesis of Temporal Causal Effects . . . . . . . . . . . . . . . . . . . . . . . . . 76
5.4. Bias-Reduced Doubly Robust Estimation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 77
5.4.1. Marginal Treatment Effects . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 77
5.5. The Augmented Synthetic Control Method . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 79
5.6. Panel Data Regression for Dynamic Causal Effects . . . . . . . . . . . . . . . . . . . . . . . . . 81
5.6.1. Panel Experiments and Dynamic Causal Effects . . . . . . . . . . . . . . . . . . . . . . . 81
5.6.2. Identifying Dynamic Causal Effects with Panel Data: Applications . . . . . . . . . . . . . 82
6. Advanced Topics 84
6.1. Testing for Unstable Root in Structural ECMs . . . . . . . . . . . . . . . . . . . . . . . . . . . . 84
6.1.1. Testing Cointegration Rank when some Cointegrating Directions are Changing . . . . . . 86
6.2. Empirical Likelihood Estimation Approach . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 87
6.2.1. High Dimensional Generalized Empirical Likelihood Estimation . . . . . . . . . . . . . . 88
6.3. High Dimensional VARs with Common factors . . . . . . . . . . . . . . . . . . . . . . . . . . . 89
6.4. Fast Algorithm for Detection of Breaks in Large VAR models . . . . . . . . . . . . . . . . . . . . 91
6.4.1. Illustrative Examples . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 91
6.4.2. Main Asymptotic theory results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 92
6.4.3. Large VAR Models: Model Formulation . . . . . . . . . . . . . . . . . . . . . . . . . . . 94
6.4.4. A Block Segmentation Scheme Based Algorithm . . . . . . . . . . . . . . . . . . . . . . 95
6.4.5. Consistency of the BSS Estimator . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 97
6.5. Sieve Bootstrap for Functional Vector Autoregressions . . . . . . . . . . . . . . . . . . . . . . . 98
6.5.1. The bootstrap procedure . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 98
7. Conclusion 101
3
B Algebraic Theory of Identification in Structural Models 110
B1. Relation between Equivalence and Orbit . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 110
B2. Iterative Schemes for Large Linear Systems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 110
References 111
4
1. Introduction
This set of lecture notes are meant to discuss fundamental concepts on the mechanisms of structural
vector autoregressive models with many economic theory driven examples which are drawn from sev-
eral seminal studies presented in the existing literature. The particular perspective we follow permits
to present several mathematical and theoretical concepts in an informal way while discussing important
concepts by presenting how some long-standing problems in the applied macroeconometrics literature
have been tackled. Moreover, our discussion of selected studies from the literature in the form of ex-
amples that present key aspects in the identification and estimation of SVAR models, main theoretical
econometric concepts such as matrix representations and asymptotic theory convergence results; are
motivated from a research-based pedagogical approach that encourages the reader to seek a compre-
hensive understanding of these concepts while fostering further discussions on current state-of-the art
methodological approaches and debates in the literature. Note that a full treatment of these topics are
presented in the "Time Series Analysis" book of Hamilton (1994). Additional material can be found in
the books of Gourieroux and Monfort (1997), Lütkepohl (2005) and Kilian and Lütkepohl (2017).
More specifically, Stock and Watson (2001) describe the key job tasks of applied macro-econometricians,
which shall also serve as the main learning objectives motivating the preparation of this set of lecture
notes. There are summarized as below:
• To be able to describe/summarize and infer about the presence of stochastic trends, cointegration
dynamics, unit root behaviour as well as structural breaks in macroeconomic time series.
• To be able to recover the structure of the macroeconomy based on data and theoretical-driven
questions using structural analysis (i.e., identification of structural shocks) as well as impulse
response analysis and construction of related test statistics and confidence intervals.
• To be able to advise macroeconomic policy-makers (e.g., using counterfactual analysis and related
methods).
Vector Autoregressive models are a popular econometric tool for estimating the dynamic interaction
between the variables included in the system. Moreover, several test statistics are constructed using VAR
models such that: Wald tests for Granger Causality, impulse response functions (IRFs) and forecast error
variance decomposition (FEVDs). Therefore, inference on these statistics is typically based on either
on first-order asymptotic approximations or on bootstrap methods. However, the deviation from i.i.d
innovations as in the case of conditional heteroscedasticity, invalidates a number of standard inference
procedures such that the application of these methods may lead to conclusions that are not in line with
the true underlying dynamics. Consequently, in many VAR applications there is need for inference
methods that are valid when innovations are only serially correlated but not independent. Nevertheless,
the VAR representation of the economy is suitable for estimation/testing as well as forecasting purposes
but is not adequate for policy analysis which is the purpose of this course (thus, our interest in SVARs).
5
1.1. The Identification Problem
Roughly speaking, the identification problem of structural parameters in linear simultaneous equations
models are closely related1 . Seminal papers discussing the identification problem include among others
Sargan (1983) and Dufour (2003). According to Hausman and Taylor (1983), necessary and sufficient
conditions for identification with linear coefficient and covariance restrictions are developed in a limited
information context. In particular, imposing covariance restrictions facilitate identification iff they imply
that a set of endogenous variables is predetermined in the equation of interest - which generalizes the
notion of recursiveness for structural learning and causal recovery. Under full information, covariance
restrictions imply that residuals from other equations are predetermined in a particular equation, and,
under certain conditions, can facilitate system identification. This implies that in the general case, FIML
first order conditions show that if a system of equations is identifiable as a whole, covariance restrictions
cause residuals to behave as instruments.
Furthermore, imposing exclusion restrictions and the normalization βii = 1, then the classical structural
econometric specification for the simultaneous equation model Y B′ + ZΓ′ = U, where the vector yi , for
i ∈ {1, ..., G}, includes G jointly dependent random variables. A key result for system identification
purposes is the fact that the relative triangularity of equations (i, j) is precisely equivalent to a zero in
the (i, j)−th position of B−1 , denoted as B−1
i j . As a result, equations (i, j) are relatively recursive if and
only if there are no paths by which a shock to u j can be transmitted to yi .
Lemma 1. Zero restrictions on (B, Σ1 ) are sufficient for identification if and only if they induced the
equivalence relation:
Corollary 1. If yi is predetermined in the first equation, then every endogenous variable in the i−th
structural equation (y j ) for which σ j1 = 0 is predetermined in the first structural equation.
Remark 1. Notice that in the case of a diagonal disturbance covariance matrix, an endogenous variable
is predetermined in the first equation implying a special structure. In particular, if yi is predetermined
in the first equation, every endogenous variable in the i−th equation is also predetermined in the first
equation. Consequently, every endogenous variable in their respective equations is predetermined in the
first equation. In other words, in the case of a diagonal variance matrix Σ, a relatively recursive situation
is one in which a set of endogeneous variables determines itself independently from the first equation
variable, and thus all elements of the set are predetermined in the first equation.
Next, we consider the equivalence between instrumental variables and covariance restrictions.
1 The notion of coefficient restrictions was originally extended to show the equivalence relationship between identifiability
and instrumental variables estimation, that is, the restrictions required for identification give rise to instrumental variables
required for estimation.
6
Lemma 2 (Rank). The parameters of the first structural equation are identifiable if and only if the 2SLS
estimator is well-defined, using W as the matrix of instruments.
Notice that rank conditions for identification are commonly used as identifying conditions for static and
dynamic factor models as well as cointegrating regression models (e.g., VECM forms).
A well-known problem to the literature of structural VAR models is that imposing the assumption of
Gaussian errors renders non-identifiable structural parameters due to the nonlinearities and complex
dependence structure of the system. Additional identifying restrictions are needed to ensure robust
identification and estimation of SVARs. Based on these concerns, the framework of Lanne et al. (2017)
shows that in practice the Gaussian case is an exception in that a SVAR model whose error vector
consists of independent non-Gaussian components is, without any additional restrictions, identified and
leads to essentially unique responses. The identification scheme employs the non-Gaussianity of struc-
tural shocks while an MLE estimator for the parameters of the non-Gaussian SVAR model is consistent
and asymptotically normally distributed. Moreover, the scheme allows to impose additional identifying
restrictions which can be formulated into testable hypotheses. Further extensions of the non-Gaussian
identification strategy includes cases such as when the system includes only strictly positive compo-
nents (see, Nyberg and Rauhala (2022)), cases when a GMM estimation approach is employed (see,
Lanne and Luoto (2021) and Keweloh (2021)) as well as specific cases when structural shocks are lep-
tokurtic (e.g. see Lanne et al. (2023)) or heavy-tailed (e.g., see Anttonen et al. (2023)).
Furthermore, while the reduced-form VAR model can be seen as a convenient description of the joint
dynamics of a number of time series that facilitates forecasting, the SVAR model is more appropriate
for answering economic questions of theoretical and practical interest. The main tools in analyzing the
dynamics in SVAR models are the impulse response function and the forecast error variance decompo-
sition. Moreover, usually when fitting a SVAR model, the joint distribution of the error terms is almost
always (either explicitly or implicitly) assumed to follow a multivariate Gaussian distribution which
might render the system non-identifiable especially under the presence of non-linearities, higher-order
dynamics and heavy-tails. The positive news is that the joint distribution of the reduced-form errors is
fully determined by their covariances only but this which implies that the structural errors of the system
cannot be identified unless we impose additional linear restrictions (see, Lanne et al. (2017)). However
the use of suitable orthogonal transformations or even permutations (due to being exchangeable random
variables), can be shown to induce observationally equivalent SVAR processes.
7
Furthermore, we discuss some key identification strategies in the following sections. Generally, speak-
ing understanding all the available tools in our disposal is important when considering relevant empirical
and theoretical driven questions in relation to the impact of structural shocks on economic outcomes.
Therefore, in this set of lecture notes we present some key concepts and recent developments in the
literature with focus on the implementation of econometric and statistical tools for the identification and
estimation of structural vector autoregressive models, (SVAR). In particular, we discuss several different
approaches that are commonly employed for the structural identification of these interdependent systems
in relation to economic phenomena such as business cycle fluctuations and macroeconomic conditions
in order to obtain policy recommendations. On the other hand, exogenous variation to macroeconomic
conditions such as the adverse effect of climate change across interconnected economies requires novel
econometric methods to ensure robust identification and estimation2 (e.g., see Pesaran et al. (2000)).
The stability of cointegrated VAR processes can be assessed with respect to the long-run relations as
well as with respect to the short-run parameters. Usually the long-run equilibrium relations are based
on economic theory and so we expect them to be relatively more stable after a perturbation of struc-
tural shocks (such as a change in the monetary policy framework), than the short-run coefficients which
are data-driven. Consider for example the study of the transmission of monetrary policy shocks to
macroeconomic variables using structural cointegrated VAR models. Roughly speaking the identifica-
tion of SVAR models is achieved by imposing restrictions on the covariance matrix of the residuals of a
reduced-form VAR to provide an economic interpretation of shocks while a structural cointegrated VAR
model is identified by imposing restrictions on the cointegrating space. Consequently, based on a struc-
tural identification scheme of the covariance matrix then metrics such as impulse-responses, variance-
decompositions and historical decompositions can be constructed.
In particular, Rubio-Ramirez et al. (2010) verifies the identification for linear Gaussian models by check-
ing whether two parameter points are observetionally equivalent iff they have the same reduced-form
representation. Moreover, a global identification can be verified by checking whether the model is
identified at different points in the parameter space prior to the estimation step. On the other hand, a
well-know fact in the literature of nonstationary autoregressive models, asymptotic distribution discon-
tinuities arise in the case when the parameter is at the boundaries of the parameter space.
Definition 1. Consider a SVAR(p) process with restrictions imposed by R. Then, the SVAR(p) is ex-
actly identified iff, for almost any reduced-form parameter point (B, Σ), there exists a unique structural
parameter point (A0 , A+ ) ∈ R such that g(A0 , A+ ) = (B, Σ).
2 Inparticular, the omission of such an important component as the climate when considering the structural analysis of
time series models could lead to non-causality, leading to inaccurate model estimates and forecasts (see, Lütkepohl (1982)).
Further properties on Noncausal vector AR processes are presented by Davis and Song (2020).
8
Generally, it is well-known that identification in parametric models implies that a specific model struc-
ture is identifiable if there is no other structure which is observationally equivalent (see, also Sargent
(1976)). For instance, the identification scheme proposed by Lanne et al. (2017) requires the construc-
tion of observationally equivalent SVAR processes when comparing two MA(∞) representations of the
model. In conjecture with conventional rank conditions that ensure the valid identification of SVAR
models under non-Gaussianity the identification of structural parameters can be extended into more
complex model configurations regardless of the presence of nonlinear or nonstationary dynamics.
More specifically, the main intuition of the identification methodology of Rubio-Ramirez et al. (2010),
is to establish general rank conditions for global identification of identified and exactly identified models.
Moreover, Bognanni (2018) highlights that the key conceptual properties which ensure SVAR tractabil-
ity include: (i) observational equivalence of points (A, F) and (AQ, FQ) for Q ∈ On in the structural
model, and (ii) the ability to reparametrize the structural model from (A, F) to (B, H, Q) and separate
identifiable elements (B, H) and non-identifiable elements Q. Thus, the above definitions imply the
existence of a unique mapping of reduced-form parameters which is invariant to scaling and reordering.
Another issue discussed by Forni and Gambetti (2014) is the presence of sufficient information to ensure
that the conditions of existence and uniqueness are ensure for identification purposes. Specifically
Forni and Gambetti (2014) present necessary and sufficient conditions under which a VAR contains
sufficient information to estimate the structural shocks. Based on this theoretical result the authors
propose a test statistic for detecting information deficiency and a procedure to amend a deficient VAR
model. A relevant example is the bivariate VAR with unemployment and labour productivity which
is found to be informentionally deficient. The building block of this framework is a representation
of the economy where q mutually orthogonal structural shocks affect macro variables through square-
summable impulse-response functions (see, also Kocikecki and Kolasa (2018, 2023)).
where ut is a q−dimensional white noise vector of structural macroeconomic shocks and F(L) is an
(n × q) matrix of square-summable linear filters in the non-negative powers of the lag operator L. The
above formulation also corresponds o the steady-state equilibrium of a DSGE model.
Assumption 2 (Information Set). We shall say that the information set Xt ∗ is given by the closed linear
space spanned by present and past values of the variables in xt∗ , where
In practice the number of observable variables n is very large, so that the econometrician needs to reduce
it in order to estimate a VAR. The VAR information set is then spanned by an s−dimensional sub-vector
of xt∗ , or, more generally, an s−dimensional linear combination of xt∗ , say zt∗ = Cxt∗ .
9
2. Linear Vector Autoregressions
Example 1. Consider the model yt = Ayt−1 + et , which implies that (see, Lai and Wei (1982))
n
yt = An y0 + ∑ An− j e j , t = 1, ..., n. (2.1)
j=1
!
α1 ... αk−1 αk
A= (2.2)
I k−1 0
which is equivalent to the matrix coefficient of a VAR(1) process. Moreover, expressing the matrix A
into its Jordan form we obtain A = CDC−1 , where D = diag D1 , ..., Dq
zj 1 0 ... 0
0 zj 1 ... 0
Dj =
.. .. . . . . . .
(2.3)
. . 0
0 0 ... ... zj
(m j ×m j )
with z j being the root of ϕ (z) with multiplicity m j and C a nonsingular matrix. Moreover, let M =
max j m j , then it holds that
q
n n −1 −1
kA k = CD C ≤ kCk C ∑ Dnj = O(nM−1 ). (2.4)
j=1
In particular, when the roots of the characteristic polynomial lie strictly inside the unit circle, then with
the independent white noise et , the system is stable. Moreover, by allowing the roots of the characteristic
polynomial to lie on the unit circle, we can formulate unstable but non-explosive systems related to the
ARIMA models. Notice the analysis of time series regression models includes unit-roots and partially
nonstationary (cointegrated) models.
An important property of I(1) variables is that there can be linear combinations of these variables that are
I(0). If this is so then these variables are said to be cointegrated. Notice that econometric cointegration
analysis can be used to overcome difficulties associated with stochastic trends in time series and applied
to test whether there exist combinations of non-stationary series that are themselfs stationary.
Definition 2. Suppose that yt is I(1). Then yt is cointegrated if there exists an N × r matrix β , of full
column rank and where 0 < r < N, such that the r linear combinations, β ′ yt = ut , are I(0). We say that
the dimension r is the cointegration rank and the columns of β are the cointegrating vectors.
10
• Testing for cointegrating relations that economic theory predicts should exist, implies that the
null hypothesis of noncointegration is not rejected. However, under the presence of breaks there
is a need for implementing tests of the null hypothesis of non-cointegration, against alternatives
allowing cointegrating relations subject to breaks.
• Relevant asymptotic theory results on the efficient estimation of the parameter path in unstable
time series models can be found in the study of Müller and Petalas (2010). Further issues include
the use of exogenous regressors (e.g., see Pesaran et al. (2000)) in VECM representations.
A VAR has several equivalent representations that are valuable for understanding the interactions be-
tween exogeneity, cointegration and economic policy analysis. To start, the levels form of the s−th
order Gaussian VAR for x is
s
xt = Kqt + ∑ A j xt− j + εt , εt ∼ N (0, Σ) . (2.5)
j=1
Example 2. Let xt be an I(1) vector of n components, each with possibly deterministic trend in mean.
Suppose that the system can be written as a finite-order vector autoregression:
Therefore, we get the following system equation representation π (L)xt = µ + ε t , t = 1, ..., T , where
π (L) = (1 − L)I n − ∑k−1 i k
i=1 Γi (1 − L)L − π L , and Γi = −I n + π 1 + π 2 + ... + π i , for i = 1, ..., k.
11
Remark 2. Within a cointegrated analysis context, it has been proved to be advantageous for both the-
oretical and practical purposes to separate the long-run behaviour of the system from the more transient
dynamics by using the error correction form of the model which is useful when measuring the impact
of structural shocks to economic outcomes such as monetary policy, fiscal policy and macro aggregates.
Thus when employing vector autoregressions to model the comovements of macroeconomic aggregates
time series stationarity is assumed (e.g., by using a first-difference transformations). Moreover, cointe-
gration dynamics can be correctly specified using a Cointegrated VAR representation. Our main interest
is the structural analysis of VAR models under time series stationarity as we explain below.
Example 3. Consider the following data generating process as below
k−1
∆Xt = αβ ′ Xt−1 + ∑ Γi ∆Xt−i + εt , for t = 1, ..., T, (2.7)
i=1
where {εt } is i.i.d with mean zero and full-rank covariance matrix Ω, and where the initial values
X1−k , ..., X0 are fixed. We are interested in the null hypothesis H0 : β = β0 . Thus, when β is a known
(p × r) matrix of full column rank β0 , the subspace spanned by β and β0 are identical.
Example 4 (Wage formation with Cointegrated VAR, see Petursson and Slok (2001)). A Gaussian
VAR(k) model is used can be rewritten in the usual error correction form in terms of stationary variables
k−1
∆xt = ∑ Γ j ∆xt− j + αβ ⊤xt−1 + Φ∆ + εt , t = 1, ..., T (2.8)
j=1
Using the cointegrated VAR model we can characterize the cointegrating relations based on a model
of wage bargaining between trade unions and firms. Thus the economic theory implies that: (i) the
marginal productivity condition for labour (downward sloping demand curve of firms product), and
(ii) real wage relation derived from the bargaining between trade unions and firms over wage. In other
words, two stationary combinations of the non-stationary data exist and that these can be identified based
on the prior knowledge of these two economic relations. The associated imposed parameter restrictions
are not rejected by the data and simplified analysis is achieved via the partial VAR formulation which
corresponds to the conditional system for ∆x1t given the past and ∆x2t .
! !
Ir 0 α11 α12
Remark 3. Let α and β be defined as below β = and α = , then the
−β2 Ik−r 0 α2
cointegration matrix becomes as below
!
α11 α12
Π = βα , (2.9)
−β2 α11 −β2 α12 + α22
where all the above parameter coefficients are unrestricted. Moreover, according to Kleibergen and Van Dijk
(1994) the behaviour of the likelihood is due to the nonidentifiedness of certain parameters, which occurs
when the model is a difference stationary one. In other words, flat priors are informative in cointegration
models because difference stationary models are infinitely favoured.
12
Example 5. Imposing long-run restrictions is a commonly used approach for the identification of struc-
tural shocks. However, usually identification assumptions for cointegrating vectors can be complicated,
however following the approach of Zha (1999), which applies block recursive assumptions to the struc-
tural VECM with long-run restrictions. The block recursive system (see, also Keweloh et al. (2023)) is
well developed in structural VAR and VECM models with short-run restrictions but not so developed in
the structural VECM with long-run restrictions (see, Hecq et al. (2000)).
Consider the structural VAR model as below
p
B(L)xt = µ + ut , B(L) = B0 − ∑ B j L j
j=1
Suppose that xt is a vector of cointegrated time series where ut is an (n × 1) vector of serially uncorre-
lated structural disturbances with a mean zero and a covariance matrix Σu and εt is an (n × 1) vector of
serially uncorrelated linear forecast errors with a mean of zero and a covariance matrix Σε .
In addition, assume that all variables are cointegrated such that there is a reduced number of com-
mon stochastic trends (k = n − r) and a number of transitory components r. Moreover, these common
trends are assumed to be generated by permanent shocks such that the vector of error terms ut is de-
composed into ut ≡ (u′kt , u′rt )′ , where ukt is a k−dimensional vector of permanent shocks and urt is an
r−dimensional vector of transitory shocks.
• Empirical evidence: The authors investigate the effects of contractionary shock to the monetary
policy on economic variables in a seven-variable VECM with long-run restrictions. In particular,
the permanent shocks include a shock that affects the long-run level of real exchange rates (a
real-exchange-rate shock) and a US monetary policy shock that affects the long-run level of US
prices. Thus a Japanese monetary policy shock can be considered as a transitory shock since the
model does not include the Japanese price, while it can be considered as a permanent shock it it
affects the long-run level of Japanese interest rates.
• Theoretical Evidence: When the monetary policy and exchange rates are simultaneously deter-
mined contemporaneously, then the recursive ordering assumptions are not satisfied. On the other
hand, identification with long-run restrictions does not suffer from this simultaneity problem be-
cause no zero restrictions are imposed on the structural parameter B0 .
13
Example 6 (The causal effects of fiscal policy shocks, see Caldara and Kamps (2017)). More recently
attention has been paid in the role of fiscal policy for stabilizing business cycles. Since empirical studies
have not reached a consensus about the effects of fiscal policy on macroeconomic variables to assess
the effects of fiscal policy the SVAR methodology is commonly used (see also Boiciuc (2015)). The
structural analysis of VARs allows to estimate the effects of fiscal policy shocks on economic activity.
Consider the structural representation of a VAR model is given by
To estimate the SVAR the reduced form is given by xt = C(L)xt−1 + ut where ut = A−10 Bεt . The relation
between structural shocks and reduced form shocks is A0 ut = Bεt . Recall that a SVAR(p) model with
the following representation
Example 7. Consider the study of Bilgili (2012) who attempts to reveal explicitly whether or not
biomass consumption can mitigate carbon dioxide (CO2 ) emissions. In other words, in order to cor-
rectly capture the underline features in the data (and produce unbiased and efficient estimators), a coin-
tegrating regression specification with regime shifts (structural breaks) are essential to understand the
long-run equilibrium of CO2 emissions with biomass consumption as well as fossil fuel consumption.
Their main findings include the presence of a statistical positive impact of fuel’s consumption and a sta-
tistical negative impact of biomass consumption on CO2 emissions. A climate-economic related event
is a change in the policy of a Country’s Energy Authority by the introduction of a policy act to introduce
measures for diminishing CO2 emissions (e.g., an increase in biomass consumption). In particular, this
can be detected in the data by identifying (dating) the presence of a regime shift using the cointegration
model with structural breaks. Furthermore, a statistical negative impact on CO2 emissions (or equiva-
lently a statistical positive impact in CO2 emissions reductions), might be expected to increase/decrease
through possible government incentives for research and development on biomass plants (assuming that
the magnitudes of other parameters, such as population growth and growth in demand for energy, will
not increase beyond the expectations) (see, also Phillips et al. (2020)).
14
Example 8 (see, Moon and Schorfheide (2002)). Suppose that φ = 1 and µ = 0. Define with y1,t = Ct
and y2t = [Wt , It ]. According to the permanent-income model all three variables are integrated of order
one I(1) which implies the following cointegration regression model
" # " #
y1,t A′
= y + ut , (2.11)
y2t I 2 2t−1
In particular, the distribution theory of estimators of the unrestricted cointegration vector A is well-
developed in the literature which typically have a T −convergence rate. Moreover, the MLE and FM-
OLS estimators of Phillips (1991) and Phillips and Hansen (1990) respectively have a mixed-Gaussian
limit distribution with a random covariance matrix. In addition, several studies concerning the estima-
tion of the restricted cointegration vectors are also presented in the literature. In particular, Saikkonen
(1995) extends the analysis for the estimation of cointegration vectors with linear restrictions to the
case in which the restriction function is nonlinear and twice differentiable. More precisely, he provides
stochastic equicontinuity conditions to make the conventional Taylor approximation approach valid.
Even if the income process is stationary such that 0 ≤ φ < 1 and µ > 0, both consumption and wealth are
I(1) processes under the optimal consumption choice. Therefore, the optimal decision rule creates re-
strictions between parameters that are associated with long-run relationships and parameters that control
the short-run dynamics (see, also Blanchard and Quah (1988)). Define with y1,t = Ct , y2,t = [∆Wt , It ]′ ,
x1,t = Wt−1 , x2,t = [1, It−1]′ , yt = [y1,t , y2,t ]′ , xt = [x1,t , x2,t ]′ . Thus, the consumption model is nested in
the following general specification
" # " #" # " #
y1,t A′11 A′21 x1,t u1,t
= ′ + , (2.12)
y2t 0 A22 x2,t u2,t
Define with ai j = vec(Ai j ). Then the unrestricted parameter vector a = [a′11 , a′21 , a′22 ]′ and b = [r, µ , φ ]′
is composed of the structural parameters. Assume that the partial sum process of ∆Wt converges to a
vector Brownian motion such that
⌊Tr⌋
1
√
T
∑ ∆Wt ⇒ B(r) ≡ BM(Ω), (2.13)
t=1
3 Relatedasymptotic theory with examples can be found in Katsouris (2023b) who present unit-root dynamics and weak
convergence arguments to a suitable topological space. Moreover, asymptotic theory results for a cointegrating predictive
regression model with near units roots when modelling the term structure of interest rates is presented by Lanne (2000).
Notice that Moon and Schorfheide (2002) show the consistency of the estimator using a Skorohod representation of the
weakly converging objective function and derive the limit distribution of the MD estimator for smooth restriction functions.
15
Example 9 (A simple climate-economic system, see Pretis (2021)). Climate and economic variables
are observed over time and space. Denote with yt = y′1t , y′2t ) denotes the relevant climate and socio-
economic variables. Denote with Y ji = (yi , ..., y j ) for i ≤ j, such that YT1 = (y1 , ..., yT ). Then, the model
of interest can be characterized as
T
fY YT1 |Y0 , θ = ∏ fy yt |Yt−1 , θ , θ ∈ Θ ⊂ Rn , (2.15)
t=1
where fy yt |Yt−1 , θ denoting the sequentially-conditioned, joint-density for yt , with (n × 1) parameter
vector θ lying in parameter space Θ. A vector autoregression process within a cointegration framework
due to the fact that economic and climate time-series are pre-dominantly non-stationary time series due
to the presence of stochastic trends and structural breaks. Therefore, climate-economic systems can
be well-approximated by cointegrated econometric models, although in addition we are interested to
measure the weather shocks into the macro-economy (see, Pretis (2020, 2021)).
Example 10. Consider the Energy Balance System studied by Pretis (2020) (see, also Carrion-i Silvestre and Kim
(2021) and Bruns et al. (2020)), such that the law of motion is captured by the following system of dif-
ferential equations
dTm
Cm = −λ Tm + F − γ (Tm − Td ) (2.16)
dt
dTd
Cd = γ (Tm − Td ) (2.17)
dt
The feedback parameter λ is of particular interest as it determines the equilibrium response of surface
temperatures Tm to a change in the forcing F (e.g., from increased CO2 concentrations). Furthermore,
focusing on a cointegration analysis which allows to examine whether there exists a combinations of
time-series that are themselves stationary, it can be shown that the system of differential equations of the
two-component EBM is equivalent to a cointegrated systen with restrictions on the parameters. Thus,
a climate-economic system is formulated as a cointegrated vector-autoregression, where the regressand
is decomposed into two variables, that is, yt = (et , ct ), such that et represents a univariate economic
variable and ct represents a univariate climate variable.
s
yt = ∑ A j yt− j + µ + εt , εt ∼ N (0, Σ) (2.18)
j=1
where yt = (et , ct )′ , εt = (εe,t , εc,t ) and ∆yt = yt − yt−1 . Thus, the full system can be written as
" # " # " # " #" # " # " #
∆et α1 h i e
t−1 Γ11 Γ12 ∆et−1 µe εe,t
= β1 β2 + + + (2.20)
∆ct α2 ct−1 Γ21 Γ22 ∆ct−1 µc εc,t
A relevant discussion on the interaction between unit roots and exogeneity can be found in Hendry
(1994) (see, also Engle et al. (1983)).
16
Therefore, the above economic and climate variables approximate the full climate-economic system
with the links between climate and the economy given by both the short-run parameters Γ and the
equilibrium relationship ht given by the cointegrating vector β ′ yt such that:
" #
h i e h i
t
ht = β1 β2 = β1 et β2 ct . (2.21)
ct
The cointegrating relation is an equilibrium one, and does not necessarily reflect purely a climate-impact
function, but rather an equilibrium between the two series, to which each series adjusts. The statistical
properties of the two-stage least squares estimator under cointegration can be found in Hsiao (1997).
Consider a general autoregressive distributed lag ARDL (p,q) model where a series, yt , is a function of
a constant term, α 0 , past values of itself stretching back pperiods, contemporaneous and lagged values
of an independent variable, xt , of lag order q, and indepenendent, identically distributed error term:
p q
yt = α 0 + ∑ α i yt−i + ∑ β j xt− j + ε t , (2.22)
i=1 j=0
Example 11. A commonly used model is the ARDL (1,1) model given by
17
Consider a demand equation system in which qt is a measure of the quantity of oil purchased, pt is a
measure of the real price of oil, and yt is a measure of the real income such that
Moreover, the demand structural system describes the behaviour of oil producers and the determinants
of income such that (the order matters)
′ , y′ , ...., y′
′
where xt−1 := 1, yt−1 t−2 t−p is a vector consisting of a constant term and p lags of each of
the three variables with yt = (qt , yt , pt )′ .
Ayt = Bxt−1 + ut ,
1 −δ −α qt b′s uts
y
−ε 1 −β yt = b′y xt−1 + ut .
1 −ζ −γ pt b′d utd
| {z }
A
We assume that these structural shocks have mean zero and are serially uncorrelated as well as uncorre-
lated with each other such that
′ D, for t = s,
E ut ut =
0, for t 6= s.
where D is a diagonal covariance matrix. Then, by pre-multiplying the structural VAR representation
with the matrix A−1 , we obtain the corresponding reduced-form VAR equation which has the follow-
ing dynamic structural model form yt = Πxt + εt , E (εt εt′ ) = Ω. Therefore the above reduced form
specification can be expressed as below
18
Then, the above reduced-form can be employed to construct impulse-response functions Ψs = ∂ yt+s /∂ εt′ .
Ψ̂s = Φ̂1 Ψ̂s−1 + Φ̂2 Ψ̂s−2 + ... + Φ̂m Ψ̂s−m , for s = 1, 2, 3, ...
Remark 4. Following Kilian and Lütkepohl (2017) we assume that the short-run income and price elas-
ticities of supply as well as the contemporaneous coefficient relating oil prices to economic activity are
all zero, although such conditions might need to be revisited especially within a climate-economy equi-
librium setting without resorting to non-equilibrium dynamic states. Nevertheless, using the Cholesky
factorization4 of ΩMLE we can obtain the estimated effects on yt+s of a one-standard-deviation increase
in one of the structural shocks at date t (see, also Hamilton (1994)).
where η and A1 , A2 , ..., A p , are a constant vector and constant matrices, respectively.
which is considered to be a stationary process provided all roots of det [A(z)] = 0, lie outside the unit
circle. Then, the process admits a causal VMA(∞) respresentation such that
∞
yt = ω + ∑ Ck ε t−k . (2.30)
k=0
Although in Section 2. we consider suitable formulations of cointegrated VAR models and their ap-
plications from related economic studies, one needs to revise the concepts presented in Chapter 10 of
Hamilton (1994) on Covariance-Stationary processes before proceeding to the material of Section 3.
Open Problems Various open problems remain in the literature such as the use of the cross-section as a
mechanism for beliefs elicitation and identification of structural shocks. Related studies from the finance
literature where the cross-section is used for improving the predictability of portfolio returns include
among others Chinco et al. (2019), Freyberger et al. (2020) and Kyle et al. (2023), but in the case of
modelling macroeconomic fundamentals using cross-sectional information is still a growing literature.
Relevant studies in this direction include Forni et al. (2017), Kong et al. (2019), Yamamoto and Horie
(2023) and Hannadige et al. (2023) among others.
4 Using a Cholesky decomposition is a standard approach in the literature and several studies employ such techniques
although alternative parametrizations are of particular interest within a unified framework.
19
3. Vector Autoregressions: Prediction and Granger Causality
where yt ∈ Rd is a d−dimensional vector. Then, the above VAR(p) process can be expressed as a
VAR(1) process using a companion matrix as below
A1 A2 A3 . . . Ap εt
yt I yt−1
d 0 0 ... 0 0
yt−1 y
. = 0 Id 0 . . . 0 t−2
. . .. . . .. .. + 0 (3.2)
. .. . . ... . . ..
.
yt−p+1 .. yt−p
0 0 . Id 0 0
such that Y t = A Y t−1 +U t , which implies that a unique solution can be determined using
t
Y t = A t Y 0 + ∑ A t− jU j . (3.3)
j=1
Assumption 3. (a) any non-zero eigenvalue λ of A satisfies the following stability condition
!
p
1
det I d − ∑ j A j = 0. (3.4)
j=1 λ
(b) The spectral radius of the companion matrix A is smaller than 1, ρ (A ) < 1.
Lemma 3. Consider the homoscedastic VAR(p) process defined in (3.1), where the roots of the charac-
teristic polynomial satisfy the conditions of Assumption 3.
(i). Consider the process yt ∼ VAR(p), and suppose that the spectral radius of the companion matrix
A satisfies ρ (A ) < 1. Then, since Y 0 = ∑∞j=0 A jU t− j almost surely then under the covariance
′
stationarity assumption it holds that Γy ( j) = E Y t Y t− j .
(ii). An equivalent expression for the VAR(p) process using the lag operator is as below
I d − A1 L − A2 L2 − ... − A p L p yt = ε t . (3.5)
20
Consequently, the following are equivalent: (i) yt is a stable process; (ii) ρ (A ) < 1; (iii) all the
roots of the characteristic polynomial det {Φ(z)} = 0 lie outside the unit circle {z ∈ C : |z| < 1}
such that det {Φ(z)} = 0 iff |z| > 1.
(iii). Hence, based on the results given in (i) and (ii) above, the VAR(p) process admits a linear process
representation such that
∞
i.i.d
yt = ∑ C j ε t− j , almost surely, with ε t ∼ (0, Σε ), (3.7)
j=0
for a sequence of C j j≥0
satisfying the summability condition ∑∞j=0 C j < ∞.
Consider that these interelated equations can be written in the following form:
Therefore, by pre-multiplying the above dependent variable with B0 we can obtain a VAR representation
given by the following expression:
In other words, we can view the above representation as a special case of the Dynamic Structural system
equation since we eliminate the interelations of the dependent variable to a reduced form of a VAR.
!
z1t
Example 14. Let zt = be an n−dimensional vector stochastic process, where z1t is an (n1 × 1)
z2t
and z2t is an (n2 × 1) such that n = n1 + n2 . Assume a linear dynamic model (see, Lastrapes (2005))
zt = A−1 −1 −1
0 A1 zt−1 + ... + A0 A p zt−p + A0 ut
zt = B1 zt−1 + ... + B p zt−p + ε t .
such that the structural errors have a variance-covariance matrix E (ε t ε t′ ) = Ω. Therefore, the MA
representation of the structural model is
zt = A0 − A1 L − ... − A p L p ut ≡ D0 + D1 L + D2 L2 + ... ut = D(L)ut . (3.11)
21
Similarly, the reduced form MA has the following representation
zt = I − B1 L − ... − B p L p ε t ≡ I +C1 L +C 2 L2 + ... ut = C(L)ε t . (3.12)
In other words, the parameters of interest are the structural dynamic multiplies or impulse response
∂z
functions such that ∂ t+k
ut = Dk . The common practice in the applied macroeconometrics literature is
to obtain estimates for B(L) and Ω and then impose suitable restrictions on the underline structure to
identify the model parameters.
Remark 6. Notice that while reduced-form parameters A j for j = 1, ..., p and the covariance matrix of
residuals Σu can be estimated consistently, the structural parameters collected in the matrix B are not
identified without further theoretical or statistical assumptions. One possible solution is to restrict the
matrix B to be a lower-triangular imposing a recursive causal structure among the model variables.
xt = C(L)ut , (3.13)
∞
where C(L) = ∑ C j L j is an one-sided polynomial in the lag operator L in infinite order. These shocks
j=0
are orthogonal white noises such that ut ∼ (0, Σu ), where Σu is diagonal. Furthermore, notice that a VAR
model based on a subset of the variables implies a different MA(∞) representation than its counterpart
based on the joint distribution. We begin by considering various examples regarding the identification
and estimation of SVAR models using properties of structural shocks.
Definition 4 (Fundamentalness in Systems). Given a covariance stationary vector process xt , the repre-
sentation xt = C(L)ut is fundemental if
(ii). C(L) has no poles of modules less or equal than unity, i.e., no poles inside the unit disc.
(iii). det[C(z)] has no roots of modules less than unity, i.e., all its roots are outise the unit disc
If the roots of det[C(z)] are outside the unit disc, we have invertability in the past, that is, the inverse
representation depends only on non-negative powers of L, and we have fundamentalness. If at least one
of the roots of det[C(z)] is inside the unit disc, then invertability and non-fundementalness holds.
22
Example 15 (Partially nonstationary multivariate autoregressive AR(p)). Consider the partially nonsta-
tionary multivariate autoregressive AR(p) given by
!
p
Φ(L)Y t ≡ Im − ∑ Φ jL j Y t = εt (3.15)
j=1
where {Y t } is an m−dimensional process and for ε t it holds that E(ε t ) = 0 and cov(ε t ) = Ω. Moreover,
it is assumed that det Φ(L) = 0 has d < m unit roots and the remaining roots are outside the unit circle
and that rank Φ(1) = r, where r = (m − d) > 0 (see, Samuelson (1941) and Dickey et al. (1986)).
Therefore, each component of the first difference Wt = (Yt −Yt−1 ) is assumed to be stationary.
Then the model has the following error-correction form below
p−1
∗
Φ (L)(1 − L)Y t = CY t−1 + ε t , ∆Y t = CY t−1 + ∑ Φ∗ ∆Y t− j + ε t (3.16)
j=1
p
Moreover, the Jordan Canonical form of ∑ j=1 Φ j is given by
!
p
P−1 ∑ Φj P = diag I d , Λr (3.18)
j=1
One can define with Zt = [Z ′1t , Z ′2t ] = QY t , where Z 1t = Q′1Y t and Z 2t = Q′2Y t .
The identification of structural vector autoregressions is a crucial step before proceeding with construct-
ing test statistics and forecasts (see Section 4.). A suitable approach found in the literature, is the
method of recursive identification5 (see, Chen and Zhao (2014)) which implies exact identification of
the structural parameter B−1 0 (e.g., see Kim and Roubini (2000)). Moreover, the presence of stochastic
singularity are crucial on whether a system is identified or unidentifiable. Based on these considerations
Komunjer and Ng (2011) employ restrictions implied by observational equivalence to establish condi-
tions for avoiding the presence of stochastic singularity. To ensure that these structural econometric
models can be still partially identified (see, also Phillips (1989)) regardless of the possible presence of
stochastic singularity, then the econometrician can exploit the existence and uniqueness of local identi-
fication of DSGE models based on the corresponding linearized solution. Lastly, given the prominent
role of expectations in news aspects such as the information flow and anticipated shocks can affect the
identification of structural shocks (e.g., optimal inter-temporal decisions discount future tax obligations,
see Leeper et al. (2013) and Mertens and Ravn (2013)).
5 The theoretical analysis of recursive identification methods is given by Soderstrom et al. (1978) and Chen (2010)
23
3.2.2. Weak Exogeneity in I(2) VAR Systems
The notion of weak exogeneity is important when considering the structural analysis of cointegrating
regression models. Under the assumption of weak exogeneity estimation of the cointegration parameters
in conditional models is accordingly influenced. In particular, for the VAR model allowing for I(1) vari-
ables under the assumption of Gaussian errors implies the use of a cointegrated VAR specification which
is suitable for modelling long-run equilibrium dynamics for multivariate time series. Within this stream
of literature one is interested to analyze the conditions under which a subset of equations is weakly
exogenous with respect to the cointegration parameters (see, Paruolo (1997, 2000), Paruolo and Rahbek
(1999), Dolado (1992) and Pesaran et al. (2000)). Moreover, Tchatoka and Dufour (2013) propose uni-
fied exogeneity test statistics and examine the pivotality property under strict exogeneity (see, also
White and Pettenuzzo (2014)) and characterize the finite-sample distributions of the statistics under H0 ,
including when identification is weak and the errors are possibly non-Gaussian. Assume that the vectors
U t = [ut ,V t ]′ for t = 1, ..., T , have the same nonsingular covariance matrix defined below
" #
σ 2 δ′
E U t U t′ = Σ = u′ > 0, t = 1, ..., T, (3.19)
δ ΣV
where ΣV has dimension G (see, also Khalaf and Urga (2014)). Then, the covariance matrix of the
reduced-form disturbances W t = [vt ,V t′ ]′ = [u +V β ,V ] takes the following form
" #
σu2 + β ′ ΣV β + 2β ′ δ β ′ ΣV + δ ′
Ω := (3.20)
ΣV β + δ ΣV
The matrix P is the Cholesky factor of Ω−1 , so P is the unique lower triangular matrix.
" #
P11 0
P= (3.22)
P21 P22
24
3.3. Local Projections
This example is based on the study of Mertens and Ravn (2010) who consider the consequences of
anticipation effects for VAR-based estimates of the impact of government spending shocks. We consider
a bivariate time series representation of a vector consisting of control variables zt , and government
spending gt . We assume transitory government spending shocks, such that µg (L) is a stable polynomial.
The identification is facilitated by assuming that information can arrive with any anticipation horizon
between 1 and q periods. Thus, the vector of observables vt = [gt , zt ]′ is formulated as below
" #" # " #
0 1 − µg (L) kt 1 Lq
vt = + Σe et , (3.25)
φzk φzk gt 0 φz,1 Θ(L)
" # " #
q−1 q−2 q−2 q−1 1 0 eg0,t
Θ(L) = ω +ω L + ... + ω L +L , Σe ≡ σg , et ≡ g (3.26)
0 λ eq,t
Example 17. Consider that the vector yt = [GOVt , GDPt ,CONt ], where all three macroeconomic vari-
ables are in real terms and in logarithms. Then, the VECM is given by the following expression
where D = Y (0)Σe B(0) and ε t B(L)−1 et , where et contains the structural shocks of interest. Notice that
despite the presence of permanent fiscal shocks, the variables in yt cointegrate since the investment-
output ratio is unaffected by the level of government spending in the long-run. Moreover, the unantic-
ipated government spending shock is allowed to affect the level of government spending immediately,
while the anticipated government spending shock is assumed not to affect government spending within
one quarter. However, the two shocks are restricted to have the same long run impact on the level of gov-
ernment spending (see, Mertens and Ravn (2010)). Overall, the estimation procedure aims to uncover
the response to an anticipated fiscal shock. On the other hand, when the anticipation rate is high which
implies that the anticipated shocks are relatively important, then biased estimates for the unanticipated
shock are obtained in small samples. Further examples can be found in Mertens and Ravn (2013).
25
Two popular methods to estimate impulse responses when a shock has already been identified include:
local projections and distributed lag models. In particular, one can show that the dynamic response
from a VAR with the shock embedded as an endogenous variable is equivalent to that of a VAR with
the shock included as an exogenous variable only when that shock has no serial correlation (e.g., see
Plagborg-Møller and Wolf (2021), Montiel Olea et al. (2021)). The result follows because a VAR with
a shock as an exogenous variable (VAR-X) can be seen as a multivariate generalization of a DLM. A
comparison of these cases can be found in Alloza et al. (2020). Moreover, Lusompa (2023) propose
an efficient estimation method for Local projections when residuals are autocorrelated. Lastly, the
simultaneous impact of monetary and fiscal policy is discussed by Bruneau and De Bandt (2003) while
the relation of fiscal policy and persistent inflation is studied by Bianchi et al. (2023).
Example 18 (see, Alloza et al. (2020)). Consider the following data generating mechanism:
Then, the process can be formulated as a SVAR of the form A0Y t = BY t−1 such that
" #" # " #" # " #
1 0 xt γ 0 xt−1 εx
= + ty (3.30)
−δ0 1 yt δ1 ρ yt−1 εt
zt = µ z + Γε t + Σ1/2 η t (3.32)
Without further restrictions, the augmented SVAR model is only identified up to orthogonal rotations
of the form B = B̃Q. According to Alloza et al. (2020) the presence of serial correlation can be thought
as cross-sectional persistence as in the case of narrative shocks, which are usually serially correlated.
This aspect can directly affect the identification and estimation of their dynamic effects. In particular,
when estimating the dynamic response of some variable to a serially correlated shock, some part of
the persistence may be spillover on the impulse response function. We discuss the use of a suitable
parametrization for modelling persistence in macroeconomic data in Section 4.6..
Recently, it has been shown that the method of local projections has advantages in drawing inference
of impulse responses, especially with respect to uniform validity (e.g., see Inoue and Kilian (2020)).
We demonstrate an example of the estimation method from the framework of Xu (2023). The VAR(p)
model is of the form yt = A1 yt−1 + ... + A pyt−p + ut , t = 1, ..., n, where ut is serially uncorrelated shock.
26
To measure the responses of the first endogenous variable y1t to shocks after h propagation periods,
where h ≥ 1, the LP method runs the regression below (see, Xu (2023)).
p−1
y1,t+h = β1 (h)yt + ∑ θ1 j (h)′yt− j + ξ1t (h) (3.33)
j=1
p−1
y2,t+h = β2 (h)yt + ∑ θ2 j (h)′yt− j + ξ2t (h) (3.34)
j=1
.. .
. = .. (3.35)
p−1
yk,t+h = βk (h)yt + ∑ θk j (h)′yt− j + ξkt (h) (3.36)
j=1
Notice that by definition it holds that θ1,p (h) = 0. Recall that for a matrix x, we denote its Frobe-
nius norm, such that |x| = [trace(x′ x)]1/2 . Suppose that {b
ut (h) : t = p∗ , ..., n − h} are OLS residuals of
the VAR(p∗ − 1) regression, using the data {yt : t = p∗ , ..., n − h}. In addition, related distributional
assumptions for the error term of the model can be imposed.
The OLS estimator is given by
!−1 !
n−h n−h
βb1LA (h) − β1 (h) = ∑∗ ubt (h)but (h)′ ∑ ubt (h)ξ1t (h) . (3.41)
t=p t=p∗
Notice that these local projection estimates correspond to an in-sample estimates for the VAR(p) model
based on p lags. On the other hand, if we are interested to obtain out-of-sample forecast sequences
then we need to implement a forecasting scheme such as a moving window which is used to obtain
pseudo-out-of-sample forecasts. The window size should be larger than the number of lags to have
good statistical properties for the forecast sequences. Relevant applications of local projections include
the construction of impulse response functions.
27
3.4. Impulse Response Functions
′
Moreover, let g1 (β ), ..., gm(β ) be a continuously differentiable function with values in m−dimensional
∂ gi ∂ gi
Euclidean space and ′ = , for i ∈ {1, ..., m}. Then, it holds that (see, Lütkepohl (1990))
∂β ∂βj
√
d ∂g ∂ g′
T g(β̂ ) − g(β ) → N 0, ′ Σβ . (3.43)
∂β ∂β
and the ut are independent, identically distributed (i.i.d) with bounded forth moments. This implies that
the usual OLS estimators have asymptotic covariance matrix given by
n ′ o
−1
Σa = Γ ⊗ Σu , Γ = E yt yt−1 ... yt−p+1 ⊗ yt yt−1 ... yt−p+1 (3.45)
In addition, if yt is Gaussian, then α̂ and σ̂ are asymptotically independent. Interval forecasts based on
predictive distributions are studied by Chatfield (1993) while the reliability of local projection estimators
of impulse response functions is examined by Kilian and Kim (2011). The asymptotic distributions of
IRFs and FEVD for VARs are established by Lütkepohl (1990) (see, also Lanne and Nyberg (2016)).
28
Remark 8. Regarding the construction of Impulse Response Functions and FVED, it can be shown that
Impulse Responses that are obtained from unrestricted VARs with roots near unity have long period
estimated impulse responses that are inconsistent. Moreover, FEVD are also inconsistent in unrestricted
VAR models with near unit roots. The inconsistency problem of IRFs from nonstationary SVAR mod-
els is discussed by Phillips (1998) (see, also Lütkepohl and Saikkonen (1997)). Specifically, the OLS
estimation approach of Phillips (1998) shows that in nonstationary VAR models with some roots at or
near unity the estimated impulse response matrices are inconsistent at long horizons and tend to random
matrices rather than the true impulse responses. In this stream of literature another issue is the consistent
estimation of the number of cointegrating relations in a system of equations (e.g., see, Barigozzi et al.
(2021b) and Barigozzi and Trapani (2022)).
The residuals of ut are the one-step ahead forecast errors associated with the VECM representation.
However, tracing the marginal effects of a change in one component of ut through the system may not
reflect the actual responses of the variables since in practice an isolated change in a single component
of ut is likely to occur if the component is correlated with the other components.
Remark 9. Identifying restrictions are imposed in order to identify the transitory and permanent shocks
in the system. For example, Castelnuovo and Surico (2010) present impulse response analysis evidence
that the VAR model is robust to two different identification strategies based on zero restrictions and
the sign restrictions when evaluating the response of prices on monetary policy shocks. Additional to
an impulse response analysis as a mechanism for studying the impact of structural shocks, the VAR
framework is often used for forecasting based on suitable forecasting schemes (e.g., see, Ng and Young
(1990), Bhansali (2002) and Mariano (2002)). Two commonly used forecasting approaches:
(i). Direct Multiperiod Forecasting (e.g., see Kang (2003), Greenaway-McGrevy (2013, 2020)).
yt+h = yt + β ′ Xt + εt (3.47)
Suppose that the we have an auxiliary equation for the Xt such that Xt = AXt−1 + ut . Then, estimating
the vector autoregression for the regressors by OLS we obtain an estimate  and thus a forecast of
Xt+ j is  j Xt and a forecast of the quantity (yt+h − yt ) is given by ∑h−1
j=0 Â j Xt . Moreover, an example, of
constructing one-step ahead forecasted values based on information from the cros-section is presented
by Katsouris (2021). An interesting extension would be to use such cross-sectional shrinkage approach
within the aforementioned econometric environment.
29
3.4.2. Forecast Error Variance Decomposition
Generally, the Forecast Error Variance Decomposition gives the proportional contribution of each of the
structural shocks to the forecast error variance of each variable at different horizons. Recall that the
h−period forecast error of yt is given by
Hence, the h−period ahead forecast error of the j−th component of yt is given by
h−1 K
∑ θ j1,i w1,t+h−i + ... + θ jK,i wK,t+h−i = ∑ θ jk,0 wk,t+h + ... + θ jk,h−1 wk,t+1 .
i=0 k=0
Moreover, the variance of the h−period forecast error of the j−th component of yt is given by
" 2 #
K K
2 2
E ∑ θ jk,0 wk,t+h + ... + θ jk,h−1 wk,t+1 ≡ ∑ θ jk,0 + ... + θ jk,h−1 .
k=1 k=1
n o
which holds since Σw = IK . Notice that the sum of the terms θ 2jk,0 + ... + θ 2jk,h−1 , is thus the k−th
shock to the forecast error variance of the h−period forecast of the j−th variable. Then, we have that
As an example of application of the above metrics see Kilian and Park (2009) (and Kilian (2009)) who
consider the effects of three structural shocks (oil supply shock, aggregate demand shock, and oil-
specific demand shock), in the US stock market. Another useful metric is the construction of historical
decompositions which facilitate the quantification of the contribution of a given structural shock to the
historically observed fluctuations in the variables included in the SVAR model. In particular, one might
be interested to find to what extent the fiscal policy shock explains recessionary conditions (e.g., see
Cevik and Miryugin (2023)) during uncertain times (e.g., see Bloom (2009)). This is useful because we
can concentrate on explaining a historical event which we know a priori might have induced a structural
change to the macroeconomic system, instead of using for example the average contribution of the fiscal
policy shock to the business cycle fluctuations as given by the FEVD (see, Kilian and Lee (2014)).
Suppose that yt is weakly stationary, and we have observations from 1 to t, such that for any t ∈ N
t−1 ∞
yt = ∑ Θswt−s + ∑ Θswt−s.
s=0 s=t
The MA coefficients Θs converge to zero the further to the past, so ŷt = ∑t−1
s=0 Θs wt−s approximates yt .
30
4. Identification of Structural Vector Autoregression Models
While theory-based identification through sign restrictions guarantees economic interpretation of the
identified shocks (e.g., see Antolín-Díaz and Rubio-Ramírez (2018) and Mavroeidis (2021)), data-based
identification offers diagnostic tools for testing the otherwise just identifying structural model specifi-
cations, while remain silent about the economic interpretation of the resulting structural shocks.
Example 21 (Partially Identified SVARs). Consider the dynamic structural models of the following
form (see, Baumeister and Hamilton (2015))
′ ′ , y′ , ..., y′
′
where xt−1 = yt−1 t−2 t−m , 1 is a (k × 1) vector with k = mn + 1 containing a constant and m
lags of y. Then, the MLE estimates of the reduced-form parameters are given by
! !−1
T T
bT = ′ ′ 1 T
Φ ∑ yt xt−1 ∑ xt−1xt−1 and Ω̂T = ∑ ε̂t ε̂t′ ,
T t=1
(4.3)
t=1 t=1
with ε̂t = yt − Φ̂T xt−1 . Moreover, Baumeister and Hamilton (2019) examines aspects related to the
structural analysis of Vector Autoregressive models with incomplete identification based on the Bayesian
approach rather than the classical frequentest approach. Generally to uniquely identify the structural pa-
rameter in a SVAR model implies of having K(K − 1)/2 restrictions. In addition when rank conditions
hold, then the model is assumed to be exactly identified (e.g., see Rubio-Ramirez et al. (2005)).
Example 22 (Identification using Stability Restrictions). According to Magnusson and Mavroeidis (2014),
often identification depends on the distributional assumptions imposed on xt . Suppose that xt is a policy
variable determined according to an underline stochastic process such that
xt = ρ xt−1 + (1 − ρ )φ yt + ηt (4.4)
Furthermore, under a deterministic rational expectations equilibrium, the dynamics of yt and xt are as
yt = β xt−1 + uyt and xt = ϕ xt−1 + uxt , where uyt and uxt are innovation sequences. Moreover, Vlaar
(2004) consider the asymptotic distribution of impulse responses in SVARs with long-run restrictions
(see, also Mittnik and Zadrozny (1993)).
31
Recursive Identification Usually recursive identification with B−1 0 being a lower diagonal matrix was
particularly popular in the earlier SVAR literature which implies exact identification. In particular, when
B−1
0 is lower triangular, so it is the matrix B0 and thus the SVAR is recursive.
Example 23. Consider the following trivariate recursive SVAR(1) model as below
b11,0 0 0 y1t b11,1 b11,2 b13,1 y1,t−1 η1t
b12,0 b22,0 0 y2t = b21,1 b22,1 b23,1 y2,t−1 + η2t
b31,0 b32,0 b33,0 y3t b31,1 b32,1 b33,1 y3,t−1 η3t
A relevant study considers the identification of monetary policy shocks using a trivariate SVAR(1) model
where yt = (xt , ∆pt , πt )′ . However, a drawback of the recursive identification approach is that a differ-
ent ordering of the variables can yields a different SVAR model with different identified structural
shocks (i.e., does not produce uniquely identified structural shocks). On the other hand, it can be eas-
ily estimated since it involves estimating the reduced-form model using OLS or MLE, computing the
covariance matrix and the computing the lower-triangular Cholesky decomposition.
• Provided that the order condition is satisfied, the rank condition for identification can be checked
by assessing sequentially whether each of the equations can be estimated by the instrumental
variables regression, using the variables excluded from each equation as instruments.
• Once an equation has been verified as being identified, then the corresponding structural shock is
identified and thus it can be used as instrument since the shocks are uncorrelated.
32
4.1.1. Dynamic Structural Factor Models with Overidentifying Restrictions
As an example of formal statistical testing for overidentifying restrictions we present the framework
proposed by Han (2018) who develops a new estimator for the impulse response functions (IRFs) in
structural factor models with a fixed number of over-identifying restrictions (see, also Chapter 12 in
Dhrymes (2013)). The proposed identification scheme nests the conventional just-identified recursive
scheme as a special case. The first step before formulating the test statistic is to present the identification
of structural shocks in the structural factor models setting (see, also Koistinen and Funovits (2022)).
Consider the following structural factor model for t = 1, ..., T such that X t = ΛF t + et ,
p
Ft = ∑ Φ j F t− j + Gη t , η t = Aζ t , (4.7)
j=1
⊤
where X t = X1t , ..., Xnt is an n−dimensional vector, F t is an r−dimensional static factor, Λ is an
⊤
(n × r) factor loading matrix and et = e1t , ...., ent is an n−dimensional idiosyncratic error term.
Moreover, G is an (r × q) matrix of rank q, η t is the q−dimensional reduced-form shock and ζ t is
the q−dimensional structural shock. We also have that E(η t η t⊤ ) = I q and A is a q × q nonsingular
matrix. Setting q ≤ r, then the proposed framework allows to to obtain dynamic factors. Furthermore,
by assuming the stationarity of F t , we have that
−1 ∞
I r − Φ1 L − ... − Φ p L p = Ir + ∑ Ψ jL j, (4.8)
j=1
where Ψ j is the coefficient matrix in the vector moving average representation. Moreover, denote with
⊤ ⊤
⊤
F t = F t−1 , ..., F t−p and Φ = Φ1 , ..., Φ p . Rearranging equations gives
X t = ΠF t + Θη t + et ≡ ΠF t + Θη t + et (4.9)
where Π = ΛΦ, Θ = ΛG and Γ = ΘA. In other words, for the i−th cross-section, the factor model can
be represented as X i = F λi + ei , where X i = [Xi1, ..., XiT ]⊤ and F = [F1 , ..., FT ]⊤ and ei = [ei1 , ..., eiT ]⊤ .
Furthermore, let F̂t denote the standard PCA estimator for Ft , such that T1 ∑t=1 T
F̂t F̂t⊤ = I r .
Remark 10. Suppose that the Ft and et are uncorrelated, then Σ = ΛΛ′ + Ψ, where Σ = Var(yt ). By
taking the variance of the static factor model we obtain the above expression for the variance, but we
also need to impose necessary and sufficient conditions for the identification. Even if ΛΛ′ is identified,
Λ is only identified up to an (r × r) rotation. Suitable conditions under which Λ can be identified imply
resolving the rotational indeterminacy (local identification). The factor loading Λ is locally identified
at Λ∗ if in a small open neighborhood of Λ∗ , then Λ∗ is the only matrix that satisfies all the identify-
ing constraints (see, Bai and Wang (2014)). According to Han (2015), the existing literature has not
addressed the overidentification problem in FAVAR models. Unlike the conventional structural VAR
analysis where the number of restrictions for identifying structural shocks, the FAVAR models tend to
involve a large number of identifying restrictions.
33
4.2. Identification via Conditional Heteroscedasticity
The presence of structural change in a sample can be exploited as an identification mechanism within
a structural system context, especially when the heteroscedasticity of structural variances across the
two regimes is exploited (see, Hodoshima (1988), Choi (2002), Lanne and Lütkepohl (2008a) as well
as Bacchiocchi and Fanelli (2011)). In other words, the presence of heteroscedasticity in the error
terms εt of the system can be employed for identifiability purposes. Although, the conditional het-
eroscedastiticy approach often requires to use a pre-testing procedure before estimating the model as
in Meitz and Saikkonen (2021), Lütkepohl et al. (2021), Bertsche and Braun (2022) and Guay (2021);
we follow the study of Brüggemann et al. (2016) who proposed a framework for the identification of
SVARs with conditional heteroscedasticity of unknown form.
Let (ut ,t ∈ Z) be a K−dimensional white noise sequence defined on a probability space (Ω, F , P),
such that each ut = (u1t , ..., u2t )′ is assumed to be measurable with respect to Ft , where (Ft ) is a
sequence of increasing σ −fields of F . Suppose that we observe a data sample (y−p+1 , ..., y0, y1 , ..., yT )
of sample size T plus p pre-sample values from the following DGP for the K−dimensional time series
yt = (y1t , ..., yKt )′ such that
−1
βb = ZZ ′ Z ⊗ I K y. (4.12)
Assume that the process yt is stable, then it has a vector moving-average representation (VMA) s.t.
∞
yt = ∑ Φ j ut− j , t ∈ Z, (4.13)
j=0
i
Φ0 = I K and Φi = ∑ Φi− j A j , i = 1, 2, ... (4.14)
j=1
34
b = vech(Σ
Moreover, we set σ = vech(Σu ) and σ bu ). The vech-operator is defined to stack column wise
the elements on and below the main diagonal of the matrix. A particular useful way to consider the
development of the asymptotic theory is to obtain the joint limiting distribution using an unconditional
central limit theorem as below
!
√ βb − β d
T 2 2
→ N (0,V ) . (4.15)
b −σ
σ
Remark 11. Notice that we can also write the following expression
∞
V (2,2) = Var(ut2 ) + ∑ Cov ut2 , ut−h
2
(4.20)
h=−∞
such that ut2 = vech(ut ut′ ). Hence, V (2,2) has a long-run variance representation in terms of ut2 that
captures the (linear) dependence structure in the underline stochastic sequence. In addition if the errors
are i.i.d then we have that V (2,2) = Var(ut2 ) = LK τ0,0,0 L′K − σ σ ′ .
Next, we focus on the residual-based moving block bootstrap resampling method. Various studies in
the literature have demonstrated that block bootstrap methods are suitable for capturing dependencies
in time series data. Specifically, we are interested in applying the moving block bootstrap technique for
the residuals obtained from a fitter VAR(p) model to approximate the limiting distribution of
′ ′
√
T b
β − β , (σ ′
b −σ) . (4.21)
b1 , ...., A
Step 1. Fit a VAR(p) model to the data to get A b p, and compute the residuals ubt = yt − A
b1 yt−1 −
bp yt−p , for t = 1, ..., T .
A
35
Step 2. Choose a block length ℓ < T and let T = ⌊T /ℓ⌋ be the number of blocks needed such that ℓN ≥ T .
Moreover, define (K × ℓ)−dimensional blocks such that
Bi,ℓ = (b
ui+1 , ..., ubi+ℓ) , i ∈ {0, ..., T − ℓ} (4.22)
such that i0 , ..., iN−1 be i.i.d random variables uniformly distributed on the set {0, 1, ..., T − ℓ}.
Moreover, we lay blocks Bi0 ,ℓ , ..., BiN−1,ℓ end-to-end together and discard the last Nℓ − T values to
get bootstrap residuals ubt∗ , ..., ub∗T .
Step 3. Centering this sequence of residuals ubt∗ , ..., ub∗T based on the following rule
T −ℓ
1
u jℓ+s = ubjℓ+s − E∗ ubjℓ+s = ub∗jℓ+s − ∑ ubs+r
T − ℓ + 1 r=0
for s ∈ {1, ..., ℓ} and j ∈ {0, 1, 2, ..., N − 1} and unconditional mean, E∗ (ut∗ ) = 0, for all t = 1, ..., T .
Step 4. Set bootstrap pre-sample values y∗−p+1 , ..., y∗0 equal to zero and generate the bootstrap sample
y∗1 , ..., y∗T according to the following
b1 yt−1
yt∗ = A ∗ bp yt−p
+ ... + A ∗
+ ut∗ . (4.23)
∗
b b b ∗ ∗′ −1 ∗
β = vec A1 , ..., A p = Z Z Z ⊗ I K y∗ (4.24)
1 T ∗ ∗′
Σ∗u = ∑ ubt ubt (4.25)
T t=1
b∗ y∗ − ... − A
where ubt∗ = yt∗ − A b∗p yt−p
∗ are the bootstrap residuals obtained from the VAR(p) fit.
1 t−1
∗ ∗
We set σ b
b = vech(Σu ).
Theorem 1 (Residual-based MBB Consistency, Brüggemann et al. (2016)). Under regularity conditions
and if ℓ3 /T → ∞ as T → ∞, it holds that
√ ∗ ′ √ ′
sup P ∗ b b ′ b∗ ′
b) ≤ x − P
T (β − β ) , (σ − σ b b ′ b −σ
T (β − β ) , (σ ′
b) ≤ x → 0 (4.26)
x∈RK̃
in probability, where P∗ denotes the probability measure induced by the residual-based MBB.
36
4.3. Identification via Non-Gaussianity of Structural Shocks
Example 24. Consider a standard d−dimensional VAR(p) process given by (see, Petrova (2022))
p
Y t = µ0 + ∑ B0,1Y t−1 + ε t ≡ µ0 + B0 ⊙ Y t−1 + ε t , ε t ∼ (0, Ω0 ). (4.27)
i=1
⊤
⊤ ⊤ is an d p × 1 vector containing the lags of
where B0 = B0,1 , ..., B0,p and Y t−1 = Y t−1 , ...,Y t−p
the vector Y t and ⊙ denotes the inner product between two vectors of the same dimension. Denote
with Ft = σ ε t , ..., ε 1 the natural filtration of the innovation sequence and with EFt the conditional
expectation operator. Recall that covariance stationarity of Y t yields a vector MA(∞) representation of
the form Y t = ∑+∞
j=−∞ Φ j ε t− j .
Assumption 5 (see, Petrova (2022)). Suppose that the following conditions hold:
(i). The VAR process is stable, so that all roots of the polynomial
!
p
ψ (z) = det I d − ∑ z j B0,i , lie outside the unit circle. (4.28)
j=1
Based on the conditions above (stability condition of the system), a suitable estimation method is the
quasi-maximum likelihood method which is robust to distributional misspecifications. The QML esti-
mators of the model parameters B0,µ = [µ , B0 ] are given as below:
!−1 !
T T
⊤ ⊤ 1 T
b0,µ =
B ∑ X t−1 X t−1 ∑ Y t−1X t−1 , Ω̂T = ∑ ε t ε t⊤,
T t=1
(4.30)
t=1 t=1
b0,µ X t−1 . Estimating the VAR model without an intercept after demeaning leads to
where ε t = Y t − B
the same QML estimators for B b0,µ and ε t , which is a simple consequence of the Frisch-Waugh-Lovell
theorem. Moreover, the asymptotic theory analysis can be developed using the Hessian matrix, denoted
with H , such that
∂ 2 ℓ(Y t ; θ )
H (Y t ; θ ) = . (4.31)
∂θ∂θ′
37
Remark 12. The result of distributional misspecification6 is that Bayesian methods leads to asymptot-
ically invalid posterior inference for the intercept and the volatility matrix and, consequently, invalid
posterior credible sets for quantities such as impulse responses, variance decompositions and density
forecasts. Therefore, a Bayesian procedure which delivers asymptotically correct posterior credible
sets regardless of distributional assumptions is desirable for constructing the posterior distribution of
quantities such as impulse response functions. (see, Petrova (2022)).
Generally, exploiting the presence of non-Gaussianity especially in a multivariate system which is more
likely to include non-linearities and complex dependencies of unknown forms between variables is
an identification method which uniquely determines the structural parameter. However, identification
up to signed-permutations reduces the identified set and thus to achieve point identification it is nec-
essary to have a mechanism for selecting a specific permutation as in Lanne et al. (2017) (see, also
Hallin and Mehta (2015)). Recently, Chan et al. (2023) propose a fast algorithm for identifying the
exact permutation scheme based on multiple sign and ranking restrictions.
Example 25 (see, Lanne et al. (2017)). Consider the structural VAR (SVAR) model as below
def
det(A) = I n − A1 z − ... − A p z p =6 0, |z| ≤ 1 (4.33)
If we left multiply by the inverse of B we obtain an alternative formulation of the SVAR model such that
A0 yt = µ ⋆ + A⋆1 yt−1 + ... + A⋆p yt−p + ε t , where we have that A0 = B−1 , µ ⋆ = B−1 µ , A⋆j = B−1 A j , j ∈
{1, ..., p}. Consider the moving average representation of the model as below:
∞
yt = ν + ∑ Ψ j Bε t− j , Ψ0 = I n , (4.34)
j=0
where ν = A(1)−1 µ is the expectation of yt and the matrices Ψ j , for j ∈ {0, 1, ..., } are determined by
the power series Ψ(z) = A(1)−1 ≡ ∑∞j=0 Ψ j z j . An appropriate identification is needed to make the two
factors in the product Bε t , and hence the impulse responses Ψ j B, unique. The uniqueness property of
the identification scheme is determined only up to a linear transformation of the parameter matrices. The
reduced form residuals are assumed to be a linear combination of d independent unobserved variables,
that is, the structural shocks. Thus, the identification implies the consistent estimation of structural
shocks given the reduced-form error terms (a nonprametric approach is proposed by Braun (2023)).
6 Petrova (2022) mentions that inaccurately imposing Gaussian distributional assumptions in standard multivariate time
series models does not affect inference on the autoregressive coefficients but distorts both classical and Bayesian inference
on the volatility matrix whenever the true error distribution has excess kurtosis relative to the multivariate normal density.
38
Assume that the matrix B ≡ SC where C ∈ Rd×d is orthogonal and S ∈ Rd×d is such that Σu = SS′ .
Thus, given a large number of time series observations {ut }t=1,...,n, then the statistical problem is to
consistently estimate the matrix product SC and determine the distribution of the structural shocks ε t .
Since, the matrix S can be directly estimated by matching from the covariance matrix Σu , using the
Cholesky decomposition, then the identification problem of the structural parameter matrix, boils down
to the consistent estimation of the orthogonal matrix C ∈ Rd×d . When structural shocks ε t are assumed
to be Gaussian, one can only identify the space of the mixed signals without separating the individual
components. The only way to ensure consistent identification is by assuming that structural shocks are
non-Gaussian and then proceeding by consistently estimating the orthogonal matrix C. To consistently
estimate the individual components of the structural shocks we assume joint non-Gaussianity and that
the orthogonal matrix C is consistently estimated. The main intuition of exploiting the non-Gaussianity
of structural shocks as an identification mechanism as in Lanne et al. (2017) can be seen via the concept
of time-reversibility (see, Findley (1986), Tong and Zhang (2005), Chan et al. (2006) and Lawrance
(1991)7 ) in time series analysis. The notion of time-series reversibility was proposed by economists and
financial engineers to address the asymmetry property of business cycles. The presence of asymmetric
signals can induce enough skewness to reject the null hypothesis of time-reversibility. Business cycle
asymmetry can be tested with respect to whether macroeconomic fluctuations are time irreversible and
can also be exploited for identification purposes as in Virolainen (2020) (see, Nyberg and Savva (2023)).
Theorem 2 (Uniqueness result). Suppose that xt is a non-Gaussian, strictly stationary, zero mean time
series whose r−th moment E(xtr ) are finite for all r = 1, 2, .... Suppose also that the spectral density of
xt is positive almost everywhere. Then, xt can have at most one mean square convergent representation
xt = ∑∞j=−∞ c j ηt− j , in which ηt is a sequence of i.i.d distributed zero mean random variables with finite
moments, ignoring changes of scale and shifts in the time origin of the series ηt (see, Chan et al. (2006)).
Example 26 (Chan et al. (2006)). Consider the univariate non-Gaussian linear process representation
∞
Xt = ∑ ′
A′j εt− j, t ∈ N (4.35)
j=−∞
There exists an integer m and a matrix B such that, for all t, (i) At ≡ Am−t B and (ii) Bεt is distributionally
equivalent to εt . This uniqueness of a non-Gaussian process implies that, if X admits representation
(4.35) and Yt = ∑∞j=−∞ A′j εt−
′ , then Y is distributionally equivalent to X if and only if the previous
j
conditions hold, with the equality concerning the innovations interpreted as equivalence in distribution.
Proof of Proposition 1 in Lanne et al. (2017) Suppose that the following two conditions hold:
(i). The error process ε t = ε1,t , ..., εd,t is a sequence of i.i.d random vectors with each component
εi,t , for i = {1, ..., n}, having zero mean and finite positive variance σi2 .
7 Professor Tony Lawrance is an Emeritus Professor at the Department of Statistics, University of Warwick, UK.
39
(ii). The components of ε t = ε1,t , ..., εd,t are (mutually) independent and at most one of them has a
Gaussian marginal distribution.
A key step is to demonstrate the existence and uniqueness of an equivalent SVAR representation. In par-
ticular, one needs to show that the structural shocks of the original model can be determined as a function
of the structural shocks of the observetionally equivalent SVAR process multiplied by an unknown per-
mutation matrix M such that ε t = M ε t∗ , M = B−1 B∗ . To ensure a valid identification scheme, we need
to examine the properties of the elements of this permutation matrix. Thus, we assume that at most one
column of M may contain more than one nonzero element (see, Lemma A.1 in Lanne et al. (2017) and
Kagan et al. (1973)8 Suppose say that the k−th column of M has at least two nonzero elements, mik and
m jk such that i 6= j. Then, it holds that
∗ ∗
εi,t = mik εk,t + ∑ miℓ εℓ,t (4.36)
ℓ=1,....n;ℓ6=k
∗ ∗
ε j,t = m jk εk,t + ∑ m jℓ εℓ,t (4.37)
ℓ=1,....n;ℓ6=k
where the random variable εk,t ∗ is assumed to be Gaussian with positive variance. Moreover, for all
ℓ ∈ {1, ..., n; ℓ 6= k}, it must hold that mik m jℓ = 0 because only the k−th element of M could have more
than one nonzero element. Notice that the purpose of the identification scheme is to distinguish or
identify the structural shocks. Due to the independence assumption of the error terms and by the fact
that each row of M has exactly one non-zero element, then we can conclude that there exist a permutation
matrix P and a diagonal matrix D = diag {d1 , ..., dn} with non-zero diagonal elements such that M = DP.
Rearranging terms implies that ε t = B−1 B∗ ε t∗ and to have observational equivalence it should hold that
△
ε t ≡ B−1 BDP ∗ ∗ ′ −1
|{z} ε t , where ε t = P D ε t (4.38)
B∗
since D is a diagonal matrix then D−1 is also a diagonal matrix and PP′ = I where P represents a
permutation matrix within the same equivalence class.
Proposition 1 in Lanne et al. (2017) characterizes a class of observationally equivalent SVAR processes
that differ only with respect to the ordering and scaling of the structural shocks in the vector εt . A
(rescaled) error vector of these set of permuted structural shocks, consists of exactly the same elements,
whose ordering varies and thus these SVAR processes induces asymptotically equivalently impulse
responses functions. Thus, the invariance transformation property allows to identify structural shocks
using an observationally equivalent class of SVAR processes based on mutual independence.
8 For the interested reader, the property is due to the theory of locally compact abelian groups which implies that any
finitely generated group can be expressed as a direct product of cyclic groups. As a result, the duality result allows to
construct a distributional equivalent topological space using properties from non-Gaussian topological spaces. Specifically,
any compact Hausdorff abelian group can be completely described by the purely algebraic properties of its dual group.
Corollary 3 in Morris (1979) defines the observetionally equivalent class of topologically isomorphic groups. These theorems
remain agnostic regarding the exact properties model estimators are assumed to have ex-ante to the imposed specification,
which implies that using a non-Gaussian distributional assumption is sufficient.
40
4.3.1. Testing for Independence of Structural Shocks
The independence property of structural shocks εt can serve to uniquely identify the structural parameter
B it at most one component of εt is Gaussian. Moreover, in non-Gaussian frameworks, one has to
consider the entire dependence structure to implement a robust structural analysis. Moreover, although
the original identification scheme of Lanne et al. (2017) is based on the assumption of at least one
Gaussian shock for full identification, the particular result is extended by Maxand (2020) who show
that the same identification scheme is valid in the case of multiple Gaussian structural shocks while
the remaining structural shocks are non-Gaussian. In addition some further studies as in Crucil et al.
(2023) and Bruns and Keweloh (2023) consider the case of non-Gaussian proxy SVARs. In particular,
Crucil et al. (2023) extending the approach of Schlaak et al. (2023) employs an instrumental regression
to estimate the structural shock of interest while staying agnostic for the remaining structural shocks.
Testing for the independence of the identified structural shocks obtained from orderings (e.g., a set
of permutations of all shocks) of the VAR residuals is an important validation step for the underline
assumptions. When a suitable test statistic of independence for the identified shocks doesn’t find evi-
dence against the hypothesis of independence then this is an indication that using a non-Gaussian MLE
approach should render valid model estimates. Motivated from the presence of such non-Gaussian and
heavy-tailed error terms in SVAR models Davis and Ng (2023) study the role of uncertainty in economic
fluctuations. Specifically, disaster shocks are shown to have short term economic impact arising mostly
from feedback dynamics. In particular, the authors show that the coefficient estimate on the infinite vari-
ance regressor would be consistent but not asymptotically normal. As a result, the coefficients of impulse
response functions whether there are computed from the VAR or by local projections are likely not to
be asymptotically normally distributed. A possible drawback of the particular identification scheme is
that exploiting the non-Gaussianity property precludes cases of non-identification. Recently, there is
a growing interest in testing not only the distributional assumptions of structural shocks but also indi-
rectly testing the absence of non-identification. In particular, Drautzburg and Wright (2023) mentiones
that criteria such as an inverted GMM test statistic based on co-dependence of higher moments, can be
employed to rule out incorrect identification schemes when the data are non-Gaussian. When the data
are Gaussian, there is no co-dependence of higher moments and the estimator becomes close to the plug-
in estimator in Moon and Schorfheide (2012) (see, also Granziera et al. (2018)). Thus, the approach of
Drautzburg and Wright (2023) exploits independence for identification, but does so allowing for weak
identification, which indeed remains valid in the Gaussian limit where there is complete lack of identifi-
cation. Further issues related to identification via non-Gaussianity is the modelling of climate-economic
systems which is relevant to the structural analysis of climate shocks to macroeconomic variables.
Remark 13. Under the assumption of non-Gaussianity, the fact that ε t are white noise does not imply
that they are serially independent. This is an important aspect of a non-Gaussian VAR model because
even if Corr(εit , εit+1 ) = 0 for i = 1, ..., l, it is still in principle possible that εit+1 is statistically dependent
on εit . Thus, in this latter case εit+1 would not be an actual innovation term, since it can be predicted by
yt (see, Moneta et al. (2013)).
41
4.4. Identification via External Instruments
Example 27. Following the framework proposed by Mittnik and Zadrozny (1993), let θ be the param-
eter vector that collects the coefficients of the original SVAR representation.
• Suppose we have the model A(L)yt = B(L)ε t , with α := vec (A0 , ..., A p ), β := vec (B0 , ..., B p ) and
′
r := vec(R), then θ = α ′ , β ′ , r′ . Moreover, let φ denote the vector of parameters which are
estimated, that is, the structural parameters of the system.
• Under the assumption that sufficient restrictions have been imposed to identify φ , we say that
these identifying restrictions induce a mapping from φ 7→ θ such that the linear map satisfies
θ = F(φ ). Let φ̂ denote an estimate of φ . Then, under stationarity, invertibility and appro-
priate regularity conditions on the linear map F(·), φ̂ will be normally distributed such that
√ d
T φ̂ − φ 0 → N (0, Ξ).
Overall, SVARs are used in the applied macroeconometrics literature to estimate the dynamic causal
effects of structural shocks. Parameters in these models have traditionally been point-identified using
zero or long-run restrictions. A commonly used approach in empirical macroeconometrics research is to
construct time series that capture exogenous changes affecting the macroeconomy. The dynamic effects
of these time series can be often estimated using distributed lag regressions. However, these variables
are correlated with a particular structural shock and uncorrelated with other structural shocks (e.g.,
target versus non-target structural shocks). In particular, external instruments impose second moment
restrictions that identify SVAR shocks and associated impulse response coefficients.
A major known issue of SVAR-IV models is the weak identification problem which appears in the case
that instruments are only weakly correlated with the variable of interest. Specifically, building on the
IV regression literature, Stock and Watson (2021) proposed weak-instrument robust inference methods
for impulse response coefficients. Their framework demonstrates how an external instrument can be
used to identify the structural shock of interest, its impulse response coefficients, historical effect on
the variables in the VAR and contribution to forecast error variances. A relevant identification scheme
combines sign restrictions with external instruments is proposed by Braun and Brüggemann (2023).
Remark 14. Recall that the invertability property implies that structural shocks can be recovered from
current and past (but not future) values of the observed macro variables. On the other hand, it is likely
that external shock measures are contaminated by substantial measurement errors, causing attenuation
bias. In other words, external instruments can be employed as an informative tool regarding shock
importance. For example, an unrestricted linear MA model using IV can be employed to construct
related statistics. In such settings a key identifying assumption is the availability of external instruments
that are correlated with the shock of interest but are otherwise dynamically uncorrelated with all other
macro variables (see, Plagborg-Møller and Wolf (2022)). Identification of SVAR models with external
instruments is also studied by Montiel Olea et al. (2021). Lastly, a framework for identifying SVARs
with weak proxies is proposed by Angelini et al. (2024).
42
4.5. Identification via Sign-Restrictions and Set-Identified SVARs
Several studies consider the set-identification of SVAR models based on a combination of economic
theory and data-driven approaches such as sign restrictions. The key idea behind sign restrictions
is to characterise a shock through imposing restrictions on the responses of some variables, while
remaining agnostic about others. In particular, Inoue and Kilian (2013, 2016) study the construc-
tion of confidence sets for set-identified parameters. Moreover, Giacomini and Kitagawa (2021) and
Giacomini et al. (2022) consider robust inference in Proxy-SVARs with external instruments. Sign re-
strictions usually imply partial identification (see, Dufour and Taamouti (2005)). Additionally, it has
several drawbacks including the fact that different structural shocks in the model might be character-
ized by the same set of sign restrictions. Nevertheless, to illustrate the implementation of the particular
identification scheme we follow the framework proposed by Gafarov et al. (2018) who contribute to the
literature of set-identified SVARs by proposing a novel delta-method interval for the coefficients of the
IRF.
Consider the n−dimensional SVAR with p−lags such that Y t = A1Y t−1 + ... + A pY t−p + Bε t . Then,
the object of interest is the k−th period ahead structural impulse response function of variable i to a
particular shock j such that λk,i, j (A, B) ≡ e′iCk (A)B j , where B j ≡ Be j and ei and e j denote the i−th and
j−th column of I n . A second object of interest is the vector of reduced-form VAR parameters
′
µ ≡ vec(A)′ , vec(B)′ ∈ M ⊂ Rd , A ≡ vec (A1 , A2 , ..., A p ) . (4.39)
Overall, a common practice in empirical macroeconomics literature is to set equality and inequality
restrictions to set-identify the structural IRFs. Specifically, an example of an equality restriction in a
monetary VAR is that prices do not react contemporaneously to monetary policy shocks. An example
of an inequality restriction is that contractionary monetary policy shock cannot increase prices. Let
R(µ ) ⊂ Rn be the set of values of B j that satisfy the inequality and equality restrictions such that
n ′ ′
R(µ ) ≡ B j ∈ R Z(µ ) B j = 0mz ×1 and S(µ ) B j ≥ 0ms ×1 , (4.40)
where Z(µ ) is a matrix of dimension (n × mz ) that collects the equality restrictions specified by the
researcher and S(µ ) is a matrix of dimension (n × ms ) that collects the inequality restrictions.
Corollary 2 (Directional Differentiability, see Gafarov et al. (2018)). Consider the auxiliary function
νk,i, j (µ ; r(µ )), such that νk,i, j (µ ; r(µ )) 6= 0, then the function νk,i, j (µ ; r(µ )) is differentiable with re-
′
spect to vec(A)′ , Σ)′ with derivative given by
43
Theorem 3 (see, Gafarov et al. (2018)). Consider a delta-method interval of the following form
" #
σb(k,i, j) σb(k,i, j)
CST (1 − α ) ≡ ν k,i, j (θb T ) − z1−α /2 √ , ν k,i, j (θb T ) + z1−α /2 √ (4.43)
T T
b b b
where θ T ≡ vec(AT ), vec(ΣT ) is the OLS estimator such that
! !−1
bT ≡ 1 T 1 T 1 T
′
A ∑ Y t X t′
T t=1 ∑ X t X t′
T t=1
bT ≡
, Σ ∑ ηb t ηb t , ηb t ≡ Y t − AbT X t
T − np − 1 t=1
(4.44)
Therefore, Gafarov et al. (2018) suggested to exploit the specific form of the directional derivative in
the SVAR framework via the following expression
n o1/2
σb(k,i, j),T ≡ max b ′b b
ν̇ k,i, j (θ T ; r) ΩT ν̇ k,i, j (θ T ; r) . (4.45)
bT )
r∈R(µ
Evaluating the posterior of sign-identified VAR Models Consider the d−variate reduced-form VAR(p)
model such that (see, Inoue and Kilian (2013))
i.i.d
for t = 1, ..., T , where et ∼ N (0n×1 , Σ) and Σ is positive definite. We write the above econometric
model such that
Y = XB + e (4.47)
′ ′ ′ ] and B = [c B B ... B ]′ . Moreover,
where Y = y1 y2 ... yT T ×k , X = [X1 X2 ... XT ], Xt = [1 yt−1 ... yt−1 1 2 p
we follow the conventional approach of specifying a normal-inverse Wishart prior distribution for the
reduced-form VAR parameters.
vec(B)|Σ ∼ N vec(B̄0 ), Σ ⊗ N0−1 , (4.48)
Σ ∼ IWn ν0 S0 , ν0 , (4.49)
44
4.6. Identification in Proxy SVARs under Possible Nonstationarity
Suppose that {Yt : t = −p + 1, ..., T } be a (d × 1) vector of observed variables that follows a SVAR
p
process Y t = µ + ∑ j=1 A jY t− j + ut , ut = Bε t , where ut is the reduced-form error terms and A j are
(d × d) coefficient matrices for j ∈ {1, ..., p}. Therefore, to study the IRFs with respect to the structural
shocks εit , we study the estimation of the i−th column of the structural parameter B denoted by B =
(b1 , ..., bd )′ with external instruments Zt = (z1t , ..., zkt )′ ∈ Rk : t = 1, ..., T .
Example 28. Let zt denote an external instrument for a structural shock of interest εkt , k ∈ {1, ..., K}.
Based on conventional assumptions from the instrumental variable estimation literature, the potential
instruments must satisfy the following conditions:
(i) zt is relevant for the underlying structural shock εkt such that E (εkt zt ) = φ 6= 0.
(ii) zt is exogenous from other structural shocks in the system such that E ε jt zt = 0, ∀ j ∈ {1, ..., K}\ {k}.
Then, based on the above two conditions it follows that up to scale φ , the population covariance between
the instrument and VAR residuals obtains the k−th column of B, denote by Bk such that
E (ut zt ) = Bk E (εkt zt ) = φ Bk .
Thus, the link between the structural shock and reduced form-error provides the condition to ensure the
instrument validity and instrument relevance. These econometric specifications are also particularly im-
portant for the identification of technology shocks as in Feve and Guay (2010) and Pagan and Pesaran
(2008) (see, also Lovcha and Perez-Laborda (2021)). However, according to Cheng et al. (2022) long-
run restrictions can lead to the problem of weak identification due to sensitivity to low-frequency cor-
relations between variables. Furthermore, due to the fact that yt often includes components which are
integrated and cointegrated it is important to develop a robust identification methodology for Proxy
SVAR models which is robust to weak identification and possibly nonstationary regressors.
Example 29 (Nearly nonstationary SVAR, see Gospodinov (2010)). Suppose that we have a bivariate
vector autoregressive process denoted with ỹt = (y1t , y2t ) such that
p
where the matrix Φ contains the largest roots of the system and Ψ(L) = I − ∑i=1 Ψi Li and where
e
A(L) = Ψ(L)(I − ΦL). Consider the reduced form VAR model ỹt = A1 ỹt−1 + ... + A p+1 ỹt−p−1 + ut . Pre-
" #
(0)
1 −b21
multiplying both sides by the matrix B0 = (0) , yields the structural VAR model B(L)yt = ε t ,
−b21 1
where B(L) = B0 A(L) and ε t = B0 ut denote the structural shocks with are assumed to be orthogonal with
variances σ12 and σ22 . Thus, writing the reduced model in the DF form
p
yt = A (1)yt−1 + ∑ A⋆⋆
⋆
j ∆yt− j + ut , (4.51)
j=1
45
e p
such that A⋆ (L) = L−1 I − A(L) , A⋆ (1) = I + T1 Ψ(1)C and A⋆⋆ ⋆
i = − ∑ j=i+1 A j . This implies that,
p p
" # " # c a⋆j,11 a⋆j,12 " # " #
∆y1t ψ12 (1) y2t ∑ ∑
0 T j=1 j=1 ∆y1t− j u1t
= c + p p + . (4.52)
∆y2t 0 ψ12 (1) y2t ∆y2t− j u2t
T ∑ a⋆j,21 ∑ a⋆j,22
j=1 j=1
Notice that for this example following Gospodinov (2010), the author considers a bivariate system
of equations and propose an overidentifying restriction which using the VECM representation of the
model reduces to the assumption that b12 is known and is equal to zero. The relationship between the
endogenous variable and the instrument is given by the second equation of the reduced VAR model as
c
∆ỹ2t = ψ22 (1) y2t−1 + u2,t , t = 1, ..., n, (4.53)
T
Moreover, based on the long-run restrictions and the model assumptions, it can be proved
Z 1 q Z 1
ρ Jc (s)dV2 (s) + 1 − ρ 2 Jc (s)dV1 (s)
(0) (0) σ1 0 0
b̂12 − b12 ⇒ Z 1 Z 1 . (4.54)
ω2 2
c Jc (s) ds + Jc (s)dV2(s)
0 0
Long-run restrictions are a popular identification method of SVAR models since the seminal contribution
of Blanchard and Quah (1988), which are commonly used to investigate the hours debate9 of Gali (1999)
as well as the impact of technology shocks on productivity and output (see, also Chari et al. (2008)).
Example 30. Following Chaudourne et al. (2014), consider a nearly stationary persistent process, which
is obtained using the local-to-unity parametrization as in Phillips (1987) who consider nearly unit-root
process to investigate the asymptotic power of the unit-root tests under a sequence of local alternatives.
Therefore, since we are interested to model a highly persistent process which is asymptotically sta-
tionary, we consider a sequence of local alternatives such that the process is locally nonstationary but
asymptotically stationary and persistent.
c
1 − ρ L ∆xt = ut − 1 − √ ut−1 , t = 1, ..., n. (4.55)
n
Notice that as n → ∞ and for a high value of ρ , this process becomes a stationary persistent process
whereas, in a finite sample, the process is characterized by a unit root.
9
In particular, the hours debate is concerned with the short-run effect of a technology shock on hours initiated by
Gali (1999) and Christiano et al. (2003). Using a bivariate SVAR framework Y1t corresponds to log productivity and Y2t
corresponds to log hours. Further studies which investigate the impact of technology shocks to business cycles include
Francis and Ramey (2005), Fève and Guay (2009) and Chaudourne et al. (2014). For example, Fève and Guay (2009) im-
pose a long-run identifying restriction which implies that the composite non-technology shock has no long-run effect on
labor productivity. This means that the upper triangular element of A(L) in the long-run must be zero, that is A12 = 0, so
to uncover this restriction from the estimated VAR(p) model, the matrix A(1) is obtained by the Choleski decomposition of
C(1)ΣC(1).
46
Thus, this process is locally nonstationary but asymptotically stationary and persistent. Such a process
appears to be particulary well suited to characterize the dynamics of hours worked because it implies
a unit root in a finite sample but is asymptotically stationary and persistent. This is typically the case
for per capita hours worked which are included in SVARs. Consider a bivariate representation where
X t := (∆X1t , X2t ) for t = 1, ..., n, such that the variable X1t contains an exact unit root and X2t is a highly
persistent variable. Notice that both variables in the vector Xt are asymptotically second order stationary
and they admit asymptotically the following Wold representation
∞
X t = C(L)ε t , C(L) = ∑ C j ε t− j , (4.56)
j=0
′
where η t = η1t , η2t is the vector of orthogonal shocks with E (η t η t′ ) = Ω a diagonal matrix.
A common identification assumption is that Ω = I 2 which implies that the variance of the structural
shocks is normalized to one. Then, we have that the error terms ε t from the reduced form are related to
the structural error terms η t as follows: ε t = A0 η t which implies that Σ = A0 A′0 . Therefore, the iden-
tification scheme, should ensure the unique identification of the structural parameter A0 and we focus
to the case of long-run restrictions as proposed by Blanchard and Quah (1988). Specifically, we use
the long-run variance-covariance matrix of the reduced form and the structural form which are related
by C(1)ΣC(1)′ = A(1)A(1)′ and A0 = C(1)−1 A(1). Typically a lower triangular structure is imposed to
the long-run impact matrix A(1) which can be easily obtained using a Choleski decomposition of the
long-run covariance matrix C(1)ΣC(1)′ . In other words, in the case where two variables are included in
X t , the first structural shock is the only one shock that can have a permanent effect on the first variable.
Consider for example, a finite sample of size T observations a structural characterization of the highly
persistent variable X 2t as a nearly stationary persistent process such that
c
∆X2t = a21 (L)∆η1t + a22 (L) 1 − 1 − √ L η2t
T
≡ a21 (L)∆η1t + ã22 (L)η2t (4.58)
where ã22 (L) = a22 (L) 1 − 1 − √cT L . Moreover, using the Beveridge-Nelson decomposition, the
above model can be rewritten as below
√
such that ã22,T (1) = a22,T (1)c/ T and ã∗22,T (L)(1 − L) = ã∗22,T (L) − ã∗22,T (1).
47
Putting everything together, it implies that the SMA bivariate representation contains a difference sta-
tionary process ∆X1t and a highly persistent process such that
" # " #" #
∆X1t a11 (L) ã12 (L) η1t
= . (4.60)
∆X2t a21 (L)(1 − L) ã22,T (L) η2t
Estimation and inference We now consider estimation and inference with the proposed specifications
of the SVAR model. The reduced-form moving average representation is retrieved by performing a finite
order VAR on the data. Suppose that the structural moving average representation can be characterized
or approximated in a small sample by a finite VAR of order p. Consider the reduced form VAR(p):
p p
(i) (i)
1 − ∑ p d11 Li −∑ d12 Li
i
D(L)X t = ε t , D(L) = I − ∑ Di L = i=1
p
i=1
p (4.61)
(i) (i) i
i=1 − ∑ d12 Li 1 − ∑ d22 L
i=1 i=1
!
(0)
1 −b12
By multiplying both sides by a matrix of the form B0 = (0) = A−1
0 , we obtain the VAR as
1 −b21
a function of the structural shocks such that B(L)X t = η t with B(L) = B0 D(L).
p p
" # 1 − d (i) Li − d (i) Li
1 −b12
(0)
∑ 11 ∑ 12
i=1 i=1
(0) p p
−b21 1 (i) (i)
− ∑ d12 Li 1 − ∑ d22 Li
!i=1 i=1
!
p p p p
(i) (0) (i) (i) (0) (i)
1− ∑ ∑ d11 Li
−∑ + b12 d12 Li
1 − ∑ d22 Li d12 Li −b12
=
i=1
p
! i=1
p
i=1
p
i=1
p
! .
(0) (i) (i) (0) (i) (i)
−b21 1 − ∑ d11 Li − ∑ d12 Li b21 ∑ d11 Li + 1 − ∑ d22 Li
i=1 i=1 i=1 i=1
48
Moreover, imposing the structural long-run impact matrix A(1) to be lower triangular implies that
B0 D(1) is also lower triangular by A(1) = D(1)−1 A0 . The long-run multiplier of the variable X 2t on
∆X 1t is then zero. Imposing this constraint yields that
!
p p p−1
(i) (0) (i) (0) (i)
∆X1t = ∑ d11 Li − b12 ∑ d21 Li ∆X1t + b12 ∆X2t + ∑ b̃12 Li∆X2t + η1t
i=1 i=1 i=1
(0)
= b11 (L)∆X1t−1 + b12 ∆X2t + b̃12 (L)∆X2t−1 + η1t .
(0)
where the element b12 = −d12 (1)/ (1 − d22 (1)).
(0)
IV Estimation The IV estimator of b12 with X2t−1 as instrument is given by the following expression
!−1 ! !−1 !
1 T 1 T 1 T 1 T h i
(0) (0)
b12 = ∑ X2t−1 ∆X̃2t
T t=1 ∑ X2t−1 ∆X̃1t
T t=1
= ∑ X2t−1∆X̃2t
T t=1 ∑ X2t−1 b12 ∆X̃2t + η1t
T t=1
(0)
We consider the estimation of the structural parameter b21 . Since η1t and η2t are orthogonal, the
(0)
residuals η b1t = ∆X̃1t − b
b12 ∆X̃2t can be used as instrument for the endogenous variable ∆X̃1t . Thus, it
(0) (0)
holds that η b1t − η1t = b b12 − b12 ∆X̃2t . Moreover, we shall define with zt = (ηb1t , X2t−1 )′ and xt =
′ (0)
∆X̃1t , X̃2t−1 . Then, the IV estimator of b12 is given by
!−1 !
1 T 1 T
βb = ∑ zt xt′ ∑ zt ∆X2t′ . (4.65)
T t=1 T t=1
Impulse Response Function Analysis Consider the following VAR(1) model as below
(0)
∆X1t = b12 ∆X2t + η1t (4.66)
(0)
X2t = b21 ∆X1t + b22 X2t−1 + η2t (4.67)
Notice that the highly persistent variable represents hours worked. Generally, it can be proved that
estimated impulse responses from LSVAR and SDVAR models are biased in a finite sample if the
measure of productivity is contaminated by low frequency movements in hours. Moreover, although
these estimators are asymptotically consistent they display a nonstandard limiting distribution.
49
Example 31 (SVAR reduced-form representation, Cheng et al. (2022)). Suppose that we have a vector
of macroeconomic variables Y t∗ such that Y t∗ := µ +Y t . Consider the SVAR model formulation
p
Y t∗ ∗
= δ + ∑ Φ jY t− j + ut , ut = Bε t , t = −p + 1, ..., T. (4.68)
|{z} j=1
(k×1)
Thus, by taking the demeaned vector of regressors Y = (Y t∗ − µ ) we proceed by considering the fol-
lowing reduced-form representation. In other words, we allow the SVAR process to consist of a set of
near-unit root and cointegrated variables.
Specifically, Cheng et al. (2022) employ the reduced-form:
C
Y 1t = I k1 + Y 1,t−1 + u1t , u1t ∈ Rk1 , (4.69)
n
Y 2t = ΠY 1t + u2t , u2t ∈ Rk2 , (4.70)
k3
Y 3t = u3t , u3t ∈ R (4.71)
′ ′
where it holds that Ψ(L)ut = et such that ut = u′1t , u′2t , u′3t and et = e′1t , e′2t , e′3t such that
(i). The roots of Ψ(L) has all roots outside the unit circle.
∞
2
(ii). Ξ0 = I K , where Ξ(1) has full rank, where ∑ j2 Ξj < ∞.
j=0
(iii). Suppose that ēt = [et′ , vt′ ]′ is an i.i.d(r + k) × 1 vector with mean zero, E ēt ēt′ = Σ is positive
definite, fourth moments of ēt are finite, and et is homoscedastic conditional on vt .
50
′ ′
Denote with X t = Y ∗1,t−1 − µ 1 , 1, Dt′ , then the original SVAR process (which includes a model
intercept) and the above reformulation based on the demeaned original series has an one-to-one trans-
formation. Moreover, this equivalent representation implies that the least-squares residual ût from the
VAR process is numerically equivalent to that obtained from regressing ∆Y t on X t . Therefore, we shall
use the conventional SVAR representation for practical estimation purposes such as to obtain the OLS
residuals of the model whereas we use the above reformulation for the asymptotic theory analysis of the
corresponding model estimator under the presence of both near-unit and cointegrated regressors.
Next, for the near unit root regressors we denote with J c (s) an (k1 × 1) vector OU process such that it
# equation dJ c (s) = CJ c (s) + dBu (s). Denote with D n the
satisfies the following "stochastic differential
√
nI k1 0
diagonal matrix D n = . Moreover, it holds that
0 I pr−k1 +1
! Y ∗ − µ
1,t−1 1
1 n n
∗ ′ −1
D −1
n ∑ X t X t′ D −1
n = D −1
n ∑ 1
′
Y 1,t−1 − µ 1 , 1, Dt D n
n t=1 t=1
Dt
n n
′
∑ D −1
n Y ∗1,t−1 − µ 1 Y ∗1,t−1 − µ 1 D −1
n ∑ D −1
n Y ∗1,t−1 − µ 1 D −1
n 0
t=1 t=1
n ′
=
∑ D −1
n Y ∗1,t−1 − µ 1 D −1
n 1 0
t=1
0 0 ΓDD
∆Y t∗ = A1 + A2 (Y ∗1,t−1 − µ 1 ) +A3 Dt + ut ,
| {z }
Y1,t−1
51
Our goal is to demonstrate that the data generating process with three equations representing the case of
both near unit roots and cointegrated regressors, is equivalent to the original SVAR formulation.
C
Y 1t = I k1 + Y 1,t−1 + u1t , u1t ∈ Rk1 ,
n
Y 2t = ΠY 1t + u2t , u2t ∈ Rk2 ,
Y 3t = u3t , u3t ∈ Rk3
where the above regressors correspond to their demeaned counterparts. In other words we start by
rewritting the regressors specific represenation to the equivalent general SVAR process.
C
∆Y ∗1t n 0 0 ∆Y ∗ − µ I 0 0
1 k
∗ 1t−1 1
C
∆Y 2t = ∗ −1
Π I k1 + −I k2 0 ∆Y 2t−1 − µ 2 + Π I k2 0 Ψ(L) et , (4.76)
∗
∆Y 3t n ∗
∆Y 3t−1 − µ 3 0 0 I k3
0 0 −I k3 | {z }
| {z } P
M
Recall that we have that ut := Pet which implies by the change of basis rule
u1t I k1 0 0 e1t I k1 ⊗ e1t
u2t := Π I k2 0 e2t ≡ Π ⊗ e2t + I k2 ⊗ e2t .
u3t 0 0 I k3 e3t e3t
∗ − µ + PΨ(L)−1 e . Multiplying both sides by
Thus in matrix notation we have that ∆Y t∗ = M Y t−1 t
−1
PΨ(L)P , we obtain
PΨ(L)P−1 ∆Y t∗ = PΨ(L)P−1 M Y t−1
∗
− µ + Pet . (4.77)
Remark 16. In summary, Cheng et al. (2022) show that that in the presence of stationary regressors,
cointegration relationships, or more than one lag variables in the VAR system, the estimation error
from the nonstationary component is asymptotically negligible. Specifically, the authors prove that the
asymptotic variance of the IRF only depends on the stationary component, but a consistent covariance
matrix estimator is available even without knowing which series are stationary. Although the approach
of Cheng et al. (2022) shows that for a proxy SVAR standard asymptotic normal inference remains valid
under a general form of nonstationarity in the VAR system, their theoretical result is not uniform over
the entire parameter space of the roots of the autoregressive model as in the spirit of Mikusheva (2007).
A relevant framework that demonstrates the robustness of SVAR identification and estimation to the
presence of unit root and cointegration dynamics is proposed by Chevillon et al. (2020).
52
4.7. Identification with Occasionally Binding Constraints
At the Zero Lower Bound (e.g., see Eggertsson et al. (2014) and Aruoba et al. (2022)) the ability of
policy makers to react to macroeconomic shocks is limited and therefore the response of the economy
to shocks will change when the policy instrument hits the ZLB. Thus, according to Mavroeidis (2021)
since the difference in the behaviour of the macroeconomic variables across the ZLB and non-ZLB
regimes is only due to the impact of monetary policy, the switch across regimes provides information
about the causal effects of policy. Therefore, the ZLB identifies the causal effect of policy during the
unconstrained regime. However, if monetary policy remains partially effective during the ZLB regime,
for example, through the use of unconventional policy instruments, then the difference in the behaviour
of the economy across regimes will depend on the difference in the effectiveness of conventional and un-
conventional policies. In this case, we obtain only partial identification of the causal effects of monetary
policy, but we can still get information bounds on the relative efficiency of unconventional policy.
From the econometric perspective a major challenge when the restriction mechanism involves OBCs is
that it generates censoring of one of the dependent variables. Specifically, once the censoring mech-
anism is triggered then one can allows for some of the coefficients for the remaining variables to
change. According to Aruoba et al. (2022) the specification of a dynamic multivariate model with
censoring and regime-dependent coefficients faces two challenges: parsimony and the existence of a
unique reduced form. These models distinguish between a shadow rate, y∗1t , and the actual interest rate
y1,t = max y∗1t , c by defining the endogenous variable st = 1 y1,t > c . This implies that the regime-
dependency of the coefficients is equivalent to capturing nonlinearities in decision rules that arise in
DSGE models with occasionally-binding constraints. Moreover, Duffy et al. (2022) mentions that in an
autoregressive model, an occasionally binding constraint naturally give rise to nonlinearity in the form
of links between unit roots and stochastic trends. A prototypical model from the VECM literature is
k−1
∆yt = g β ′ yt−1 + ∑ Γ j ∆yt−1 + ut , (4.78)
j=1
Causal ordering provides a mechanism for causal effect justification in identification and estimation
of SVARs. In particular, Moneta et al. (2011) studies statistical methodologies for causal inference in
SVAR models while other authors propose a graph-theoretic method for causal analysis in vector au-
toregressions which allows to select the causal (contemporaneous) order of a SVAR model. Thus, a
key role to the theory of causal discovery plays the framework for recursive causal models proposed by
Kiiveri et al. (1984). The recursive causal modelling approach utilizes the notion of conditional inde-
pendence (pioneered by Wright (1934) and Wold (1960)), for system identification purposes. Moreover,
Swanson and Granger (1997) discuss the estimation of impulse response functions based on a causal
approach to residual orthogonalization in vector autoregressions and propose a search algorithm as sta-
tistical mechanism for non-recursive structural models while Hausman and Taylor (1983) discuss the
use of IV regression as an identification mechanism for systems of simultaneous equations
53
In this direction, Silva and Shimizu (2017) examine structural learning using instrumental variables un-
der the non-Gaussianity assumption. Their IV causal discovery method provides a set of potential causal
effects, from which a unique causal discovery is attainable if at least two instrumental variables (under
the same conditioning set) are present in the true underline graph. Moreover, to facilitate computa-
tional flexibility this methodology focuses on completeness, that is, on the characterization of relation
as causals when instrumental variables satisfy the given Graphical criteria, rather than the identifica-
tion of all causal effects induced from faithfulness assumptions. A graphical approach to estimating
high-dimensional VAR models is presented by Fragetta and Melina (2011) and Bertsche et al. (2023).
Example 32. Consider the standard linear structural VAR model as below
p
yt = ∑ A j yt− j + A0 ε t , ut = A0 ε t (4.79)
j=1
We shall demonstrate the above ideas with based on the SVAR process by considering the independence
properties of structural shocks. These are summarized as below:
• If (εt,1 , ..., εt,K ) are mutually independent and non-Gaussian then the error equation is an ICA
model and the matrix A0 can be identified up to a post-multiplication of a generalized permutation
matrix (see, Eriksson and Koivunen (2004), Gourieroux et al. (2017) and Lanne et al. (2017)).
• The recursiveness property of A−1 0 , implies that a recursive causal structure on the contemporane-
ous variables of the structural VAR model hold: if the assumption is true, one can re-order the vari-
ables entering in yt in a "Word Causal Chain" (see, Wold (1960)), so that each variable yt,i causes
yt, j and no variable yt, j causes yt,i for i < j (lower triangular elements) where i, j ∈ {1, ..., K}.
• Since the structural parameter A−10 is a triangular matrix, a convenient way to represent the error
vector ut is with a directed acyclic graph (DAG) (see, discussion in Moneta et al. (2013)).
ut,1 = εt,1
ut,2 = αεt,1 + εt,2
ut,3 = β εt,1 + γεt,2 + εt,2
In other words, when the true DAG is known, then we are able to add zero restrictions (just-identifying or
over-identifying) on the structural parameter A−1
0 and recover the independent shocks εt (see, Cordoni et al.
(2023)). Thus, knowing the causal structure is key for identification in structural econometric models.
Notice that the ICA approach works well in the setting of identifying structural shocks because it pre-
serve the causal ordering and causality of variables.
54
Example 33. Consider the formulation of the multivariate time series model with respect to contempo-
raneous and lagged coefficients
The structural parameter B has a zero diagonal by definition and ε t is a (k × 1) vector of error terms.
The model can be equivalently expressed in the standard SVAR form
where Γ0 = (I − B). Moreover, in the standard SVAR model it is assumed that the covariance matrix
Σε = E (ε t ε t′ ). However, since the variables (y1t , ..., ykt ) are endogenous, these models cannot be directly
estimated without biases. We can thus derive the reduced-form VAR model as below
yt = Γ−1 −1
0 Γ1 yt−1 + ... + Γ0 Γ p yt−p + ε t (4.82)
yt = A1 yt−1 + ... + A p yt−p + ut (4.83)
where ut is a vector of zero-mean white noise processes with covariance matrix such that Σu = E (ut ut′ ),
which in general will not be diagonal. Although the VAR parameters can be easily directly estimated
from the observed data, this is not sufficient to recover the parameters of the SVAR equation, whose
number is larger than the number of parameters in the VAR equation. In other words, the reduced-
form VAR model is only adequate for estimation and forecasting purposes and not sufficient for policy
analysis (see, Moneta et al. (2013)).
Consider the Wold Moving Average (MA) representation of the VAR model such that
∞ ∞ ∞
yt = ∑ Φ j ut− j = ∑ Φ j Γ−1
0 Γ0 ut− j = ∑ Ψ j ε t− j . (4.84)
j=0 j=0 j=0
where the Φ j for j = 0, 1, 2, ... are the MA coefficient matrices and Ψ j for j = 0, 1, 2, ... represent the
impulse responses of the elements of yt to the shocks of ε t− j for j = 0, 1, 2, ...
i i
Φ0 = I, Φi = ∑ A j Φi− j , Ψ0 = Γ−1
0 , Ψi = ∑ A j Ψi− j . (4.85)
j=1 j=1
where A j = 0 for j > p. Therefore, we can see that the impulse response coefficients Ψ j can be obtained
from the reduced-form VAR parameters A j only if we know the SVAR coefficient matrix Γ0 = (I − B).
Since the impulse responses are crucial for policy analysis, it is clear that we need to recover the SVAR
representation for this purpose. The problem is that any invertible unit-diagonal matrix Γ0 is compatible
with the coefficient matrices we obtain by estimating the VAR reduced form. In other words, we shall
find the correct Γ0 which produces the right transformation Γ0 ut = ε t of the VAR error terms ut . The
equivalence between covariance restrictions and IV was first discussed by Hausman and Taylor (1983).
55
Example 34 (Impulse Response Functions). The impulse response functions are calculated considering
the system in levels. Specifically, the forecast error of the h−step forecast of Yt is given by
Yt+h −Yt (h) = ut+h + Φ1 ut+h−1 + ... + Φh−1 ut+1 . (4.86)
Notice that the Φi are obtained from the Ai recursively by the following expression
i
Φi = ∑ Φi− j A j , i = 1, 2, ... (4.87)
j=1
with Φ0 = Ik . Since vt = (I − B0 )ut , then the forecast-error of the h−step ahead forecast of Yt is
Yt+h −Yt (h) = Θ0 vt+h + Θ1 vt+h−1 + ... + Θh−1vt+1 . (4.88)
where Θi = Φi (I − B0 )−1 . Notice that the element ( j, k) of Θi represents the response of the variable y j
to a unit shock in the variable yk , i periods ago.
Example 35 (Uncertainty Shocks). A growing steam of literature considers the identification of uncer-
tainty shocks especially in relation to financial uncertainty and macroeconomic conditions. Since the
work of Bloom (2009) (see, also Jurado et al. (2015) and Baker et al. (2016)) a vast literature studies
on the role of uncertainty shocks for macroeconomic fluctuations. In particular, measuring uncertainty
is important for correctly measuring the impact it has on macroeconomic variables through structural
analysis (e.g., see Trung (2019)). We can show that the identification procedure proposed in the study
of Forni et al. (2023) is equivalent to the Proxy-SVAR methodology (see also Carriero et al. (2023a)).
Proof. We can show that the OLS estimates are identical to those in the study of Mertens and Ravn
(2013) if the number of lags of yt included in the regression of the exogenous instruments zt is equal to
the number of lags of the VAR for yt . Since A(L)yt = µ + εt
Suppose that k < p such that k ∈ {1, ..., p}, then using the stack vectors forms we obtain
′ ′
y p+1−k ε p+1−k
y′ ε ′
Yk = p+2−k , X = 1 Y1 ...Yp , E = p+2−k (4.90)
... ...
y′n−k ′
εn−k
Moreover, let Y = Y0 . Hence, the VAR equation can be written as Y = X A + E (proof to be completed).
Further studies which consider the identification of SVARs under the presence of policy uncertainties
include Mumtaz and Surico (2018) (see, also Ramey (2011)).
56
4.9. Identifying Aggregate Fluctuations
A particular stream of literature focuses on the identifying aggregated fluctuations through sectoral
and firm-specific dynamics. Specifically, Gabaix (2011) proposed a framework for investigating the
granular origins of aggregate fluctuations. In addition, Gabaix and Koijen (2023) illustrate when the
market is sufficiently concentrated, then one can use the collection of idiosyncratic shocks to individual
micro units, at each time period, as an instrument for endogenous aggregate variables. Moreover, a
full-information approach to Granular Instrumental variables is proposed by Baumeister and Hamilton
(2023) while a framework for heterogeneity-robust granular instruments is presented by Qian (2023).
Lastly, Banafti and Lee (2022) establish inferential theory for GIVs in high dimensional settings.
Example 36. Suppose we are willing to assume that the shock to country j′ s demand consist of a global
demand shock fct that affects all consumers across economies in the same way and a purely idiosyncratic
shock ηc jt . In other words, the GIV approach provides a systematic way of constructing instruments
from suitably weighted idiosyncratic shocks (using observational datasets) and use them as instruments
for aggregate endogenous variables. Then, the framework of Banafti and Lee (2022) aims to generalize
the method of granular instrumental variable.
The global market clearing condition is given by ySt = dt , where ySt := S′ y·t = ∑N
i=1 Si yit , where S is the
N
(N × 1) vector of shares that are normalized such that ∑i=1 Si = 1.
Consider the following panel simultaneous equations model with factor error structure
Asymptotic distribution for structural parameters in factor augmented regressions in time series and
panel models is already well-developed in the literature. When both N and T are large, then it can be
shown that the sampling error from estimating the high dimensional precision matrix, the factors, as well
as the instrument is negligible in the asymptotic distribution of the structural parameters. Moreover,
in the case of supply elasticity, the estimator will additionally depend on the estimated (potentially
high dimensional) precision matrix. Therefore, the contribution of Banafti and Lee (2022) to the GIV
literature focuses around relaxing the homogenous loadings assumption which is overly restrictive and
propose constructing an estimate of the instrument using PCA or iterative OLS-PCA methods.
57
4.10. Extensions
Suppose that the k−dimensional process {Y t : t ∈ Z} has the following VARMA representation
p q
Yt − ∑ ΦhYt−h = εt − ∑ Θh εt−h , (4.94)
h=1 h=1
where {εt } is the innovation process, a sequence of independent and identically distributed random
variables with zero-mean and an invertible covariance matrix Σ (see, also Gouriéroux et al. (2020)).
The following representation also holds
! !
p q
j j
Ik − ∑ Φ̃0 Φ̃ j B Yt = Ik − ∑ Φ̃0 Θ̃ j B εt . (4.95)
j=1 j=1
Theorem 4 (see, Mélard et al. (2006)). Let Φ(1) be a (k ×k) matrix such that rank {Φ(1)} = (k −d) = r,
where 0 < d < k. Let P1 be any (k ×d) matrix such that Φ(1)P1 = 0 and P2 be any (k ×r) matrix such that
′
P = [P1 , P2 ] is invertible and the columns of P2 are orthogonal to those of P1 . Let P−1 ≡ Q = Q′1 , Q′2
stands for the inverse of P, with Q1 and Q2 having d and r rows, respectively then the matrices P1Q1
and P2 Q2 are uniquely determined. The singular value decomposition of Φ(1) can be written as Φ(1) =
UDV ′ , where U and V are (k × k) orthogonal matrices.
(Sketch Proof)
• The null space of Φ(1), that is, the set of vectors x such that Φ(1)x = 0, is the space generated by
the columns of the matrix G = [vr+1 , ..., vk ]. Now let P1 and P1∗ by any (k × d) matrices such that
Φ(1)P1 = 0 and Φ∗ (1)P1 = 0.
• Since P1 and P1∗ are full rank matrices, there exist full rank (d × d) matrices α and α ∗ such that
P1 = Gα and P1∗ = Gα ∗ . Hence, we have that P1∗ = P1 α −1 α ∗ .
• Similarly, it can be proved that P2∗ = P2 β −1 α ∗ . Then, using the inverse matrix representations, it
can be proved that P1∗ Q∗1 = P1 Q1 and P2∗ Q∗2 = P2 Q2 , which demonstrates that indeed the matrices
P1 Q1 and P2 Q2 are uniquely defined (see, Mélard et al. (2006)).
Forecasting in VARMA models Modelling and forecasting a large set of macroeconomic variables
can be done using VARMA representations. However, to ensure a robust estimation, the identification of
uniquely parametrized VARMA models requires imposing restrictions on the parameter space to ensure
that the model is uniquely identified (see, Dias and Kapetanios (2018)).
A0Y t = A1Y t−1 + ... + A pY t−p +B0 ut + B1 ut−1 + ... + Bq ut−q (4.96)
| {z } | {z }
AR component MA component
58
Define the lag polynomials such that
A(L) = A0 − A1 L − A2 L2 − ... − A pL p
B(L) = B0 − B1 L − B2 L2 − ... − Bq Lq
We shall say that the model is unique if there is only one pair of stable and invertible polynomials A(L)
and B(L), respectively which satisfies the canonical MA representation (see, also Poskitt (2006))
∞
Yt = A(L)−1 B(L) ≡ Θ(L)ut = ∑ Θ j ut− j .
j=0
Assumption 7. Consider the VARMA(p1 , q1 ) model defined by the stochastic difference equations
! !
p1 p1
I k − ∑ A j L j Xt = I k + ∑ B j L j εt (4.97)
j=1 j=1
lie outside the unit ball in C, then (4.97) is a stable solution (see, Hallin and Paindaveine (2004)).
Remark 17. Notice that in contrast to the reduced form VAR models, setting A0 = B0 = Ik is not a
sufficient condition (only ensures existence) to ensure a unique VARMA representation. The uniqueness
of the VARMA(p, q) representation is guaranteed by imposing restrictions on the pair of stable and
invertible polynomials A(L) and B(L) (see, also Deistler (1983)).
Remark 18. A second important aspect to emphasize is that the framework proposed by Hallin and Paindaveine
(2004) presents with a mathematical rigorous way the construction of an optimal-rank test statistic for
testing the adequacy of the lag orders of a VARMA model. Generally, any test statistics associated with
any parameter vectors from SVAR processes, should have power functions with suitable properties (such
as the monotonicity property) to detect departures from the null hypothesis. Notice that optimal-rank
based tests remain valid under arbitrary elliptically symmetric innovation densities, including those with
infinite variance and heavy-tails. In particular, the authors show that these optimal-rank based tests are
uniformly more powerful than those based on cross-covariances. Moreover, based on the classical LAN
result it allows to derive testing procedures that are locally and asymptotically optimal under a given
innovation density f , based on a non-Gaussian form of cross-covariances.
Suppose that our aim is to construct a statistic for testing the null hypothesis θ = θ 0 against the alterna-
tive θ 6= θ 0 . Choosing p0 < p1 and q0 < q1 allows one to test the adequency of the specified VARMA
coefficients in θ 0 , while contemplating the possibility of possibly higher-order VARMA models.
59
Therefore, this null hypothesis is thus invariant under the group of affine transformation ε t 7→ M ε t if
and only if MAi M −1 = Ai for all i = 1, ...p0 and MB j M −1 = B j for all j = 1, ..., q0, that is, iff each Ai
and each B j commutes with any invertible matrix M, which holds true iff they are proportional to the
(k × k) identify matrix (see, Hallin and Paindaveine (2004)).
(n) (n) ′
Write H (n) (θ 0 , Σ, f ) for the hypothesis under which an observation X (n) := X 1 , ..., X n is gener-
ated by the VARMA(p0 , q0 ) model with parameter value θ 0 satisfying Assumption A and innovation
S S
process satisfying Assumption B. Our objective is to test H (n) (θ 0 ) := Σ f H (n) (θ 0 , Σ, f ) against
S (n) (θ ). Consequently, Σ and f play the role of nuisance parameters; note that the unions, in
θ 6=θ 0 H
the definition of H (n) (θ 0 ), are taken over all possible values of Σ and f .
Let A(L) and B(L) be such that Ai = 0 for i = p0 + 1, ..., p1, and Bi = 0 for i = q0 + 1, ..., q1, and consider
the sequences of linear difference operators
p1
(n)
A(n) (L) := I k − ∑ Ai + n−1/2 γ i Li (4.98)
i=1
p1
(n)
B(n) (L) := I k + ∑ Bi + n−1/2 δ i Li (4.99)
i=1
is such that supn (τ(n) )′ τ(n) < ∞. These operators define a sequence of VARMA models such that
Testing for Invertibility A key assumption for valid identification, estimation and forecasting with
VARMA models (see, Gouriéroux et al. (2020), Sims (2012) and Lütkepohl (2002)), is the invertability
of underline polynomial representations. In particular, Chen et al. (2017) propose a test for invertabil-
ity or fundamentalness of structural vector autoregressive moving average models generated by non-
Gaussian independent and identically distributed structural shocks. Moreover, it can be proved that in
these models under certain regularity conditions the Wold innovations are a martingale difference se-
quence if and only if the structural shocks are fundamental. Thus, the particular representation provides
a mechanism for testing the presence of invertability. In other words, Chen et al. (2017) convert the
statistical problem of polynomial invertability of VARMA processes into a statistical testing problem
for the martingale difference sequence property of the Wold innovations. An MLE approach for non-
invertible ARMA models is studied by Meitz and Saikkonen (2013). Further issues on identification
and estimation of non-invertible SVARMA models are presented by Funovits (2020).
60
4.10.2. Extension 2: Non-Linear Dynamic System
A more challenging task is the econometric identification in nonlinear systems. In addition, for a non-
linear dynamic system excited by a non-Gaussian disturbances, the system equation governing the prob-
ability distribution of the system response contains an infinite number of terms which makes identifica-
tion and estimation considerable complex. Therefore, a pertrubation scheme is proposed to obtain an
approximate solution to the system, thereby the effects of excitation non-Gaussianity can be evaluated
(see, Cai and Lin (1992), Grigoriu (1995) and Sengupta and Kay (1989)).
Consider the unit-root VAR model, A(z)yt = ε t , which crucially depends on A(z)−1 (L), thus we address
the issue of of the inversion of A(z) around a pole z = z0 . Denote with
K
A(z) = ∑ A jz j, A j 6= 0, (4.102)
j=0
be a matrix polynomial of order n and degree K and z0 such that deg A(z) = 0. Therefore,
K
1 ( j)
A(z) = ∑ A (z − z0 ) j (4.103)
j=0 j!
+∞ −1
A(z)−1 = ∑ N j (z − z0 ) j = ∑ N j (z − z0 ) j + M(z). (4.104)
j=−m j=−m
where m is the order of the pole and the first term above represents the principal part while M(z) =
∑+∞ j j
j=0 N (z − z0 ) , represents the regular part. Using the fact that A(z)A(z)
−1 = I
n
! ! !
+∞ K +∞ m+k
1
In = ∑ N j (z − z0 ) j ∑ k! A(k) (z − z0)k = ∑ ∑ N k− j A( j) (z − z0 )k (4.105)
j=−m k=0 k=−m j=0
!
+∞ h
1 ( j)
= ∑ ∑ A N h−m− j (z − z0 )h−m (4.106)
h=0 j=0 j!
Notice that the Unit-Root VAR model is a linear non-homogeneous difference-equation system in matrix
form whose solution can be formally written as below
where the first term is a particular solution of the non-homogeneous equation and the second is the so-
called complementary solution. However, both depend on the operator A−1 (L), and thus on the matrix
A−1 (z) since the algebras of the polynomial functions of the lag operator L and of the complex variable
z are isomorphic. These expressions are useful when deriving statistical tests for invertability.
61
4.10.3. Extension 3: Automatic Inference for Finite Order VARs
Selecting lag length is useful when p is unknown, which can directly impact the estimation of the effects
of structural shocks. Notice that we only discuss lag selection procedures for stationary settings10.
Roughly speaking the selection of the optimal lag length is achieved via the use of information criteria
such as the AIC, SIC and HQC (e.g., see Ivanov and Kilian (2005) and Choi and Kurozumi (2012)).
In this section, we focus on the framework proposed by Kuersteiner (2005) who consider automatic
inference for infinite order vector autoregressions. In particular, infinite order dynamics in cointegration
models has been examined by several studies. Specifically, yt has an infinite order VAR representation
∞ ∞
yt = µ + ∑ ξ j yt− j + vt , C(L) = ∑ C jL j. (4.108)
j=1 j=0
!
∞
where µ = C(1)−1 µy and C(L)−1 = ξ (L) such that ξ (L) = I − ∑ ξ j L j . Therefore, the approximate
j=1
model with VAR coefficient matrices ξ1,p , ..., ξ p,p has the following formulation
Definition 5. The general-to-specific procedure proposed by Ng and Perron (2001) is based on:
(i). p̂n = p if, at significance level α ∈ (0, 1), W (p, p) is the first statistic in the sequence W ( j, j)
such that { j = pmax , ..., 1}, which is significantly different than zero or
(ii). p̂n = 0 if, W ( j, j) is not significantly different from zero for all { j = pmax , ..., 1}, where pmax is
such that pmax /n → 0 and n1/2 ∑∞j=pmax +1 ξ j → 0 as n → ∞.
′
Lemma 4. Let Yt = yt′ − µy′ , yt−1 ′ − µy′ , ..., yt−p+1
′ − µy′ , and let
1 n−1 1 n−1
Γ1,p = ∑ Yt,p (yt+1 − µy )′
(n − p) t=1
Γp = ∑ Yt,pYt,p′
(n − p) t=p
(4.111)
10 A relevant framework for model selection in explosive time series regressions is presented by Tao and Yu (2020).
62
Denote with ξ ( p̂n )′ = Γ′1, p̂n Γ′p̂n such that ξ ( p̂n ) is given by Definition 5 above, then it holds that
Proof. Choose δ such that 0 < δ < 13 , and pick a sequence p∗min such that
!−1
∞
pmin = max p∗min , pmax − n1/2 ∑ ξj . (4.113)
j=p∗min +1
Moreover, since pmin ≤ pmax , and by changing sign of the expression above it follows that
!−1 !
∞ ∞
−1
min pmax − p∗min , n1/2 ∑ ξj = pmax − pmin ≤ n1/2 ∑ ξj (4.114)
∗
j=p +1
j=p∗ +1
min min
such that (pmax − pmin ) → ∞ and (pmax − pmin ) = O p (nδ ). Furthermore, we have that
(n−1) (n−1)
1 ′ 1
Γ̂1, p̂n − Γo1, p̂n = ∑ n Yt, p̂ ( µy − ȳ) + 1 p̂n ⊗ ( µy − ȳ) ∑ (yt+1 − µy )′
(n − p̂n ) t= p̂n (n − p̂ )
n t= p̂n
(n − 1 − p̂n )
+ 1 p̂n ⊗ (µy − ȳ) (yt − µy )′ .
(n − p̂)
1 (n−1) n ′ o
≤ O p (nδ ) ∑ trace E yt+1 − µy yt+1 − µy
(n − pmax )2 t,s=pmin
= O p (n−1+δ ).
63
4.10.4. Extension 4: Locally Trimmed OLS Procedure for VARs
A particular methodology which has been proposed in the literature to deal with the problem of out-
liers when estimating vector autoregression models, is the least-trimmed squares procedure. In par-
ticular, Agullo et al. (2008) proposed a multivariate least-trimmed squares estimator and prove Fisher-
consistency under the assumption of elliptically symmetric error distributions. Moreover, Cavaliere and Georgiev
(2009) consider robust methods for estimation and unit root testing in autoregressions with infrequent
outliers whose number, size, and location can be random and unknown11.
Let {yt |t ∈ Z} be a d−dimensional stationary time series. The vector autoregressive model of order p,
denoted by VAR(p), is given by the following expression
with yt a d−dimensional vector, the intercept parameter B0 a vector in Rd and the slope parameters
B1 , ...., B p are matrices in Rd×d . Throughout the paper, B′ will stand for the transpose of the ma-
trix B. Moreover, the d−dimensional error terms εt are assumed to be independently and identically
distributed with a density function of the form
g (u′ Σu)
fε t = , (4.116)
(detΣ)1/2
with Σ a positive definite matrix, called the scatter matrix, and g a positive function. If the second
moment of εt exists, Σ will be proportional to the covariance matrix of the error terms. Existence of
a second moment, will not be required for the robust estimator. We focus on the unrestricted VAR(p)
model, where no restrictions are put on the parameters B1 , ...., B p . Suppose that the multivariate
time series yt is observed for t = 1, ..., T . Then, the vector autoregression model has the following
multivariate regression representation
yt = B ′ xt + εt (4.117)
′ , ..., y′
′ q
for t ∈ {p + 1, ..., T } with xt = 1, yt−1 t−p ∈ R , where q = pd + 1. Moreover, the matrix B =
(B1 , ...., B p)′ ∈ Rq×d contains all the unknown matrices of coefficients. In matrix notation, we have that
′ ′
X = x p+1 , ..., xT ∈ Rn×q is the matrix of regressors and Y = y p+1 , ..., yT is the matrix of predictants,
where n = T − p is the sample size after excluding the estimation for the lag components of the model.
The classical least squares estimator for the regression parameter B is estimated by
−1
bOLS = X ′ X
B X ′Y , (4.118)
11 In particular, the authors show that standard inferene based on ordinary least squares estimation of an augmented Dickey-
Fuller (ADF) regression may not be reliable, because (a) clusters of outliers may lead to inconsistent estimation of the
autoregressive parameters, and (b) large outliers induce a jump component in the asymptotic distribution of UR test statistics.
64
and the scatter matrix Σ is estimated by the following expression
bOLS = 1 b
′
b
Σ Y − X BOLS Y − X BOLS . (4.119)
n−d
Remark 19. Outliers can affect parameter estimates, model specification and forecasts based on the
selected model. Therefore, a careful examination of the nature of outliers is necessary to ensure robust
statistical inference. Moreover, outliers can be of different nature, the most well known types being
additive outliers and innovational outliers. More specifically, for the vector autoregressive model yt is
an additive outlier if only its own value has been affected by contamination. On the other hand, an outlier
is said to be innovational if the error term εt is contaminated. Innovational outliers will therefore have
an effect on the next observations as well, due to the dynamic structure of the series. Additive outliers
have an isolated effect on the time series, but they still may seriously affect the parameter estimates.
Remark 20. Notice that for the Residual Autocovariance (RA) estimators are considered to be affine
equivariant version of the estimators of LI and Hui (1989). The particular estimators are considered to be
resistant to outliers and therefore we can consider employing them for in the case of constructing a robust
estimator to outliers for the Vector Autoregressive Model. For this purpose, we use the Multivariate
Least Trimmed Squares (MLTS) estimator proposed by Agullo et al. (2008). The specific estimator is
defined by minimizing a trimmed sum of squared Machalanobis distances, and can be computed by a
fast algorithm. Moreover, the procedure also provides a natural estimator for the scatter matrix of the
residuals, which can then be used for model selection criteria.
The Multivariate Least Trimmed Squares Estimator The unknown parameters of the VAR(p) will
be estimated via the multivariate regression model. To allow for the robust estimation of the model
under the presence of outliers the Multivariate Least Trimmed Squares (MLTS) estimator is employed,
which is based on the idea of the Minimum Covariance Determinant estimator. The way that the MLTS
estimator works is that it selects the subset of h observations which minimize the determinant of the
covariance matrix of residuals based on a least squares fit from the particular subsample.
Consider the dataset Z = {(xt , yt ) ,t = p + 1, ..., T } ⊂ Rd+q . We denote with H the collection of all
subsets of size h, such that:
H = H ⊂ {p + 1, ...., T } | N {H} = h (4.120)
Then, for any subset H ∈ H , we denote with B bOLS the classical least squares estimator based on the
observations of the particular subset given by the following expression
−1
bOLS (H) = X ′ XH
B XH′ YH , (4.121)
H
where XH and YH are submatrices of X and Y respectively, consisting of the rows of X and Y that
corresponds to the subsample H.
65
Then, the corresponding scatter matrix estimator computed from this subset is constructed as below
bOLS = 1 b
′
b
Σ YH − XH BOLS (H) YH − XH BOLS (H) . (4.122)
h−d
The MLTS estimator is now defined as below
h i
bMLTS (k) = B
B bOLS ĥ , where Ĥ = argmin det ΣbOLS (H) (4.123)
H∈H
Moreover, the associated estimator of the scatter matrix of the error terms is given by
bMLTS (H) = cα Σ
Σ bOLS Ĥ (4.124)
where cα is the correction factor to obtain consistent estimation of Σ for the error distribution of the
model, and α the trimming proportion for the MLTS estimator, that is, α ≈ 1 − nh . In the case of
multivariate error terms it has been shown that cα = (1 − α ) Fχ 2 (qα ) , where Fχ 2 is the cumulative
d+2 d+2
distribution function of a χ 2 distribution with q degrees of freedom, and qα = χq,1−
2
α is the upper α
quantile of this distribution.
Equivalent characterizations of the MLTS estimator are given by Agullo et al. (2008). The particular
study shows that any B e ∈ Rd×q minimizing the sum of the h smallest squared Mahlanobis distance of
its residuals, subject to the condition that det [Σ] = 1 is a solution of 4.123. More formally,
( )
h
e = argmin
B 2
∑ ds:n (B, Σ) such that |Σ| = 1 (4.125)
(B,Σ) j=1
where d1:n (B, Σ) , ..., dn:n (B, Σ) is the ordered sequence of the residual Mahalanobis distances is
′ 1/2
ds (B, Σ) = yt − B ′ xt Σ−1 yt − B ′ xt (4.126)
for some B ∈ Rd×q . Notice that the MLTS estimator minimizes the sum of the h smallest squared
distances of its residuals, and is therefore the multivariate extension of the Least Trimmed Squares (LTS)
estimator of Rousseeuw (1984). The particular MLTS estimator has been found to have low efficiency
so the literature has proposed the reweighted version of the estimator to improve the performance of the
MLTS estimator. Therefore, since the efficiency of the MLTS estimator is quite low, the re-weighted
version of the particular estimator is employed to improve the performance of the MLTS.
The Reweighted Multivariate Least Trimmed Squares (RMLTS) estimates are defined as
bRMLTS = B
B bOLS (J) and Σ bRMLTS = cδ ΣbOLS (J) (4.127)
n o
2 b b
J = j ∈ {1, ..., n}|d j BMLT S , ΣMLT S ≤ qδ (4.128)
2
and qδ = χq,1−δ.
66
Since outliers tend to have large residuals withrespect to the initial robust MLTS estimator, then it
2
implies a large residual Mahalanobis distance d j B bMLTS , Σ
bMLTS . If the later is above the critical value
qδ , then the observation is flagged as an outlier. The final RMLTS is then based on those observations
not having been detected as outliers. For instance, one can set δ = 0.01 and take as trimming proportion
for the initial MLTS estimator α = 25%.
Determining the autoregressive order To select the order p of a vector autoregressive model, infor-
mation criteria are computed for several values of p and an optimal order is selected by minimizing the
criterion. Most information criteria are in terms of the value of the log likelihood ℓk of the VAR(p)
model. Using the model assumption for the distribution of the error terms, we obtain that
T n
ℓk = ∑ g εt Σ−1 εt − log det Σ , (4.129)
t=k+1 2
with n = T − k. When error terms are multivariate normal the above leads to the following expression
n np 1 T
ℓk = −
2
det Σ − log (2π ) −
2 ∑ εt′ Σ−1 εt
2 t=k+1
(4.130)
The log likelihood will depend on the autoregressive order via the estimate of the covariance matrix of
the residuals. In particular, for the classical least squares estimator we have that
T
bOLS = 1
Σ ∑ bεt (k)bεt′(k)
n − p t=k+1
(4.131)
where bεt′ (k) are the residuals corresponding to the VAR(p) model. Using trace properties, the last term
of the log-likelihood function corresponds to (n − p)p/2 for the OLS estimator. In order to prevent
that outliers might affect the optimal selection of the information criteria, we estimate Σ by the RMLTS
estimator given by the following expression:
bRMLTS = cδ
εt (k)′,
m(k) − p ∑
Σ t ∈ J(k)b
εt (k)b (4.132)
with J(k) as above and m(k) the number of elements in J(k). In this case, the last term of the log-
likelihood function equals now (m(k) − p) p (2cδ ).
Open Problems Some possible extensions of the particular approach is to propose a framework which
employes the LTS method that accounts for outliers in multivariate time series models, specifically to
the setting of structural vector autoregressions.
67
4.10.5. Extension 5: Time-Varying VAR Models
The identification and estimation of time-varying VAR models is useful when modelling dynamically
macroeconomic events such as monetary policy regimes (see, Primiceri (2005), Koop and Korobilis
(2013)) and Chan et al. (2016) as well as time-varying instrumental variable estimation as in Kapetanios et al.
(2019), Petrova (2019) and Giraitis et al. (2021). A relevant setting is the framework of factor aug-
mented SVAR models (e.g., see Eickmeier et al. (2015)).
Financial Connectedness
The study of network topology and financial connectedness is used for modelling systemic risk, spillover
effects and financial contagion using the VAR framework (see, discussion in Katsouris (2023b)). Esti-
mating systemic risk measures implies identifying shocks which are transmitted via system-wide con-
nectedness. In particular, Diebold and Yilmaz (2012); Diebold and Yılmaz (2014) proposed a frame-
work for modelling the network topology via the forecast error variance decomposition (FEVD) of
an estimated VAR process, as a transformation method that summarizes connectedness between a set
of institutions. Thus this approach allows to construct measures of directional connectedness which
are analogous to bilateral imports and exports for each of a set of N countries. Recent implementa-
tions are presented by Barigozzi and Hallin (2017), Baruník and Křehlík (2018), Demirer et al. (2018),
Korobilis and Yilmaz (2018), Geraci and Gnabo (2018), Rehman et al. (2018), Yoon et al. (2019), Barigozzi et al.
(2021a), Laborda and Olmo (2021), Mensi et al. (2023) and Baruník and Ellington (2024).
Consider, an N dimensional covariance-stationary series Yt = (y1,t , ..., yN,t ), which is generated from a
VAR(p) process for t = 1, ..., T . Then, based on the Wold decomposition theorem we have that
∞
Φ(L)Y t = εt , with Φ(L) = ∑ ΦhLh ,
h=0
such that |Φ(z)| lie outside the unit-circle, then the VAR process has the MA(∞) representation expressed
as Yt = Θ(L)εt , where Θ(L) is an (NxN) matrix of infinite lag polynomials and ε t is considered to
be a white-noise generated by the covariance matrix Σ. Therefore, the H−step generalized variance
gH
decomposition12 matrix DgH = [di j ] is given by
σi−1 H−1 ′
j ∑h=0 (ei Θh Σe j )
2
di j = ′ , (4.133)
∑H−1 ′
h=0 (ei Θh ΣΘh ei )
′
where e j are selection matrices e.g., e j = (0, ...1, ...0) , Θh is the coefficient matrix evaluated at the h-
lagged shock vector and σi j is the jth diagonal element of Σ. Modelling the asymmetric responses of
the economy to monetary policy shocks over the business cycle due to the impact of the monetary trans-
mission mechanism (e.g., see Eickmeier et al. (2015), Morley and Piger (2012) and Veldkamp (2005)),
requires to employ specifications which allow for time-varying effects (see, Paul (2020)).
12 See
also, Lanne and Nyberg (2016) who employ a GFEV decomposition for linear and nonlinear multivariate models.
Moreover, Katsouris (2023a) develops methods for the identification and estimation of financial networks.
68
5. Dynamic Causal Effects
The two main types of identification presented so far in the literature consists of point-identification
and set-identification. Specifically, point-identified SVARs are interpreted as achieving identification
through internal instruments. Consequently, in these models, structural shocks are the interventions
of interest, and thus the goal is to estimate the dynamic causal effect of these shocks on macroeco-
nomic outcomes. In contrast, modern microeconometric identification stategies rely heavily on external
sources of variation that provide quasi-experiemnts to identify causal effects. Such external variation
might be found, for example, in firm-specific characteristics that introduce as-if randomness in the vari-
able of interest (the treatment). The use of such external instruments in microeconometrics has proven
highly productive and has yielded compelling estimates of causal effects. Therefore, the use of external
instruments has opened a new and rapidly growing research avenue in macroeconometrics, in which
credible identification is obtained using as-if random variation in the shock of interest that is distinct
from - external to - the macroeconomic shocks hitting the economy (Stock and Watson (2018)).
Following the paper of Stock and Watson (2018), a starting point for formulating the theory on identi-
fication and estimation of dynamic causal effects in macroeconomics, is that the expected difference in
outcomes between the treatment and control groups in a randomized experiment with a binary treatment
is the average treatment effect. Roughly speaking, if a binary treatment X is randomly assigned, then
all other determinants of Y are independent of X , which implies that the (average) treatment effect is
E Y |X = 1 − E Y |X = 0 (5.1)
In the linear model Y = α + β X + u, where β is the treatment effect, random assignment implies that
E (u|X ) = 0 so that the population regression coefficient is the treatment effect. When randomization
is conditional on covariates W , then the treatment effect for an individual with covariates W = w is
estimated by the outcome of a random experiment on a group of subjects with the same value of W ,
E Y |X = 1,W = w − E Y |X = 0,W = w . (5.2)
Example 37. Usually the path of observed macroeconomic variables arising from current and past
shocks and measurement error are collected into the (m × 1), εt error vector εt . Therefore, the (n × 1)
vector of macroeconomic variables Yt can be written in terms of current and past innovation terms
Yt = Θ(L)εt , (5.3)
69
Then the variance matrix of the error terms Σ = E[ε t ε t ] is assumed to be positive-definite to ensure the
existence of a non ill-conditioned covariance matrix. Moreover these disturbance terms (shocks) are
assumed to be mutually uncorrelated. The expression Yt = Θ(L)εt , corresponds to the structural moving
average representation of Yt . Specifically, the coefficients of Θ(L) are the structural impulse response
functions, which are the dynamic causal effects of the shocks (see, also Bacchiocchi et al. (2018)).
Under the null of invertability, the following SVAR representation applies:
Therefore, under invertability, the structural moving average is Y t = C(L)Θ0 ε t , where C(L) = A(L)−1 ,
such that Θ(L) = C(L)Θ0 . The null and alternative hypotheses are then
Recall that the SVAR can be also written in state-space form as below:
Y t = BX t (5.6)
X t = AX t−1 + Gε t , (5.7)
⊤
where X t = Y t⊤ ,Y t−1
⊤ , ....,Y ⊤
t−p+1 , such that A is the companion matrix and B = I n 0 · · · 0 is a
selection matrix. Then, the local projection regression equation is written as below:
Remark 21. The study of Stock and Watson (2018) verifies important insights regarding the identifica-
tion of dynamic causal effects. In particular, under the assumption of Gaussian errors, every invertible
model has multiple observationally equivalent non-invertible representations, which imply that to iden-
tify a unique representation some external information regarding the system is required. If we assume
that the structural shocks are independent and non-Gaussian then using information from higher-order
restrictions the causal structure of the system can be identified. In practice, external instruments can be
employed to estimate dynamic causal effects directly without using an indirect VAR identification step.
70
In particular, time variation in the parameters of the model allows to: (i) capture changes in the behaviour
of individuals as they respond to public health conditions and (ii) to include shifts in the transmission
and clinical outcomes of the pandemic.
We write our SVAR model such that
In addition, we assume that the vector ε t , conditional on past information and the initial conditions
y0 , ..., y1−p , is Gaussian with mean zero and covariance matrix I n . Without loss of generality, we assume
that the first equation of the SVAR characterizes the policy rule. This implies that
• α +,1 denotes the first column of A+ for ℓ ∈ {0, ..., p} and αs,i j denotes the (i, j) entry of As , where
s ∈ {0, +}, and describes the systematic component of the policy rule.
Restricting the systematic component of the policy rule is equivalent to restricting αs,i j and identifying
a policy shock that we call the stringency shock. Recall that the dynamic causal effect of a unit inter-
vention in ε j,t ∈ εt on zt+h is given by E[zt+h |ε j,t = 1; εt−1 , ....] − E[zt+h|ε j,t = 0; εt−1 , ....] . Then, the
structural impulse response functions can be derived by ∂ zt+h /∂ ε j,t for h = 0, 1, 2, ....
Example 39 (ECB monetary policy and bank default risk, see Soenen and Vander Vennet (2022)).
In this example, the primary focus is the impact of ECB monetary policy on bank risk. There is an on-
going debate on the effect of accommodative monetary policy on bank risk taking and financial stability
in general. Specifically, one concern relates to the potential increase in risk taking and the possible
under-pricing of risk. The relevant question is whether or not monetary policy causes excessive risk
taking by banks, since this could hamper financial stability which can be reflected in higher bank CDS
spreads. Next, regarding the choice of the variable of interest: we cannot use the policy rate because
of the zero lower bound constraint and, similarly, we cannot use the ECB balance sheet because some
important monetary policy measures did not affect the balance sheet.
71
Therefore, when assessing the causal impact of monetary policy, we decide to employ a structural VAR
because incorporating a broad set of financial market indicators allows us not only to identify actual
ECB monetary policy decisions, but also to capture anticipation effects and instances in which financial
markets judge that monetary policy actions were insufficient, given the prevailing market conditions.
In other words, we can estimate a time series of exogeneous monetary policy shocks by modelling a
set of relevant financial market variables in a structural VAR model which is given by the following
econometric specification:
Remark 22. A macroeconomic puzzle relevant to the optimal economic policy decision-making with
respect to macro-prudential restructuring and debt collection is the aspect of the tax policy conducted
over the Business Cycle. In particular, it is well known that government spending has typically been
procyclical in developing economics but exhibit acyclical or countercyclical behaviour in industrial
economies. Moreover, an interesting avenue of further research is the counterfactual distribution of the
cyclical behaviour of tax rates, as opposed to tax revenues which are endogenous to the business cycle
and hence cannot shed light on the cyclicality of tax policy (see, Vegh and Vuletin (2015)).
Therefore the identification and estimation of dynamic treatment effects allows to investigate the VAT
behaviour of firms in relation to imposed by the government tax rates. The intuition goes as follows.
Governmental taxation is the pillar of fiscal policy and decisions on the optimal tax policy are usually
taken in relation to their effectiveness on various macroeconomic outcomes, such as output fluctuations
as well as structural changes in agent’s behaviour over the business cycle (aggregate fluctuations). From
the econometrics perspective to facilitate inference requires to derive the asymptotic behaviour of the
dynamic treatment effect estimator across a panel of economics which have undergone a period of fiscal
adaption to new tax policies. In other words, one would be interested to obtain the asymptotic behaviour
of these estimators based on the trajectory of the control against the experimental groups in relation to a
common taxation policy as well as to obtain counterfactual outcomes when a control group has adapted
an updated fiscal policy across the business cycle. In this direction, we are interested to incorporate
tax policy variables such as taxation revenues which are often considered as policy outcomes. Further-
more, due to the fact that these variables might be positive correlated during the business cycle, suitable
instrumental variables that can ensure consistent model estimation and inference procedures can be em-
ployed. In particular, Vegh and Vuletin (2015) use the VAT rates across economic during the business
cycle as a valid instrumental variable. Moreover, the authors use a novel tax database which combines
both corporate and personal income tax rates as well as data on value-added taxes. A relevant study
which discusses the effects of global tax reform is presented by Gomez Cram and Olbert (2023).
72
5.2. Counterfactual Analysis
In particular, the parameter of interest is the ATE := E [Y1 −Y0 ]. Moreover, the Conditional Average
Treatment Effect (CATE) defined by
captures how the average treatment effect changes across sub-populations with covariate value X = x.
Furthermore, under the unconfoundness assumption τ (x) can be nonparametrically identified by
τ (x) := E [Y |D = 1, X = x] − E [Y |D = 0, X = x] (5.13)
(ii) X has a compact support Supp(X ) and there exists c > 0 such that c < Π(x) < 1 − c for all
x ∈ Supp(X ).
Remark 23. Notice that the above assumptions are fairly standard in the causal inference literature and
we assume that they hold throughout the remaining of the paper unless otherwise states. In particular,
Assumption 1 (i), corresponds to the unconfoundedness condition while Assumption 1 (ii) corresponds
to the common support condition for the set of covariates in the model.
Consider the k−step ahead forecast of Yt given the information available up to time t ∗ which is given by
the following expression
b Yt ∗ +k (0)|Ft ∗ ≡ X ′∗ βb + E
Yt ∗ +k|t ∗ (0) := E b Zt ∗ +k |Ft ∗ . (5.14)
t +k
Notice that by definition, Yt ∗ +k (0) is the potential outcome expected at time t ∗ +k in case the intervention
does not occur. Therefore, an estimator of the point causal effect can be defined as
τbt ∗ +k := Yt ∗ +k (1) −Yt ∗ +k|t ∗ (0) = τt ∗ +k + Xt ∗ +k β − βb + Zt ∗ +k − E
b Zt ∗ +k |Ft ∗
Let D ∈ {0, 1} be a binary treatment indicator and Yd be the potential outcomes for d = 0, 1 in the
status quo environment (see, also Hsu et al. (2022)). In particular, Y1 is the outcome if an individual
is exogenously assigned to the treatement (D = 1) and Y0 is the outcome in the absence of treatement
(D = 0). Then, the actual observed outcome is
73
We observe a d−dimensional vector of pretreatment covariates X = (X1 , ..., Xk) in the status quo envi-
ronment and X ⋆ = X1⋆ , ..., Xk⋆ in the counterfactual environment which is of the same dimension as the
set of covariates X .
The corresponding treatment indicator, outcome, and potential outcomes in the counterfactual environ-
ment are denoted as D∗ ,Y ∗ and Yd∗ for d = 0, 1, respectively with
However, notice that since the treatment has not been implemented in the counterfactual environment
yet, neither D∗ nor Y ∗ is observed in our model (see, Hsu et al. (2022)). We adopt the former approach
such that a closed-form solution exists
1 n h i
δb∗ = ∑ Eb Y1 |X = X ∗
j − b
E Y0 |X = X ∗
j (5.17)
n j=1
b Yd |X = X ∗ is the Nadaraya-Watson estimator such that
where E j
∑ Yi1 {Di = d} Kx,h (Xi − x)
b Yd |X = X = ∗ i=1
E j n (5.18)
∑ 1 {Di = d} Kx,h (Xi − x)
i=1
The asymptotic properties for δb∗ can be derived under weaker regularity conditions stated below. Fur-
thermore based on regularity conditions it can be proved that
√ b∗
d
n δ − δ ∗ → N 0, σδ2∗ . (5.19)
Therefore, compared to the semiparametric efficiency bound of the ATE estimator we have that
Var(Y1 |X ) Var(Y0 |X ) 2
E + + E Y1 −Y0 |X − E (Y1 −Y0 ) . (5.20)
p(X ) 1 − p(X )
74
5.3. Time Series Experiments and Causal Effects
Time series causal effects are defined as a comparison between the potential outcomes at a fixed point in
time. Thus, the primary object of interest is the temporal average of these causal effects. Therefore, the
only source of randomness in our formulation is the randomization of the treatment. Once we begin the
experiment, the randomization reveals a particular outcome path; however, this does not change the set
of potential outcomes, and therefore, does not alter what we would have observed had we administered
a different treatment. Thus, in this sense, we are treating the potential outcomes as predetermined, and
the only source of randomness comes from which path we observe.
Definition 6 (General Causal Effects). For paths w1:t and w′1:t , the t−th causal effect is given by
τ (w1:t , w′1:t ) = E Yit (w1:t ) −Yit (w′1:t ) . (5.21)
Then, the temporal average treatment effect of the paths w1:T and w′1:T is
1 T
τ (w1:t , w′1:t ) = ∑ τt w1:t , w′1:T . (5.22)
T t=1
Moreover, we define the potential outcomes just intervening on treatment the last j periods as
Yit (xt− j:t ) = Yit Xi,1:t− j−1, xt− j:t (5.23)
In other words, this "marginal" potential outcome represents the potential or counterfactual level of
regime A in country i, if we let welfare spending run its natural course up to t − j − 1 and just set
the last j lags of spending to xt− j:t . Thus, we can define an important quantity of interest, that is, the
contemporaneous effect of treatment (CET) of Xit on Yit such that:
τc (t) := E Yit Xi,1:t−1, 1 −Yit Xi,1:t−1 , 0 ≡ E Yit (1) −Yit (0) .
Furthermore, following Bojinov and Shephard (2019) specifically in longitudinal studies and survey
studies with a follow-up wave of cross-sectional datasets, we are particularly interested to investigate
the treatment effect of survey participants as well as their membership in certain endogenously gen-
erated groups. This implies that we are modelling the presence of time-varying exposure to dynamic
causal estimands. Therefore, to construct a feasible structural econometric framework we consider the
treatment and potential outcome path with time-varying effects. At each time step, t = 1, ..., T , we as-
sume that each individual (sampling unit), is exposed to either treatment, Wt = 1, or control, Wt = 0, and
then we measure the corresponding outcome variable of interest.
75
Although we focus on binary treatments, our results can be generalized to multiple treatments as well.
Thus, the random "treatment path" is given by W1:t = W1 , ...,Wt . As a result, the potential path for the
treatment path w1:t can corresponds to the following sequence
Y1:t (w1:t ) = Y1 (w1 ),Y2 (w1:2 ), ...,Y2(w1:t ) . (5.24)
Following the framework proposed by Bojinov and Shephard (2019), to access whether the treatment
has a statistically significant effect, we consider the sharp null of no temporal causal effects
In particular, this will null hypothesis will be tested against a portmanteu alternative. Thus, invoking the
sharp null hypothesis implies that Yt (wobs ′ ′
1:t ) = Yt (w1:t ) for all w1:t . Moreover, the particular formulation
of the null hypothesis means "the null of no temporal causality at lag p ≥ 0 of the treatment on the
outcome" holds, such that H0,p : τt,p = 0, for all t = 1, 2, ..., T.
In the case where p = 0, then τbt,0 estimation error can be written as below
1 {Wt = 1} 1 {Wt = 0}
ut,0 = − Yt (wobs
1:t ) (5.26)
pt (1) pt (0)
Yt (wobs
1:t )
ER ut,0 |FT,t−1 = 0 and Var ut,0 |FT,t−1 = . (5.27)
pt (1)pt (0)
In particular, whenever pt (1) = pt (0) the estimator of the upper bound to the variance for the contem-
poraneous causal effect, is equal to the variance under the sharp null of no treatment effect. Then, the p
lagged error variance VarR ut−p,p|FT,t−p−1 equals to
2 1 1 1
ν p2 = Yt (wobs
1:t ) 2p ∑ + . (5.28)
2 w∈{0,1} p
pt (1, w) pt (0, w)
76
Consider the conditional independence property which implies that Yit (1),Yit (0) ⊥ X i |U i . By the law
of iterated expectations this implies that
E Yit (x)|Ci = 1 = E E Yit (x)|U i ,Ci = 1 Ci = 1 (5.30)
for x = 0, 1.
According to Vermeulen and Vansteelandt (2015) the proposed approach may more generally lend itself
better to small-sample inference. For instance, suppose that interest lies in the marginal causal effect
τ = E[Y (1) − Y (0)]. Because the proposed estimation strategy does not require acknowledging the
uncertainty of the estimated nuisance parameters (up to first order), we foresee that it may potentially
lend itself better to randomization inference. How such randomization inference could be accomplished
and how it performs in small to large samples is indeed relevant questions for further investigation.
Consider i.i.d data {Z i = (Yi , Ai , X i ) , i = 1, ..., n}, where Yi is the outcome of interest, Ai is a dichotomous
treatment taking values zero and one and X i is a sufficient set of covariates to control for confounding
of the treatment effect, in the sence that Y (a) ⊥ A|X for a ∈ {0, 1}. Specifically, Y (a) denotes the
counterfactual outcome for treatment level a ∈ {0, 1}, which is linked to the observed data through the
consistency assumption (i.e., Y (a) = Y iff A = a).
Assumption 9 (Covariates-Treatment Independence). Let X j,t be a vector of covariates that are predic-
tive of the outcome of unit j, for all t ∈ {t ∗ + 1, ..., T } such covariates are not affected by the intervention,
such that, X j,t (1) = X j,t (0).
Definition 7. Let t ∗ be the time of the intervention and define with k ∈ {1, ..., K} such that t ∗ + K = T .
For any k, the point, cumulative and temporal average causal effects on unit j at time t ∗ + k are defined,
respectively, such that
77
Notice that the selection on observables assumption says that policies are independent of potential out-
comes after appropriate conditioning. Moreover, the conditioning set includes the unknown error vector
of the model, εt , such that Yt,1 (d),Yt,2(d), ..., ⊥ εt |zt . This is because, conditional on the set of instru-
ments zt , randomness in Dt is due exclusively due to randomness induced from the unknown εt vector.
Therefore, in order to convert the above assumption into a testable hypothesis which has an identifiable
parameter space, we consider the sharp null hypothesis such that Yt, j (d ′ ) = Yt, j (d) = Yt+ j .
Then, if we substitute the observed with the potential outcomes we obtain the following testable condi-
tional independence assumption:
Notice that Dt is not necessarily a binary treatment status but can be a set of different taxation policies
which implies that we have a continuum of treatment status.
Remark 24. A relevant example is presented by Menchetti et al. (2023) who consider estimating the
effect of policy interventions in observational time series settings in the absence of untreated units, by
combining counterfactual outcomes and ARIMA models for prediction purposes. Specifically, when
focusing on observational time series data, commonly used methodologies for policy evaluation within
the Rubin Causal Model (RCM) framework include Differences-in-Differences (DiD) and synthetic
control methods. However, a main drawback of these two approaches is that they require the presence
of controls that did not experience the treatment, and therefore finding untreated units is a challenging
task (Menchetti et al., 2023). Moreover, the particular method is not guaranteed to reveal relevant causal
effects especially due to the presence of unobserved factors. For example, if our main goal is to estimate
the causal effect of a new intervention policy such as due to the adoption of new tax rates, as well as
its indirect impact on a common fiscal policy across similar economies (e.g., VAT policy), then a causal
time series modelling approach can be employed to obtain the causal estimands of interest. However,
one may find that these causal effects have no negative or positive impact on the spending trajectory of
similar economies that did not adopt new fiscal policies, suggesting that unobserved factors may drive
spending decision making more than taxation policy. Other relevant counterfactual frameworks can be
constructed (see, Kilian and Lewis (2011)), around issues from the international finance literature such
as the modelling sovereign default risk (e.g., see Hatchondo et al. (2016)).
78
5.5. The Augmented Synthetic Control Method
The synthetic control method (SCM) is a popular approach for estimating the impact of a treatment on a
single unit in panel data settings (see, Ben-Michael et al. (2021)). In other words is a weighted average
of control units that balances the treated units pre-treatment outcomes and other covariates as closely as
possible. We define the treated potential outcome as Yit = Yit (0) + τit , where the treatment effects τit are
fixed parameters. Since the first unit is treated, the key estimand of interest is
Remark 25. Notice that pre-treatment outcomes serve as covariates in SCM, we use Xit , for t ≤ T0 , to
represent pre-treatment outcomes.
• We are interested on a potential outcome system as a foundational framework for analyzing dy-
namic causal effects on outcomes in observational time series settings. We consider settings in
which there is a single unit observed over time. At each time period t ≥ 1, the unit receives a
vector of assignments Wt , and an associated vector of outcomes Yt are generated. The outcomes
are causally related to the assignments through a potential outcome process, which is a stochastic
process that describes what would be observed along counterfactual assignment paths.
• A dynamic causal effect is generally defined as the comparison of the potential outcome process
along different assignment paths at a fixed point in time.
Assumption 11 (Assignment and Potential Outcome). The assignment process {Wt }t≥1 satisfies
Wt ∈ W := ×dk=1
w
Wk (5.38)
The potential outcome process, is for any deterministic sequence {ws }s≥1 with ws ∈ W for all s ≥ 1,
Yt {ws }s≥1 t≥1 , where the time−t potential outcomes satisfies Yt {ws }s≥1 ∈ Y ⊂ Rdy .
Remark 26. The simplest case is when the assignment is scalar and binary W = {0, 1}, in which
case Wt = 1 corresponds to "treatment" and Wt = 0 is "control". Moreover, the potential outcome
Yt {ws }s≥1 may depend on future assignments {ws }s≥t+1 . The next assumption considers the de-
pendence structure we allow in the proposed econometric framework, restricting the potential outcome
to only depend on past and contemporaneous assignments.
79
Assumption 12 (Non-anticipating Potential Outcomes). For each t ≥ 1, and all deterministic {wt }t≥1
and {wt′ }t≥1 with wt , wt′ ∈ W ,
Yt w1:t , {ws }s≥t+1 = Yt w1:t , w′s s≥t+1
almost surely. (5.39)
Remark 27. The above assumption is a stochastic process analogue of non-interference. Furthermore,
it still allows for rich dependence on past and contemporaneous assignments. Under Assumption 2, we
drop references to the future assignmnets in the potential outcome process and write
Yt {ws }s≥1 t≥1
= Yt w1:t t≥1
(5.40)
Then, the set Yt w1:t : w1:t ∈ W t collects all the potential outcomes at time t.
n o
Assumption 13 (Output). The output is {Wt ,Yt }t≥1 = Wt ,Yt W1:t t≥1 . The {Yt }t≥1 is called the
outcome process.
Furthermore, we assume that the assignmemnt process is sequentially probabilistic, meaning that any
assignment vector may be realized with positive probability at time t given the history of the observable
stochastic process up to time t − 1. Let {Ft }t≥1 denote the natural filtration generated by the realized
{wt , yt }t≥1 .
Assumption 14 (Sequentially probabilistic assignment process). The assignment process satisfies 0 <
P Wt = w Ft−1 < 1 with probability one for all w ∈ W . Moreover, the probabilities are determined
by a filtered probability space of Wt , Yt (w1:t ), w1:t ∈ W t t≥1 .
Remark 28. Various further applications of the aforementioned frameworks can be considered starting
with the pioneered work of Aigner and Balestra (1988) who consider the optimal experimental design
for error correction models, the econometric framework of De Chaisemartin and d’Haultfoeuille (2020)
which corresponds to the two-way fixed effects estimators with heterogeneous treatment effects as well
as the novel framework of Lewis and Syrgkanis (2020) which corresponds to a time series setting based
on machine learning and statistical learning techniques. However, we consider all aforementioned set-
tings as more advanced topics and beyond the scope of our current discussion.
80
5.6. Panel Data Regression for Dynamic Causal Effects
Following the framework proposed by Bojinov et al. (2021) we consider a design-based framework
which provides a generalization of the finite population literature in cross-sectional causal inference and
time series experiments to panel experiments.
1 N
τ̄.t (w, w̃; p) := ∑ τi,t (w, w̃; p).
N i=1
(5.41)
T
1
τ̄i. (w, w̃; p) := ∑
T − p t=p+1
τi,t (w, w̃; p). (5.42)
T N
1
τ (w, w̃; p) := ∑ ∑
N(T − p) t=p+1 i=1
τi,t (w, w̃; p). (5.43)
Example 40. Suppose that we aim to generate the potential outcomes for the panel experiment using
an autoregressive model formulated as below
Yi,t = φi,t,1Yi,t−1 (wi,1:t−1 ) + ... + φi,t,t−1Yi,1 (wi,1 ) + βi,t,0wi,t + ... + βi,t,t−1wi,1 + εi,t , ∀ t > 1. (5.44)
In a more compact form the above specification can be written as Yi,1 (wi,1 ) = βi,1,0wi,1 + εi,1 , where
φ , for s = 1 β , for s = 0
φi,t,s = and βi,t,s (5.45)
0, for s > 1. 0, for s > 0.
Within a simulation setting we can vary the choice of φ , which governs the persistence of the process,
and β , which governs the size of the contemporaneous causal effects. Varying degrees of persistence in
conjecture with the size of contemporaneous causal effects might cause size distortions for values close
to the boundary of the parameter space. A robust method that accommodates for these the presence of
these features in the aforementioned econometric environment is an interesting avenue for further re-
search. In addition, we can vary the probability of treatment pi,t−p (w) = p(w) as well as the distribution
of the errors εi,t , which can be either sampled from a standard normal or a Cauchy distribution.
81
5.6.2. Identifying Dynamic Causal Effects with Panel Data: Applications
Example 41 (see, Goes (2016)). Consider the following panel SVAR(1) model defined as below
where yi,t ≡ [ci,t , ki,t ]′ is a bi-dimensional vector of stacked endogenous variables, such that ci,t is the log
of GDP per capita and ki,t is the proxy for institutional quality, f i is a diagonal matrix of time-invariant
individual-specific intercepts. Moreover, A(L) = ∑ pj=0 A j L j is a polynomial of lagged coefficients, A j is
a matrix of coefficients, and ei,t is a vector of stacked residuals, and B is a matrix of contemporaneous
coefficients. However, since f i is correlated to the error terms, estimation through OLS leads to biased
coefficients. As proposed by Baltagi (2008), a strategy to obtain consistent parameters and eliminate
individual fixed-effects when N is large and T is fixed, is to apply first-differencing and use lagged
instruments. We consider the GMM/IV technique using a system of m = 2 equations. Each equation
in the system has the first difference of an endogenous variable on the left hand side, p lagged first
differences of all m endogenous variables on the right hand side, and no constant.
p p
j j
∆y1,i,t = ∑ γ11 ∆y1,i,t− j + ... + ∑ γ1m ∆ym,i,t− j + e1,i,t (5.47)
j=1 j=1
.. .
. = ..
p p
j j
∆ym,i,t = ∑ γm1∆y1,i,t− j + ... + ∑ γmm ∆ym,i,t− j + em,i,t (5.48)
j=1 j=1
Moreover, the model has an equivalent vector moving average (V MA) representation which implies that
the Panel SVAR model can be formulated as follows
∞ ∞
Byi,t = Φ(L)ei,t , Φ(L) := ∑ Φ jL j ≡ ∑ A1j L j (5.49)
j=0 j=0
shocks of the model are assumed to be uncorrelated, such that, ui,t u′i,t = I m , we identify the matrix B
by decomposing the variance-covariance matrix into two triangular matrices. Therefore, to identify the
model we impose one restriction in order to orthogonalize the contemporaneous responses.
In particular, using the Cholesky ordering, and based on the variables of interest, institutional quality is
set to have no contemporaneous effect on GDP per capita while the latter is allowed to contemporane-
ously impact the former. Overall, the study of Goes (2016) investigates the relation between institutions
and economic growth using a structural panel VAR(1) model. Next, we focus on the approach to recover
the IRs from the coefficient matrices.
82
Specifically, we take the following VMA representation of the Panel SVAR such that
Example 42. We briefly discuss the econometric specification proposed by Huber et al. (2023) which is
suitable for dealing with cross-country heterogeneity in panel VARs using finite mixture models. Specif-
ically, a global econometric identification and estimation model (see, also Feldkircher et al. (2020)) is
yit = β i + Ai1 yit−1 + ... + Aip yit−p + Bi1 y−it−1 + ... + Bip y−it−p + ε it . (5.51)
Under the Gaussian "state-of-the-world" one can impose Gaussian assumptions regarding the error co-
variances across countries. Thus, we can stack the country-specific errors ε it in a K−dimensional vector
′
ε t = ε ′1t , ..., ε ′Nt and assume that ε t ∼ N 0, Σt , where Σt is a (K × K) dimensional covariance ma-
trix. Furthermore, according to Huber et al. (2023), there are three important dimensions of model
uncertainty which have been identified in the literature as summarized below:
• The first source of model uncertainty is concerned with modelling contemporaneous relations
across the shocks in the system (static interdependancies).
• The second source of model uncertainty focuses on the question whether coefficients associated
with lagged domestic variables are homogeneous across countries (homogeneity restrictions). In
particular, if such domestic coefficients are similar, then the so-called homogeneity restrictions
might be imposed, effectively introducing the same set of coefficients for various countries and
thus reducing the number of free parameters.
• The third source of model uncertainty deals with the question whether to allow for lagged depen-
dencies between countries (dynamic interdependencies). Therefore, if we are interested to capture
the international effects of climate shocks in globalized markets then capturing cross-country in-
terdependencies on the macroeconomic level is crucial.
• From the empirical contribution perspective one can focus on how climate shocks impact country-
specific macroeconomic fundamentals. Then the empirical analysis of the novel framework pro-
posed by demonstrate that climate shocks are have a substantial impact on short-term interest rates
and inflation, which is the primary conventional tool and target variable of central banks, and to a
lesser degree on output and exchange rates (see, Huber et al. (2023)).
Example 43 (Macro-finance Application). A causal effect question of interest from the macro-finance
literature is the linkage between financial crises and economic contraction (see, Huber (2018), Romer and Romer
(2017) and Baron et al. (2021)). In particular, Mei et al. (2023) employ a local projection panel data es-
timator with fixed effects to study the link between financial distress and economic conditions deteriora-
tion. Thus, an econometric issue of interest is to overcome the so-called Nickell bias when considering
classical FE estimators. Consider an h−period ahead panel LP model
(h) (h)
yi,t+h = µi + β (h) xi,t + εi,t+h , for t = 1, ...., n − h, h = 0, 1, ..., H. (5.52)
83
6. Advanced Topics
We follow the framework proposed by Boswijk (1994). Consider the single-equation error correction
model of a time series {yt } conditional upon the (k × 1) vector time series {zt }, for t ∈ {1, ...T } as
p−1
∆yt = β0′ ∆zt ′
+ λ yt−1 − θ + ∑ γ j ∆yt− j − β j′ ∆zt− j + vt , (6.1)
j=1
where {vt } is an innovation process relative to zt , yt− j , zt− j , j = 1, 2, ... with positive variance ω 2 .
Notice that θ and β j are (k × 1) parameter vectors such that j ∈ {1, ..., p − 1}, such that θ defines the
long-run equilibrium relation y = θ ′ z, the deviations from which lead to a correction of yt by a proportion
of λ , the adjustment or error correction coefficient.
The conditional model is said to be stable if all roots of the characteristic equation
!
p−1
ϕ (ζ ) = (1 − ζ ) 1 − ∑ γ jζ j − λ ζ = 0, (6.2)
j=1
are outside the unit circle. In other words, stability of the model implies that the disequilibrium error
(yt − θ ′ zt ) is a stationary process, even though zt and yt are integrated of order one and hence nonsta-
tionary. Specifically, if the model is stable, then xt = (yt , zt′ )′ is cointegrated with cointegrating vector
(1, −θ )′ . Thus, the purpose of this exercise is to develop a class of tests for the null hypothesis that the
characteristic equation has a unit root, so that the model is unstable, against the alternative hypothesis of
stability. Furthermore, the single-equation conditional model can be seen as a special case if a structural
error correction model. This is a system of Error Correction Equations for a (g × 1) vector of time series
{yt } conditional upon {zt } such that (see, Boswijk (1994)):
p−1
Γ0 ∆yt = B0 ∆zt + Λ Γyt−1 + Bzt−1 + ∑ Γ j ∆yt− j + B j ∆zt− j + vt , (6.3)
j=1
where {vt } is an innovation process with a positive-definite covariance matrix Ω such that the above
expression corresponds to a parametrization of a conditional model of yt given zt (if Γ0 6= I g ). Next
we consider the identification of these matrix parameters. In particular, the above model implies g
cointegrating relationships Γy + Bz = 0, provided that it is stable, such that the characteristic equation
is expressed as below
!
p−1
ϕ (ζ ) = (1 − ζ ) Γ0 − ∑ Γ jζ j − ΛΓζ = 0. (6.4)
j=1
84
In other words, unless the parameters are restricted in some way, the structural model is not identified.
Specifically, we identify the long-run relations by imposing restrictions of the usual form such that
where Γi and Bi denote the i−th row of Γ and B, respectively, and where Ri is a known matrix of
approximate order. Thus, the rank condition for identification of the i−th long-run relation
′
rank Ri Γ B = (g − 1) (6.6)
Then, the remaining parameters are identified by the normalization Γ0,ii = 1, and the restriction that Λ =
diag λ1 , ..., λg , that is, where Λ is a diagonal matrix. In practice, this means that only the disequilibrium
error of the i−th long-run relation appears in the i−th structural error correction equation which means
that the i−th equation is (over)-identified if the i−th long-run equation is.
Remark 29. The reason for restricting the error correction matrix Λ to be diagonal is twofold. Firstly, it
allows for an interpretation of these separate equations as representing economic behaviour of a group
of agents, whose target consists in a particular long-run relationship, such as a money demand relation
or a consumption function. Furthermore, notice that the possibility that all endogenous variables are
affected by each disequilibrium error is not excluded. However, this is considered to be a property of
the reduced form of the system rather than the structural form. Moreover, imposing a symmetric matrix
Λ = diag λ1 , ..., λg , facilitates the implementation and interpretation of a test for instability. Notice that
we assume that the number of stable relationships is equal to the number of cointegrating relationships.
Consequently, under this null hypothesis, there is no error correction in the i−th equation, which sug-
gests that the i−th row of the system Γy + Bz = 0 is not a cointegrating relationship. Specifically, let zt
be generated by the following expression:
p−1
∆zt = α Γyt−1 + Bzt−1 + ∑ A1 j ∆yt− j A2 j ∆zt− j + ε t , (6.7)
j=1
Assumption 1, can be interpreted as a form of exogeneity condition, because it states that the coin-
tegration properties of the conditional model carry over to the full VAR system. Understanding well
these concepts are essential especially if one is interested to apply advanced identification techniques in
SVARs such as the Markov switching regime (see, Lanne and Saikkonen (2002) and Lanne et al. (2010))
as well as identification in Gaussian Mixture Vector Autoregressive Models (see, Kalliovirta et al. (2016)
and Meitz et al. (2023)). In particular, Kalliovirta et al. (2016) develop a framework for Gaussian mix-
ture vector autoregression which is a nonlinear vector autoregression and is designed for analyzing time
series that exhibit regime-switching dynamics13 (see, Kalliovirta et al. (2015) and Virolainen (2020)).
Next, we consider another example motivated from the structural break literature specifically in the
context of cointegrated models (see, also Hansen and Seo (2002)).
13 The property of long memory and regime switching is examined in the study of Diebold and Inoue (2001) while a regime
switching panel data regression model is proposed by Cheng et al. (2019). For a forecasting application see Nyberg (2018).
85
6.1.1. Testing Cointegration Rank when some Cointegrating Directions are Changing
Consider the following VECM of order p with a structural break such that
p
∆X t = α 0 β ′0 X t−1 + Φ0 Dt 1 {t ≤ k0 } + α 1 β ′1 X t−1 − X k0 + Φ1 Dt 1 {t > k0 } + ∑ Γ j ∆X t− j + ε t ,
j=1
Then, the date of the break is characterized by the fraction λ0 of the sample size n such that k0 = ⌊λ0 ⌋,
where λ0 ∈ [λ , λ̄ ] and 0 < λ < λ̄ < 1. Moreover, the VECM is assumed to satisfy the stability condition
over the two regimes as given by the Assumption below.
Assumption 16. Suppose that {εt } is a vector martingale difference sequence with respect to Ft−1 such
that the variance-covariance matrix Σε = E[εt εt′ |Ft−1 ], for t = 1, 2, ....
Remark 30. Following Andrade et al. (2005), we emphasize that the assumption above does not allow
for a shift in the innovation covariance matrix Σε across the two regimes. This would split the model into
two separate ones if the short-term coefficient matrices were also shifting with the break and would thus
require a separate analysis over the two sub-samples. Similarly we do not allow for a different number
of cointegration relationships across the two two regimes. In particular, we consider the following cases:
Case 1. The break (structural change) does not affect the loading factors. In this case their estimation
requires an identification constraint. In other words, α0 and α1 span the same vector space and
thus a natural identification scheme for the two regimes is given by α ′ Σε α = Ir .
Case 2. The break affects the cointegration vectors and the loading factors. This implies that the spaces
spanned by α0 and α1 differ across the two regimes.
Theorem 6 (Andrade et al. (2005)). Consider the model defined by the first expression with α0 = α1 =
α and Assumption 2. Under H0 , when T → ∞ it holds that
(Z Z λ −1 Z λ )
λ0
d 1 0 0
ξT (λ0 ) → trace dW 0 G′0 G0 G′0 du G0 dW ′0
λ0 0 0 0
(Z Z 1 −1 Z 1 )
1 1
+ trace dW 0 G′1 G1 G′1 du G1 dW ′0
(1 − λ0 ) λ0 λ0 λ0
86
6.2. Empirical Likelihood Estimation Approach
Example 44. Consider the following constant coefficient autoregressive model with time-varying vari-
ances as below:
Based on the aforementioned assumptions the estimation of the unknown parameter vector β o based on
the OLS estimator β̂ is given by
!−1 !
√ 1 T ⊤ 1 T ⊤ d
T (β − β o ) = ∑ X t−1 X t−1
T t=1 ∑ X t−1 εt
T t=1
→ N (0, Λ), (6.11)
where Λ = Ω−1 −1
1 Ω2 Ω1 , are defined as (p + 1) × (p + 1) matrices.
To construct an empirical likelihood function, the estimation equations are defined as below:
⊤
W t (b) = X t−1 · Yt − X t−1 b , (6.12)
for a generic parameter b ∈ R p+1 . By using the Lagrange multipliers method, we have that λ̂ = λ̂ (b) ∈
R p+1 is the solution of the following set of equations:
1 T W t (b)
∑
T t=1 1 + λ̂ ′ ·W t (b)
= 0. (6.13)
T
′
ℓ(b) = 2 ∑ log 1 + λ̂ ·W t (b) (6.14)
t=1
d
and it holds that ℓ(β 0 ) → χ p2 , as T → ∞.
87
6.2.1. High Dimensional Generalized Empirical Likelihood Estimation
In this section, we discuss relevant aspects to the empirical likelihood estimation for high dimensional
dependent data, which is applicable to time series regression models (see, Chang et al. (2015)). Further
frameworks related to identification and estimation of high-dimensional time series with SVAR models
are proposed by Krampe et al. (2023), Zhang (2023) and Adamek et al. (2022). In particular, denote
with θ = (θ1 , ..., θ p)′ be a p−dimensional parameter taking values in a parameter space Θ. Consider a
sequence of r−dimensional estimating equation such that
g(Xt , θ ) = g1 (Xt , θ ), ...., gr(Xt , θ ) (6.15)
for some r ≥ p, Then, the model information regarding the data and the data parameter is summarized
by moment restrictions below:
E g(Xt , θ 0 ) = 0. (6.16)
where θ0 ∈ Θ is the true parameter. Furthermore, in order to preserve the dependence structure among
the underlying data, we employ the blocking technique. Let M and L be two integers denoting the block
length and separation between adjacent blocks, respectively. Then, the total number of blocks is
(n − M)
Q=⌊ ⌋+1 (6.17)
L
Then, the EL estimator θ o is θb EL = argmax θ inΘ log L (θ ). Consequently, the maximization problem
can be carried out more efficiently by solving the corresponding dual problem, which implies that θb EL
can be obtained as below:
Q
θb EL = arg min max
θ ∈Θ
∑ log 1 + λ ⊤φM (Bq, θ ) ,
b n (θ ) q=1
λ ∈Λ
(6.18)
n o
b n (θ ) := λ ∈ Rr : λ ⊤ · φM (Bq , θ ) ∈ N , q = 1, ..., Q
Λ (6.19)
where m ≥ 1 is some constant. In particular, for conventional vector autoregressive models such that
88
Notice that these set of model parameters {A1 , ..., Am } correspond to coefficient matrices and are usu-
ally estimated based on some maximum likelihood approach (usually the MLS under Gaussianity) and
according to the distributional assumptions of the model. Moreover, suppose that ηt is the white noise
series such that
⊤ ⊤
h Yt , ...,Yt−m ; θ 0 = Yt − A1Yt−1 − ... − AmYt−m ⊗ Yt⊤ , ....,Yt−m . (6.22)
High dimensional time series analysis, requires to assume that the dimensionality of Yt is large in relation
to sample size, that is, s → ∞ as n → ∞. Within a high-dimensional environment, the number of estimat-
⊤ ⊤
ing equation and unknown parameters are both s2 m. On the other hand, if we replace Yt⊤ , ....,Yt−m
⊤
by Yt⊤ , ....,Yt−m−ℓ
⊤ for some fixed ℓ ≥ 1, then the model will be over-identified. This phenomenon
of over-parametrization in such models is well-known to the literature. Thus, to implement a consistent
estimation approach, the sparsity assumption allows to employ a penalized estimation methodology.
We follow the framework proposed by Miao et al. (2023) who study high-dimensional vector autore-
gressions (VARs) augmented with common factors that allow for strong cross-sectional dependence.
This approach allows to incorporate in a unified framework a convinient mechanism for accommo-
dating the interconnectedness and temporal co-variability that are often present in large dimensional
systems.
Consider the N−dimensional vector-valued time series {Yt } = (y1t , ..., yNT )′ , the high-dimensional
VAR model of order p with CFs given by
p
Yt = ∑ A0jY t− j + Λ0 f t0 + ut , t = 1, ..., T, (6.23)
j=1
where A01 , ..., A0p are the (N × N) transition matrices and ut is an N−dimensional vector of unobserved
idiosyncratic errors. Moreover, the analytical framework allows for both the number of cross-sectional
units N and the number of time periods T to pass to infinity. The lag length is also allowed to (slowly)
grow to infinity with (N, T ). Estimation then is a natural high-dimensional problem.
Then, the N−dimensional VAR(p) process {Yt } can be rewritten in a companion form as an N p−dimensional
VAR(1) process with common factors such that:
Yt A01 A02 . . . A0p−1 A0p Yt−1 Λ0 f t0 ut
Yt−1 IN 0 ... 0 0 Yt−2 0 0
= . ..
.. . .. ..
.
.. .. + .. + .. . (6.24)
. . . . . . . .
Yt−p−1 0 0 ... 0 0 Yt−p 0 0
| {z } | {z } | {z } | {z } | {z }
X t+1 Φ Xt Ft Ut
89
As a result, the reverse characteristic polynomial of Yt can be written as below:
p
A (z) ≡ IN − ∑ A0j z p . (6.25)
j=1
In the low-dimensional framework, the process is stationary if A (z) has no roots in and on the complex
unit circle, or equivalently the largest modules of the eigenvalues of Φ is less than 1. Therefore, to
achieve identification, we shall study the Gram or signal matrix SX = X ′ X/T and its population coun-
terpart ΣX = E (X t′ X t ). In other words, one can study the deviation bounds for the Gram matrix, under
the Gaussianity assumption and boundedness of the spectral density function.
In order to ensure that the matrix ΣX is well-behaved, we write X t+1 as a moving average process of
infinite order MA(∞) such that
∞ ∞ ∞
Xt+1 ≡ ∑ Φ j F t− j +U t− j = ∑ Φ j F t− j + ∑ Φ jU t− j (6.26)
j=0 j=0 j=0
Eigenvalue Analysis:
(f)
• First, consider Xt+1 = ∑∞j=0 Φ j F t− j , the component due to the common factors. The covariance
matrix of F t is a high-dimensional matrix with rank R0 and explosive non-zero eigenvalues. In
other words, even if the largest modules of the eigenvalues of Φ is smaller than 1, the variances
(f)
of the entries of Xt+1 are not assumed to be uniformly bounded.
(f) (f)
• Consider yit , which is the i−th entry of Xt+1 . Let e j,M be the j−th column of IM . Noting that
(f) ′ ( f ) (f)
yit = e1,p ⊗ ei,N Xt+1 , we can write yit as the MA(∞) process given below:
∞ ′ 0 ∞
(f) (f)
yit = ∑ e1,p ⊗ ei,N Φ j e1,p ⊗ ei,N ft− j≡ ∑ αiN ( j) ft−0 j , (6.27)
j=0 j=0
Assumption 17 (Miao et al. (2023)). Consider that the following conditions hold:
90
6.4. Fast Algorithm for Detection of Breaks in Large VAR models
Currently in the literature there are two approaches in terms of the identification of break-points. Firstly,
there is a so-called "top down" procedure, in the sense that one tests all the data to determine if there is at
least one change-point and iterates the procedure in the intervals immediately to the "left" and "right" of
the most recently detected change-point. However, the particular methodology has certain weaknesses
compared to a number of other thresholding procedures. A second approach is the so-called "bottom
up" procedure motivated by the observation in the presence of multiple change-points or to mitigate
the effects of inadequately controlled drift in the "baseline" mean value, it may be useful to compare
a candidate change-point at j to an appropriate "local" background ( j, k), where i < j < k. Similar
approaches are the Wild Binary Segmentation (WBS) methodology which uses a random set of possible
backgrounds and an apparently determined threshold (see, Fang et al. (2020)).
Example 46. Suppose we are interested in testing whether there is a change in the covariance matrix of
the stationary multivariate time series {Y n1 , ...,Y n1 }. Then, this implies that
(0) (1)
Y ni := Y n1 1 {i ≤ k} +Y n1 1 {i > k} , 1 ≤ i ≤ n, (6.29)
(0)
for some change-point in the interval 1 ≤ k ≤ n, where Y n1 denotes the stationary time series before and
(1)
change-point and Y n1 denotes the stationary time series after the change-point. Therefore, we assume
that it induces a change to the structure of the covariance matrix which implies the following expressions
(0) (0) (1) (1)
Σn := Var Yn,1 and Σn := Var Yn,k+1 (6.30)
Denote with Σn ([ j]) := Var Y n j , 1 ≤ j ≤ n. Testing the change-point problem at a single location
(0) (0)
H0 : Σn ([ j]) ≡ Σn ∀ j = 1, ..., n and H1 : ∃ k ∈ {2, ..., n} : Σn ([k]) 6= Σn , (6.31)
which tests the null hypothesis of stability of the covariance structure against the alternative hypothesis
of a change. An equivalent formulation of the statistical testing problem under consideration is given as
The proposed change-point detection test statistics are based are based on the partial sums of the sta-
tionary multivariate time series such that Snk = ∑i≤k Y niY ′ni , k ≥ 1. Under the assumption that the
true change-point k is proportional to the sample size, k = ⌊λ n⌋, where 0 < r < λ < 1, the unknown
break-point can be consistently estimated using the following estimator
1 ′ k
k̂n := argmax vn Snk − Snn wn (6.33)
n0 ≤k≤n n n
91
The formulation of the test statistic presented in the above example, implies that is the (smallest) time
index k ≥ n0 := ⌊nr⌋, for some small r > 0, at which the maximum in the definition of the test statistic
Cn is attained. On the other hand, in order to robustify these CUSUM-type test statistics for detecting
changes not only in the middle of the sample, the following statistic is proposed
Tn,md 1 ( j − i)
Cn,md = , with Tn,md = max √ ∑ v′n Snℓ − Snn wn (6.34)
α̂n 1≤ j≤n n 1≤ℓ≤ j n
which considers the maximal deviation from the average of the cumulated sums, the maximum being
taken over all possible subsamples Y nℓ , i ≤ ℓ ≤ j, for 1 ≤ i < j ≤ n. To be more precise, the particular
algorithm implies that each subsample induces a corresponding CUSUM statistic and then the maximal
test statistic is taken over all these subsamples. Therefore, it can be proved
d
Tn,md → sup |Bo (s) − Bo(t)|, as n → ∞. (6.35)
0<s<t<1
Further studies related to testing for structural breaks in covariance structures include Kao et al. (2018).
b be the sample covariance matrix such that Σ
In particular, Σ b = 1 ∑T yt yt′ . Then, for a given τ ∈ [0, 1],
T t=1
we define a point in time ⌊T τ⌋, and we use the subscripts τ and 1 − τ to denote quantities calculated
using the subsamples t = 1, ..., ⌊T τ⌋ and t = ⌊T τ⌋ + 1, ..., T , respectively. In particular, we consider the
sequence of partial sample estimators
⌊T τ⌋ T
bτ = 1 ∑ yt yt′ b1−τ = 1
yt yt′
T (1 − τ) t=⌊T∑
Σ and Σ (6.36)
T τ t=1 τ⌋+1
We denote with wt = vec yt yt′ and w̄t = vec yt yt′ − Σ .
Assumption 18. Let supt E kyt k2r < ∞ for some r > 2. Then, we define with
" ! !′ #
T T
1
VΣ,T = E
T ∑ w̄t ∑ w̄t (6.37)
t=1 t=1
⌊T τ⌋ 1/2
1
√
T
∑ w̄t → VΣ Wn2 (τ), (6.38)
t=1
92
In particular, if no serial dependence is present, a possible choice is the full sample estimator
1 T h i h i′
b ′ b
V Σ = ∑ wt wt − vec Σ b
vec Σ (6.39)
T t=1
Alternatively, one could use the sequence of partial sample estimators such that
h
1 T i h i′ h i h i′
b ′ b
V Σ = ∑ wt wt − τ vec Στ b b
vec Στ + (1 − τ) vec Σ1−τ b
vec Σ1−τ .
T t=1
Furthermore, to accommodate for the case in which Ψℓ ≡ E w̄t w̄t−ℓ =6 0 for some ℓ, we propose a
weighted sum-of-covariance estimator with bandwidth m, such that
m h
i
b ℓ b b ′
Ṽ Σ = Ψ0 + ∑ 1 − Ψℓ + Ψℓ , (6.40)
ℓ=1 m
T
Ψb ℓ = 1 ∑ wt − vec Σb b ′
wt−ℓ − vec Σ (6.41)
T t=ℓ+1
Theorem 8. Under the null hypothesis of no structural break in Σ then, if Assumption 18 holds, as
T → ∞, there exists a δo > 0 such that
1
sup V̂ Σ,τ −V Σ = o p . (6.45)
1≤⌊T τ⌋≤T T δo
Remark 31. Aue et al. (2009) consider the detection of structural breaks in the covariance structure of
multivariate time series under the assumption of stationarity and ergodicity. Their approach identifies
breaks in the unconditional mean of a p−dimensional multivariate time series. Moreover, Lee et al.
(2023) propose a change point test for structural vector autoregressive models via independent compo-
nent analysis. The test statistic is constructed based on the fitted residuals and the CUSUM functional.
93
Furthermore, a growing literature develops methodologies for testing structural breaks in factor models
as in Breitung and Eickmeier (2011), Baltagi et al. (2021) and Bai et al. (2024) among others (see, also
Bai and Perron (1998)). Regarding the literature relevant to structural break testing around the Vector
Autoregressive framework of interest there are two main approaches. On the one hand, the approach
of Safikhani et al. (2022) implies a fast and scalable algorithm for detection of structural breaks in
large VAR models. The particular modelling environment corresponds to a piecewise linear Vector
Autoregression model. Thus, testing for structural breaks takes the form of a time-varying VAR model
and therefore in practise the statistical problem of interest is to test for the presence of time-varying
vector autoregressive matrix coefficients across the blocks. Moreover, since the model is parametrized
in a form of a large VAR model with a corresponding companion matrix, then the focus is on the fast
and accurate break detection algorithm with respect to the dimensionality of the model. On the other
hand, Gouriéroux et al. (2020) proposed an econometric framework for identification and estimation of
the structural vector autoregressive model where the main focus is how the distributional assumptions
on the shocks affect the identification procedure. In practice, based on the non-Gaussianity assumption
of the structural shocks the model is identifiable and the estimation of model parameters is achieved
via the use of the likelihood function. Therefore, testing for structural breaks in such a setting could be
implemented using an appropriate specification which captures time-variation.
For the remaining of this section, we follow the framework proposed by Safikhani et al. (2022). Their
procedure is found to be scalable to big high-dimensional time series datasets with a computational
√
complexity that can achieve O( n), where n is the length of the time series (sample size), compared to
an exhaustive procedure that requires O(n) steps.
Suppose that there exist m0 break points 0 < t1 < ...tm0 < T , with t0 = 0 and tm0+1 = T where j ∈
{1, ..., m0 + 1}. Notice that if there are no break points, then m0 = 0 and so (t0 = 1,t1 = T ) which
corresponds to the full sample period.
Thus, we can assume that m0 ≥ 1, which implies that (t0 = 1,t2 = T ) and so j = {1, 2}. In other words,
for t j−1 ≤ t < t j (e.g., t0 ≤ t < t1 , i.e., t = 1, ..., T ) we have the following model:
1/2
yt = Φ(1, j) yt−1 + ... + Φ(q, j) yt−q + Σ j εt (6.46)
Since in the case that m0 = 1 then j ∈ {1, 2} then we have a two-regime model estimation:
1/2
M0 : yt = Φ(1,1) yt−1 + ... + Φ(q,1) yt−q + Σ1 ε t , t0 ≤ t < t1 , t ∈ {0, ...,t1} , (6.47)
1/2
M1 : yt = Φ(1,2) yt−1 + ... + Φ(q,2) yt−q + Σ2 ε t , t1 ≤ t < t2 , t ∈ {t1 , ..., T } , (6.48)
There is a single break-point and we fit a different VAR model before the break, i.e. M0 pre-break based
on the observations up to the unknown break-point, and a we fit M1 post-break based on the observations
from the unknown break-point and onward until the end of the full sample (see, Safikhani et al. (2022)).
94
• In each segment we estimate the VAR process. Let yt be a p−dimensional vector of observations
at time t and Φ(ℓ, j) be an (p × p) sparse coefficient matrix corresponding to the ℓ−th lag of a VAR
process of order q during the j−th stationary segment, and εt is a white-noise process with zero
mean and variance matrix Σ j .
• In each segment [t j−1,t j ), all model parameters are assumed to be fixed. However, the auto-
regressive (AR) parameters Φ(ℓ, j) will change values across segments, while the error covariance
across all segments is assumed to be Σ j = σ 2 I.
Based on the aforementioned setting proposed by Safikhani et al. (2022), the unknown parameters in
each segment are: the number of break-points m0 , their locations t j , j = 1, ..., m0, the VAR parameters
Φ(q, j) together with the covariance matrix. Therefore, according to , the statistical problem is then to
detect the break points t j , in a computational efficient manner that is also scalable for very large values
of T . Additionally, we are also interested to estimate accurately the VAR parameters Φ(ℓ, j) , under a
high-dimensional environment (p >> T ).
The main idea of the BSS is to partition the time points into blocks of size bn and fix the VAR parameters
within each block. To this end, define a sequence of time points q = r0 < r1 < ... < rkn = T + 1 which
play the role of the end points for the blocks. In other words, the total number of blocks is calculated as:
n
ri+1 − ri = bn , for i ∈ {0, ...., kn − 2} , kn = ⌊ ⌋. (6.49)
bn
where kn is the total number of blocks based on the (effective) sample size and bn is the length size of
each block which is kept fixed across all blocks. Notice that throughout, the dimensions of the vector yt
remains fixed having p elements as well as the (effective) sample size being n.
The theoretical results presented in Safikhani et al. (2022) deal with the high-dimensional case such
that the sparsity levels dk j increase with the sample size, T . Specifically, we define with p ≡ p(n) and
m0 ≡ m0 (n) and dk j ≡ dk j (n), where n = T − q + 1. In other words, n is the effective sample size or the
length of the time series for estimation of model parameters purposes, since the number of lags used for
constructing additional regressors affects the sample size (see, Safikhani et al. (2022)).
The minimum distance between two consecutive break points is denoted by
Let Y ∈ Rn×p and X ∈ Rn×kn pq and θ ∈ Rkn pq×p and E ∈ Rn×p . In other words, based on this
parametrization, θi 6= 0 for i ≥ 2 implies a change in the VAR coefficients. Therefore, for j ∈ {1, ..., m0},
the structural break points t j can be estimated as block-end time point ri−1 , where i ≥ 2 and θi 6= 0.
95
Then, the model is formulated as below
where πb = kn p2 q and Θ = vec(θ ) ∈ Rπb ×1 . The initial estimate of the parameter Θ is obtained:
kn i
b = arg min 1 kY − ZΘk2 + λ1,n kΘk + λ2,n ∑
Θ ∑ θj . (6.52)
2 1
Θ∈∈Rπb ×1 n i=1 j=1 1
Notice that the above high dimensional problem (convexity is preserved) uses fused lasso penalty with
two ℓ1 penalties controlling the number of break points and the sparsity of the VAR model.
Denote the set of indices of blocks with nonzero jumps and corresponding estimated change points
obtained from solving the high-dimensional problem as below:
n o
b i 6= 0, i = 1, ..., kn ,
Ibn = î1 , î2 , ..., îm̂ = i : Θ (6.53)
n F o
A = {tˆ1, tˆ2 , ..., tˆm̂} = ri−1 : i ∈ Ibn . (6.54)
The total number of estimated change points in this step corresponds to the cardinality of the set An ,
where m̂ = |An |. Moreover, the block size bn acts as a tuning parameter that regulates the number
of model parameters to be estimated. In addition to this, the algorithm applies a local screening step.
Since the set An of candidate change points tends to overestimates their numbers, a screening step to
eliminate redundant break points is required. According to Safikhani et al. (2022), the main idea goes
is to estimate the VAR parameters locally on the left and right side of each selected break point in the
first step and compare them to one VAR parameter estimated from combining the left and right of the
selected break point as one large stationary segment. Now, if the selected break point is close to a true
break point, the sum of squared errors calculated assuming stationarity around the true break point will
be much larger compared to the sum of squared errors calculated from two separate VAR parameter
estimates on the right and left of the selected break point. Therefore, we can get consistent estimates
of the number of break points by minimizing a localized information criterion (LIC) comprising of the
sum of squared errors and a penalty term on the number of break points.
Next, we describe the localized screening step in more details. Recall that A = {tˆ1, tˆ2 , ..., tˆm̂} is the
set of candidate break points selected in the first step. Then, for each subset A ⊂ An , we define the
following local VAR parameter estimates for some b ti ∈ A such that
( )
b
1 ti −1 2
an t=∑
bbti ,1 = arg min
ψ yt − ψbti ,1Yt−1 + ηbti ,1 ψbti ,1 (6.55)
b
ti ,1 2 1
b
t −a
i n
where an corresponds to the neighbourhood size in which the VAR parameters are estimated. The
procedure of Safikhani et al. (2022), can be directly implemented for identifying and dating the presence
of multiple break-points in multivariate time series. In contrast to the literature of SVAR models, the
assumption is that the model parameters are identifiable but a possible parameter instability could exist.
96
6.4.5. Consistency of the BSS Estimator
To establish the consistency properties of the BSS-based estimator relevant regularity conditions are
required which we briefly present below.
Assumption 19 (see, Safikhani et al. (2022)). For each j ∈ {1, 2, ..., m0 + 1}, the process
1/2
y( j),t = Φ(1, j) y( j),t−1 + ... + Φ(q, j) y( j),t−q + Σ j ε t (6.56)
Assumption 20 (see, Safikhani et al. (2022)). The matrices Φ(1, j) are sparse such that for all k ∈
{1, ...p} and j ∈ {1, ..., m0}, dk j << p, that is, dk j /p = o(1). There exists a positive constant MΦ > 0
such that
max Φ(·, j) ∞
≤ MΦ (6.58)
1≤ j≤m0 +1
Example 47. Consider the CUSUM series of yℓt over a generic segment [s, e] for some 1 ≤ s < e < T ,
r !
b e
ℓ 1 (b − s + 1)(e − b) 1 1
Ys,b,e =
σℓ e−s+1 ∑ yℓt − e − b ∑ yℓt ,
b − s + 1 t=s
(6.59)
t=b+1
for b = s, ..., e − 1, where σℓ denotes a scaling constant for treating all rows of the sequence yℓt such that
1 ≤ ℓ ≤ N on equal footing. Notice that if εℓt were i.i.d Gaussian random variables, the maximum like-
ℓ
lihood estimator of the change-point location for yℓt , s ≤ t ≤ e would coincide with arg max Ys,b,e .
b∈[s,e)
Moreover, Ds,ne (m) takes the contrast between the m largest CUSUM values ℓ
, 1 ≤ ℓ ≤ m and the
Ys,b,e
rest at each b, and thus partitions the coordinates into the m that are the most likely to contain a change-
point and those which are not in a point-wise manner. Then, the test statistic is derived by maximizing
the two-dimensional array of DC statistics over both time and cross-sectional indices as
Therefore, the presence of multiple break-points implies that there are multiple regimes within the full
sample, where the corresponding time series sequences appear to have structural breaks.
Theorem 9 (Fang et al. (2020)). Let (X1 , ..., Xm) be an independent sequence of normally distributed
random variables with mean µ and variance 1. Then for Zi, j,k we have for b → ∞ and m ≈ b2 ,
P max Zi, j,k ≥ b . (6.61)
0≤i< j<k≤m
97
6.5. Sieve Bootstrap for Functional Vector Autoregressions
According to Paparoditis (2018), bootstrap procedures for Hilbert space-valued time series proposed so
far in the literature, are mainly attempts to adapt, to the infinite dimensional functional framework, of
bootstrap methods that have been developed for the finite dimensional time series case. Specifically,
Paparoditis (2018) shows, that under quite general assumptions, the stochastic process of Fourier coef-
ficients obeys a so-called vector autoregressive representation and this representation plays a key role in
developing a bootstrap procedure for the functional time series at hand. To capture the essential driving
functional parts of the underlying infinite dimensional process, the first m functional principal compo-
nents are used and the corresponding m−dimensional time series of Fourier coefficients is bootstrapped
using a pth order vector autoregression fitted to the vector time series of sample Fourier coefficients.
In this way, a m−dimensional pseudo-time series of Fourier coefficients is generated which imitates
the temporal dependence structure of the vector time series of sample Fourier coefficients. Thus, using
the truncated Karhunen-Loeve expansion, these pseudo-coefficients are then transformed to functional
bootstrap replicates of the main driving, principal components, of the observed functional time series.
The main building block of the procedure proposed by Paparoditis (2018) is to generate pseudo-replicates
X1∗ , ..., Xn∗ of the functional time series at hand by first bootstrapping the m−dimensional time series of
Fourier coefficients ξt = (ξ1,t , ξ2,t , ..., ξm,t ), t = 1, 2, ...n, corresponding to the first m principal compo-
nents. This m−dimensional time series of Fourier coefficients is bootstrapped using the autoregressive
representation ξt . Then, the generated m−dimensional pseudo-time series of Fourier coefficients is then
transformed to functional principal pseudo-components by means of truncated Karhunen-Loeve expan-
sion ∑nj=1 ξ j,t v j . Adding to this, an appropriately resampled functional noise leads to the functional
pseudo-time series X1∗ , X2∗, ..., Xn∗. However, since the ξt′ s are not observed, we work with the time series
of estimates scores. The functional sieve bootstrap algorithm of Paparoditis (2018) is:
Step 1: Select a number m = m(n) of functional principal components and an autoregressive order p =
p(n), both finite and depending on n
Step 2: Let
⊤
b b
ξ = ξ j,t = hXt , vbj i, j = 1, ..., m , t = 1, ..., n, (6.62)
be the m−dimensional series of estimated Fourier coefficients, where we denote with vbj , j =
λ1 > b
1, ..., m the estimated eigenfunctions corresponding to the estimated eigenvalues b λ2 > ... >
b
λm of the sample covariance operator
n n
Cb0 = n−1 ∑ (Xt − X̄n ) ⊗ (Xt − X̄n ) , X̄n = n−1 ∑ Xt . (6.63)
t=1 t=1
98
Step 3: Let Xbt,m = ∑mj=1 ξbj,t vbj and define the functional residuals U
bt,m = Xt − Xbt,m and define the functional
residuals Ubt,m = Xt − Xbt,m , t = 1, ..., n.
Step 4: Fit a pth order vector autoregressive process to the m−dimensional time series ξbt , t = 1, ..., n, de-
note by Ab j,p (m), j = 1, ..., p, the estimates of the autoregressive matrices and by ebt,p the residuals,
p
b j,p(m)ξbt− j , t = p + 1, p + 2, ..., n.
ebt,p = ξbt − ∑ A (6.64)
j=1
p
ξt∗ = ∑ Ab j,p(m)ξt−∗ j + et∗ , (6.65)
j=1
where et∗ , t = 1, ..., n are i.i.d random vectors having as distribution the empirical distribution of the
centered residual vectors eet,p = ebt,p − eb̄n,t , for t = p + 1, p + 2, ..., n and eb̄ = (n − p)−1 ∑t=p+1
n
ebt,p .
and U1∗ , ...,Un∗ are i.i.d random functions obtained by choosing with replacement from the set of
b̄ t = n−1 ∑n U
b̄ t and U
bt,m − U b
centered functional residuals U t=1 t,m , for t = 1, ..., n.
Remark 32. Notice that X1∗ , ..., Xn∗ are functional pseudo-random variables and the the autoregressive
representation of the vector time series of Fourier coefficients is solely used as a toll to bootstrap the
m main functional principal components of the functional time series. In fact, it is this autoregressive
representation which allows the generation of the pseudo-time series of Fourier coefficients ξ1∗ , ..., ξn∗
in Step 4 and Step 5 in a way that imitates the dependence structure of the sample Fourier coefficients
ξ1 , ..., ξn. These pseudo-Fourier coefficients are transformed to bootstrapped main principal components
by means of the truncated and estimated Karhunen-Loeve expansion which together with the additive
functional noise Ut∗ , lead to the new functional pseudo-observations X1∗, ..., Xn∗ (see, Paparoditis (2018)).
Furthermore, an important issue for bootstrap-based inference methods is to investigate the validity
of the functional sieve bootstrap applied in order to approximate the distribution of some statistic
Tn = T (X1 , ..., Xn) of interest, when the bootstrap analogue Tn∗ = T (X1∗, ..., Xn∗) is used. Notice that
establishing validity of a bootstrap procedure for time series heavily depends on two issues. On the
dependence structure of the underlying process which affects the distribution of the statistic of interest
and on the capability of the bootstrap procedure used to mimic appropriately this dependence structure.
99
Recall the definition given by Paparoditis (2018) such that
m
Xt∗ = ∑ 1⊤j ξ j∗vbj +Ut∗ (6.67)
j=1
!−1
∞ p
b m,p (z) = Im + ∑ Ψ
Ψ b l,p(m)zl = b j,p(m)z j
Im − ∑ A (6.68)
l=1 j=1
converges for |z| ≤ 1. Thus, following Paparoditis (2018) we can decompose each sequence as below
∞ m
Xt∗ = ∑ ∑ 1⊤j Ψb l,p(m)et−l
∗
vbj +Ut∗ (6.69)
l=0 j=1
M−1 m ∞ m
∗
Xt,M = ∑ ∑ 1⊤j Ψb l,p(m)et−l
∗ b l,p(m)e∗ vbj +U ∗
vbj + ∑ ∑ 1⊤j Ψ t−l t (6.70)
l=0 j=1 l=M j=1
where for each t ∈ Z, e∗s,t , s ∈ Z is an independent copy of {e∗s , s ∈ Z}. Notice that,
∞ m
∗
XM ∗
− XM,M = ∑ ∑ 1⊤j Ψb l,p(m) e∗M−l − e∗M−l,M vbj (6.71)
l=M j=1
Thus, evaluating the first expectation term, we obtain that using kAk2F = tr AA⊤ and the submulti-
plicative property of the Frobenius matrix norm, that
2
∞ m ∞ 2 ∞ 2
b l,p (m)e∗ b l,p (m)Σ∗ (m)Ψ
b l,p(m) ≤ Σ 1/2
be,p b l,p (m)
E ∑∑ 1⊤j Ψ bj
M−l,M v = ∑ tr Ψ (m) ∑ Ψ ,
l=M j=1 l=M F l=M F
100
7. Conclusion
There are various statistical identification procedures for SVAR models such as imposing sign-restrictions,
short-run restrictions, long-run restrictions, identification by higher-order moments as well as iden-
tification by conditional heteroscedasticity and changes in volatility (e.g., see, Lanne and Lütkepohl
(2008a,b). More precisely, by exploiting the heteroscedasticity of the errors εt , then the identification
of the structural parameters of the system can be achieved. In particular, Lanne et al. (2010) assume
Markov switching and a smooth transition in the covariance matrix of the error term εt of the model.
Specifically, the structural analysis of VARs in this case is based on the Markov regime switching prop-
erty to identify shocks, which holds due to the fact that the reduced form error covariance matrix varies
across states. Furthermore, a different stream of literature considers identification in SVAR models by
assuming that the error terms are non-Gaussian and mutually independent (e.g., see Lanne et al. (2017)
and Gouriéroux et al. (2020)). For example, Lanne et al. (2010) assume that the errors of the model are
independent over time with a distribution that is a mixture of two Gaussian distributions with zero means
and diagonal covariance matrices, one of which is an identify matrix and the other one has positive diag-
onal elements, which for identifiability have to be distinct. Roughly speaking there are specific systems
where the use of the non-Gaussianity property is a more suitable identification strategy of structural
shocks of the SVAR system which we leave for discussion in a follow-up paper.
In this note we have reviewed key statistical properties of VAR processes, cointegrated VAR processes
and SVAR processes as well as suitable identification strategies for the structural parameters of these
structural econometric specifications (see, also Mumtaz et al. (2018)). Moreover, we have discusses
several empirical-driven applications from the applied macroeconometrics literature and we studied
the proposed identification and estimation methods employed to those cases. Then, motivated from
the literature of dynamic causal effects (see, Mertens and Ravn (2013), Stock and Watson (2018) and
Budnik and Rünstler (2023)), we have motivated the main statistical and econometric tools for identifi-
cation of structural econometric models, commonly used for counterfactual analysis purposes, and then
presented some examples of time series regression models, both important for understanding suitable
techniques for policy-making evaluation in macroeconometrics. We leave any further presentation of
other relevant topics for further research discussions. Lastly, some advanced topics were also presented
such as break testing for unstable roots in structural error correction models, estimation techniques for
the optimal lag order, the estimation of a high dimensional VAR model as well as a fast algorithmic
procedure presented in the literature for detecting structural breaks in large VAR models.
In the Appendix of this set of lecture notes we present an example of the maximum likelihood estimation
for α −stable autoregressive processes (see, Andrews et al. (2009)) as well as the main asymptotic theory
results for inference for the VEC(1) model from Guo and Ling (2023).
101
A Maximum Likelihood Theory
Following the framework of Andrews et al. (2009), we consider the MLE estimation for both causal
and non-causal autoregressive time series processes with non-Gaussian α −stable noise. In particular,
the estimators for the autoregressive parameters are n−1/α −consistent (parametric rate) and converge in
distribution to the maximizer of a random function. Although the form of this limiting distribution is
intractable, the shape of the distribution can be examined using the bootstrap procedure. Moreover, the
estimators for the parameters of the stable noise distribution have the traditional n1/2 rate of convergence
and are asymptotically normal. In addition, the behavior of the estimators for finite samples is studied
via simulation, and we use the maximum likelihood estimation to fit a noncausal autoregressive model
(see, Andrews et al. (2009) and Andrews and Davis (2013)).
Let {Xt } be the AR process which satisfies the difference equations
φ0 (B)Xt = Zt , (A.1)
Furthermore, because φ0 (z) 6= 0 for |z| = 1, the Laurent series expansion of 1/φ0 (z) is given by
+∞
1
= ∑ ψ j z j, (A.3)
φ0 (z) j=−∞
• If φ0 (z) 6= 0 for |z| ≤ 1, then ψ j = 0 for j < 0 and so {Xt } is a causal process, since Xt =
∑+∞
j=−∞ ψ j Zt− j , is a function of the past and present {Zt }.
Moreover, in the purely non-causal process case, the coefficients ψ j satisfiy
1 − φ01 z − ... − φ0pz p ψ0 + ψ−1 z−1 + ... = 1, (A.5)
−1
which, if φ0p 6= 0, implies that ψ0 = ψ−1 = ... = ψ1−p = 0 and ψ1−p = −φ0p .
102
Therefore, to express φ0 (z) as the product of causal and purely noncausal polynomials, suppose
φ0 (z) = 1 − θ01 z − ... − θ0r0 zr0 1 − θ0,r0 +1 z − ... − θ0,r0+s0 zs0 , (A.6)
which implies that θ0+ is a causal polynomial and θ0∗ is a purely noncausal polynomial. Therefore, φ0 (z)
has a unique representation as the product of causal and noncausal polynomials if the true order of the
AR polynomial φ0 (z) is less than p.
Consider the following sequence {Zt } such that
Zt = 1 − θ01 B − ... − θ0r0 Br0 1 − θ0,r0 +1 B − ... − θ0,r0+s0 Bs0 Xt (A.8)
where θ = (θ1 , ..., θ p)′ and θ = (θ01 , ..., θ0p)′ . Moreover, let η = (η1 , ..., η p+4) = (θ1 , ..., θ p, α , β , σ , µ )′ =
n
(θ ′ , τ ′ )′ . In other words, given a realization {Xt }t=1 , then the log-likelihood of η can be approximated
by the conditional log-likelihood such that
s
L (η , s) = ∑ ln f Zt (θ , s); τ + ln|θ p |1 {s > 0} , (A.12)
t=p+1
n
where the component {Zt (θ , s)}t=p+1 is computed separately.
Remark 33. The use of causal and non-causal autoregressive processes can be helpful in understanding
concepts such as identification, granger non-causality and weak exogeneity while ensuring that these
properties are not violated when learning for dynamic causal effects based on macroeconomic models.
Representations for such univariate processes when modelling explosive bubbles/regimes are studied by
Fries and Zakoian (2019), Davis and Song (2020), Cavaliere et al. (2020) and Hecq and Voisin (2021)
while related representations for multivariate processes are studied by Lanne and Saikkonen (2013),
Velasco and Lobato (2018) and Rygh Swensen (2022). A framework for optimal forecasting with non-
causal autoregressive time series is presented by Lanne et al. (2012).
103
A2. Inference for the VEC(1) Model
Consider an m−dimensional vector Y t satisfying the following equation (see, Guo and Ling (2023))
Denote with
′ −1 ′ −1
Q = β 0 , α 0⊥ Q−1 = α 0 β ′0 α 0 , β 0⊥ α 0⊥ β ′0⊥ . (A.16)
Denote with
! ! ! !
β ′0Y t Z 1,t β ′0 ut W 1,t
Zt = ≡ , Wt = ≡ (A.18)
α ′0⊥Y t Z 2,t α ′0⊥ ut W 2,t
Using Z 1,t = β ′0Y t and Z 2,t = α ′0⊥Y t we obtain the following expressions
t −1
Z 1,t ≡ β ′0Y t = β ′0C ∑ us + β ′0 α 0 β ′0 α 0 Ξ(L)β ′0 ut + β ′0CY 0 , (A.20)
s=1
t −1
Z 2,t ≡ α ′0⊥Y t = α ′0⊥C ∑ us + α ′0⊥ α 0 β ′0 α 0 Ξ(L)β ′0 ut + α ′0⊥CY 0 , (A.21)
s=1
104
! !
t t t
Z 2,t = α ′0⊥C ∑ us +Y 0 ≡ 0, I d Q ∑ us +Y 0 ≡ 0, I d ∑ γ s , (A.23)
s=1 s=1 s=1
∞ ∞
γ s = ∑ QDi ε s−i ≡ ∑ φ i ε s−i , with φ i = QDi = O p(ρ i ) for some ρ ∈ (0, 1). (A.24)
i=0 i=0
−1 −1
Theorem 10 (Guo and Ling (2023)). Let β̄ ⊥ = β 0⊥ α ′0⊥ β 0⊥ , β̄ = α 0 β ′0 α 0 and φ = Q ∑∞
i=0 Di .
Suppose that Assumptions hold and that εt has a symmetric distribution, then it follows
(a). Πb ols − Π0 β̄ →d R1 Γ−1 ,
11
(b). n Πb ols − Π0 β̄ →d R∗ Γ−1 − R∗ Γ−1 Γ21 Γ−1 ,
⊥ 2 22 1 11 22
Remark 34. Notice that in order to establish the asymptotic distribution of the OLS estimator for the
VEC(1) model with heavy tailed errors, we shall consider the weak convergence of the matrix func-
tionals based on the above limit results. In particular, the OLS related to the short-term parameters is
not consistent and converges to a functional of stable matrix-processes. Moreover, the inconsistency is
because {ut } is a series of dependent random vectors which is a critical difference from the i.i.d white
noises cases as in She et al. (2022).
105
By defining suitable functionals as above we can establish the joint weak convergence given below
! !
∞ ∞
1 n ′ 1 n
′ d
∑ Y t−iY t−
a2n t=1 j , 2 ∑ ε t Y t−k →
ãn t=1 ∑ ΨℓS1 Ψ′ℓ+i− j , ∑ Sℓ+k+1 Ψ′ℓ (A.30)
ℓ=0 ℓ=0
∞
where 1 ≤ j ≤ i ≤ p, 1 ≤ k ≤ p and Y t = ∑ Ψℓε t−ℓ.
ℓ=0
−1
Moreover, denoting with Bn = (Dn Q) and H such that
ãn −1/2
H = diag Ir , n an I m−r0 .
an 0
Then, based on the above asymptotic results and the joint weak convergence we obtain that
1 1 −2 ∞
! ′
n S
a2n 11
√ an S12n
n d
∑ CℓS1Cℓ 0
H −1 Dn ′
∑ Zt−1Zt−1 −1
Dn H = 1 1 → ℓ=0 Z .
1 −2 1
t=1 √ 2 S11 √ an S12n 0 φ P(r)P′
(r)dr φ
n an n 0
Understanding the existing methods for decomposing permanent and transitory effects in cointegrating
regressions is useful also in correctly implementing relevant identification strategies for SVAR models.
In particular, Gonzalo and Ng (2001) propose a framework for analyzing the dynamic effects of perma-
nent and transitory shocks. Moreover Hecq et al. (2000) apply the permanent-transitory decomposition
(see, Beveridge and Nelson (1981) and Quah (1992)) in VAR models with cointegration and common
cycles (see, also Garratt et al. (2006) and Myers et al. (2018)).
Specifically decomposing which shocks are temporary and which shocks have a permanent effects is
important for cointegration analysis purposes. Consider a bivariate SVAR(1) system, then if there as
at least one permanent shock a cointegrated VAR representation can be identified or a Vector Error
Correction Model while if all shocks have a permanent shock then there is no underline cointegration
dynamics and thus the model represents a VAR system in first differences. In other words, under the
absence of transitory shocks then the system involves no cointegration dynamics. On the other hand,
there should be at least one shock that is not transitory but has a permanent impact in order to ensure
that a cointegrated VAR representation holds.
106
Consider the following three-variable VEC
q
∆yt = µ + αβ ′ yt−1 + ∑ Γ j ∆yt− j + εt ,
j=1
such that there are r ≤ 2 cointegrating relationships, β contains the cointegrating vectors and µ , α , Γ j
are unknown parameters to be estimated. Moreover, the error terms εt are assumed to be serially
uncorrelated but may be contemporaneously correlated. Based on the aforementioned formulation,
Gonzalo and Granger (1995) showed that given the three variable VEC representation with two cointe-
grating vectors, yt can be decomposed into permanent and transitoru components as below
−1 −1 ′
yt := Aa′⊥ yt + Bβ ′ yt ≡ β⊥ α⊥
′
β⊥ a′⊥ yt + α β ′ α β yt
| {z } | {z }
permanent component transitory component
where A and B are loading matrices which scale the two components accordingly. Notice that the
decomposition works because the term a′⊥ yt represents the only linear combination of yt that ensures the
transitory component Bβ ′ yt has no long-run impact on yt . Furthermore, the particular decomposition
when modelling the joint behaviour of multiple time series, is also useful in applications related to
permanent and temporary income shock dynamics as well as to household consumption patterns.
Remark 35. Long-run restrictions were pioneered with the seminal study of Blanchard and Quah (1988).
Moreover, the advantage of considering cointegration dynamics is that it can be shown that the long-run
behaviour of an n dimensional system with r cointegrating relationships can be described by only (n −r)
independent stochastic trends. In other words, the presence of common trends can be identified as struc-
tural innovations to the system which implies that the remaining r structural innovations cannot have
a permanent impact to the system. Therefore, in order to conduct proper inference in these structural
models the asymptotic distribution of the parameters of interest has to be derived.
Example 48 (Multivariate Beveridge-Nelson Decomposition with I(1) and I(2) series). The study of
Murasawa (2015) considers that the consumption Euler equation implies that the output growth rate
and the real interest rate are of the same order of integration. To estimate the natural rates and gaps of
macroeconomic variables jointly, we need to develop a multivariate Beveridge-Nelson Decomposition
when some series are I(1) and other are I(2). These common sets of macro variables as below:
• inflation rate: let Pt be the price level and πt = ln(Pt /Pt−1 ) be the inflation rate. Moreover, let
r̂t := it − πt+1 be the ex-post interest rate.
• interest rate: let It be the 3-month interest rate (annual rate in percent).
• unemployment rate: let Lt be the labor force, Et be employment and Ut := −ln(Et /Lt ) be the
unemployment rate.
Assume that {∆lnYt } is I(1) without a drift, {πt } is I(1) and rt is also I(1).
107
A2.2. Automated Estimation of VECMs
A related stream of literature includes the notion of irreducible cointegrating relation, which implies that
no variable can be omitted without loss of the cointegration property (see, Davidson (1998a)) as well as
the notion of causal ordering of cointegrating relations (see, Hoover (2020)). In addition, the "minimal"
algorithm proposed by Davidson (1998a) is extended to the literature of automated cointegration models
as in Krolzig (2003) and Liao and Phillips (2015). Note that the "minimal" algorithm of Davidson
(1998a) is different from other methods since it tries to find all the smallest subsets of variables which
are cointegrating. Furthermore, his method is based on the Wald test statistic developed by Davidson
(1998b) and thus the algorithm is faster since the Wald test does not require restricted optimization.
Moreover, the approach proposed by Liao and Phillips (2015) corresponds to an automated estimation
of vector error correction model to tackle model selection and associared issues of post-model selection
inference which present well known challenges in empirical econometric research.
Following the automated estimation of vector error correction model approach proposed by Liao and Phillips
(2015), we consider the parametric VEC representation of a cointegrated system
p
∆Y t = Π0Y t−1 + ∑ B0, j ∆Y t− j + ut , (A.32)
j=1
2
n
b 0, B
(Π b0 ) = arg min ∑ ∆Y t − ΠY t−1 − ∑ B j ∆Y t− j
Π,B1 ,...,B p∈Rm×m t=1 j≤p
p m
+ n ∑ λb, j,n B j + n ∑ λr,k,n Φn,k (Π) ,
j=1 k=1
where λb, j,n and λr,k,n are tuning parameters that directly control the penalization.
First Order VECM Estimation Consider the following formulation of the VECM as below
In other words, the model based on the above formulation contains no deterministic trend and no lagged
differences. Therefore, our focus in this simplified system is to outline the approach to cointegrating
rank selection and develop key elements in the limit theory, showing consistency in rank selection and
reduced rank coefficient matrix estimation.
108
Assumption 21 (WN). Suppose that {ut }t≥1 is an m−dimensional i.i.d process with zero mean and
non-singular covariance matrix Ωu .
Under the above assumption then partial sums of ut satisfy the functional law below
1 [n·]
√ ∑ ut ⇒ Bu (·), (A.34)
n t=1
−1
Define with Q = [β0 , α 0⊥ ]′ , and note that C = β 0⊥ α ′0⊥ β ′0⊥ α ′0⊥ . Then, using the matrix Q we
obtain the following equivalent expression
Under the above assumptions, we have the following functional law as below
" # " #
1 [n·] β ′0 Bu (·) Bη1 (·)
√ ∑ η t ⇒ Bη (·) = QBη (·) = ≡ . (A.37)
n t=1 α ′0⊥ Bu (·) Bη2 (·)
b of Π0 is
Notice that the unrestricted LS estimator Π
! !−1
n n n
b unrest
Π n = arg min ∑ k∆Y t − ΠY t−1k2 = ∑ ∆Y tY t−1
′ ′
∑ ∆Y t−1Y t−1 .
Π∈Rm×m t=1 t=1 t=1
Decomposing nearly nonstationary from nearly stationary roots in the set of regressors is discussed
in Paruolo (1997) who considers that the cointegrated system X t = Pβ X t + Pβ⊥ X t in two systems of r
stationary β ′ X t and (d − r) nonstationary β⊥′ X t components (see, also Stock and Watson (1988)).
109
B Algebraic Theory of Identification in Structural Models
Following, Kociecki (2011) we shall avoid using the definition that only the structural parameters that
are in one-to-one correspondence with the reduced form parameters are identified. The reason behind
this is that this is not always the case. In fact, there are other structural parameters, which are identified,
but can not be uniquely recovered from the reduced form parameters. In other words, an one-to-one
correspondence is a necessary condition for identification but is not a unique one.
There is a close connection between equivalence class and orbit. For a number of econometric models,
equivalence classes are simply orbits. Therefore, when equivalence class is an orbit the approach to
identification based on checking local properties of the likelihood (information matrix), is rather mis-
placed. We present a relevant definition which is a fundamental tool in statistical invariance theory.
Definition 11 (Kociecki (2011)). A function f : Θ → Y is said to be invariant under some action of a
group G on Θ, if f (θ ) = f (g◦ θ ) for any g ∈ G, θ ∈ Θ. Moreover, a function f : Θ → Y is called maximal
invariant G−invariant if f is G−invariant and for any θ1 , θ2 ∈ Θ, f (θ1 ) = f (θ2 ) implies θ1 = g ◦ θ2 for
some g ∈ G, that is, θ1 and θ2 lie on the same orbit.
Example 49 (Finite Mixture Models). Suppose that the pdf of a finite mixture of two normal distribu-
tions is given by
− 21 1 2 − 21 1 2
fyt (yt ) = p1 · (2π ) exp − (yt − µ ) + p2 · (2π ) exp − (yt − µ ) . (B.1)
2 2
Definition 12. A square matrix B is irreducible if there exists no permutation matrix Q for which
" #
B 11 B 12
QBQ⊤ = (B.2)
0 B22
|βii | ≥ ∑ |βi j |, with strict inequality holding at least for one i. (B.3)
j=1, j6=i
110
References
Adamek, R., Smeekes, S., and Wilms, I. (2022). Local projection inference in high dimensions. arXiv
preprint arXiv:2209.03218.
Agullo, J., Croux, C., and Van Aelst, S. (2008). The multivariate least-trimmed squares estimator.
Journal of Multivariate Analysis, 99(3):311–338.
Aigner, D. J. and Balestra, P. (1988). Optimal experimental design for error components models. Econo-
metrica: Journal of the Econometric Society, pages 955–971.
Alloza, M., Gonzalo, J., and Sanz, C. (2020). Dynamic effects of persistent shocks. arXiv preprint
arXiv:2006.14047.
Andrade, P., Bruneau, C., and Gregoir, S. (2005). Testing for the cointegration rank when some cointe-
grating directions are changing. Journal of Econometrics, 124(2):269–310.
Andrews, B., Calder, M., and Davis, R. A. (2009). Maximum likelihood estimation for α -stable autore-
gressive processes. The Annals of Statistics, 37(4):1946–1982.
Andrews, B. and Davis, R. A. (2013). Model identification for infinite variance autoregressive processes.
Journal of Econometrics, 172(2):222–234.
Angelini, G., Cavaliere, G., and Fanelli, L. (2024). An identification and testing strategy for proxy-svars
with weak proxies. Journal of Econometrics, 238(2):105604.
Angrist, J. D. and Kuersteiner, G. M. (2011). Causal effects of monetary shocks: Semiparametric con-
ditional independence tests with a multinomial propensity score. Review of Economics and Statistics,
93(3):725–747.
Antolín-Díaz, J. and Rubio-Ramírez, J. F. (2018). Narrative sign restrictions for svars. American
Economic Review, 108(10):2802–2829.
Anttonen, J., Lanne, M., and Luoto, J. (2023). Statistically identified svar model with potentially skewed
and fat-tailed errors. Available at SSRN 3925575.
Arias, J. E., Fernández-Villaverde, J., Rubio-Ramírez, J. F., and Shin, M. (2023). The causal effects of
lockdown policies on health and macroeconomic outcomes. American Economic Journal: Macroe-
conomics, 15(3):287–319.
Aruoba, S. B., Mlikota, M., Schorfheide, F., and Villalvazo, S. (2022). Svars with occasionally-binding
constraints. Journal of Econometrics, 231(2):477–499.
Aue, A., Hörmann, S., Horváth, L., Reimherr, M., et al. (2009). Break detection in the covariance
structure of multivariate time series models. The Annals of Statistics, 37(6B):4046–4087.
Bacchiocchi, E., Castelnuovo, E., and Fanelli, L. (2018). Gimme a break! identification and estima-
tion of the macroeconomic effects of monetary policy shocks in the united states. Macroeconomic
Dynamics, 22(6):1613–1651.
Bacchiocchi, E. and Fanelli, L. (2011). New identification strategies in structural vector autoregressive
models with structural changes, with an application to us monetary policy. hypothesis, page 2.
Bai, J., Duan, J., and Han, X. (2024). The likelihood ratio test for structural changes in factor models.
Journal of Econometrics, 238(2):105631.
Bai, J. and Perron, P. (1998). Estimating and testing linear models with multiple structural changes.
111
Econometrica, pages 47–78.
Bai, J. and Wang, P. (2014). Identification theory for high dimensional static and dynamic factor models.
Journal of Econometrics, 178(2):794–804.
Baillie, R. T. and Kapetanios, G. (2013). Estimation and inference for impulse response functions from
univariate strongly persistent processes. The Econometrics Journal, 16(3):373–399.
Baker, S. R., Bloom, N., and Davis, S. J. (2016). Measuring economic policy uncertainty. The quarterly
journal of economics, 131(4):1593–1636.
Baltagi, B. H. (2008). Econometric analysis of panel data, volume 4. Springer.
Baltagi, B. H., Kao, C., and Wang, F. (2021). Estimating and testing high dimensional factor models
with multiple structural changes. Journal of Econometrics, 220(2):349–365.
Banafti, S. and Lee, T.-H. (2022). Inferential theory for granular instrumental variables in high dimen-
sions. arXiv preprint arXiv:2201.06605.
Barigozzi, M. and Hallin, M. (2017). A network analysis of the volatility of high dimensional financial
series. Journal of the Royal Statistical Society Series C: Applied Statistics, 66(3):581–605.
Barigozzi, M., Hallin, M., Soccorsi, S., and von Sachs, R. (2021a). Time-varying general dynamic factor
models and the measurement of financial connectedness. Journal of Econometrics, 222(1):324–343.
Barigozzi, M., Lippi, M., and Luciani, M. (2021b). Large-dimensional dynamic factor models: Es-
timation of impulse–response functions with i (1) cointegrated factors. Journal of Econometrics,
221(2):455–482.
Barigozzi, M. and Trapani, L. (2022). Testing for common trends in nonstationary large datasets. Jour-
nal of Business & Economic Statistics, 40(3):1107–1122.
Barnichon, R. and Brownlees, C. (2019). Impulse response estimation by smooth local projections.
Review of Economics and Statistics, 101(3):522–530.
Baron, M., Verner, E., and Xiong, W. (2021). Banking crises without panics. The Quarterly Journal of
Economics, 136(1):51–113.
Baruník, J. and Ellington, M. (2024). Persistence in financial connectedness and systemic risk. Euro-
pean Journal of Operational Research, 314(1):393–407.
Baruník, J. and Křehlík, T. (2018). Measuring the frequency dynamics of financial connectedness and
systemic risk. Journal of Financial Econometrics, 16(2):271–296.
Baumeister, C. and Hamilton, J. D. (2015). Sign restrictions, structural vector autoregressions, and
useful prior information. Econometrica, 83(5):1963–1999.
Baumeister, C. and Hamilton, J. D. (2019). Structural interpretation of vector autoregressions with
incomplete identification: Revisiting the role of oil supply and demand shocks. American Economic
Review, 109(5):1873–1910.
Baumeister, C. and Hamilton, J. D. (2021). Advances in using vector autoregressions to estimate struc-
tural magnitudes. Econometric Theory, pages 1–39.
Baumeister, C. and Hamilton, J. D. (2023). A full-information approach to granular instrumental vari-
ables. Technical report, Working Paper, UCSD.
Ben-Michael, E., Feller, A., and Rothstein, J. (2021). The augmented synthetic control method. Journal
of the American Statistical Association, 116(536):1789–1803.
112
Bertsche, D. and Braun, R. (2022). Identification of structural vector autoregressions by stochastic
volatility. Journal of Business & Economic Statistics, 40(1):328–341.
Bertsche, D., Brüggemann, R., and Kascha, C. (2023). Directed graphs and variable selection in large
vector autoregressive models. Journal of Time Series Analysis, 44(2):223–246.
Beveridge, S. and Nelson, C. R. (1981). A new approach to decomposition of economic time series
into permanent and transitory components with particular attention to measurement of the ‘business
cycle’. Journal of Monetary economics, 7(2):151–174.
Bhansali, R. J. (2002). Multi-step forecasting. A Companion to Economic Forecasting, 1:207–221.
Bianchi, F., Faccini, R., and Melosi, L. (2023). A fiscal theory of persistent inflation. The Quarterly
Journal of Economics, page qjad027.
Bilgili, F. (2012). The impact of biomass consumption on co2 emissions: cointegration analyses with
regime shifts. Renewable and Sustainable Energy Reviews, 16(7):5349–5354.
Blanchard, O. and Perotti, R. (2003). An empirical investigation of the dynamic effects of shocks to
government spending and taxes on output. Quarterly Journal of Economics, 117:1–329.
Blanchard, O. J. and Quah, D. (1988). The dynamic effects of aggregate demand and supply distur-
bances.
Bloom, N. (2009). The impact of uncertainty shocks. econometrica, 77(3):623–685.
Bognanni, M. (2018). A class of time-varying parameter structural vars for inference under exact or set
identification.
Boiciuc, I. (2015). The effects of fiscal policy shocks in romania. a svar approach. Procedia Economics
and Finance, 32:1131–1139.
Bojinov, I., Rambachan, A., and Shephard, N. (2021). Panel experiments and dynamic causal effects:
A finite population perspective. Quantitative Economics, 12(4):1171–1196.
Bojinov, I. and Shephard, N. (2019). Time series experiments and causal estimands: exact randomiza-
tion tests and trading. Journal of the American Statistical Association, 114(528):1665–1682.
Boswijk, H. P. (1994). Testing for an unstable root in conditional and structural error correction models.
Journal of econometrics, 63(1):37–60.
Braun, R. (2023). The importance of supply and demand for oil prices: Evidence from non-gaussianity.
Quantitative Economics, 14(4):1163–1198.
Braun, R. and Brüggemann, R. (2023). Identification of svar models by combining sign restrictions with
external instruments. Journal of Business & Economic Statistics, 41(4):1077–1089.
Breitung, J. and Eickmeier, S. (2011). Testing for structural breaks in dynamic factor models. Journal
of Econometrics, 163(1):71–84.
Brüggemann, R., Jentsch, C., and Trenkler, C. (2016). Inference in vars with conditional heteroskedas-
ticity of unknown form. Journal of econometrics, 191(1):69–85.
Bruneau, C. and De Bandt, O. (2003). Monetary and fiscal policy in the transition to emu: what do svar
models tell us? Economic Modelling, 20(5):959–985.
Bruns, M. and Keweloh, S. (2023). Testing for strong exogeneity in proxy-vars. Available at SSRN
4674721.
Bruns, S. B., Csereklyei, Z., and Stern, D. I. (2020). A multicointegration model of global climate
113
change. Journal of Econometrics, 214(1):175–197.
Budnik, K. and Rünstler, G. (2023). Identifying structural vars from sparse narrative instruments: Dy-
namic effects of us macroprudential policies. Journal of Applied Econometrics, 38(2):186–201.
Cai, G. and Lin, Y. (1992). Response distribution of non-linear systems excited by non-gaussian impul-
sive noise. International journal of non-linear mechanics, 27(6):955–967.
Caldara, D. and Kamps, C. (2017). The analytics of svars: a unified framework to measure fiscal
multipliers. The Review of Economic Studies, 84(3):1015–1040.
Carriero, A., Marcellino, M., and Tornese, T. (2023a). Macro uncertainty in the long run. Economics
Letters, 225:111067.
Carriero, A., Marcellino, M. G., and Tornese, T. (2023b). Blended identification in structural vars.
BAFFI CAREFIN Centre Research Paper, (200).
Carrion-i Silvestre, J. L. and Kim, D. (2021). Statistical tests of a simple energy balance equation in a
synthetic model of cotrending and cointegration. Journal of Econometrics, 224(1):22–38.
Castelnuovo, E. and Surico, P. (2010). Monetary policy, inflation expectations and the price puzzle. The
Economic Journal, 120(549):1262–1283.
Cavaliere, G. and Georgiev, I. (2009). Robust inference in autoregressions with multiple outliers. Econo-
metric Theory, pages 1625–1661.
Cavaliere, G., Nielsen, H. B., and Rahbek, A. (2020). Bootstrapping noncausal autoregressions: with
applications to explosive bubble modeling. Journal of Business & Economic Statistics, 38(1):55–67.
Cevik, S. and Miryugin, F. (2023). It’s never different: Fiscal policy shocks and inflation.
Chan, J. C., Eisenstat, E., and Koop, G. (2016). Large bayesian varmas. Journal of Econometrics,
192(2):374–390.
Chan, J. C., Matthes, C., and Yu, X. (2023). Large structural vars with multiple sign and ranking
restrictions.
Chan, K.-S., Ho, L.-H., and Tong, H. (2006). A note on time-reversibility of multivariate linear pro-
cesses. Biometrika, 93(1):221–227.
Chang, J., Chen, S. X., and Chen, X. (2015). High dimensional generalized empirical likelihood for
moment restrictions with dependent data. Journal of Econometrics, 185(1):283–304.
Chari, V. V., Kehoe, P. J., and McGrattan, E. R. (2008). Are structural vars with long-run restrictions
useful in developing business cycle theory? Journal of Monetary Economics, 55(8):1337–1352.
Chatfield, C. (1993). Calculating interval forecasts. Journal of Business & Economic Statistics,
11(2):121–135.
Chaudourne, J., Fève, P., and Guay, A. (2014). Understanding the effect of technology shocks in svars
with long-run restrictions. Journal of Economic Dynamics and Control, 41:154–172.
Chen, B., Choi, J., and Escanciano, J. C. (2017). Testing for fundamental vector moving average
representations. Quantitative Economics, 8(1):149–180.
Chen, H.-F. (2010). New approach to recursive identification for armax systems. IEEE Transactions on
Automatic Control, 55(4):868–879.
Chen, H.-F. and Zhao, W. (2014). Recursive identification and parameter estimation. CRC Press.
Cheng, T., Gao, J., and Yan, Y. (2019). Regime switching panel data models with interactive fixed
114
effects. Economics Letters, 177:47–51.
Cheng, X., Han, X., and Inoue, A. (2022). Instrumental variable estimation of structural var models
robust to possible nonstationarity. Econometric Theory, 38(5):845–874.
Chevillon, G., Mavroeidis, S., and Zhan, Z. (2020). Robust inference in structural vector autoregressions
with long-run restrictions. Econometric Theory, 36(1):86–121.
Chinco, A., Clark-Joseph, A. D., and Ye, M. (2019). Sparse signals in the cross-section of returns. The
Journal of Finance, 74(1):449–492.
Choi, I. (2002). Structural changes and seemingly unidentified structural equations. Econometric The-
ory, 18(3):744–775.
Choi, I. and Kurozumi, E. (2012). Model selection criteria for the leads-and-lags cointegrating regres-
sion. Journal of Econometrics, 169(2):224–238.
Christiano, L., Eichenbaum, M. S., and Vigfusson, R. J. (2003). What happens after a technology shock?
Cordoni, F., Doremus, N., and Moneta, A. (2023). Identification of vector autoregressive models with
nonlinear contemporaneous structure.
Crucil, R., Hambuckers, J., and Maxand, S. (2023). Do monetary policy shocks affect financial un-
certainty? a non-gaussian proxy svar approach. A Non-gaussian Proxy SVAR Approach (June 5,
2023).
Davidson, J. (1998a). Structural relations, cointegration and identification: some simple results and
their application. Journal of econometrics, 87(1):87–113.
Davidson, J. (1998b). A wald test of restrictions on the cointegrating space based on johansen’s estima-
tor. Economics Letters, 59(2):183–187.
Davis, R. and Ng, S. (2023). Time series estimation of the dynamic effects of disaster-type shocks.
Journal of Econometrics, 235(1):180–201.
Davis, R. A. and Song, L. (2020). Noncausal vector ar processes with application to economic time
series. Journal of Econometrics, 216(1):246–267.
De Chaisemartin, C. and d’Haultfoeuille, X. (2020). Two-way fixed effects estimators with heteroge-
neous treatment effects. American Economic Review, 110(9):2964–2996.
Deistler, M. (1983). The properties of the parameterization of armax systems and their relevance for
structural estimation and dynamic specification. Econometrica: Journal of the Econometric Society,
pages 1187–1207.
Demirer, M., Diebold, F. X., Liu, L., and Yilmaz, K. (2018). Estimating global bank network connect-
edness. Journal of Applied Econometrics, 33(1):1–15.
Dhrymes, P. (2013). Mathematics for econometrics. Springer.
Dias, G. F. and Kapetanios, G. (2018). Estimation and forecasting in vector autoregressive moving
average models for rich datasets. Journal of Econometrics, 202(1):75–91.
Dickey, D. A., Bell, W. R., and Miller, R. B. (1986). Unit roots in time series models: Tests and
implications. American statistician, pages 12–26.
Diebold, F. X. and Inoue, A. (2001). Long memory and regime switching. Journal of econometrics,
105(1):131–159.
Diebold, F. X. and Yilmaz, K. (2012). Better to give than to receive: Predictive directional measurement
115
of volatility spillovers. International Journal of forecasting, 28(1):57–66.
Diebold, F. X. and Yılmaz, K. (2014). On the network topology of variance decompositions: Measuring
the connectedness of financial firms. Journal of econometrics, 182(1):119–134.
Dolado, J. J. (1992). A note on weak exogeneity in var cointegrated models. Economics Letters,
38(2):139–143.
Drautzburg, T. and Wright, J. H. (2023). Refining set-identification in vars through independence.
Journal of Econometrics.
Duffy, J. A., Mavroeidis, S., and Wycherley, S. (2022). Cointegration with occasionally binding con-
straints. arXiv preprint arXiv:2211.09604.
Dufour, J.-M. (2003). Identification, weak instruments, and statistical inference in econometrics. Cana-
dian Journal of Economics/Revue canadienne d’économique, 36(4):767–808.
Dufour, J.-M. and Taamouti, M. (2005). Projection-based statistical inference in linear structural models
with possibly weak instruments. Econometrica, 73(4):1351–1365.
Eggertsson, G., Ferrero, A., and Raffo, A. (2014). Can structural reforms help europe? Journal of
Monetary Economics, 61:2–22.
Eickmeier, S., Lemke, W., and Marcellino, M. (2015). Classical time varying factor-augmented vector
auto-regressive models—estimation, forecasting and structural analysis. Journal of the Royal Statis-
tical Society Series A: Statistics in Society, 178(3):493–533.
Engle, R. F., Hendry, D. F., and Richard, J.-F. (1983). Exogeneity. Econometrica: Journal of the
Econometric Society, pages 277–304.
Eriksson, J. and Koivunen, V. (2004). Identifiability, separability, and uniqueness of linear ica models.
IEEE signal processing letters, 11(7):601–604.
Fang, X., Li, J., and Siegmund, D. (2020). Segmentation and estimation of change-point models: false
positive control and confidence regions. The Annals of Statistics, 48(3):1615–1647.
Feldkircher, M., Huber, F., and Pfarrhofer, M. (2020). Factor augmented vector autoregressions, panel
vars, and global vars. Macroeconomic Forecasting in the Era of Big Data: Theory and Practice,
pages 65–93.
Fève, P. and Guay, A. (2009). The response of hours to a technology shock: A two-step structural var
approach. Journal of Money, Credit and Banking, 41(5):987–1013.
Feve, P. and Guay, A. (2010). Identification of technology shocks in structural vars. The Economic
Journal, 120(549):1284–1318.
Findley, D. F. (1986). The uniqueness of moving average representations with independent and iden-
tically distributed random variables for non-gaussian stationary time series. Biometrika, 73(2):520–
521.
Forni, M. and Gambetti, L. (2014). Sufficient information in structural vars. Journal of Monetary
Economics, 66:124–136.
Forni, M., Gambetti, L., and Sala, L. (2023). Macroeconomic uncertainty and vector autoregressions.
Econometrics and Statistics.
Forni, M., Hallin, M., Lippi, M., and Zaffaroni, P. (2017). Dynamic factor models with infinite-
dimensional factor space: Asymptotic analysis. Journal of Econometrics, 199(1):74–92.
116
Fragetta, M. and Melina, G. (2011). The effects of fiscal policy shocks in svar models: a graphical
modelling approach. Scottish Journal of Political Economy, 58(4):537–566.
Francis, N. and Ramey, V. A. (2005). Is the technology-driven real business cycle hypothesis dead?
shocks and aggregate fluctuations revisited. Journal of Monetary Economics, 52(8):1379–1399.
Freyberger, J., Neuhierl, A., and Weber, M. (2020). Dissecting characteristics nonparametrically. The
Review of Financial Studies, 33(5):2326–2377.
Fries, S. and Zakoian, J.-M. (2019). Mixed causal-noncausal ar processes and the modelling of explosive
bubbles. Econometric Theory, 35(6):1234–1270.
Fry, R. and Pagan, A. (2011). Sign restrictions in structural vector autoregressions: A critical review.
Journal of Economic Literature, 49(4):938–960.
Funovits, B. (2020). Identifiability and estimation of possibly non-invertible svarma models: A new
parametrisation. arXiv preprint arXiv:2002.04346.
Gabaix, X. (2011). The granular origins of aggregate fluctuations. Econometrica, 79(3):733–772.
Gabaix, X. and Koijen, R. S. (2023). Granular instrumental variables. Journal of Political Economy
(forthcoming).
Gafarov, B., Meier, M., and Montiel Olea, J. L. (2018). Delta-method inference for a class of set-
identified svars. Journal of Econometrics, 203(2):316–327.
Gali, J. (1999). Technology, employment, and the business cycle: do technology shocks explain aggre-
gate fluctuations? American economic review, 89(1):249–271.
Ganics, G., Inoue, A., and Rossi, B. (2021). Confidence intervals for bias and size distortion in iv and
local projections-iv models. Journal of Business & Economic Statistics, 39(1):307–324.
Garratt, A., Robertson, D., and Wright, S. (2006). Permanent vs transitory components and economic
fundamentals. Journal of Applied Econometrics, 21(4):521–542.
Geraci, M. V. and Gnabo, J.-Y. (2018). Measuring interconnectedness between financial institutions
with bayesian time-varying vector autoregressions. Journal of Financial and Quantitative Analysis,
53(3):1371–1390.
Giacomini, R. and Kitagawa, T. (2021). Robust bayesian inference for set-identified models. Econo-
metrica, 89(4):1519–1556.
Giacomini, R., Kitagawa, T., and Read, M. (2022). Robust bayesian inference in proxy svars. Journal
of Econometrics, 228(1):107–126.
Giraitis, L., Kapetanios, G., and Marcellino, M. (2021). Time-varying instrumental variable estimation.
Journal of Econometrics, 224(2):394–415.
Goes, C. (2016). Institutions and growth: A gmm/iv panel var approach. Economics letters, 138:85–91.
Gomez Cram, R. and Olbert, M. (2023). Measuring the expected effects of the global tax reform. Review
of Financial Studies.
Gonzalo, J. and Granger, C. (1995). Estimation of common long-memory components in cointegrated
systems. Journal of Business & Economic Statistics, 13(1):27–35.
Gonzalo, J. and Ng, S. (2001). A systematic framework for analyzing the dynamic effects of permanent
and transitory shocks. Journal of Economic dynamics and Control, 25(10):1527–1546.
Gorodnichenko, Y. and Lee, B. (2020). Forecast error variance decompositions with local projections.
117
Journal of Business & Economic Statistics, 38(4):921–933.
Gospodinov, N. (2010). Inference in nearly nonstationary svar models with long-run identifying restric-
tions. Journal of Business & Economic Statistics, 28(1):1–12.
Gourieroux, C. and Monfort, A. (1997). Time series and dynamic models, volume 3. Cambridge
University Press.
Gourieroux, C., Monfort, A., and Renne, J.-P. (2017). Statistical inference for independent component
analysis: Application to structural var models. Journal of Econometrics, 196(1):111–126.
Gouriéroux, C., Monfort, A., and Renne, J.-P. (2020). Identification and estimation in non-fundamental
structural varma models. The Review of Economic Studies, 87(4):1915–1953.
Granziera, E., Moon, H. R., and Schorfheide, F. (2018). Inference for vars identified with sign restric-
tions. Quantitative Economics, 9(3):1087–1121.
Greenaway-McGrevy, R. (2013). Multistep prediction of panel vector autoregressive processes. Econo-
metric Theory, 29(4):699–734.
Greenaway-McGrevy, R. (2020). Multistep forecast selection for panel data. Econometric Reviews,
39(4):373–406.
Grigoriu, M. (1995). Linear and nonlinear systems with non-gaussian white noise input. Probabilistic
engineering mechanics, 10(3):171–179.
Guay, A. (2021). Identification of structural vector autoregressions through higher unconditional mo-
ments. Journal of Econometrics, 225(1):27–46.
Guo, F. and Ling, S. (2023). Inference for the vec (1) model with a heavy-tailed linear process errors.
Econometric Reviews, 42(9-10):806–833.
Hallin, M. and Mehta, C. (2015). R-estimation for asymmetric independent component analysis. Jour-
nal of the American Statistical Association, 110(509):218–232.
Hallin, M. and Paindaveine, D. (2004). Rank-based optimal tests of the adequacy of an elliptic varma
model. The Annals of Statistics, 32(6):2642–2678.
Hamilton, J. D. (1994). Time series analysis. Princeton university press.
Han, X. (2015). Tests for overidentifying restrictions in factor-augmented var models. Journal of
Econometrics, 184(2):394–419.
Han, X. (2018). Estimation and inference of dynamic structural factor models with over-identifying
restrictions. Journal of Econometrics, 202(2):125–147.
Hannadige, S. B., Gao, J., Silvapulle, M. J., and Silvapulle, P. (2023). Forecasting a nonstationary time
series using a mixture of stationary and nonstationary factors as predictors. Journal of Business &
Economic Statistics, pages 1–13.
Hansen, B. E. and Seo, B. (2002). Testing for two-regime threshold cointegration in vector error-
correction models. Journal of econometrics, 110(2):293–318.
Hatchondo, J. C., Martinez, L., and Sosa-Padilla, C. (2016). Debt dilution and sovereign default risk.
Journal of Political Economy, 124(5):1383–1422.
Hausman, J. A. and Taylor, W. E. (1983). Identification in linear simultaneous equations models with co-
variance restrictions: An instrumental variables interpretation. Econometrica: Journal of the Econo-
metric Society, pages 1527–1549.
118
Hecq, A., Palm, F. C., and Urbain, J.-P. (2000). Permanent-transitory decomposition in var models with
cointegration and common cycles. Oxford Bulletin of Economics and Statistics, 62(4):511–532.
Hecq, A. and Voisin, E. (2021). Forecasting bubbles with mixed causal-noncausal autoregressive mod-
els. Econometrics and Statistics, 20:29–45.
Hendry, D. F. (1994). On the interactions of unit roots and exogeneity. Econometric Reviews, 14(4):383–
419.
Hodoshima, J. (1988). Estimation of a single structural equation with structural change. Econometric
Theory, 4(1):86–96.
Hoover, K. D. (2020). The discovery of long-run causal order: A preliminary investigation. Economet-
rics, 8(3):31.
Hsiao, C. (1997). Statistical properties of the two-stage least squares estimator under cointegration. The
Review of Economic Studies, 64(3):385–398.
Hsu, Y.-C., Lai, T.-C., and Lieli, R. P. (2022). Counterfactual treatment effects: Estimation and infer-
ence. Journal of Business & Economic Statistics, 40(1):240–255.
Huber, F., Krisztin, T., and Pfarrhofer, M. (2023). A bayesian panel vector autoregression to analyze the
impact of climate shocks on high-income economies. The Annals of Applied Statistics, 17(2):1543–
1573.
Huber, K. (2018). Disentangling the effects of a banking crisis: Evidence from german firms and
counties. American Economic Review, 108(3):868–898.
Inoue, A. and Kilian, L. (2013). Inference on impulse response functions in structural var models.
Journal of Econometrics, 177(1):1–13.
Inoue, A. and Kilian, L. (2016). Joint confidence sets for structural impulse responses. Journal of
Econometrics, 192(2):421–432.
Inoue, A. and Kilian, L. (2020). The uniform validity of impulse response inference in autoregressions.
Journal of Econometrics, 215(2):450–472.
Ivanov, V. and Kilian, L. (2005). A practitioner’s guide to lag order selection for var impulse response
analysis. Studies in Nonlinear Dynamics & Econometrics, 9(1).
Johansen, S. (1995). Identifying restrictions of linear equations with applications to simultaneous equa-
tions and cointegration. Journal of econometrics, 69(1):111–132.
Jorda, O. (2005). Estimation and inference of impulse responses by local projections. American eco-
nomic review, 95(1):161–182.
Jurado, K., Ludvigson, S. C., and Ng, S. (2015). Measuring uncertainty. American Economic Review,
105(3):1177–1216.
Kagan, A. M., Linnik, I., and Rao, C. R. (1973). Characterization problems in mathematical statistics.
(No Title).
Kalliovirta, L., Meitz, M., and Saikkonen, P. (2015). A gaussian mixture autoregressive model for
univariate time series. Journal of Time Series Analysis, 36(2):247–266.
Kalliovirta, L., Meitz, M., and Saikkonen, P. (2016). Gaussian mixture vector autoregression. Journal
of econometrics, 192(2):485–498.
Kang, I.-B. (2003). Multi-period forecasting using different models for different horizons: an applica-
119
tion to us economic time series data. International Journal of Forecasting, 19(3):387–400.
Kao, C., Trapani, L., and Urga, G. (2018). Testing for instability in covariance structures. Bernoulli,
24(1):740–771.
Kapetanios, G., Marcellino, M., and Venditti, F. (2019). Large time-varying parameter vars: A nonpara-
metric approach. Journal of Applied Econometrics, 34(7):1027–1049.
Katsouris, C. (2021). Forecast evaluation in large cross-sections of realized volatility. arXiv preprint
arXiv:2112.04887.
Katsouris, C. (2023a). Estimating conditional value-at-risk with nonstationary quantile predictive re-
gression models. arXiv preprint arXiv:2311.08218.
Katsouris, C. (2023b). Limit theory under network dependence and nonstationarity. arXiv preprint
arXiv:2308.01418.
Keweloh, S. A. (2021). A generalized method of moments estimator for structural vector autoregressions
based on higher moments. Journal of Business & Economic Statistics, 39(3):772–782.
Keweloh, S. A., Hetzenecker, S., and Seepe, A. (2023). Monetary policy and information shocks in a
block-recursive svar. Journal of International Money and Finance, page 102892.
Khalaf, L. and Urga, G. (2014). Identification robust inference in cointegrating regressions. Journal of
Econometrics, 182(2):385–396.
Kiiveri, H., Speed, T. P., and Carlin, J. B. (1984). Recursive causal models. Journal of the australian
Mathematical Society, 36(1):30–52.
Kilian, L. (2009). Not all oil price shocks are alike: Disentangling demand and supply shocks in the
crude oil market. American Economic Review, 99(3):1053–1069.
Kilian, L. and Kim, Y. J. (2011). How reliable are local projection estimators of impulse responses?
Review of Economics and Statistics, 93(4):1460–1466.
Kilian, L. and Lee, T. K. (2014). Quantifying the speculative component in the real price of oil: The
role of global oil inventories. Journal of International Money and Finance, 42:71–87.
Kilian, L. and Lewis, L. T. (2011). Does the fed respond to oil price shocks? The Economic Journal,
121(555):1047–1072.
Kilian, L. and Lütkepohl, H. (2017). Structural vector autoregressive analysis. Cambridge University
Press.
Kilian, L. and Park, C. (2009). The impact of oil price shocks on the us stock market. International
economic review, 50(4):1267–1287.
Kim, S. and Roubini, N. (2000). Exchange rate anomalies in the industrial countries: A solution with a
structural var approach. Journal of Monetary economics, 45(3):561–586.
Kleibergen, F. and Van Dijk, H. K. (1994). On the shape of the likelihood/posterior in cointegration
models. Econometric theory, 10(3-4):514–551.
Kociecki, A. (2011). Algebraic theory of indentification in parametric models. National Bank of Poland
Working Paper, (88).
Kocikecki, A. and Kolasa, M. (2018). Global identification of linearized dsge models. Quantitative
Economics, 9(3):1243–1263.
Kocikecki, A. and Kolasa, M. (2023). A solution to the global identification problem in dsge models.
120
Journal of Econometrics, 236(2):105477.
Koistinen, J. and Funovits, B. (2022). Estimation of impulse-response functions with dynamic factor
models: A new parametrization. arXiv preprint arXiv:2202.00310.
Komunjer, I. and Ng, S. (2011). Dynamic identification of dynamic stochastic general equilibrium
models. Econometrica, 79(6):1995–2032.
Kong, J., Phillips, P. C., and Sul, D. (2019). Weak σ -convergence: Theory and applications. Journal of
Econometrics, 209(2):185–207.
Koo, B., Lee, S., and Seo, M. H. (2022). What impulse response do instrumental variables identify?
arXiv preprint arXiv:2208.11828.
Koop, G. and Korobilis, D. (2013). Large time-varying parameter vars. Journal of Econometrics,
177(2):185–198.
Korobilis, D. and Yilmaz, K. (2018). Measuring dynamic connectedness with large bayesian var models.
Available at SSRN 3099725.
Krampe, J., Paparoditis, E., and Trenkler, C. (2023). Structural inference in sparse high-dimensional
vector autoregressions. Journal of Econometrics, 234(1):276–300.
Krolzig, H.-M. (2003). General-to-specific model selection procedures for structural vector autoregres-
sions. Oxford Bulletin of Economics and Statistics, 65:769–801.
Kuersteiner, G. M. (2005). Automatic inference for infinite order vector autoregressions. Econometric
Theory, 21(1):85–115.
Kyle, A. S., Obizhaeva, A. A., and Wang, Y. (2023). Beliefs aggregation and return predictability. The
Journal of Finance, 78(1):427–486.
Laborda, R. and Olmo, J. (2021). Volatility spillover between economic sectors in financial crisis predic-
tion: Evidence spanning the great financial crisis and covid-19 pandemic. Research in International
Business and Finance, 57:101402.
Lai, T. L. and Wei, C. Z. (1982). Least squares estimates in stochastic regression models with applica-
tions to identification and control of dynamic systems. The Annals of Statistics, 10(1):154–166.
Lanne, M. (2000). Near unit roots, cointegration, and the term structure of interest rates. Journal of
Applied Econometrics, 15(5):513–529.
Lanne, M., Liu, K., and Luoto, J. (2023). Identifying structural vector autoregression via leptokurtic
economic shocks. Journal of Business & Economic Statistics, 41(4):1341–1351.
Lanne, M. and Luoto, J. (2021). Gmm estimation of non-gaussian structural vector autoregression.
Journal of Business & Economic Statistics, 39(1):69–81.
Lanne, M., Luoto, J., and Saikkonen, P. (2012). Optimal forecasting of noncausal autoregressive time
series. International Journal of Forecasting, 28(3):623–631.
Lanne, M. and Lütkepohl, H. (2008a). Identifying monetary policy shocks via changes in volatility.
Journal of Money, Credit and Banking, 40(6):1131–1149.
Lanne, M. and Lütkepohl, H. (2008b). A statistical comparison of alternative identification schemes for
monetary policy shocks.
Lanne, M., Lütkepohl, H., and Maciejowska, K. (2010). Structural vector autoregressions with markov
switching. Journal of Economic Dynamics and Control, 34(2):121–131.
121
Lanne, M., Meitz, M., and Saikkonen, P. (2017). Identification and estimation of non-gaussian structural
vector autoregressions. Journal of Econometrics, 196(2):288–304.
Lanne, M. and Nyberg, H. (2016). Generalized forecast error variance decomposition for linear and
nonlinear multivariate models. Oxford Bulletin of Economics and Statistics, 78(4):595–603.
Lanne, M. and Saikkonen, P. (2002). Threshold autoregressions for strongly autocorrelated time series.
Journal of Business & Economic Statistics, 20(2):282–289.
Lanne, M. and Saikkonen, P. (2013). Noncausal vector autoregression. Econometric Theory, 29(3):447–
481.
Lastrapes, W. D. (2005). Estimating and identifying vector autoregressions under diagonality and block
exogeneity restrictions. Economics letters, 87(1):75–81.
Lawrance, A. (1991). Directionality and reversibility in time series. International statistical re-
view/revue internationale de statistique, pages 67–79.
Lee, S., Lee, S., and Maekawa, K. (2023). Change point test for structural vector autoregressive model
via independent component analysis. Journal of Statistical Computation and Simulation, 93(5):687–
707.
Leeper, E. M., Walker, T. B., and Yang, S.-C. S. (2013). Fiscal foresight and information flows. Econo-
metrica, 81(3):1115–1145.
Lewis, G. and Syrgkanis, V. (2020). Double/debiased machine learning for dynamic treatment effects
via g-estimation. arXiv preprint arXiv:2002.07285.
LI, W. K. and Hui, Y. (1989). Robust multiple time series modelling. Biometrika, 76(2):309–315.
Liao, Z. and Phillips, P. C. (2015). Automated estimation of vector error correction models. Econometric
Theory, 31(3):581–646.
Lovcha, Y. and Perez-Laborda, A. (2021). Identifying technology shocks at the business cycle via
spectral variance decompositions. Macroeconomic Dynamics, 25(8):1966–1992.
Lusompa, A. (2023). Local projections, autocorrelation, and efficiency. Quantitative Economics,
14(4):1199–1220.
Lütkepohl, H. (1982). Non-causality due to omitted variables. Journal of Econometrics, 19(2-3):367–
378.
Lütkepohl, H. (1990). Asymptotic distributions of impulse response functions and forecast error vari-
ance decompositions of vector autoregressive models. The review of economics and statistics, pages
116–125.
Lütkepohl, H. (2002). Forecasting cointegrated varma processes. A Companion to Economic Forecast-
ing, Blackwell, Oxford, pages 179–205.
Lütkepohl, H. (2005). New introduction to multiple time series analysis. Springer Science & Business
Media.
Lütkepohl, H., Meitz, M., Netšunajev, A., and Saikkonen, P. (2021). Testing identification via het-
eroskedasticity in structural vector autoregressive models. The Econometrics Journal, 24(1):1–22.
Lütkepohl, H. and Saikkonen, P. (1997). Impulse response analysis in infinite order cointegrated vector
autoregressive processes. Journal of Econometrics, 81(1):127–157.
Magnusson, L. M. and Mavroeidis, S. (2014). Identification using stability restrictions. Econometrica,
122
82(5):1799–1851.
Marcellino, M., Stock, J. H., and Watson, M. W. (2006). A comparison of direct and iterated multistep
ar methods for forecasting macroeconomic time series. Journal of econometrics, 135(1-2):499–526.
Mariano, R. S. (2002). Testing forecast accuracy. A companion to economic forecasting, 2:284–298.
Mavroeidis, S. (2021). Identification at the zero lower bound. Econometrica, 89(6):2855–2885.
Maxand, S. (2020). Identification of independent structural shocks in the presence of multiple gaussian
components. Econometrics and Statistics, 16:55–68.
Mei, Z., Sheng, L., and Shi, Z. (2023). Implicit nickell bias in panel local projection. arXiv preprint
arXiv:2302.13455.
Meitz, M., Preve, D., and Saikkonen, P. (2023). A mixture autoregressive model based on student’st–
distribution. Communications in Statistics-Theory and Methods, 52(2):499–515.
Meitz, M. and Saikkonen, P. (2013). Maximum likelihood estimation of a noninvertible arma model
with autoregressive conditional heteroskedasticity. Journal of Multivariate Analysis, 114:227–255.
Meitz, M. and Saikkonen, P. (2021). Testing for observation-dependent regime switching in mixture
autoregressive models. Journal of Econometrics, 222(1):601–624.
Mélard, G., Roy, R., and Saidi, A. (2006). Exact maximum likelihood estimation of structured or unit
root multivariate time series models. Computational statistics & data analysis, 50(11):2958–2986.
Menchetti, F., Cipollini, F., and Mealli, F. (2023). Combining counterfactual outcomes and arima mod-
els for policy evaluation. The Econometrics Journal, 26(1):1–24.
Mensi, W., Aslan, A., Vo, X. V., and Kang, S. H. (2023). Time-frequency spillovers and connected-
ness between precious metals, oil futures and financial markets: Hedge and safe haven implications.
International Review of Economics & Finance, 83:219–232.
Mertens, K. and Ravn, M. O. (2010). Measuring the impact of fiscal policy in the face of anticipation:
a structural var approach. The Economic Journal, 120(544):393–413.
Mertens, K. and Ravn, M. O. (2013). The dynamic effects of personal and corporate income tax changes
in the united states. American economic review, 103(4):1212–1247.
Miao, K., Phillips, P. C., and Su, L. (2023). High-dimensional vars with common factors. Journal of
Econometrics, 233(1):155–183.
Mikusheva, A. (2007). Uniform inference in autoregressive models. Econometrica, 75(5):1411–1452.
Mittnik, S. and Zadrozny, P. A. (1993). Asymptotic distributions of impulse responses, step responses,
and variance decompositions of estimated linear dynamic models. Econometrica: Journal of the
Econometric Society, pages 857–870.
Moneta, A., Chlaß, N., Entner, D., and Hoyer, P. (2011). Causal search in structural vector autoregres-
sive models. In NIPS Mini-Symposium on Causality in Time Series, pages 95–114.
Moneta, A., Entner, D., Hoyer, P. O., and Coad, A. (2013). Causal inference by independent component
analysis: Theory and applications. Oxford Bulletin of Economics and Statistics, 75(5):705–730.
Montiel Olea, J. L. and Plagborg-Moller, M. (2021). Local projection inference is simpler and more
robust than you think. Econometrica, 89(4):1789–1823.
Montiel Olea, J. L., Stock, J. H., and Watson, M. W. (2021). Inference in structural vector autoregres-
sions identified with an external instrument. Journal of Econometrics, 225(1):74–87.
123
Moon, H. R. and Schorfheide, F. (2002). Minimum distance estimation of nonstationary time series
models. Econometric Theory, 18(6):1385–1407.
Moon, H. R. and Schorfheide, F. (2012). Bayesian and frequentist inference in partially identified
models. Econometrica, 80(2):755–782.
Morley, J. and Piger, J. (2012). The asymmetric business cycle. Review of Economics and Statistics,
94(1):208–221.
Morris, S. A. (1979). Duality and structure of locally compact abelian groups. Math Chronicle, 8.
Müller, U. K. and Petalas, P.-E. (2010). Efficient estimation of the parameter path in unstable time series
models. The Review of Economic Studies, 77(4):1508–1539.
Mumtaz, H., Pinter, G., and Theodoridis, K. (2018). What do vars tell us about the impact of a credit
supply shock? International Economic Review, 59(2):625–646.
Mumtaz, H. and Surico, P. (2018). Policy uncertainty and aggregate fluctuations. Journal of Applied
Econometrics, 33(3):319–331.
Murasawa, Y. (2015). The multivariate beveridge–nelson decomposition with i (1) and i (2) series.
Economics Letters, 137:157–162.
Myers, R. J., Johnson, S. R., Helmar, M., and Baumes, H. (2018). Long-run and short-run relationships
between oil prices, producer prices, and consumer prices: What can we learn from a permanent-
transitory decomposition? The Quarterly Review of Economics and Finance, 67:175–190.
Ng, C. N. and Young, P. C. (1990). Recursive estimation and forecasting of non-stationary time series.
Journal of Forecasting, 9(2):173–204.
Ng, S. and Perron, P. (2001). Lag length selection and the construction of unit root tests with good size
and power. Econometrica, 69(6):1519–1554.
Noh, E. (2017). Impulse-response analysis with proxy variables. Available at SSRN 3070401.
Nyberg, H. (2018). Forecasting us interest rates and business cycle with a nonlinear regime switching
var model. Journal of Forecasting, 37(1):1–15.
Nyberg, H. and Rauhala, S. (2022). A structural vector autoregression containing positive-valued com-
ponents. Available at SSRN.
Nyberg, H. and Savva, C. S. (2023). Risk-return trade-off in international stock returns: Skewness and
business cycles. Econometrics and Statistics.
Pagan, A. R. and Pesaran, M. H. (2008). Econometric analysis of structural systems with permanent
and transitory shocks. Journal of Economic Dynamics and control, 32(10):3376–3395.
Paparoditis, E. (2018). Sieve bootstrap for functional time series. The annals of Statistics, 46(6B):3510–
3538.
Paruolo, P. (1997). Asymptotic inference on the moving average impact matrix in cointegrated 1 (1) var
systems. Econometric Theory, 13(1):79–118.
Paruolo, P. (2000). Asymptotic efficiency of the two stage estimator in i (2) systems. Econometric
Theory, 16(4):524–550.
Paruolo, P. and Rahbek, A. (1999). Weak exogeneity in i (2) var systems. Journal of Econometrics,
93(2):281–308.
Paul, P. (2020). The time-varying effect of monetary policy on asset prices. Review of Economics and
124
Statistics, 102(4):690–704.
Pesaran, M. H., Shin, Y., and Smith, R. J. (2000). Structural analysis of vector error correction models
with exogenous i (1) variables. Journal of econometrics, 97(2):293–343.
Petrova, K. (2019). A quasi-bayesian local likelihood approach to time varying parameter var models.
Journal of Econometrics, 212(1):286–306.
Petrova, K. (2022). Asymptotically valid bayesian inference in the presence of distributional misspeci-
fication in var models. Journal of Econometrics, 230(1):154–182.
Petursson, T. G. and Slok, T. (2001). Wage formation and employment in a cointegrated var model. The
Econometrics Journal, 4(2):191–209.
Phillips, P. C. B. (1987). Towards a unified asymptotic theory for autoregression. Biometrika,
74(3):535–547.
Phillips, P. C. B. (1989). Partially identified econometric models. Econometric Theory, 5(2):181–240.
Phillips, P. C. B. (1991). Optimal inference in cointegrated systems. Econometrica: Journal of the
Econometric Society, pages 283–306.
Phillips, P. C. B. (1998). Impulse response and forecast error variance asymptotics in nonstationary
vars. Journal of econometrics, 83(1-2):21–56.
Phillips, P. C. B. and Hansen, B. E. (1990). Statistical inference in instrumental variables regression
with i (1) processes. The review of economic studies, 57(1):99–125.
Phillips, P. C. B., Leirvik, T., and Storelvmo, T. (2020). Econometric estimates of earth’s transient
climate sensitivity. Journal of Econometrics, 214(1):6–32.
Plagborg-Møller, M. and Wolf, C. K. (2021). Local projections and vars estimate the same impulse
responses. Econometrica, 89(2):955–980.
Plagborg-Møller, M. and Wolf, C. K. (2022). Instrumental variable identification of dynamic variance
decompositions. Journal of Political Economy, 130(8):2164–2202.
Poskitt, D. S. (2006). On the identification and estimation of nonstationary and cointegrated armax
systems. Econometric Theory, 22(6):1138–1175.
Pretis, F. (2020). Econometric modelling of climate systems: The equivalence of energy balance models
and cointegrated vector autoregressions. Journal of Econometrics, 214(1):256–273.
Pretis, F. (2021). Exogeneity in climate econometrics. Energy Economics, 96:105122.
Primiceri, G. E. (2005). Time varying structural vector autoregressions and monetary policy. The Review
of Economic Studies, 72(3):821–852.
Qian, E. (2023). Heterogeneity-robust granular instruments. arXiv preprint arXiv:2304.01273.
Quah, D. (1992). The relative importance of permanent and transitory components: identification and
some theoretical bounds. Econometrica: Journal of the Econometric Society, pages 107–118.
Ramey, V. A. (2011). Identifying government spending shocks: It’s all in the timing. The Quarterly
Journal of Economics, 126(1):1–50.
Rehman, M. U., Shahzad, S. J. H., Uddin, G. S., and Hedström, A. (2018). Precious metal returns and
oil shocks: A time varying connectedness approach. Resources Policy, 58:77–89.
Romer, C. D. and Romer, D. H. (2017). New evidence on the aftermath of financial crises in advanced
countries. American Economic Review, 107(10):3072–3118.
125
Rousseeuw, P. J. (1984). Least median of squares regression. Journal of the American statistical
association, 79(388):871–880.
Rubio-Ramirez, J. F., Waggoner, D. F., and Zha, T. (2010). Structural vector autoregressions: Theory
of identification and algorithms for inference. The Review of Economic Studies, 77(2):665–696.
Rubio-Ramirez, J. F., Waggoner, D. F., and Zha, T. A. (2005). Markov-switching structural vector
autoregressions: theory and application.
Rygh Swensen, A. (2022). On causal and non-causal cointegrated vector autoregressive time series.
Journal of Time Series Analysis, 43(2):178–196.
Safikhani, A., Bai, Y., and Michailidis, G. (2022). Fast and scalable algorithm for detection of structural
breaks in big var models. Journal of Computational and Graphical Statistics, 31(1):176–189.
Saikkonen, P. (1995). Problems with the asymptotic theory of maximum likelihood estimation in inte-
grated and cointegrated systems. Econometric Theory, 11(5):888–911.
Samuelson, P. A. (1941). Conditions that the roots of a polynomial be less than unity in absolute value.
The Annals of Mathematical Statistics, 12(3):360–364.
Sargan, J. D. (1983). Identification and lack of identification. Econometrica: Journal of the Econometric
Society, pages 1605–1633.
Sargent, T. J. (1976). The observational equivalence of natural and unnatural rate theories of macroeco-
nomics. Journal of Political Economy, 84(3):631–640.
Schlaak, T., Rieth, M., and Podstawski, M. (2023). Monetary policy, external instruments, and het-
eroskedasticity. Quantitative Economics, 14(1):161–200.
Sengupta, D. and Kay, S. (1989). Efficient estimation of parameters for non-gaussian autoregressive
processes. IEEE Transactions on Acoustics, Speech, and Signal Processing, 37(6):785–794.
She, R., Mi, Z., and Ling, S. (2022). Whittle parameter estimation for vector arma models with heavy-
tailed noises. Journal of Statistical Planning and Inference, 219:216–230.
Silva, R. and Shimizu, S. (2017). Learning instrumental variables with structural and non-gaussianity
assumptions. Journal of Machine Learning Research, 18(120):1–49.
Sims, E. R. (2012). News, non-invertibility, and structural vars. In DSGE Models in Macroeconomics:
Estimation, Evaluation, and New Developments, pages 81–135. Emerald Group Publishing Limited.
Soderstrom, T., Ljung, L., and Gustavsson, I. (1978). A theoretical analysis of recursive identification
methods. Automatica, 14(3):231–244.
Soenen, N. and Vander Vennet, R. (2022). Ecb monetary policy and bank default risk. Journal of
International Money and Finance, 122:102571.
Stock, J. and Watson, M. (2021). Inference in svars identified with external instruments. Journal of
Econometrics, 225:74–87.
Stock, J. H. and Watson, M. W. (1988). Testing for common trends. Journal of the American statistical
Association, 83(404):1097–1107.
Stock, J. H. and Watson, M. W. (2001). Vector autoregressions. Journal of Economic perspectives,
15(4):101–115.
Stock, J. H. and Watson, M. W. (2018). Identification and estimation of dynamic causal effects in
macroeconomics using external instruments. The Economic Journal, 128(610):917–948.
126
Swanson, N. R. and Granger, C. W. (1997). Impulse response functions based on a causal approach to
residual orthogonalization in vector autoregressions. Journal of the American Statistical Association,
92(437):357–367.
Tao, Y. and Yu, J. (2020). Model selection for explosive models. In Essays in honor of Cheng Hsiao,
volume 41, pages 73–103. Emerald Publishing Limited.
Tchatoka, F. D. and Dufour, J.-M. (2013). On the finite-sample theory of exogeneity tests with possibly
non-gaussian errors and weak identification.
Tong, H. and Zhang, Z. (2005). On time-reversibility of multivariate linear processes. Statistica Sinica,
pages 495–504.
Trung, N. B. (2019). The spillover effects of us economic policy uncertainty on the global economy: A
global var approach. The North American Journal of Economics and Finance, 48:90–110.
Vegh, C. A. and Vuletin, G. (2015). How is tax policy conducted over the business cycle? American
Economic Journal: Economic Policy, 7(3):327–370.
Velasco, C. and Lobato, I. N. (2018). Frequency domain minimum distance inference for possibly
noninvertible and noncausal arma models. The Annals of Statistics, 46(2):555–579.
Veldkamp, L. L. (2005). Slow boom, sudden crash. Journal of Economic theory, 124(2):230–257.
Vermeulen, K. and Vansteelandt, S. (2015). Bias-reduced doubly robust estimation. Journal of the
American Statistical Association, 110(511):1024–1036.
Virolainen, S. (2020). Structural gaussian mixture vector autoregressive model with application to the
asymmetric effects of monetary policy shocks. arXiv preprint arXiv:2007.04713.
Vlaar, P. J. (2004). On the asymptotic distribution of impulse response functions with long-run restric-
tions. Econometric Theory, 20(5):891–903.
Watson, M. W. (1994). Vector autoregressions and cointegration. Handbook of econometrics, 4:2843–
2915.
White, H. and Pettenuzzo, D. (2014). Granger causality, exogeneity, cointegration, and economic policy
analysis. Journal of Econometrics, 178:316–330.
Wold, H. O. (1960). A generalization of causal chain models (part iii of a triptych on causal chain
systems). Econometrica: journal of the Econometric Society, pages 443–463.
Wright, S. (1934). The method of path coefficients. The annals of mathematical statistics, 5(3):161–215.
Xu, K.-L. (2023). Local projection based inference under general conditions. Available at SSRN
4372388.
Yamamoto, Y. and Horie, T. (2023). A cross-sectional method for right-tailed panic tests under a mod-
erately local to unity framework. Econometric Theory, 39(2):389–411.
Yoon, S.-M., Al Mamun, M., Uddin, G. S., and Kang, S. H. (2019). Network connectedness and net
spillover between financial and commodity markets. The North American Journal of Economics and
Finance, 48:801–818.
Zha, T. (1999). Block recursion and structural vector autoregressions. Journal of Econometrics,
90(2):291–316.
Zhang, Y. (2023). Statistical inference of high-dimensional vector autoregressive time series with non-
iid innovations. arXiv preprint arXiv:2310.07364.
127
Misspecifying the distributional assumptions leads to asymptotically invalid posterior inference, impacting key aspects like impulse responses and variance decompositions. Hence, it's crucial to employ Bayesian methods robust to distributional assumptions to ensure valid credible sets for structural parameters and other quantities of interest .
The main challenge is identifying relevant causal effects due to the absence of untreated controls and unobserved factors influencing outcomes. Approaches like synthetic control methods and ARIMA models can help by creating counterfactual predictions, but they still require careful consideration of potential biases due to these unobserved factors .
Identification of technology shocks in SVAR models allows economists to isolate the impact of technological changes on macroeconomic variables, providing insights into the propagation mechanisms of such shocks and their contributions to business cycle fluctuations .
Pre-treatment outcomes serve as covariates to balance the treatment and control units, ensuring that any observed treatment effect is attributable to the treatment itself rather than pre-existing differences. This balances the treated and control units, thus increasing the validity of causal inference .
The identity matrix signifies that the reduced-form errors have a diagonal variance-covariance matrix, indicating that each error term is uncorrelated with the others. This simplifies the model's estimation and interpretation by assuming that the errors are orthogonal and have unit variance .
Fundamentalness indicates that the vector process representation is invertible, meaning the structural shocks can be recovered uniquely from the observable variables. It ensures that the past can be forecasted using present information without ambiguity, which is critical for accurate impulse response analysis in econometric modeling .
ICA helps to identify structural shocks by exploiting statistical independence rather than the more traditional Gaussian assumptions. By maintaining the causal order of shocks, ICA provides a robust framework for recovering independent shocks, crucial for interpreting complex dynamic systems .
Dynamic panel data regression identifies causal effects by exploiting time variation and within-unit changes across panel data, accounting for individual heterogeneity and evolving dynamics over time. This methodology helps differentiate between contemporaneous and lagged effects of interventions, thereby effectively capturing the causal impact .
Identifiability ensures that the model parameters can be uniquely determined from the data. In linear ICA models, it is achieved through the assumption of statistical independence and certain mathematical constraints, such as non-Gaussianity. In structural econometric models, identifiability is often achieved using causal ordering and zero restrictions on structural parameters .
Autoregressive models accommodate dynamic causal effects by leveraging lagged variables to model dependencies over time. This allows the estimation of causal relationships as they evolve, capturing both contemporaneous and historical influences of treatment effects on outcomes .