Robust Tests for Factor-Augmented Forecasts
Robust Tests for Factor-Augmented Forecasts
1 University
of Bologna
2 BI Norwegian Business School
Abstract
We present four novel tests of equal predictive accuracy and encompassing for out-of-sample fore-
casts based on factor-augmented regression. We extend the work of Pitarakis (2023a,b) to develop the
inferential theory of predictive regressions with generated regressors which are estimated by using
Common Correlated Effects (henceforth CCE) - a technique that utilizes cross-sectional averages of
grouped series. It is particularly useful since large datasets of such structure are becoming increasingly
popular. Under our framework, CCE-based tests are asymptotically normal and robust to overspec-
ification of the number of factors, which is in stark contrast to existing methodologies in the CCE
context. Our tests are highly applicable in practice as they accommodate for different predictor types
(e.g., stationary and highly persistent factors), and remain invariant to the location of structural breaks
in loadings. Extensive Monte Carlo simulations indicate that our tests exhibit excellent local power
properties. Finally, we apply our tests to a novel EA-MD-QD dataset by Barigozzi et al. (2024b), which
covers Euro Area as a whole and primary member countries. We demonstrate that CCE factors offer a
substantial predictive power even under varying data persistence and structural breaks.
1 Introduction
Nowadays, forecasting by using so-called diffusion indexes has become increasingly popular thanks to
the growing availability of large datasets. In fact, this parsimonious approach allows forecasters to extract
the predictive content of many potential predictors into a reduced number of indexes, also referred to as
latent factors, when the data follows an approximate factor structure proposed by Chamberlain (1983).
Because the factors are unobserved, they must be estimated in the first step before augmenting the smaller
forecasting model (possibly of the autoregressive nature). The dominant method to estimate the factors is
Principal Components (PC) (see e.g. Bai and Ng, 2002). The simple yet powerful motivation behind this
setup is that the estimated factors have been shown to enhance the accuracy of economic forecasts em-
pirically (see Stock and Watson, 1999; Forni et al., 2003; Ludvigson and Ng, 2011 or Ciccarelli and Mojon,
2010 from an extensive literature). Although the rich content of predictive information embedded in
the factors represents an opportunity to make more accurate forecasts, it is essential to to evaluate their
1
predictive power formally. Fundamentally, the answer boils down to out-of-sample forecast evaluation
between the factor-augmented model and its nested baseline.
In this article, we build on recent advances in the forecasting literature: both methodological and data
advancements. We develop novel tests of forecast accuracy and encompassing for nested regressions with
latent factors. Our contribution is to extend the novel and extremely useful inferential theory of Pitarakis
(2023a,b) to the case of generated regressors. Specifically, we establish the conditions under which the
asymptotic normality of the proposed statistics hold even if the regressors are estimated, and many em-
pirical regularities hold: 1) persistent series, 2) presence of structural breaks, or 3) uncertainty about the
number of factors. We achieve this by revisiting the methodology Stauskas and Westerlund (2022), who
were the first to abandon the usual PC approach and instead adopted the Common Correlated Effects
(CCE) approach by Pesaran (2006) in forecast comparison setup.1 In contrary to extracting the eigenvec-
tors corresponding to the largest eigenvalues as factors, CCE aims to average out the noise in the data
and settle on the so-called low-rank component of the data, which is the factors identified up to a linear
transformation. In other words, the estimated CCE factors correspond to the cross-section averages (CAs)
formed from blocks of predictors in the large panel - a number of which is usually substantial.
Indeed, the target variable to forecast is typically a macroeconomic variable, such as inflation, pro-
duction, or a financial variable, such as excess returns on stocks or bonds. In these cases, it is standard
to organize the predictors into blocks of variables, such as consumption, money aggregates, prices, ex-
change rates and so on, each of which contains a large number of series. In fact, Eickmeier and Ziegler
(2008) consider 52 studies that use factor augmented-augmented regression models to forecast inflation
and/or output. The data sets employed in these studies all have the same structure with series being
organized in blocks of variables. Importantly, such structure is becoming increasingly popular, where a
standout example is FRED-MD (monthly) database by McCracken and Ng (2016) and its quarterly coun-
terpart FRED-QD (see McCracken and Ng, 2020). Recently, Barigozzi et al. (2024b) presented their Euro-
pean counterpart - a novel EA-MD-QD dataset. It covers the Euro Area as a whole and the main member
countries with 118 series organized in 11 blocks. Given that such datasets offer a large number of esti-
mated factors and cover a long time horizons, a robust procedure to evaluate their predictive content is
the key.
The study of Stauskas and Westerlund (2022) explores the statistics of Clark and McCracken (2001)
and McCracken (2007), which solve the degeneracy problem that occurs when the seminal tests of West
(1996) are implemented in the nested environment (see also a survey in Diebold (2015)). However,
Clark and McCracken (2001) and McCracken (2007) come at the expense of highly non-standard limit-
ing distributions that are based on functionals of stochastic integrals of Brownian motion and depend
on the relative growth rate of in-sample versus out-of-sample observations. Although simulation-based
approaches exist for estimating asymptotically valid critical values (see Clark and McCracken, 2012; or
Hansen and Timmermann, 2015), their practical implementation remains challenging. While the fact that
these test work under CCE-estimated factors and - very elegantly - that the degrees of freedom of the
asymptotic distributions correspond to the number of CAs (known) and not the true number factors
(typically unknown) is shown by Stauskas and Westerlund (2022), the non-standard inference and imple-
mentation issues remain inherited.
The key feature of tests of encompassing and forecast accuracy in Pitarakis (2023a,b) is that they put an
end to the joint problem of variance degeneracy and nonstandardness of asymptotics in nested models.
In fact, the proposed studentised statistics are not only well-defined under the null in the limit, but also
asymptotically distributed according to a standard normal free from nuisance parameters. However, the
inferential theory is established under the assumption that the predictors are observed and is silent on
the case where they are generated regressors. Therefore, we fill in the gap in the literature in the context
1 CCE has been used in forecasting before. For example, Karabiyik and Westerlund (2021) and Massacci and Kapetanios
2
of generated regressors estimated using the CCE methodology. When combined with the framework of
Pitarakis (2023a,b), CCE mechanics solve many issues that arise in Stauskas and Westerlund (2022). In
stark contrast to all existing tests of equal accuracy for factor-augmented regressions, our user-friendly
theoretical framework can flexibly accomodate for: 1) different degrees of persistency in the regressors, 2)
structural breaks in the factor loadings in the full sample and, perhaps most importantly, 3) overestima-
tion of the number of factors. To the best of our knowledge, this is the first result in the CCE literature,
where the redundant CAs do not have any asymptotic effect, irrespective of the expansion rates of N and
T. In spirit, this result resembles Moon and Weidner (2015) in the PC context. Furthermore, extensive
Monte Carlo experiments on size and local power exhibit excellent finite sample properties of our tests
under a plethora of realistic economic scenarios, thus making our methodology relevant and applicable
in many empirical applications. Ultimately, this study is the first to examine the predictive capacity of
factors in the EA-MD-QD dataset by Barigozzi et al. (2024b) in a nested setup. They offer powerful pre-
dictive insight in the context of macroeconomic forecasting especially in the presence of large structural
breaks such as the Covid pandemic.
We also highlight that this work goes in parallel to, and draws some interim results from, the PC-
based theory of Margaritella and Stauskas (2024), where the core focus is the impact of loading weakness.
Throughout we focus on one-step ahead forecasts so that we can impose the smallest set of regularity
conditions to accommodate all of proposed tests, but h-step ahead forecasts are possible for some tests
under additional conditions. Finally, our results hold under both the null hypothesis of equal predictive
ability as well as under the local alternative.
The rest of the paper is organized as follows. In Section 2 we introduce our forecasting setup, the
CCE approach and the tests under the observed predictors together with a list of assumptions. Section 3
dives into the CCE-based statistics and presents the main results. In Section 4, we conduct an extensive
Monte Carlo study. The notation adopted in the paper is as follows. Scalars, vectors and matrices are
denoted bypa ∈ R, a ∈ R p and A ∈ R c× p , respectively. For any matrix A, the Frobenius norm is defined
as kAk = tr(A′ A) where tr(.) is the trace operator, and kA − Bk signifies A = B + o p (1) where B is a
second matrix. Moreover, rk(A) and A+ represent rank and Moore-Penrose inverse, respectively, while
λmin/max (A) is the minimum/maximum eigenvalue of A. Next, ⌊ x⌋ represents the integer part of x, sup
denotes the supremum, k0 represents the in-sample observations, while T − k0 = n are the out-of-sample
observations. With respect to the notion of convergence, → and → p denote the limit and convergence in
probability, respectively. Finally, e
a and b
a represent the unfeasible and feasible estimators.
2 Econometric setup
Introduce a vector of predictors zt = (w′t , f′t )′ ∈ R q+r such that we have the following forecasting models:
where wt represents a small set of known predictors. While we may choose the observables wt based on
economic theory or our experience as forecasters, it is important to remember that nowadays we face an
abundance of potential predictors that may or may not improve our forecasts. In order to exploit the
information embedded in large dimensional datasets, one could extract the potential predictive content
of the many regressors into just a few series of latent factors ft . This situation is reflected in (2.2). Because
ft is typically unobserved, we define the infeasible least squares (LS) estimator of δ for each t = 1, . . . , T as
! −1
t −1 t −1
e
δt = ∑ zs z′s ∑ zs ys +1 ,
s =1 s =1
3
′
where the out-of-sample forecast error is defined as ue2,t+1 = yt+1 − e
δ zt (e θt and ue1,t+1 are defined anal-
ogously). Throughout, zt can be stationary or near stationary. In this study, we are not interested in
specific time series properties of zt . However, we want to vary its degree of persistence and its effect
on the forecast comparison problem. Large databases, such as FRED-MD or EA-QD-MD, offer “recipes”
on how to standardize the long series (see e.g. Appendix A in Barigozzi et al., 2024b). However, some
persistence may remain, and this can be flexibly modeled via modederate integration in the spirit of
Magdalinos and Phillips (2009). The implication of this setup is that T 11+τ ∑tT=1 ft f′t converges to a positive
definite matrix for τ ∈ (0, 1).
To extract ft , we assume that our panel of M predictors can be divided into a finite number m of blocks
exogenously, while later in simulation exercises we will propose methods for grouping the series. The
economic rationale underlying this approach lies in the idea that variables with similar economic features
are expected to be driven fundamentally by a series of (possibly unobserved) common shocks (see e.g.
Hallin and Liška, 2011; Moench et al., 2013; Ando and Bai, 2016) and, as a result, are to be grouped into
one block. The FRED-MD database of the Federal Reserve (we refer to Ludvigson and Ng (2011) and
McCracken and Ng (2016) for more details) is among the primary examples, and it is comprised of 128
monthly time series spanning from 1959Q1 to the present, which are broadly classified into eight groups:
(i) output and income, (ii) labor market, (iii) housing, (iv) consumption, orders and inventories, (v) money
and credit, (vi) interest and exchange rates, (vii) prices, and (viii) stock market. It is then a reasonable be-
lief to assume that every series in each block is, due to their very characteristics, sensitive to same changes
in the economic environment.
For illustration purposes, let us assume that every block has the same number of predictors N = M/m.
By letting Ib = {1b , . . . , Nb } be the set of indexes of the series contained in each block for b = 1, . . . , B, we
denote the predictor panel data variable as xi,t = [ xi1 ,t , . . . , xim ,t ]′ ∈ R, where xib ,t can be seen as the i-th
predictor in the b-th block. The data generating process is assumed to be
where ft ∈ Rr is a vector of common factors, Λi ∈ Rr ×m is the matrix of factor loadings, and ei,t ∈ R m are
the idiosyncratic components. In this work, we estimate the latent factors by using the methodology of
common correlated effects (CCE) developed by Pesaran (2006) as
bft = xt = Λ′ ft + et , (2.4)
where A = N −1 ∑iN=1 Ai for any matrix Ai . Note that for every fixed t = 1, . . . , T, we have ket k =
O p ( N −1/2 ) under many empirically relevant assumptions. Therefore, given that rk(Λ) = r, we obtain
′
ft = (Λ )+bft + O p ( N −1/2 ), (2.5)
which justifies using averages to approximate the factors up to a transformation. Clearly, if m = r, then
+ −1
Λ = Λ , which is equivalent to knowing the number of factors r. Therefore, the overall set of feasible
zt = (w′t , bf′t )′ ∈ R q+m , which means that the feasible estimator is given by
predictors is b
! −1
t −1 t −1
b
δt = ∑ b z′s
zs b ∑ bzs ys+1, (2.6)
s =1 s =1
′
where ub2,t+1 = yt+1 − b
δb zt .
4
′
Note that ket k = O p ( N −1/2 ) implies that 1
T 1+ τ ∑tT=1 bftbf′t is asymptotically singular, because Λ 1
T 1+ τ ∑tT=1 ft f′t Λ
is, unless m = r.
Consider the case when m > r, which is equivalent to overspecifying the number of factors. Indeed,
based on the discussion above, FRED-MD would provide us with m = 8 estimated factors, while the
average number detected in the literature is smaller.2 In order to proceed with the analysis of (2.6), we
need to re-define our target object since we are now estimating f0t = [f′t , 01×(m−r ) ]′ . Hence, similarly to
Stauskas and Westerlund (2022), we introduce a so-called rotation matrix H ∈ R m×m , such that
′ ′ ′ ′ ′
H bft = H Λ ft + H et = f0t + H et . (2.7)
√
Next, we let D N = diag(Ir , NIm−r ), such that
Now, the transformed error term is e0t = [e0r,t′ , e0−′ r,t ]′ , where e0r,t = O p ( N −1/2 ), but e0−r,t = O p (1). This
−1
leads to 11+τ ∑tT=1 bf0bf0′ = O p (1), whose inverse exists. Note that if m = r, then H is equivalent to Λ and
T t t
− 1′
D N = Ir , and so (2.8) boils down to ft + Λ et . We leave the explicit definition of H in the Appendix.
Similarly to Stauskas and Westerlund (2022), the device in (2.8) provides analytical means to track the
contribution of the redundant m − r averages on b δ t and, in turn, ub2,t+1 which is the key object in the
tests. Note that based on (2.6) (in stacked notation, where we suppress dependence on t) and by the
Frisch-Wough-Lovell Theorem:
′
b (b
F MW bF ) −1 b
F′ MW y (b
F′ MW bF ) −1 b
F′ MW y
δt = = , (2.9)
(W′ MbF W)−1 W′ MbF y (W′ MbF0 W)−1 W′ MbF0 y
where MA = I − A(A′ A)−1 A′ is a projection matrix. It is essential that projecting onto bft is equivalent to
projecting onto bf0t since HD N is a full rank matrix. In the next section we introduce the tests under zt , and
then give their representations under b zt .
5
consider the test for forecast nesting from Pitarakis (2023a). Assuming that the factors are observed, the
test statistic for forecast nesting from Pitarakis (2023a) is given by
" #!
1 1 T −1 2 1 n 1 k0 + m0 −1 n 1 T −1
√ ∑ ue1,t+1 −
2 m0 n t∑ n − m0 n t=k∑
s f ,1 (m0 ) = √ ue1,t+1 ue2,t+1 + √ ue1,t+1 ue2,t+1 ,
e1
ω n t= k0 = k0 + m
0 0
(2.10)
where m0 = ⌊nµ0 ⌋ = ⌊( T − k0 )µ0 ⌋ is a cut-off point to split the average for µ0 ∈ (0, 1), µ0 6= 1/2 and is
the estimated standard deviation of the limiting distribution of the test statistic. However, the true factors
are not observed and must be estimated instead, so the subscript ” f ” indicates that the test is infeasible.
Second, Pitarakis (2023b) proposed two additional tests of forecasting accuracy whose corresponding
infeasible statistics are:
k0 + l10 −1 0 k0 + l20 −1
1 n 1 l 1
s f ,2 (l10 , l20 ) = √ ∑ ue21,t+1 − 10 √ ∑ ue22,t+1 (2.11)
e 2 l10
ω n t= k0 l2 n t = k 0
and
n
1 1
s f ,3 (ν0 , λ02 ) = ∑ s f ,2 (l1 , ⌊nλ02 ⌋). (2.12)
e 3 n(1 − ν0 ) l
ω
1 =⌊ nν0 ⌋+1
where l 0j = ⌊nλ0j ⌋ for j = 1, 2, with λ0j ∈ (0, 1) representing two portions of the out-of-sample period.
That is, the forecast MSE loss differentials of both models are computed over partially overlapping out-of-
e 2j for j = 1, 2, 3 are the corresponding variance estimators.
sample segments, i.e. l10 > l20 or vice-versa, and ω
Note that if the two segments were fully overlapping then the estimator variance would be asymptotically
degenerate as in Clark and McCracken (2001), so l10 = l20 is ruled out to avoid incurring into this exact
issue. Further, observe from the formulation of (2.12) that, once l20 (or λ02 ) is fixed, new test statistics can be
obtained by averaging (2.11) over some chosen feasible set of l1 based on the tuning parameter ν0 ∈ (0, 1).
As shown by the analysis of Margaritella and Stauskas (2024), among other averaging possibilities is also
a fourth statistic of the form
n
1 1
s f ,4 (λ01 , ν0 ) = ∑ s f ,2 (⌊nλ01 ⌋, l2 ), (2.13)
e 4 n(1 − ν0 ) l
ω
2 =⌊ nν0 ⌋+1
which is obtained by fixing l10 (or λ01 ) and averaging (2.11) over some chosen feasible set of l2 . Finally,
Margaritella and Stauskas (2024) also show for j ∈ {1, 2, 3, 4} that the variances of the four statistics share
a common element, i.e.
!
n n
1 1
ω e2 = ∑ ue22,t+1 − ∑ ue22,t+1
e i2 ∝ φ (2.14)
n t =1 n t =1
2.2 Assumptions
Throughout our analysis, we employ the following set of assumptions.
Assumption 1. E (ut+1 |Ft ) = 0, E (u2t+1 |Ft ) = σ2 , and E (u4t+1 ) < ∞ for all t, where Ft is the sigma-algebra
generated by {zt , zt−1 , ...}.
Assumption 2.
6
(a) {zt } is a mildly integrated process as defined in Magdalinos and Phillips (2009). In particular,
C
zt = Rz,T zt−1 + uz,t , Rz,T = Iq+r + , τ ∈ (0, 0.5), C < 0 (diagonal ),
Tτ
where uz,t is a zero-mean linear process.
(b) If τ = 0, then |λmin (C)| < 2 in order to ensure that Rz = Iq+r + C is inside the unit circle, and uz,t is
such that {zt } has absolute summable autocovariances.
(c) For a given τ ∈ [0, 0.5) and some κ ∈ (0, 1), we have
⌊κT ⌋
1 Σf f Σ′w f
sup ∑ zs z′s − κΣ ZZ = o p (1), Σ ZZ =
κ ∈(0,1) T 1+ τ s =1
Σw f Σww
as T → ∞.
Assumption 3.
(a) (i) If τ ∈ (0, 0.5), then {ei,t } is uncorrelated over time with E (ei,t ) = 0m×1 , E (ei,t e′i,t ) = Σee,i,t
positive definite and E (kei,t k4 ) < ∞.
(ii) If τ = 0, then we let ei,t = Ci ( L)ǫi,t = ∑∞ j=0 C i,j ǫi,t− j , where ǫi,t is independent across t with
E (ǫ i,t ) = 0m×1 , E (ǫ i,t ǫi,t ) = Σǫǫ,i,t positive definite, E (kǫ i,t k4 ) < ∞, and ∑∞
′
j=0 j
1/2 kC k < ∞.
i,j
(b) For τ ∈ [0, 0.5) and η ∈ (0, 1), we have the following limit as N → ∞:
N N ⌊ ηT ⌋
1
sup ∑∑ ∑ ei,k e′j,k − ηΣee = o p (1)
η ∈(0,1) NT i=1 j=1 k=1
(c) ξ t = vec ( Ne t e′t − E ( Net e′t )) is strong mixing with coefficients of size −bd(b − d) with b > 4 and
b > d > 2, E (kξ t kb ) < ∞, and lim T →∞ T −1 ∑tT=1 ∑sT=1 E (ξ t ξ ′s ) is positive definite.
A note on assumptions. Assumption 1 is standard in the forecasting literature, and it implies that we
focus on one-step-ahead forecasts (this assumption is made in Campbell and Yogo, 2006, Hjalmarsson,
2010, or Breitung and Demetrescu, 2015). While some tests in Pitarakis (2023a) and Pitarakis (2023b) al-
low for h-step-ahead forecasts, we focus on the smallest set of conditions to accommodate all of them.
Assumption 2 (a) flexibly regulates the persistence of the predictors zt in the manner of moderately in-
tegrated process via the parameter τ, whereas (b) makes sure that the process is stationary when τ = 0.
Note that Magdalinos and Phillips (2009) have τ ∈ (0, 1). We need to limit the speed at which the process
approaches ordinary unit root in order to still be able to eliminate factor estimation error in recursive
samples. Hence, we rule out the possibility of more persistent processes given by τ ∈ (0.5, 1). Later, we
7
also make a conjecture on the effects of a moderately explosive system of variables (C > 0) on our testing
framework. Part (c) ensures that the second empirical moment converges in recursive samples, as well.
Assumption 3 (a) is important, as it reveals that we cannot maintain both idiosyncratics correlated in an
arbitrary way and persistent predictors simultaneously. The intuition is as follows. As customary in the
CCE literature, we require that terms of the form of T 11+τ ∑tT=1 zt e′i,t are negligible for all i. When τ ∈ (0, 0.5)
and {ei,t }tT=1 is serially correlated, they shrink to zero at a rate that is too slow for our purposes, unless
E (ei,t e′i,s ) = 0m×m . Under Assumption 2 (b), the rate is suitable even if idiosyncratics exhibit more general
time-dependence. Overall, this can be seen as a price for conducting “comparative statics” by altering τ,
as this restriction is inherent in the process of Magdalinos and Phillips (2009).3 Assumption 3 (a) allows
for weak cross-section dependence together with unconditional time and cross-section heteroskedastic-
ity. For example, a spatial dependence structure employed in Stauskas and Westerlund (2022) is a special
case. Assumption 3 (c) is similar to the one in Stauskas and Westerlund (2022), and it regulates depen-
dence so that {ξ s }ts=1 obeys the invariance principle and is needed to make sure that (the feasible) δbt does
not dominate the asymptotic theory. Assumption 4 is standard, but it can be relaxed at the expense of
higher moment requirements, while Assumption 5 treats the factor loadings as fixed parameters. They
can be made stochastic as in, for instance, Pesaran (2006). It also ensures that the averages are informative
about the factors. Finally, Assumption 6 adapts the local alternative parameterization necessary for our
statistics to possibly moderately integrated environments.
b m0 n t∑ n − m0 n t=k∑
+ √ ue1,t+1 (ue2,t+1 − ub2,t+1 ) + √ ue1,t+1 (ue2,t+1 − ub2,t+1 )
2ω = k0 0 + m0
(3.1)
e2
ω
s fb,2 (l10 , l20 ) = s f ,2 (l10 , l20 ) + s f ,2 (l10 , l20 ) −1
b2
ω
k0 + l20 −1 k0 + l20 −1
1 n 2 1
+ √ ∑ ue2,t+1 (ue2,t+1 − ub2,t+1 ) − √ ∑ (ue2,t+1 − ub2,t+1 )2 , (3.2)
b 2 l20
ω n t= k0 n t= k0
e3
ω 1 n n − ⌊nν0 ⌋
s fb,3 (ν0 , λ02 ) = s f ,3 (ν0 , λ02 ) + s f ,3 (ν0 , λ02 ) −1 +
b3
ω ωb 3 l20 n(1 − ν0 )
k0 + l20 −1 k0 + l20 −1
2 1
× √ ∑ ue2,t+1 (ue2,t+1 − ub2,t+1 ) − √ ∑ (ue2,t+1 − ub2,t+1)2 , (3.3)
n t= k0 n t= k0
and
e4
ω
s fb,4 (ν0 , λ02 ) = s f ,4 (ν0 , λ02 ) + s f ,4 (ν0 , λ02 ) −1
b4
ω
3 In the Appendix, we provide additional Monte Carlo evidence which reveal that correlation in ei,t is not really harmful for
statistical power of the tests.
8
!
n k 0 + l2 − 1 k 0 + l2 − 1
1 n 2 1
+ ∑ ωb 4 l2 √ ∑ ue2,t+1 (ue2,t+1 − ub2,t+1 ) − √ ∑ (ue2,t+1 − ub2,t+1 )2 .
n(1 − ν0 ) l =⌊ nν ⌋
n t= k0 n t= k0
2 0
(3.4)
It is clear from the expansions of feasible statistics that ue2,t+1 − ub2,t+1 is the central component in the
upcoming analysis. It can be shown that this difference admits the following representation:
where each of the components brings a distinct contribution to the difference between feasible and infea-
sible forecast. For instance, I I I takes into account the usual factor estimation error, which stems from the
r factors, while I I additionally involves an error from estimating α under the observed ft . The component
0 ′
I is the most subtle one. In particular, here e δt = [e θt ]′ ∈ R m , which is e
α ′t , 01×(m−r ) , e δt appended with extra
m − r zeros in place of the redundant averages. With δ bt ∈ R being the LS estimator of the model param-
m
eters with bzt , we have Q N = diag(HD N , Iq ) transforming bft into bf0 inside of (2.6). In effect, I tracks the
t
asymptotic effect of the factor overestimation error via the m − r redundant averages. Lemma 1 below
formalizes its behavior uniformly over t.
D T (Q− 1b e0
N δt − δt ) = o p (1)
1 τ 1 τ
uniformly in t, where D T = diag( T 4 + 2 Ir , T 1/4 Im−r , T 4 + 2 Iq ).
There are significant differences between our Lemma 1 and its counterpart in Stauskas and Westerlund
(2022). Firstly, we undertake our analysis under non-stationary predictors zt . The subtle detail is that
the feasible predictors are b
zt with the excess averages driven by stationary idiosyncratics. Therefore, we
effectively have a mixed integration order of the regressors, which is reflected by the normalization matrix
0
D T . Strikingly, the appropriately normalized Q−1 δbt − δet is asymptotically negligible even if m > r. The
N
intuition behind these results is as follows. Let τ = 0, such that D T = T 1/4 Iq+m , to demonstrate that this
difference does not originate from the persistence of the regressors. Then, we can show that
0r
T 1/4 (Q− 1b e0 −1/4 vt + o p ( T −1/4 ) = O p ( T −1/4 ),
N δt − δt ) = T (3.6)
0q
where
! −1
t −1 t −1
1 1
vt = ∑ e0−r,s e0−′ r,s √ ∑ e0−r,s us+1 ∈ Rm−r
T s =1 T s =1
and kvt k = O p (1) which is the same martingale difference process as in Stauskas and Westerlund (2022).
Representation in (3.6) immediately implies that
√ 0r
−1 b 0
e
T ( Q N δ t − δ t ) = v t + o p ( 1) , (3.7)
0q
and so we are√back to Stauskas and Westerlund (2022). The reason why we only need to scale up by
T 1/4 and not T is because the former characterizes the local power of the tests in Pitarakis (2023a)
and Pitarakis (2023b). In particular, the asymptotic distribution of his infeasible statistics is generated
9
by {u2t+1 − σ2 }, whereas the components involving (functions of) {zt ut+1 } - which are the distribution
generators in Clark and McCracken (2001) - are negligible. In other words, “the weight” is placed on the
former martingale difference instead of the latter. Naturally, the same happens with the components of
the feasible statistics, and so the influence of vt becomes negligible, as well. In general, Lemma 1 suggests
that the redundant m − r cross-section averages should not interfere with the asymptotic normality of the
feasible statistics. In practice, it means that practitioners can use all the available blocks in FRED-MD or
EA-MD-QD datasets and stay agnostic about the true number of factors, as long as m ≥ r. This stands in a
sharp contrast to the CCE literature, where the redundant m − r averages result in an asymptotic bias, un-
less TN −1 → 0, which is a substantial restriction (see e.g. Karabiyik et al., 2017; or De Vos and Stauskas,
2024). To the best of our knowledge, this is the first result in the CCE literature when the redundant m − r
CAs stay completely harmless even if TN −1 → c > 0.
because T −τ/2b z0t is uniformly bounded by the results in Lemma 3.1 of Magdalinos and Phillips (2009).4
Notice that here the requirement of τ ∈ (0, 0.5) does not play a role, since the first component in (3.10)
is negligible for all τ ∈ (0, 1). However, the restriction occurs when we analyze squared sums of (3.10).
These arguments can subsequently be used in the expansions (3.1) - (3.4) which is formalized by Lemma
2.
k0 + l 0 −1
(b) √1 ∑ t = k 02 (ue2,t+1 − ub2,t+1 )2 = o p (1)
n
k0 + l 0 −1
(c) √1 ∑ t = k 02 ue2,t+1 (ue2,t+1 − ub2,t+1 ) = o p (1)
n
1 n √1 k + l −1
(d) n ∑ l2 l2 n
0
∑t= k0
2
(ue2,t+1 − ub2,t+1 )2 = o p (1)
1 n √1 k + l −1
(e) n ∑ l2 l2 n
0
∑t= k0
2
ue2,t+1 (ue2,t+1 − ub2,t+1 ) = o p (1)
e2 − φ
(f) |φ b2 | = o p (1).
Notice that a direct consequence of the last part of Lemma 2 is that the estimators of the standard
deviations of the test statistics (3.1)-(3.4) are all consistent. In fact, it follows from the workings of Pitarakis
(2023a), Pitarakis (2023b) and Margaritella and Stauskas (2024) that we have Corollary 1:
4 Note z0t ≤ T −τ/2 zt + T −τ/2 e0t as idiosyncratics are stationary.
that actually T −τ/2 zt = O p (1), however T −τ/2 b
| {z }
o p ( 1)
10
Corollary 1 Suppose that the conditions in Lemma 1 and Lemma 2 hold. Then, as ( N, T ) → ∞, we have
e 2j − ω
ω e2 − φ
b 2j = θ j φ b2 = o p (1),
for each j, where θ j is a function of tuning parameters specific to the particular statistic.
As a result, it follows by the Continuous Mapping Theorem that ω b j = o p (1) for all j ∈ (1, 2, 3, 4)
ej − ω
ej
ω
so that the respective expressions ωb j − 1 are all negligible. By using Lemma 2 and Corollary 1, we can
show that s fb,1 (m0 ), s fb,2 (l10 , l20 ), s fb,3 (ν0 , λ02 ) and s fb,4 (ν0 , λ01 ) admit the asymptotic representations reported
below in Theorem 1:
Theorem 1 Suppose that the conditions in Lemma 1, Lemma 2 hold. Then, as ( N, T ) → ∞, and using the results
in Corollary 1, we have
Since the decompositions in expressions (3.1)-(3.4) are functions of the quantities in Lemma 2 and Corol-
lary 1, the results in Theorem 1 follow readily.
The direct implication of Theorem 1 is that the asymptotic distributional theory of our four test statis-
tics follows through directly from Pitarakis (2023b) and Pitarakis (2023a). It is noteworthy to make two
remarks:
1. These asymptotic expansions hold without restrictions on the relative expansion rate of N and T so
long as they both diverge to infinity, thus making our testing procedure relevant in and applicable
to many large-dimensional macroeconomic and financial datasets. However, note from the rate of
convergence in Lemma 1 that the m − r redundant averages still decay quite slowly. As a direct
consequence, in order to be valid, our setup requires us to impose the restriction τ ∈ [0, 0.5) and
thus to exclude those processes with less moderate deviations from unity, i.e. when τ ∈ [0.5, 1).
2. While we make the assumption that the regressors wt and ft are stationary or near-stationary mildly
integrated (C < 0), the limit theory still holds and is equally favourable with moderately explosive
regressors (C > 0). We restrict our discussion to the case of C with distinct diagonal elements - i.e.
ci = c j for all i, j - in order to rule out singularity issues arising from the case of moderately explosive
cointegration among the variables (C without distinct diagonal elements or, equivalently, ci = c j for
some i 6= j). Say that Rz,T from Assumption 2 (a) can be partitioned into two diagonal r × r and
q × q matrices, i.e. Rz,T = diag(R f ,T , Rw,T ). Under the theory of moderately explosive systems, we
require that there exists an increasing sequence (k T ) T ∈N such that
−k T kT −T
Rz,T → 0, T τ Rz,T →0
2
so that E u1 < ∞ instead of bounded fourth moments. With respect to how Lemma 1 changes, let
1
T . Since they show
us define D T = diag(D F , DW ) with D F = diag( T τ R Tf,T , T 4 Im−r ) and DW = T τ Rw,T
11
−T ′
that T −2τ Rz,T Z ZR− T
z,T = O p (1) and T
− τ R − T Z ′ V = O (1) for any error V independent from Z,
z,T p
we conjecture that we can use these results to demonstrate that (3.6) still holds as the properties of
vt do not change.
While Theorem 1 demonstrates a significant applicability of the new statistics, their value in empirical
settings can be boosted even more via the power enhancement and their robustness to structural breaks
in the factor loadings.
Proposition 1. Let sPower b j for j = 2, 3, 4 be a power-enhanced statistic with the enhancement term △
= s fb,j + △ b j.
b f ,j
Then, under Assumptions 1-6 as ( N, T ) → ∞
sPower
fb,j
= sPower
f ,j + o p ( 1)
for j = 2, 3, 4.
Statistics from Pitarakis (2023b) involve several tuning parameters and therefore consequences on their
(small sample) performance may be incurred. As a solution, they can be power-enhanced, where △ j is
b j , which, according to
based on ŭ22,t+1 = ue22,t+1 − (ue1,t+1 − ue2,t+1 )2 . Clearly, we only have the feasible △
Proposition 1, obeys △ b j = △ j + o p (1). Due to the asymptotic equivalence, we will focus on the power-
enhanced statistics in our Monte Carlo experiments.
Proposition 2. Let the loadings break at some time point D = ⌊φT ⌋ with φ ∈ (0, 1), such that
Proposition 2 further increases applicability of the tests. It states that under a single break in factor load-
ings, the starting model can be reformulated into a model with 2r factors (similarly, with d breaks, we can
reformulate it into a model with 2d × r factors). The result is similar to the findings in Stauskas and Westerlund
(2022). However, in their case, the condition of D ∈ (1, k0 ] is necessary. In other words, the effect of
the structural break must be subsumed in the initial estimation of the parameters before going into the
out-of-sample analysis. Yet, it does not play any role in the current setup. The intuition behind this is
the following. The asymptotic distributions of statistics in Clark and McCracken (2001) are highly non-
standard. If D ∈ (k0 , T − 1] (the out-of-sample period), there is a break in the asymptotic variance of the
statistic. While it is typically possible to simulate the non-standard distribution, the break in variance
makes the simulation infeasible as the break date is unknown. In our Proposition 1, the break location
does not matter, because all the components of the asymptotic expansions are negligible even if breaks in
their variances occur (see the Appendix for a formal explanation).
Remark 1 If τ = 0 (stationary setup), the rank condition in Assumption 5 can be tested. In particular, this can
be achieved by using a rank classifier in De Vos et al. (2024). Next, let g be the number of CAs used in practice
(specifically, r ≤ g ≤ m). In applications, we may like to take an even further step and attempt setting g = r, which
tantamounts to estimating the number of factors in the PC literature (see e.g. Ahn and Horenstein, 2013). This can
be achieved by using Information Criterion from De Vos and Stauskas (2024) (see their Proposition 3):
!
1 N ′
NT i∑
IC ( P) = log det Xi MbFq P Xi + g · m · p N,T (3.11)
=1
12
where P is a combination of column indices of X, and q P picks the corresponding g = cols(q P ) averages in practice.
Note that here we stacked Xi over t = 1, . . . , k0 to stress the assumption that r does not change over time. Hence,
we would implement the selection once in the in-sample period.
Let accordingly P0 denote the set of averages from X such that rk(Λq P ) = r, cols(q P0 ) = r, and p N,T → 0 is a
penalty term in function of the panel dimensions N, T. This leads to the following selector for the CAs such that
r = g:
where P denotes the index set of all possible combinations of averages. Note that if the realized selection was g > r
(over-selection tends to occur in practice) the results of our Theorem 1 would still hold, which is a major advantage
for practitioners. We note, however, that such selection can be performed only when τ = 0 (based on unpublished
technical report by Ditzen and Stauskas, 2025).
We examine a large number of modelling specifications of Λi,t , ǫi,t , ut , δ, ρi , β and τ. For conciseness,
we summarise our proposed DGPs - labelled (1) to (9) - in Table 1. To start with Equation (4.1), we explore
different distributional assumptions for ut , which is drawn from N (0, 1), t(10) or N (0, (1 − φ) + φu2t−1 )
for φ = 21 across the various DGPs. These three modelling choices allow us to investigate the effect dif-
ferences of a standard normal distribution, fat tails and ARCH effects of first order on our tests. Further,
Equation (4.2) describes the data-generating process of the factor model. Notice here that the loadings
specification encompasses both the time-invariant, nonzero mean case as well as the time-varying cases in
DGP (8), which incorporates a structural shift in the loadings mean from 1 to 2 at time T/2 as in Chen et al.
(2014). Notice that the break is in the out-of-sample period in order to test the location invariance pre-
dicted by our theory. On the contrary, DGP (9) has time-invariant loadings with mean zero, which implies
13
that Λ is asymptotically zero and CCE breaks down. Finally, we choose in Equation (4.3) a common au-
toregressive specification5 for the time dynamics of the factors (see Bai and Ng, 2006, Gonçalves et al.,
2017, and Stock and Watson, 2002, for instance). In a similar manner, in Equation (4.4), we allow the
panel idiosyncratic components ei,t to be weakly dependent in i and t (similarly to Bai and Ng, 2002,
Banerjee et al., 2008, Boivin and Ng, 2006, and Chen et al., 2014), meaning that they exhibit serial correla-
tion via the coefficient ρi as well as cross-sectional correlation of spatial type by means of the β coefficient
in Equation (4.5). Note that the shocks ǫi,t in this equation are drawn from N (0m×1 , Σǫǫ,i ), N (0m×1 , Σǫǫ,i,t )
or t(10), meaning that the first distribution acts as a benchmark, whereas the second and third distribu-
tions introduce time-varying volatility and fat tails, respectively. Here, both covariance matrices Σǫǫ,i and
Σǫǫ,i,t are simply diagonal and their non-zero elements are generated from U (0.5, 1.5) so to accomodate
for cross-sectional heteroskedasticity, as permitted under CCE. Lastly, remember that we are working un-
der the theory of moderately integrated systems à la Magdalinos and Phillips (2009), whereby we assume
that the parameters θ and δ can be written as θ = 1 − θ0 /T τ with kθ0 k < ∞ and δ = 1 − δ0 /T τ with
kδ0 k < ∞. Hence, we consider τ ∈ {0.1, 0.2, 0.3, 0.4, 0.49} in the next section.
4.2 Results
In this section, we report size and local power experiments run on 10,000 MC replications across different
DGP and parameter settings. The significant level is set to 5% and the critical values are obtained from the
usual standard normal table. In detail, we present four results for the statistic s fb,1 in the form of two tables
and two figures: 1) the first table shows the sensitivity of test size and local power to different choices of
µ and ( N, T ), 2) the second table investigates the evolution of local power across all DGPs under various
( N, T ) combinations, 3) the first figure illustrates the local power of the proposed DGPs across different
values of the parameter α, and finally 4) the second figure provides insight into the effect of misspecifica-
tion in the number of blocks on size and local power.
5 Notice that if δ = 0 then the factors can be considered as a ’static’ sequence of i.i.d. shocks. Hence, even if α 6 = 0, the
probability of rejecting the null hypothesis when it is false (test power) should converge to the desired significant level (as for
the test size) because the factors are expected to be uninformative.
14
Table 2 reports the results of size on the left-hand side and local power on the right-hand side under
DGP (2). This choice of DGP assume moderate dependence in the idiosyncratic components and so it
could be considered as an unsophisticated but realistic representation of economic data under a factor
model structure. Hence, it should be regarded as a benchmark for all other MC settings. Starting from
the case of α = 0, observe on the one side that the test size converges to the significant level as T → ∞
but it is seems to be insensitive to different levels of N. On the other side, whereas, it is evident that size
deteriorates as µ0 approaches 0.45, as anticipated by the theory of Pitarakis (2023a). In essence, reliable
test size is obtained with smaller µ0 and larger T. Next, we turn our attention to the case of α = 1 under
the local alternative. Here, the local power improves mildly as N → ∞ but strongly T → ∞. Contrary to
test size, the local power converges to 1 as µ0 approaches 0.45. We conclude that the choice of µ0 imposes
a trade-off between size and local power, but such trade-off becomes materially unimportant as the time
dimension of the data diverges.
15
α = 1 (local power), µ0 = 0.45
N T DGP (1) (2) (3) (4) (5) (6) (7) (8) (9)
Table 3 reports the local power results across DGPs (1) to (9) for the triplet α = 1, and µ = 0.45.
Given our theoretical results, it is unsurprising that the nondegenerate DGPs, namely (2) to (8), all ex-
hibit satisfactory local power as the time dimension grows to infinity, but we should point out that the
results are not excellent in the case of DGPs (6) and (7). In fact, these statistics are constructed based on
averages of the (squared) predictive regression residuals so it is natural to think that their estimation is
affected by ARCH effects and fat tails. Nonetheless, our findings seem to indicate that these issues are
only of finite sample nature and vanish asymptotically with the time dimension. Similarly to Table 2, in
addition, we see that the local power improves with N given T fixed. This is due to the fact that factor
uncertainty - which is inherent of CCE estimation - disappears with the number of predictors in each
CCE block. Loosely speaking, our tests should not discern between estimated and observed factors for
sufficiently large block sizes, meaning that the theoretical framework of Pitarakis (2023a,b) is virtually
restored. With respect to the first DGP, whereas, the values are basically flat in N and fall towards the
significant level. The reason is that the panel data has no factor structure nor any kind of dependence, so
that the CCE-estimated factors should not improve upon the baseline specification and offer additional
predictive power. Finally, DGP (9) models the rank violation in CCE estimation so statistical inference is
not valid in this case and, indeed, the local power is not satisfactory under any combination of N and T.
Figure 1 reports the power curves for all DGPs under N = 200, T = 600 and µ0 = 0.45. The story is
nearly identical to the discussion about Table 3, with the local power improving as α moves away from 0.
Once again, but perhaps unsurprisingly, DGP (6) performs significantly worse across the group in terms
of size and power, whereas DGP (7) exhibits excellent test size but its local power seems to match that of
the former DGP for higher levels of α. Similarly to before, the local power of DGP (1) should be inter-
preted as test size, while that of DGP (9) should not be relied upon.
16
As anticipated by our results in Section 3, the asymptotic distributions of the tests are invariant under
a misspecification of the number of factors. Our MC simulations highlight this theoretical finding very
clearly in Figure 2. In this experiment run on DGP (2) with N = 200 T = 600 and µ = 0.45, the true num-
ber of factors is 1 but we consider three possible choices for the number of blocks, namely m ∈ {1, 2, 3},
so to account for a correct specification as well as overspecification. It is clear from the figure that the
power curves corresponding to these three choices are almost undistinguishable numerically, especially
for the higher values of α. Peculiarly, more blocks also yield a slightly better local power in the vicinity
of 0. However, we notice that each additional block costs as a marginal size distortion when α = 0, but
we believe that it is a price that we are willing to accept for the insensitivity of the local power to an
overspecification in the number of blocks. That is, we prefer the redundancy of including an irrelevant
regressor in our predictive equation than the risk of omitting a relevant one.
Finally, it remains to investigate the effect of nonstationarity in our tests. For this purpose, we illustrate
in Figure 3 the power curves of DGP (2) under varying levels of persistency as τ ∈ {0, 0.1, 0.2, 0.3, 0.4, 0.49}.
As before, we set N = 200, T = 600, h = 1 and µ0 = 0.45. In line with the previous findings, our test
exhibits excellent local power even in the presence of more persistent regressors. It should however be
pointed out that our results deteriorate mildly with increasing persistence, with the case of τ = 0.49 being
close to the boundary and performing the worst amongst the peer group. In addition, unreported simu-
lations show that the local power drops severely and size distortions are observed when τ ∈ [0.5, 1).
In the interest of space, we report additional MC simulations for the other three statistics in the Ap-
pendix. However, we would like to point out that the analysis on test size, local power and the effect of
regressor persistence for s fb,2 , s fb,3 and s fb,4 still shows similar, if not better, results.
1
DGPs
0.9 (1)
(2)
0.8 (3)
(4)
(5)
0.7 (6)
(7)
Local power
0.6 (8)
(9)
0.5
0.4
0.3
0.2
0.1
0
0 0.1 0.2 0.3 0.4 0.5 0.6
Figure 1: Evolution of local power for different DGPs and values of α. Setting: (N, T) = (200,600), µ0 =
0.45.
17
1
0.9
0.8
0.7
Local power
0.6
0.5
0.4
0.3
Figure 2: Evolution of local power for different values of m and α. Setting: DGP = (2), (N, T) = (200,600),
µ0 = 0.45. The true number of factors is r = 1.
0.9
0.8
0.7
Local power
0.6
0.5
0.4
Persistence
0.3 =0
=0.1
0.2 =0.2
=0.3
0.1 =0.4
=0.49
0
0 0.1 0.2 0.3 0.4 0.5 0.6
Figure 3: Evolution of local power for different values of τ ∈ [0, 21 ) and α. Setting: DGP (2), (N, T) = (200,
600), h = 1, ν0 = 0.8, µ0 = 0.45.
18
To summarise, our CCE-based tests perform quite well under a plethora of empirically relevant eco-
nomic scenarios. More importantly, our MC simulations seem to support two important theoretical con-
tributions of this work to the factor model literature. First, our tests seem to be invariant to uncertainty in
the number of factors and, second, they can accommodate for moderately integrated systems of variables.
5 Empirical results
In this section, we provide empirical evidence on the predictive power of CCE factors in macroeconomic
forecasting. For this purpose, we use the novel, publicly available6 EA-MD-QD dataset of Barigozzi et al.
(2024b). The dataset is comprised of quarterly and monthly macroeconomic time series for both the Euro
Area (EA) as a whole and its ten primary member countries, and it is updated on a monthly basis and
constantly revised. To the best of our knowledge, this is the first article to use this dataset in the context
of forecasting and out-of-sample (factor-based) model evaluation. In detail, the EA dataset consists of
118 time series - which include 71 and 47 variables collected at quarterly and monthly frequencies, re-
spectively - that have been recorded starting from 2000:1 to the last available observations. Similarly to
the FRED-MD dataset (McCracken and Ng, 2016), the variables are classified into 11 blocks: (1) National
Accounts, (2) Labor Market Indicators, (3) Credit aggregates, (4) Labor Costs, (5) Exchange Rates, (6) In-
terest Rates (7) Industrial Production and Turnover, (8) Prices, (9) Confidence Indicators, (10) Monetary
Aggregates and (11) Others. The country-level datasets, have a comparable number of time series and the
same classification groups, with the exception of Monetary Aggregates since it only recorded at EA level.
We treat the data according to the thorough suggestions of Barigozzi et al. (2024b). First, we aggregate
the data at the quarterly level to address the mixed-frequency nature of the variables. Next, while our
testing framework is able to accomodate for different degrees of persistency in the time series, it is neces-
sary to address the issue of missing values and outliers beforehand. To do so, we induce stationarity in the
data according to the transformations recommended for each variable in Appendix A of Barigozzi et al.
(2024b). Subsequently, we impute the missing values using the EM algorithm of Stock and Watson (2002)
and follow standard screening for outliers as detailed by McCracken and Ng (2016). Completion of these
steps results in a first dataset with fully treated stationary data. By reversing all first-differencing trans-
formations, we also recover a second dataset with fully treated persistent data.
In the forecasting exercises, we choose a collection of different time series as target variables in our
regressions so to display the predictive power and usefulness of estimated factors. The selected target
variables include: real GDP (GDP), industrial production index (IPMN), harmonised index of consumer
prices excluding energy and food (HICPNEF), total unemployment (UNETOT), 3-month interest rates
(IRT3M), 6-month interest rates (IRT6M), long-term interest rates (LTIRT), wages and salaries (WS), gen-
eral government total financial liabilities (GGFB), non-financial corporations financial liabilities (NFCLB),
households total financial liabilities (HHLB), real exchange rate against 42 industrial countries (REER42),
residential property prices (HPRC), and share prices (SHIX). We then extract the factors from stationary
and persistent datasets (after removing the target variable to be forecasted) and use them to augment an
AR(1) benchmark. Note that richer specifications can be attained by including, among other possibilities,
other exogenous regressors or additional lags in the factors. While doing so can improve the predictive
accuracy, we do not explore it in what follows. In addition, in real life, a forecaster may decide to choose
subsets of the blocks so to achieve a more parsimonious forecasting model. For instance, Gonçalves et al.
(2017) investigates the predictive content of factors for the equity premium associated with the S&P 500
Composite Index by comparing all permutations of models that include a single lag of at least one of the
6 Freely available at the page: [Link]
19
eight estimated macroeconomic factors. However, we are not excessively concerned with this considera-
tion since our testing framework is invariant under the use of redundant blocks after all.
To guarantee the validity of some of the assumptions in Section (2.2), we run some diagnostics on the
estimated factor model. In the interest of space, we report the results when the EA real GDP is the target
variable in the forecasting equation, but similar conclusions for all other variables can be drawn based on
additional results in the last section of the Appendix. The first assumption that we examine is that the
number of blocks is at least as large as the number of true factors, i.e. m ≥ r. To estimate the number of
factors, we use the eigenvalue-ratio (ER) criterion by Ahn and Horenstein (2013) with rmax = 8 and find
one factor in the stationary data, which is in line with the findings of Barigozzi et al. (2024b). Following
Barigozzi et al. (2024a), we also borrow this result in the context of the persistent data7 since, as argued by
Onatski and Wang (2021), the number of factors in levels should coincide with the number of factors in
first differences provided that no spurious effects are estimated. Barigozzi et al. (2024b) also implement
the test of Bai and Ng (2002), which select seven factors. Given that Ahn and Horenstein (2013) is often
parsimonious in small sample while the information criterion of Bai and Ng (2002) tends to overselect the
number of factors, a researcher should probably accept that the true number of factors in this dataset lies
between 1 and 7. As discussed next, this number may also change as the result of structural breaks.
Another assumption that we make is that the factor loadings are constant over time. In light of the fact
that, since the inception of the euro currency and the start of the dataset, the economies of the European
Union (EU) members have experienced systemic shocks such as the Great Financial Crisis, a sovereign
debt crisis and the Covid-19 pandemic, we raise the question about the possibility of structural breaks in
the factor model. In fact, the presence of d breaks in the sample requires us to reformulate Equation (2.3)
into a model with 2d × r factors, meaning that the number of CAs needs to satisfy m ≥ 2d × r unequiv-
ocably. In order to address our concerns about large shocks, we use the sup-LM statistic-based test of
Chen et al. (2014) to detect big structural breaks in these loadings at unknown dates. In particular, there is
evidence against the no break null hypothesis at the timestamps 2009:1 and 2020:2 at the 5% significance
level. Unsurprisingly, these dates exactly coincide with the Great Recession in Europe and the national
lockdowns imposed by EU countries since these events had systemic and material repercussions on the
entire Euro area. With this finding in mind, we then confront a consequential concern and inspect if r
changes due to these structural breaks in the factor loadings. Indeed, this concern is not new in the litera-
ture as it was previously raised, for instance, by Breitung and Eickmeier (2011) and Bai and Ng (2007) in
the context of US macroeconomic data. Further, Barigozzi et al. (2024a) argues that the effect of the Covid
shock is pervasive in the time series of most real macroeconomic variables and should therefore be treated
as an additional latent common factor. In fact, the Covid shock induced a large shift not only in the levels
(Maroz et al., 2021; Ng, 2021) but also in the volatility (Carriero et al., 2022; Lenza and Primiceri, 2022) of
macroeconomic variables. Hence, we apply once more the ER criterion to the subsamples between the
structural breaks and find: 1) r = 1 from 2000:2 to 2008:4, 2) r = 2 from 2009:1 to 2020:1, and 3) r = 3
from 2020:2 to the end of the sample. It should be noted that even if r = 3 throughout the entire sam-
ple, the number of CAs is still not smaller than 2r so we conclude that we have sufficient estimated factors.
Before turning to our equal predictive ability and encompassing tests, one more diagnostic step re-
mains. As a matter of fact, another assumption in our setup is that the idiosyncratic components in
the Equation (2.3) are at most weakly dependent cross-sectionally with stationary and persistent factors
but serially uncorrelated with non-stationary factors. To illustrate that the assumption on cross-sectional
correlation holds, we apply the sparse estimation methodology via adaptive thresholding of Cai and Liu
7 To our knowledge, Peña and Poncela (2006) is the only paper in the literature that proposes a test for the number of non-
20
(2011) to the covariance matrix of the panel residuals in the stationary and persistent cases. For the former
case, we find that the estimated covariance matrix is almost exactly diagonal with average (resp. average
absolute) value of the cross-correlation of 0.002 (resp. 0.007), which suggests that the estimated station-
ary factors are able to explain most of the variation in the co-movements of the cross-sectional units. For
the latter case, whereas, the estimated matrix is non-diagonal but sparse with average (resp. average
absolute) value of the cross-correlation of 0.005 (resp. 0.036). We regard this level of cross-sectional de-
pendence as very weak overall. On the other hand, serial correlation in the idiosyncratics seems to be
more assiduous according to standard Ljung-Box Q-testing. In the stationary dataset, we find average
(resp. average absolute) serial correlations of order 1 equal to -0.002 (resp. 0.184). Conversely, the cor-
responding value is 0.558 (resp. 0.559) in the persistent case. While our findings are not of concern in
the context of stationary data, it raises apprehension when it comes to persistent data since we do not
allow for serial correlation in association with non-stationary regressors. To alleviate some distress, we
report additional Monte Carlo simulations as a benchmark in the Appendix where we explore the effect
of strong serial correlation in DGP (3) with τ ∈ (0, 0.5) on size and local power of all four statistics. These
additional results seem to indicate for all four statistics that, in spite of our rates of convergence, test size
is fundamentally unaffected while local power suffers a small but more than acceptable loss. As a result,
we also suspect that our convergence rates may possibly be conservative in this scenario. Since the time
and cross-sectional dependence in DGP (3) is stronger than the one observed in the persistent data, we
trust that our equal forecasting ability and encompassing tests are still able to convey practical insight
into the predictive power of the estimated non-stationary factors. With all these pre-testing analysis in
mind, we finally dedicate our attention to testing predictive accuracy.
Upon estimation, we apply our four novel tests of equal predictive accuracy and encompassing to
assess whether the predictive content of CCE-estimated factors improves upon the forecasting ability of
the autoregressive baseline with the selected target variables. The results are reported in Table 4.
Stationary Persistent
EA s fb,1 s fb,2 s fb,3 s fb,4 s fb,1 s fb,2 s fb,3 s fb,4
Table 4: p-values of CCE-based equal predictive accuracy and encompassing test statistics on EA quarterly
data. Setting: h = 1, ν0 = 0.8, µ0 = 0.45, λ01 = 1, λ02 = 0.65. The significance level is chosen at 10%.
21
First and foremost, it is paramount to highlight that the predictive content inherently embedded in
our CCE-estimated factors is particularly strong with target variables that play a key role in the design
of political, fiscal or monetary policies. These variables are economic growth, inflation (excluding energy
and food), unemployment and wages. Since monitoring inflation is at the core of the European Central
Bank’s mandate, it is very valuable to observe from a policymaking perspective that the estimated factors
have formidable forecasting power for consumer prices under all tests and types of persistence. Focusing
on the individual tests, we see that s fb,1 exhibits statistical significance across the largest collection of tar-
get variables, with the p-values almost always consistently below the 1% level (especially with stationary
data). Perhaps surprisingly, our statistics detect better overall forecast accuracy using stationary data,
which is against the popular belief that stationary data incurs in a material loss of information content
compared to non-stationary data. However, this finding was already anticipated by our theoretical re-
sults and further corroborated by our Monte Carlo simulations in light of the fact that the convergence
rates of our statistics slow down with increasing persistency in the data. Further, while the ultimate ob-
jective of our tests is to uncover the predictive specifications where the estimated factors improve upon
the (autoregressive) baseline, it is also important to identify those situations where the factors are not
informative about the future and should therefore be excluded from our forecasting considerations, as
is rightfully and conventionally the case for the changes in, and the level of, the long-term interest rate,
the exchange rate and the stock market returns. In fact, these three variables are notoriously difficult, if
not impossible, to predict because - assuming the efficient market hypothesis holds - the expectations of
rational, informed investors must already be priced into financial markets.
Nevertheless, we acknowledge that our results may inadvertedly be affected by structural breaks in
the coefficients of the factor-augmented regression. Although we have dedicated considerable attention
to breaks in the factor loadings, there exist no statistical procedure to detect structural instability à la
Corradi and Swanson (2014) under the CCE framework. In fact, these authors propose a test for the
joint hypothesis of structural stability of both factor loadings and regression coefficients in the factor-
augmented forecasting model in a PC setting. As previously argued, the period associated with the
Covid-19 pandemic should be treated by including an additional factor in the specification, which will by
construction capture a portion of the shift in the level of the target variables. Hence, overestimating the
number of CAs makes us robust to: 1) the possibility of breaks in the factor loadings, and consequently
2) an increase in the number of factors resulting from a large shock, but possibly also 3) a large break in
the factor-augmented regression coefficients. However, it is undeniably not unreasonable to think that all
concerns about the last point have been dispelled once and for all.
Next, we repeat the same forecasting exercise with data on Germany, France, Italy and Spain. Notice
that the selected target variables are the same, with the exception of the 3-month and 6-month interest
rates since they are not available at the country level in the dataset of Barigozzi et al. (2024b). The re-
sults of our equal predictive accuracy and encompassing tests are reported in Tables 5 to 8. Inasmuch
as the economies and institutions of these countries are structurally distinct, the results of our tests will
naturally be heterogeneous across countries. That is, the factors extracted from each country’s dataset
are informative about the future evolution of different variables. In short, it seems from our results that
macroeconomic factors are still able to offer consistent forecasting insight into economic growth, wages
and (possibly) consumer prices across the four countries, with the notable exceptions of Spain and Italy.
With respect to the former, the factors appear to show quite weak predictive ability while, for the latter,
they anticipate the bulk of economic activity across most variables, excluding the stock market. Similarly
to the EA data, among our four statistics, s fb,1 detects the greatest number of specifications where our
macroeconomic factors yield forecasting power.
22
Stationary Persistent
Germany s fb,1 s fb,2 s fb,3 s fb,4 s fb,1 s fb,2 s fb,3 s fb,4
Table 5: p-values of CCE-based equal predictive accuracy and encompassing test statistics on German
quarterly data. Setting: h = 1, ν0 = 0.8, µ0 = 0.45, λ01 = 1, λ02 = 0.65. The significance level is chosen at
10%.
Stationary Persistent
France s fb,1 s fb,2 s fb,3 s fb,4 s fb,1 s fb,2 s fb,3 s fb,4
Table 6: p-values of CCE-based equal predictive accuracy and encompassing test statistics on French
quarterly data. Setting: h = 1, ν0 = 0.8, µ0 = 0.45, λ01 = 1, λ02 = 0.65. The significance level is chosen at
10%.
23
Stationary Persistent
Italy s fb,1 s fb,2 s fb,3 s fb,4 s fb,1 s fb,2 s fb,3 s fb,4
Table 7: p-values of CCE-based equal predictive accuracy and encompassing test statistics on Italian quar-
terly data. Setting: h = 1, ν0 = 0.8, µ0 = 0.45, λ01 = 1, λ02 = 0.65. The significance level is chosen at 10%.
Stationary Persistent
Spain s fb,1 s fb,2 s fb,3 s fb,4 s fb,1 s fb,2 s fb,3 s fb,4
Table 8: p-values of CCE-based equal predictive accuracy and encompassing test statistics on Spanish
quarterly data. Setting: h = 1, ν0 = 0.8, µ0 = 0.45, λ01 = 1, λ02 = 0.65. The significance level is chosen at
10%.
24
All things considered, we conclude that - although the common correlated effects method may appear
deceivingly simplistic at first glance - regression equations augmented with CCE-estimated factors offer
powerful predictive insight in the context of macroeconomic forecasting in both levels and differences,
especially in the presence of big structural breaks such as the Covid pandemic.
6 Conclusion
In this paper, we revisited a study of Stauskas and Westerlund (2022), who were the first to analyze tests
of equal forecasting accuracy and nesting for factor-augmented regressions under the CCE-estimated fac-
tors. While their results in context of tests by Clark and McCracken (2001) bring important practical impli-
cations, merging CCE mechanics with a very novel class of tests by Pitarakis (2023a) and Pitarakis (2023b)
opens new possibilities. These include observed predictors and latent factors that are non-stationary, or
robustness to structural breaks in loadings, in a sense that their location in the sample does not mat-
ter. Most importantly, the overspecification of the number of factors has as asymptotically negligible
influence, as well. This goes against the usual trend of CCE literature, where the excess cross-section
averages result in bias even asymptotically. This eliminates the need to know the number of factors, as
we only need to bound the number of factors with the number of the averages from above to preserve
the asymptotic normality. In practice, the suggested upper bound is often substantial, as can be seen from
our empirical application on the novel EA-MD-QD dataset. Ultimately, Monte Carlo simulations reveal
an excellent small sample performance across a plethora of empirically relevant and challenging data
generating processes.
25
7 Appendix
In the auxiliary lemmas and Lemma 1, we will use the stacked notation, such that A ∈ R (t−1)× p stacks
p × 1 vectors over time in a recursive fashion.
1
Z′ E = O p ( N −1/2 T −τ/2 ) (A.1)
T 1+ τ
as ( N, T ) → ∞.
Proof. To begin with, we can write T 11+τ Z′ E = T 11+τ ∑tT=1 zt e′t . For this and the upcoming auxiliary results,
we use the total sample for t = 1, . . . , T, however, by Assumption 2 (c), the same rates hold in a recursive
setup for each t. Then for some positive ǫ,
! !
1 T 1 T
zt e′t > ǫ ≤ ǫ−1 E z e′
T 1+τ t∑ 1+ τ ∑ t t
P
=1 T t =1
!
1 1 T √ ′
T 1+τ t∑
= √ E zt ( Net )
ǫ N =1
!
1 1 T √
T 1+τ t∑
≤ √ E kzt k ( Net )
ǫ N =1
1 1 T √
= √ ∑
ǫ N T 1+ τ t = 1
E (k z t k) E ( Ne t )
!1/2 !
1 1 T 1 T √ 2 1/2
≤ √
ǫ NT τ/2 T 1+τ t=1
∑ E(kzt k)2 T t∑
E ( Net )
=1
| {z }
O ( 1)
−1/2 − τ/2
= O( N T ), (A.2)
because
T T q 2
1 2 1
∑ E(kzt k) ≤ ∑ E (tr(zt z′t ))
T 1+ τ t =1 T 1+ τ t =1
T
1
=
T 1+ τ ∑ E(tr[zt z′t ]) = O(1) (A.3)
t =1
by the concavity of function appearing in Jensen’s inequality. Note that this result gives a different rate
than for similar terms in Pitarakis (2023a), which is brought down exactly by time dependence in {ei,t }.
Also note that independence between the predictors and idiosyncratics is not strictly necessary, and we
use it for simplicity. Indeed,
! !
1 T 1 1 T √
zt e′t > ǫ ≤ √ E
T 1+τ t∑ T 1+τ t∑
P kzt k ( Net )
=1 ǫ N =1
1 1 T 1/2 √
2 1/2
2
≤ √ ∑ E k zt k
ǫ N T 1+ τ t = 1
E Net
26
1/2
√ 2
supt E Net 1/2
2
− τ/2
≤ √ sup E T zt
ǫ NT τ/2 t
= O( N −1/2 T −τ/2 ), (A.4)
2 1/2
because supt E T −τ/2 z t = O(1) by Lemma 3.1 in Magdalinos and Phillips (2009).
Corollary A.1 Under Assumptions 1-6, but ei,t is uncorrelated over time, we have that
1
Z′ E = O p ( N −1/2 T −(1+τ )/2 ) (A.5)
T 1+ τ
as ( N, T ) → ∞.
Proof. Assume that ei,t is uncorrelated over time. Then by Markov’s inequality we obtain
!
2
T T
1 1
zt e′t > ǫ ≤ ǫ−2 E 1+τ ∑ zt e′t
T 1+τ t∑
P
=1 T t =1
" #!
T T
−2 1 1 ′ ′
1 + τ t∑ ∑ zt e t e s zs
=ǫ E tr
T 1+ τ =1 s =1
" #!
T T
1 1
= ǫ−2 1+τ tr 1+τ ∑ ∑ E zs zt e′t es ′
T T t =1 s =1
" #!
T T
1 1
= ǫ−2 1+τ tr 1+τ ∑ ∑ E z′t zs E e′t es
T T t =1 s =1
" #!
T
1 1
= ǫ−2 1+τ tr 1+τ ∑ E z′t zt E e′t et
T T t =1
" #!
−2 1 1 T ′
′
T 1+τ t∑
=ǫ tr E zt zt E Net et
NT 1+τ =1
" #!
−2
supk0 ≤t≤ T E ( Ne′t et ) 1 T ′
T 1+τ t∑
≤ǫ tr E zt zt
NT 1+τ =1
" #
− 1 − 1− τ 1 T ′
≤ O( N T ) × tr 1+τ ∑ E zt zt
T t =1
| {z }
O ( 1)
− 1 − 1− τ
= O( N T ), (A.6)
1
Z′ u = O p ( T −(1+τ )/2 )
T 1+ τ
as T → ∞.
27
Proof. By using similar steps as in Lemma A.1, we get
! !
1 T −2 1 1 T T ′
T 1+τ t∑ T 1+τ t∑ ∑ zt zs u t +1 u s +1
P zt u t +1 > ǫ ≤ ǫ E
=1 T 1+ τ =1 s =1
!
−2 1 1 T ′ 2
T 1+τ t∑
=ǫ E zt zt E (ut+1 |Ft−1 )
T 1+ τ =1
T
σ2 1
= ∑ tr E (zt z′t )
ǫ − 2 T 1+ τ T 1+ τ t =1
−(1+ τ )
= O( T ), (A.7)
which implies that T 11+τ ∑tT=1 zt ut+1 = O p ( T −(1+τ )/2 ). Note that this rate is effectively assumed in
Pitarakis (2023a) and Pitarakis (2023b). However, it naturally follows from our more primitive assump-
tions.
because it determines the object that is being estimated. The way we approach this issue is by introducing
the following rotation matrix, which is chosen such that ΛH = [Ir , 0r ×(m−r ) ] and that is going to play the
−1
same role as Λ under m = r:
" #
−1 −1
Λr − Λr Λ −r
H= = [ Hr , H −r ] ∈ R m × m , (A.9)
0(m−r )×r I m −r
− 1′ ′ − 1′
where Hr = [Λr , 0r ×(m−r ) ]′ ∈ R m×r and H−r = [−Λ−r Λr , Im−r ]′ ∈ R m×(m−r ). If m = r, we define
−1 −1 √
H = Hr = Λr = Λ . We further introduce D N = diag(Ir , NIm−r ) ∈ R m×m with D N = Im if m = r.
′
By pre-multiplying bft by D N H , we obtain
′ ′ ′ ′
D N H bft = bf0t = D N H Λ ft + D N H et = f0t + e0t , (A.10)
′ −1 √ ′ − 1′
where f0t = [f′t , 0′(m−r )×1 ]′ ∈ R m and e0t = [er,t Λr , N (e−r,t − Λ−r Λr er,t )′ ]′ = [e0r,t′ , e0−′ r,t ]′ ∈ R m with
−1′b
′
er,t ∈ Rr and e−r,t ∈ R m−r being the partitions of et = [er,t , e′−r,t ]′ . If m = r, then bf0t = Λ = ft ft , f0t
− 1′
and e0t= Λ et , and so we are back in (A.8). Hence, since ker,t k = O p ( N −1/2 ) and ke−r,t k = O p (1),
0 0
when m > r we are no longer estimating ft but rather [f′t , e0−′ r,t ]′ .8 The fact that ft is included in this object
suggests that asymptotically, CCE should be able to account for the unknown factors even if m > r. By
ensuring the existence of H, Assumption 5 makes this possible. However, we also note that because of
the presence of e0−r,t , the asymptotic distribution theory will depend on whether m = r or m > r.
√ ′
scaling by N in D N and hence in e0−r,t is necessary, for without it H bft is estimating [ f′t , 0′( m−r )×1] ′ , whose second
8 The
28
7.3 Proof of Lemma 1
To prove Lemma 1, let us begin from the decomposition in Stauskas and Westerlund (2022) (proof of
Lemma 1, Appendix A):
D T (Q− 1b
N δt − δt )
e0
′
b 0 ′ b 0 + b 0 ′ (F MW F)+ F′ MW b 0 δ0
FD ( F M W F ) F M W − Z
= 0(m−r )×(t−1)
1 τ 0
T 4 + 2 Iq (W′ MbF0 W)+ W′ MbF0 − (W′ MF W)+ W′ MF Z b0δ
′
b 0 ′ b 0 + b 0 ′ (F MW F)+ F′ MW 0
DF (F MW F ) F MW − 0(m−r )×(t−1)
Er α
−
1
+ τ ′ + ′ ′ + ′
0
T 4 2 Iq (W MbF0 W) W MbF0 − (W MF W) W MF Er α
′ + F′ M
b 0 ′ b 0 + b 0 ′ ( F M W F ) W
+ DF (F MW F ) F MW − 0(m−r )×(t−1)
u
1 τ
T 4 + 2 Iq (W′ MbF0 W)+ W′ MbF0 − (W′ MF W)+ W′ MF u
I III V
= − + , (A.11)
II IV VI
1 τ 1 τ
where D T = diag(DF , T 4 + 2 Iq ) with DF = diag( T 4 + 2 Ir , T 1/4 Im−r ) and the definitions of I − VI are im-
plicit. Assume that ei,t is uncorrelated over time.
Now we can work out the order of each of the terms by applying simple adjustments to the proofs
of and by using the results in Stauskas and Westerlund (2022). Indeed, starting with I, and recalling that
′ ′ 1 τ
b 0 δ0 = b
Z F0 (D N H Λ )+′ α + Wθ, α = T − 4 − 2 α0 , MW W = 0(t−1)×n , and the fact that b F 0′ M W b
F0 is a full rank
matrix wp1, then
F 0′ M W b
DF (b F 0′ M W Z
F0 ) + b b 0 δ0
′ ′
F 0′ M W b
= DF (b F 0′ M W b
F0 ) + b F 0′ M W b
F0 (D N H Λ )+′ α + DF (b F0′ MW Wθ
F0 ) + b
′ ′
= (D N H Λ )+′ α0 = [α0′ , 01×(m−r )) ]′ , (A.12)
′ ′ ′ ′ 0 ′ ′
which holds wp1. Further recall that (F′ MW F)+ F′ MW = (D N H Λ )′ (F0 MW F0 )+ F0 MW , E (D N H Λ )′ α =
1 τ 0
T 4 + 2 Er α0 and the fact that F′ MW F is full rank and positive definite wp1. Then
t − 1− τ F ′ M W F = t − 1− τ F ′ F − t − 1− τ F ′ W ( t − 1− τ W ′ W ) + t − 1− τ W ′ F
→ p Σ f f − Σ′w f Σ− 1
ww Σ f w = Σ f .w , (A.13)
which is positive definite by assumption because ΣZZ is (Assumption 2). In addition, the rank of T −1−τ F′ MW F
is invariant to changes in T → ∞, so
( t − 1 − τ F ′ M W F ) + = O p ( 1) . (A.14)
0
In a similar manner, (t−1−τ W′ W)+ = O p (1). Turning now to T −1−τ F′ MW Er ,
0 0 0
T − 1 − τ F ′ M W Er α 0 ≤ T − 1− τ F ′ E r α0 − T − 1− τ F ′ W ( T − 1− τ W ′ W ) + T − 1 − τ W ′ Er α0
1 1 τ
= O p ( N − 2 T − 2 − 2 ). (A.15)
29
By using these results,
DF (F′ MW F)+ F′ MW Z
b 0 δ0
1 τ 1 τ 0
= T 4 + 2 α + T 4 + 2 ( F ′ M W F ) + F ′ M W Er α
0
= α0 + ( T − 1− τ F ′ M W F ) + T − 1− τ F ′ M W E r α0
1 1 τ
= α0 + O p ( N − 2 T − 2 − 2 ). (A.16)
whereas using the definition and properties of F0 together with previous results gives
1 τ
T 4 + 2 (W′ MF W)+ W′ MF Z
b 0 δ0
1 τ 0 1 τ
= T 4 + 2 ( W ′ M F W ) + W ′ M F Er α + T 4 + 2 θ
0 1 τ
= ( T −1−τ W′ MF W)+ T −1−τ W′ MF Er α0 + T 4 + 2 θ. (A.19)
( T − 1− τ W ′ M F W ) + − s − 1− τ Σ +
w. f = o p (1). (A.22)
( T − 1 − τ W ′ M F W ) + = O p ( 1) . (A.23)
Further, since
0
T − 1 − τ W ′ M F Er α 0
0 0
≤ T − 1− τ W ′ E r α 0 − T − 1− τ W ′ F ( T − 1− τ F ′ F ) + T − 1− τ F ′ E r α0
1 1 τ
= O p ( N − 2 T − 2 − 2 ), (A.24)
30
implying that
0
( T − 1− τ W ′ M F W ) + T − 1− τ W ′ M F E r α0
0 1 1 τ
≤ ( T − 1− τ W ′ M F W ) + T − 1− τ W ′ M F E r α0 = O p ( N − 2 T − 2 − 2 )
so that
1 τ
T 4 + 2 (W′ MF W)+ W′ MF Z
b 0 δ0
0 1 τ 1 τ 1 1 τ
= ( W ′ M F W ) + W ′ M F E r α0 + T 4 + 2 θ = T 4 + 2 θ + O p ( N − 2 T − 2 − 2 ) (A.25)
In detail,
− 1 0′ 0 −1 T − 1− τ F ′ F 0r ×(m−r )
G F F G = , (A.29)
0(m−r )×r 0(m−r )×(m−r )
" #
−1 1 τ −1
− 1 0′ 0 −1 T − 1− τ F ′ E r Λ r N 2 T − 1− 2 ( F ′ E − r − F ′ E r Λ r Λ − r )
G F E G = , (A.30)
0(m−r )×r 0(m−r )×(m−r )
" −1 ′
#
− 1 0′ 0 −1 T − 1 − τ ( Λ r ) ′ Er F 0(m−r )×r
G E F G = 1 τ ′ ′ −1 ′ , (A.31)
N 2 T −1− 2 (E−r F − Λ−r (Λr )′ Er F) 0(m−r )×(m−r )
and
" 0′ 0 τ 0′ 0
#
− 1 0′ 0 T − 1− τ E r E r T − 1− 2 E r E − r
G E E G −1 = τ 0′ 0 0′ 0 (A.32)
T − 1− 2 E − r E r T − 1 E − r E − r
where
0′ 0 −1 ′ −1
Er Er = ( Λr ) ′ Er Er ( Λr )
0′ 0 1 −1 ′ −1 ′ −1
Er E−r = N 2 ((Λr )′ Er E−r − (Λr )′ Er Er Λr Λ−r )
0′ 0 1 ′ −1 ′ −1 ′ −1
E −r Er = N 2 ( E − r E r Λ r − Λ − r ( Λ r ) ′ E r Er Λ r )
0′ 0 ′ ′ −1 ′ −1 ′ ′ −1 ′ −1
E −r E −r = N ( E −r E −r − E −r Er Λr Λ −r − Λ −r ( Λr ) ′ Er E −r + Λ −r ( Λr ) ′ Er Er Λr Λ −r )
31
τ 1 1 0 0′
Now notice that T −1− 2 F′ E is of order O p ( N − 2 T − 2 ), which implies that G−1 F0′ E G−1 and G−1 E F0 G−1
1 0′
are both of order O p ( T − 2 ) by symmetry arguments by Corollary A.1. However, by Lemma A.1 G−1 E F0 G−1 =
0′ 0
O p (1) if the idiosyncratics are serially correlated. Finally, it remains to address the terms in T −1 E E :
−1 ′ −1 −1 2 ′
T − 1− τ ( Λ r ) ′ E r E r ( Λ r ) ≤ Λr T − 1 − τ Er E r
= O p ( N −1 T − τ ), (A.33)
′ −1
by virtue of the fact that T −1 E E is O p ( N −1 ) while Λr and Λr are O p (1). In turn, (A.33) implies
1 τ −1 ′ −1 ′ −1 1 −1 τ ′
N 2 T −1− 2 ((Λr )′ Er E−r − (Λr )′ Er Er Λr Λ−r ) ≤ N 2 Λr T − 1− 2 E r E − r
1 −1 2 τ ′
+ N 2 Λr T − 1− 2 E r E r Λ −r
1 τ
= Op(N− 2 T− 2 ) (A.34)
1 τ ′ −1 ′ −1 ′ −1 1 τ
By symmetry, it also follows that N 2 T −1− 2 (E−r Er Λr − Λ−r (Λr )′ Er Er Λr ) = O p ( N − 2 T − 2 ). More-
over,
′ ′ −1 ′ −1 ′ ′ −1 ′ −1
NT −1 (E−r E−r − E−r Er Λr Λ−r − Λ−r (Λr )′ Er E−r + Λ−r (Λr )′ Er Er Λr Λ−r )
′ ′ −1 −1 ′
≤ N T − 1 E − r E − r + N T − 1 E − r Er Λr Λ −r + N Λr Λ −r T − 1 Er E − r
−1 2 2 ′
+ N Λr Λ −r T −1 E −r E −r
= O p ( 1) . (A.35)
rk[G−1 b
F 0′ b
F0 G−1 ] → rkG−1 SG−1 (A.38)
almost surely. Following Karabiyik et al. (2017), (A.37) and (A.38) lead to the key result
1 τ 1
( G −1 b F0 G−1 )+ − (G−1 SG−1 )+ = O p ( N − 2 T − 2 ) + O p ( T − 2 ).
F 0′ b (A.39)
( G −1 b
F 0′ b
F 0 G − 1 ) + = O p ( 1) . (A.40)
32
In turn, it follows by using (A.77) in the Supplement of Stauskas and Westerlund (2022) and (A.41) that
so that
and
0 0
( W ′ M F0 W ) + W ′ M F0 E r ≤ ( T − 1− τ W ′ M F 0 W ) + T − 1− τ W ′ M F 0 E r
1 1 τ
= O p ( N − 2 T − 2 − 2 ). (A.47)
33
0 1 1 τ
For III, recall from (A.16) that ( T −1−τ F′ MW F)+ T −1−τ F′ MW Er α = O p ( N − 2 T − 2 − 2 ). Further, we use the
0 1
fact that DF G−1 (G−1 F b 0′ M W b
F0 G −1 ) + G −1 b
F0′ MW Er α = O p ( N − 2 T −τ/2 ). Indeed, given that DF G−1 =
1
T − 4 Im , we have
1 0 1 τ 0
T − 4 G −1 b
F 0′ M W E r α = T − 2 − 2 G −1 b
F 0′ M W E r α0
1 τ 0 1 τ 0
≤ T − 2 − 2 G −1 b
F 0′ E r α0 + T − 1− τ b
F 0′ W ( T − 1− τ W ′ W ) + T − 2 − 2 G −1 W ′ E r α0
0 1 τ ′ 0
≤ T −(1+τ ) F′ Er α0 + T − 2 − 2 G−1 E Er α0
1 τ 0
+ T − 1− τ b
F 0′ W ( T − 1− τ W ′ W ) + T − 2 − 2 G − 1 W ′ Er α 0
0 τ ′ 0
≤ T −(1+τ ) F′ Er α0 + MT − 2 T −1 E Er α0
1 τ 0
+ T − 1− τ b
F 0′ W ( T − 1− τ W ′ W ) + T − 2 − 2 G − 1 W ′ Er α0
1 τ τ 1 1 τ
= O p ( N − 2 T − 2 ) + O p ( N −1 T − 2 ) + O p ( N − 2 T − 2 − 2 )
1 τ
= O p ( N − 2 T − 2 ). (A.49)
With respect to (G−1 b F 0′ M W Fb0 G−1 )+ , we proceed analogously to (A.36). Let SW be defined as
" #
F′ MW F 0r ×(m−r )
S MW = 0′ 0
0(m−r )×r E−r MW E−r
" # " #
F′ F 0r ×(m−r ) F′ W(W′ W)+ W′ F 0r ×(m−r )
= 0′ 0 − 0′ 0
0(m−r )×r E−r E−r 0(m−r )×r E −r W ( W ′ W ) + W ′ E −r
= S − SW . (A.50)
and exploiting the definition of MW yields the decomposition
G −1 b
F 0′ M W b
F0 G −1 = G −1 b
F 0′ b
F0 G −1 − G −1 b
F 0′ W ( W ′ W ) + W ′ b
F0 G −1 . (A.51)
Thus,
G −1 b
F 0′ M W b F0 G −1 − G −1 S MW G −1
≤ G −1 F b 0′ F
b0 G−1 − G−1 SG−1 + G−1 b F 0′ W ( W ′ W ) + W ′ b
F0 G −1 − G −1 S W G −1
1 τ 1
= O p ( N − 2 T − 2 ) + O p ( T − 2 ), (A.52)
0 0 1 τ 0′ 0
using the fact that T −1−τ F0′ E and T −1−τ W′ E are both of order O p ( T − 2 − 2 ), while T −1 E E is of
order O p (1), which ensures that the first term on the right-hand side of (A.52) drives the order (i.e. the
second one does not dominate). Indeed the slowest decaying component of the second term is
0
T −1−τ F′ W( T −(1+τ ) W′ W)+ T −1−τ/2 W′ E−r (A.53)
0
≤ T −1/2 T −(1+τ ) F′ W ( T −(1+τ ) W′ W)+ T −(1+τ )/2 W′ E−r
= O p ( T −1/2 ). (A.54)
Therefore, we obtain
1 τ 1
G −1 b F0 G −1 = G −1 S MW G −1 + O p ( N − 2 T − 2 ) + O p ( T − 2 ).
F 0′ M W b (A.55)
Since G−1 SG−1 and G−1 SW G−1 are positive definite wp1, then so are G−1 SMW G−1 and G−1 b
F 0′ M W b
F0 G −1 .
Further, the rank of G SMW G is m, which is also the rank of G b
− 1 − 1 − 1 F MW b
0 ′ 0 − 1
F G . It then follows that as
N, T → ∞,
rk[G−1 b
F 0′ M W b
F0 G−1 ] → rkG−1 SMW G−1 (A.56)
34
almost surely. Following Karabiyik et al. (2017), (A.55) and (A.56) lead to the key result
1 τ 1
( G −1 b
F 0′ M W b
F0 G −1 ) + − ( G −1 S MW G −1 ) + = O p ( N − 2 T − 2 ) + O p ( T − 2 ). (A.57)
( G −1 b
F 0′ M W b
F 0 G − 1 ) + = O p ( 1) . (A.58)
as required. Finally, we are able to obtain the order III by using the definition of α and DF G−1 = Im T −1/4 :
′ + F′ M
′ + ′ ( F M W F ) W 0
III = DF (b F MW b
0
F ) b
0 0
F MW − Er α
0(m−r )×(t−1)
−(1+τ ) ′ + T −(1+ τ ) F ′ M
− 1 b 0′ − + − − 1
− τ
′ 0 ( T F M W F ) W 0
≤ (G F MW b F G ) G T 2 2b
0 1 1 0
F M W Er α + 0
E r α0
0(m−r )×(t−1)
1 1
= O p (( NT )− 2 ) + O p ( N − 2 T −τ ). (A.60)
where SMW is defined as in (A.50). Beginning from the third term on the right-hand side of (A.61),
" # " 0′ #
− 1− τ F ′ M F ) +
+ 0′ −1 ( T W 0r ×(m−r ) −1 Er M W u
D F S MW E M W u = D F G 0′ 0 G 0′
0(m−r )×r ( T −1 E −r M W E −r ) + E −r M W u
" 3 τ 0′
#
( T − 1− τ F ′ M W F ) + T − 4 − 2 E r M W u
= 0′ 0 3 0′
( T −1 E −r M W E−r ) + T − 4 E −r M W u
1
= O p ( T − 4 ), (A.62)
35
Finally, the order of the remaining term in (A.61) can be shown to be
− 1 0′
D F G −1 ( G −1 b
F 0′ M W b
F0 G −1 ) + − ( G −1 S + −1
MW G ) G F M W u
1
≤ ( G −1 b
F 0′ M W b
F0 G −1 ) + − ( G −1 S + −1
MW G ) T − 4 G − 1 F 0′ M W u
3 τ
≤ ( G −1 b
F 0′ M W b
F0 G −1 ) + − ( G −1 S + −1
MW G ) T − 4 − 2 F 0′ u
3 τ
+ ( G −1 b
F 0′ M W b
F0 G −1 ) + − ( G −1 S + −1
MW G ) T − 1− τ F 0′ W ( T − 1− τ W ′ W ) + T − 4 − 2 W′ u
1 1 τ 3
= O p ( N − 2 T − 4 − 2 ) + O p ( T − 4 ). (A.65)
Hence, putting together the results in (A.62), (A.64) and (A.65) gives the order of V:
1
V = O p ( T − 4 ). (A.66)
Using (A.77) in the Supplement of Stauskas and Westerlund (2022), note that
1 τ
T − 2 − 2 W′ (MbF0 − MF0 )u
1 τ 0 0′ 0 0′ 0 0′ 0 0′
≤ T − 2 − 2 W ′ ( E − r ( E − r E − r ) + E − r + E r ( F ′ F ) + E r + Er ( F ′ F ) + F ′ + F ( F ′ F ) + E r + b F 0′ b
F0 [(b F 0′ ) u
F0 ) + + S + ]b
τ 0 0′ 0 1 0′ τ 0 1 0′
≤ T − 1− 2 W ′ E − r ( T −1 E −r E −r ) + T − 2 E − r u + T − τ T − 1− 2 W ′ E r ( T − 1− τ F ′ F ) + T − 2 Er u
τ τ 0 1 τ τ 1 0′
+ T − 2 T − 1− 2 W ′ E r ( T − 1− τ F ′ F ) + T − 2 − 2 F ′ u + T − 2 T − 1− τ W ′ F ( T − 1− τ F ′ F ) + T − 2 Er u
1 τ
+ T − 2 − 2 W′ b
F0 G −1 ( G −1 b
F 0′ F
b0 G−1 )+ − (G−1 SG−1 )+ G −1 b
F 0′ u
1 τ 1 1 τ 1 τ 1 τ 1
= O p ( T − 2 ) + O p ( N −1 T − 2 ) + O p ( N − 2 T − 2 − 2 ) + O p ( N − 2 T − 2 ) + O p ( N − 2 T − 2 ) + O p ( T − 2 )
1 1 τ
= O p ( T − 2 ) + O p ( N − 2 T − 2 ), (A.68)
by (A.39) and the independence of W and E. By putting this latter result together with ( T −1−τ W′ MbF0 W)+ −
1 τ 1
( T −1−τ W′ MF0 W)+ = O p ( N − 2 T − 2 ) + O p ( T − 2 ), we can show that (A.67) reduces to
1 1 τ
VI ≤ T − 4 ( T −1−τ W′ MbF0 W)+ − ( T −1−τ W′ MF0 W)+ T − 2 − 2 W′ (MbF0 − MF0 )u
1 1 τ
+ T − 4 ( T −1−τ W′ MbF0 W)+ − ( T −1−τ W′ MF0 W)+ T − 2 − 2 W ′ M F0 u
1 1 τ
+ T − 4 ( T −1−τ W′ MbF0 W)+ T − 2 − 2 W′ (MbF0 − MF0 )u
3 1 1 τ
= O p ( T − 4 ) + O p ( N − 2 T − 4 − 2 ). (A.69)
36
7.4 Proof of Lemma 2
We now prove Lemma 2. Consider (a). Beginning from the definitions of ue1,t+1 and ub2,t+1 and the fact
′ 0′
that δet zt = δet Q′N b αt ′ e0r,t :
z0t − e
′
ue2,t+1 − ub2,t+1 = δbt′ b
zt − δet zt
= δet′ b
zt − δet0′ Q′N b
zt + eαt ′ e0r,t
= δbt′ Q−1′ b
z0t − δet0′ b
N αt ′ e0r,t
z0t + e
0
= (Q− 1b
N δt − δet )′ b αt ′ e0r,t ,
z0t + e
0
= (Q− 1b e ′ z0t + (e
N δt − δ t ) b α t − α)′ e0r,t + α′ e0r,t (A.71)
the first term of which Lemma 1 is concerned with. With respect to the last term on the right-hand side of
1
(A.71), whereas, recall that e0r,t = O p ( N − 2 ) and observe that
αt − α
e ≤ ( t − 1− τ F ′ M W F ) − 1 t − 1− τ F ′ M W u
1 τ
= O p ( T − 2 − 2 ). (A.72)
αt ′ e0r,t
e = (eαt − α)′ e0r,t + α′ e0r,t
≤ (eαt − α)′ e0r,t + α′ e0r,t
1 1 τ
= Op(N− 2 T− 2 − 2 ) (A.73)
which holds uniformly in t, thus meaning that the last term on the right-hand side of (A.71) is negligible.
et − θ)′ wt +
Let us consider the last two summations in (3.3). First, using the definition ue1,t+1 = ut+1 − (θ
′
α ft and the result in (A.71) gives
1 k0 + m0 −1
√ ∑ ue1,t+1(ue2,t+1 − ub2,t+1 )
n t= k0
1 k0 + m0 −1 1 k0 + m0 −1 e
= √ ∑
n t= k0
ut+1 (ue2,t+1 − ub2,t+1 ) − √ ∑ (θt − θ)′ wt (ue2,t+1 − ub2,t+1)
n t= k0
1 k0 + m0 −1 ′
+ √ ∑ α ft (ue2,t+1 − ub2,t+1)
n t= k0
1 k0 + m0 −1 e0 ′ z0 + (e 1 k0 + m0 −1 e
= √ ∑ ut+1 ((Q− 1b
N δt − δ t ) bt α t − α)′ e0r,t + α′ e0r,t ) − √ ∑ (θt − θ)′ wt (ue2,t+1 − ub2,t+1 )
n t= k0 n t= k0
1 k0 + m0 −1 ′
+ √ ∑ α ft (ue2,t+1 − ub2,t+1)
n t= k0
1 k0 + m0 −1 −1 b e 0 ′ 0 1 k0 + m0 −1 ′ 0 1 k0 + m0 −1 ′ 0
= √ ∑ t +1 N t t t
n t= k0
u ( Q δ − δ ) b
z + √ ∑ t
n t= k0
(eα − α ) e u
r,t t+1 + √ ∑ α er,t ut+1
n t= k0
1 k0 + m0 −1 e 1 k0 + m0 −1 ′
− √ ∑ (θt − θ)′ wt (ue2,t+1 − ub2,t+1 ) + √ ∑ α ft (ue2,t+1 − ub2,t+1)
n t= k0 n t= k0
= I + II + III − IV + V, (A.74)
37
q
T √ 1
where the definitions of I-V are implicit. Begining from II, and using the fact that n = 1− π0
+
O( T −1/2 ), we have
1 k0 + m0 −1
√ ∑ (eαt − α)′ e0r,t ut+1
n t= k0
1 1 1 k0 + m0 −1 √
=√ √ 1/2+τ/2 ∑ αt − α)′ ( Ne0r,t )ut+1
T 1/2+τ/2 (e
N n T t= k0
1 T 1 k0 +m0 −1 1/2+τ/2 √
= T −τ/2 √ √ ∑ T (eαt − α)′ ( Ne0r,t )ut+1
N nT T t=k0
r
1 T 1 k0 + m0 −1 √ 0
≤ T −τ/2 √ sup T 1/2+τ/2 (e
T t∑
αt − α) |( Ner,t )ut+1 |
N n t = k0
1 τ
= O p ( N − 2 T − 2 ), (A.75)
1 τ
because supt kT 2 + 2 (e
αt − α)k = O p (1) as in (38) and (46) of Pitarakis, 2023a, and the term with the sum-
mation is of the same order. In a similar manner, the order of III is obtained as follows:
1 k0 + m0 −1 ′ 0
√ ∑ α er,t ut+1
n t= k0
r
1 T 0 1 k0 + m0 −1 √
≤ √ 1 τ
α √ ∑ ( Ne0r,t )ut+1
NT 4 + 2 n T t= k0
1 1 τ
= Op(N− 2 T− 4 − 2 ) (A.76)
1 k0 + m0 −1 e0 ′
√ ∑ u t +1 ( Q − 1b −1 0
N δt − δt ) D T D T b
zt
n t= k0
− 14 1 k0 + m0 −1 −1 b 0 1
=T √ ∑ (Q N δ t − δet )′ D T T 4 D− 1 0
T bzt u t +1
n t= k0
τ 1 1 k0 + m0 −1 h −1 b 0
i
= T 2−2 √ ∑ (Q N δ t − δet )′ D T T 1/4 ( T 1/4 D− 1
T )T
− τ/2 0
b
zt u t +1
n t= k0
1
= o p (T − 4 ) (A.77)
h i
by noting that the summand, (Q− 1b
δ t − e0t )′ D T T 1/4 ( T 1/4 D−1 ) T −τ/2 b
δ z0t ut+1 , is a heterogeneous martin-
N T
0
gale difference process by Assumption 1. In particular, (Q− 1b e ′
N δt − δ t ) D T T
1/4 = O p (1) by Lemma 1,
while
as supt T −τ/2 zt = O p (1) based on Magdalinos and Phillips (2009). For IV, whereas, we can start by
showing that, as T → ∞, we have
1 τ 1 τ
T 4 + 2 (θ
et − θ) = T 4 + 2 ( T −1−τ W′ W)−1 T −1−τ W′ (Fα + u)
1 τ
= (t−1−τ W′ W)−1 t−1−τ W′ Fα0 + ( T 1+τ t−1−τ ) T 4 + 2 (t−1−τ W′ W)−1 T −1−τ W′ u
38
1 1 τ
= (t−1−τ W′ W)−1 t−1−τ W′ Fα0 + T − 4 ( T 1+τ t−1−τ )(t−1−τ W′ W)−1 T − 2 − 2 W′ u
1
= (t−1−τ W′ W)−1 t−1−τ W′ Fα0 + O p ( T − 4 )
→ p Σ− 1 0
ww Σ w f α , (A.79)
1 τ
because of the result in (A.21) and the fact that (t−1−τ W′ W)−1 T − 2 − 2 W′ u = O p (1). Hence, we have
1 τ
et − θ) = O p (1). To be more precise, we have that sup T 41 + τ2 (θ
that T 4 + 2 (θ et − θ) = O p (1). Further,
t
substituting (A.71) in IV yields
1 k0 + m0 −1 e
√ ∑ (θt − θ)′ wt (ue2,t+1 − ub2,t+1)
n t= k0
1 k0 + m0 −1 e e0 ′ z0t + (e
=√ ∑ (θt − θ)′ wt ((Q− 1b
N δ t − δt ) b αt − α )′ e0r,t + α′ e0r,t )
n t= k0
1 k0 + m0 −1 e e0t )′ b 1 k0 + m0 −1 e
=√ ∑ (θt − θ)′ wt (Q−
N
1b
δ t − δ z 0
t + √ ∑ (θt − θ)′ wt (eαt − α)′ e0r,t
n t= k0 n t= k0
1 k0 + m0 −1 e
+√ ∑ (θt − θ)′ wt α′ e0r,t .
n t= k0
(A.80)
1 k0 + m0 −1 e
√ ∑ (θt − θ)′ wt α′ e0r,t
n t= k0
k0 + m0 −1
1 τ 1 1 τ 1 τ 1 τ
= T− 4 − 2 √ ∑ T 4 + 2 (θet − θ)′ T − 4 − 2 wt T 4 + 2 α′ e0r,t
n t= k0
1 τ 1 τ 1 k0 + m0 −1 − 1 − τ
≤ T − 4 − 2 α0 sup T 4 + 2 (θ
et − θ) √ ∑ T 4 2 wt e0r,t′
k0 ≤ t ≤ k0 + m0 −1 n t= k0
r k0 + m0 −1
1 τ T 1
0 4+2 (θet − θ) e0r,t
≤ α sup T
n T 1+ τ ∑ wt
k0 ≤ t ≤ k0 + m0 −1 t= k0
r ! 12 ! 21
k0 + m0 −1 k0 + m0 −1
− τ2 1 τ T 1 2 1 2
0 4+2 (θet − θ) e0r,t
≤T α sup T
n T 1+ τ ∑ wt
T ∑
k0 ≤ t ≤ k0 + m0 −1 t= k0 t= k0
1 τ
= O p ( N − 2 T − 2 ). (A.81)
1 k0 + m0 −1 e
√ ∑ (θt − θ)′ wt (eαt − α)′ e0r,t
n t= k0
k0 + m0 −1
1 τ 1 1
= T− 4 − 2 √ ∑
τ
et − θ)′ T − 21 − τ2 wt T 21 + τ2 (e
T 4 + 2 (θ αt − α)′ e0r,t
n t= k0
1 1 τ 1 τ 1 k0 + m0 −1 − 1 − τ
≤ T− 4 sup T 2 + 2 (e
αt − α) sup T 4 + 2 (θ
et − θ)√ ∑ T 2 wt e0r,t′
k0 ≤ t ≤ k0 + m0 −1 k0 ≤ t ≤ k0 + m0 −1 n t= k0
r
1 1 τ 1 τ T 1 k0 + m0 −1
≤ T− 4 T 2 + 2 (e + e wt e0r,t
n T 1+τ t∑
sup αt − α) sup T 4 2 (θt − θ)
k0 ≤ t ≤ k0 + m0 −1 k0 ≤ t ≤ k0 + m0 −1 = k0
39
r ! 12
k0 + m0 −1
− 14 − 2τ 1 τ 1 τ T 1 2
2+ 2 4+2 (θet − θ)
≤T sup T (eαt − α) × sup T
n T 1+ τ ∑ wt
k0 ≤ t ≤ k0 + m0 −1 k0 ≤ t ≤ k0 + m0 −1 t= k0
! 21
k0 + m0 −1
1 2
× ∑ e0r,t
T t= k0
1 1 τ
= O p ( N − 2 T − 4 − 2 ). (A.82)
1 k0 + m0 −1 1 + τ e 1 τ
e0
= √ ∑ T 4 2 (θt − θ)′ T − 4 − 2 wt b
z0t ′ D− 1 −1 b
T D T (Q N δt − δt )
n t= k0
r k0 + m0 −1
1 τ 0 T 3 τ
≤ sup T 4+2 (θet − θ) sup D T (Q− 1b
N δt − δet ) ∑ T − 4 − 2 wtb
z0t ′ D−
T
1
k0 ≤ t ≤ k0 + m0 −1 k0 ≤ t ≤ k0 + m0 −1 n t= k0
r k0 + m0 −1
1 1 τ τ 0 T 1
≤ T 4 D−
T
1
sup T 4+2 (θet − θ) sup T 2 D T (Q− 1b
N δt − δet ) ∑ z0t ′
wtb
k0 ≤ t ≤ k0 + m0 −1 k0 ≤ t ≤ k0 + m0 −1 n T 1+ τ t= k0
≤ O ( 1) ×
1 τ
4+2 (θet − θ)
τ
D T (Q− 1b e0
sup T sup T N δt − δ t )
2
k0 ≤ t ≤ k0 + m0 −1 k0 ≤ t ≤ k0 + m0 −1
! 21 ! 21
k0 + m0 −1 k0 + m0 −1
1 2 1 2
× ∑ wt ∑ z0t
b
T 1+ τ t= k0
T 1+ τ t= k0
1 τ 1
= O p ( N − 2 ) + O p ( T 2 − 4 ), (A.83)
which is negligible if τ ∈ 0, 21 . Hence, IV reduces to
1 k0 + m0 −1 e 1 τ 1
√ ∑ (θt − θ)′ wt (ue2,t+1 − ub2,t+1 ) = O p ( N − 2 ) + O p ( T 2 − 4 ). (A.84)
n t= k0
Finally, the order of V can be obtained by employing an argument similar IV. That is,
1 k0 + m0 −1 ′
√ ∑ α ft (ue2,t+1 − ub2,t+1)
n t= k0
1 k0 + m0 −1 ′ e0 ′ z0t + √1
k0 + m0 −1
= √ ∑ α ft ( Q − 1b
N δ t − δt ) b ∑ α′ ft (e
αt − α)′ e0r,t
n t= k0 n t= k0
1 k0 + m0 −1 ′ ′ 0
+ √ ∑ α ft α er,t
n t= k0
1 τ 1
= O p ( N − 2 ) + O p (T 2 − 4 ) (A.85)
because
1 k0 + m0 −1 ′ 0
√ ∑ α ft (Q−N1 δbt − δet )′bz0t
n t= k0
1 k 0 + m 0 − 1 1 + τ ′ − 1 − τ 0′ − 1
= √ ∑ T 4 2 α T 4 2 ft bzt DT
n t= k0
40
r k0 + m0 −1
0 T 3 τ
≤ α0 sup D T (Q− 1b
N δt − δet ) ∑ T − 4 − 2 ft b
z0t ′ D−
T
1
k0 ≤ t ≤ k0 + m0 −1 n t= k0
! 21 ! 21
k0 + m0 −1 k0 + m0 −1
τ 0 1 2 1 2
≤ O ( 1) × α 0 sup T D T (Q−
2 1b
N δt − δet ) ∑ ft ∑ z0t
b
k0 ≤ t ≤ k0 + m0 −1 T 1+ τ t= k0
T 1+ τ t= k0
− 12 τ 1
2−4
= Op(N ) + O p (T
), (A.86)
which is negligible if τ ∈ 0, 21 similarly to (A.83), whereas
1 k0 + m0 −1 ′
√ ∑ α ft (eαt − α)′ e0r,t
n t= k0
1 k0 + m0 −1 1 + τ ′ − 3 − τ 1 τ
= √ ∑ αt − α)′ e0r,t′
T 4 2 α T 4 ft T 2 + 2 (e
n t= k0
1 τ 1 k0 + m0 −1 − 3 − τ 0
≤ α0 sup T 2 + 2 (e √
α t − α) ∑ T 4 ft er,t
k0 ≤ t ≤ k0 + m0 −1 n t= k0
r
1 1 τ T 1 k0 + m0 −1
≤ T − 4 α0 T 2 + 2 (e ft e0r,t
n T 1+τ t∑
sup αt − α )
k0 ≤ t ≤ k0 + m0 −1 = k0
r ! 12 ! 21
k0 + m0 −1 k0 + m0 −1
1 τ 1 τ T 1 2 1 2
≤ T − 4 − 2 α0 sup T 2 + 2 (e
αt − α ) 1+ τ ∑ ft ∑ e0r,t
k0 ≤ t ≤ k0 + m0 −1 n T t= k0
T t= k0
− 12 − 14 − τ2
= Op(N T ), (A.87)
and
1 k0 + m0 −1 ′ ′ 0
√ ∑ α ft α er,t
n t= k0
1 k0 + m0 −1 1 + τ ′ − 1 − τ 1 τ
= √ ∑ T 4 2 α T 2 ft T 4 + 2 α′ e0r,t
n t= k0
21 k0 + m0 −1 − 1 − τ 0
≤ α0 √ ∑ T 2 ft er,t
n t= k0
r
0 2 T 1 k0 + m0 −1
ft e0r,t
n T 1+τ t∑
≤ α
= k0
r ! 21 ! 12
τ 2 T 1 k0 + m0 −1 2 1 k0 + m0 −1
2
≤ T − 2 α0 e0r,t
n T 1+τ t∑ ∑
ft
= k0
T t= k0
1 τ
= O p ( N − 2 T − 2 ). (A.88)
1 k0 + m0 −1 1 τ 1
√ ∑ ue1,t+1(ue2,t+1 − ub2,t+1) = O p ( N − 2 ) + O p (T 2 − 4 ).
n t= k0
(A.89)
which is negligible if τ ∈ 0, 21 .
41
0
Now, let us turn our attention to Lemma 2 (b) by recalling that ue2,t+1 − ub2,t+1 = (Q− 1b e ′ z0 + (e
N δt − δt ) bt αt −
′ 0 ′ 0
α) er,t + α er,t and proceeding as follows:
k0 + l20 −1
1 3 k0 + m0 −1 e0t )′ b 3 k0 + m0 −1
√ ∑ (ue2,t+1 − ub2,t+1 )2 ≤ √ ∑ ((Q−
N
1b
δ t − δ z 0 2
t ) + √ ∑ ((eαt − α)′ e0r,t )2
n t= k0 n t= k0 n t= k0
3 k0 + m0 −1 ′ 0 2
+ √ ∑ (α er,t )
n t= k0
= 3(U1 + U2 + U3 ). (A.90)
Starting from U1 ,
k0 + l20 −1
1 e0 ′
U1 = √ ∑ ((Q− 1b −1 0 2
N δ t − δt ) D T D T b
zt )
n t= k0
k0 + l20 −1
0 2 1 2 2
≤ sup D T (Q− 1b
N δt − δet ) √ ∑ D−
T
1
z0t
b
k0 ≤ t ≤ k0 + m0 −1 n t= k0
r k0 + m0 −1
2 τ 0 2 T 1 2
≤ T 1/4
D−
T
1
sup T 2 D T (Q− 1b
N δt − δet ) ∑ z0t
b
k0 ≤ t ≤ k0 + m0 −1 n T 1+ τ t= k0
r k0 + m0 −1
τ
e0 2 T 1 2
≤ O p ( 1) × sup T 2 D T (Q− 1b
N δt − δ t ) ∑ z0t
b
k0 ≤ t ≤ k0 + m0 −1 n T 1+ τ t= k0
−1 τ − 12
= Op(N ) + O p (T ), (A.91)
which is negligible τ ∈ 0, 21 . Next,
k0 + l20 −1
1 1 τ 1 τ
U2 = √ ∑ ( T 2 + 2 (eαt − α)′ T − 2 − 2 e0r,t )2
n t= k0
r k0 + m0 −1
− 21 − τ 1 τ 2 T1 2
2+2 e0r,t
≤ T sup T (eαt − α )
nT ∑
k0 ≤ t ≤ k0 + m0 −1 t= k0
− 21 − τ
= O p ( N −1 T ). (A.92)
Moving on to U3 ,
k0 + l20 −1
1 1 τ 1 τ
U3 = √ ∑ ( T 4 + 2 α′ T − 4 − 2 e0r,t )2
n t= k0
r k0 + m0 −1
−τ 0 2 T1 2
≤ T sup α ∑ e0r,t
k0 ≤ t ≤ k0 + m0 −1 nT t= k0
−1 − τ
= Op(N T ). (A.93)
Combining (A.91)-(A.93) yields
k0 + l20 −1
1 1
√ ∑ (ue2,t+1 − ub2,t+1 )2 = O p ( N −1 ) + O p ( T τ − 2 ). (A.94)
n t= k0
Then we show that the term in (c) is negligible. Indeed, using the definition ue2,t+1 = ut+1 − (δet − δ)′ zt
and the result in (A.71) gives
k0 + l20 −1
1
√
n
∑ ue2,t+1 (ue2,t+1 − ub2,t+1 )
t= k0
42
k0 + l20 −1
1 0
=√ ∑ ut+1 ((Q− 1b
N δ t − δt ) b αt − α )′ e0r,t + α′ e0r,t )
e ′ z0t + (e
n t= k0
k0 + l20 −1
1 0
+√ ∑ (δet − δ )′ zt ((Q− 1b e ′ z0 + (e
N δt − δt ) bt αt − α)′ e0r,t + α′ e0r,t )
n t= k0
k0 + l20 −1 k0 + l20 −1 k0 + l20 −1
1 0 1 1
=√ ∑ (Q− 1b
N δt − δet )′ b
z0t ut+1 +√ ∑ (eα t − α)′ e0r,t ut+1 +√ ∑ α′ e0r,t ut+1
n t= k0 n t= k0 n t= k0
k0 + l20 −1 k0 + l20 −1
1 e0 ′ z0t + √1
+√ ∑ (δet − δ )′ zt (Q− 1b
N δt − δt ) b ∑ (δet − δ)′ zt (eαt − α)′ e0r,t
n t= k0 n t= k0
k0 + l20 −1
1
+√ ∑ (δet − δ )′ zt α ′ e0r,t (A.95)
n t= k0
Note that the first three terms in this expression are the same as in (A.74), so we only address the other
three terms. Before looking at the fourth term, observe that
δet − δ ≤ ( t − 1− τ Z ′ Z ) − 1 t − 1− τ Z ′ u
1 τ
= O p ( T − 2 − 2 ). (A.96)
due to fact that t−1−τ Z′ Z → p Σ ZZ which is positive definite wp1 (Assumption 2). Hence,
k0 + l20 −1
1 e0 ′ z0
√ ∑ (δet − δ)′ zt (Q− 1b
N δ t − δt ) bt
n t= k0
k0 + l20 −1
1 1 τ 1 τ
e0 ′ z0t
= √ ∑ T 2 + 2 (δet − δ)′ T − 2 − 2 zt D− 1 −1 b
T D T (Q N δt − δt ) b
n t= k0
r k0 + l20 −1
1 τ τ 0 T 1
≤ D−
T
1
sup T 2+2 (δet − δ) sup T 2 D T (Q− 1b
N δt − δet ) ∑ zt z0t
b
k0 ≤ t≤ k0 + l20 −1 k0 ≤ t≤ k0 + l20 −1
n T 1+ τ t= k0
r ! 12
k0 + m0 −1
− 41 1 τ τ 0 T 1 2
≤ MT sup T 2+2 (δet − δ ) sup T D T (Q−
2 1b
N δt − δet ) ∑ zt
k0 ≤ t≤ k0 + l20 −1 k0 ≤ t≤ k0 + l20 −1
n T 1+ τ t= k0
! 21
k0 + m0 −1
1 2
× ∑ z0t
b
T 1+ τ t= k0
1 1 τ 1
= O p ( N − 2 T − 4 ) + O p ( T 2 − 2 ), (A.97)
which is negligible for all values of τ ∈ (0, 1). With respect to the fifth term in (A.95), by using the results
in (A.72) and (A.96), we have that
k0 + l20 −1
1
√ ∑ (δet − δ)′ zt (eαt − α )′ e0r,t
n t= k0
k0 + l20 −1
1 1 τ 1 τ
= √ ∑ T 2 + 2 (δet − δ )′ T −1−τ zt T 2 + 2 (e
αt − α)′ e0r,t
n t= k0
r k0 + l20 −1
− 12 1 τ 1 τ T 1
≤T sup T 2+2 (δet − δ ) sup T 2+2 (eαt − α ) ∑ zt e0r,t
k0 ≤ t≤ k0 + l20 −1 k0 ≤ t≤ k0 + l20 −1
n T 1+ τ t= k0
43
r ! 12
k0 + m0 −1
− 12 − 2τ 1 τ 1 τ T 1 2
≤T sup T 2+2 (δet − δ ) sup T 2+2 (eαt − α ) ∑ zt
k0 ≤ t≤ k0 + l20 −1 k0 ≤ t≤ k0 + l20 −1
n T 1+ τ t= k0
! 21
k0 + m0 −1
1 2
× ∑ e0r,t
T t= k0
1 1 τ
= O p ( N − 2 T − 2 − 2 ). (A.98)
k0 + l20 −1
− 34 − τ 1 1 τ 1 τ
=T √ ∑ T 2 + 2 (δet − δ)′ zt T 4 + 2 α′ e0r,t
n t= k0
r k0 + l20 −1
− 14 1 τ T 1
=T α 0
sup T 2+2 (δet − δ ) ∑ zt e0r,t
k0 ≤ t≤ k0 + l20 −1
n T 1+ τ t= k0
r ! 12 ! 21
k0 + m0 −1 k0 + m0 −1
− 14 − 2τ 1 τ T 1 2 1 2
≤T α 0
sup T 2+2 (δet − δ ) ∑ zt ∑ e0r,t
k0 ≤ t≤ k0 + l20 −1
n T 1+ τ t= k0
T t= k0
1 1 τ
= Op (N− 2 T− 4 − 2 ) (A.99)
Parts (d) and (e) follow directly based on the logic in Pitarakis (2023b), and in particular, the passage from
(A.6) to (A.8), which demonstrates that the remainder still vanishes inside of the integral. That is
k 0 + l2 − 1
1 n
n 1 1 k 0 + l2 − 1 2 1
n
n
∑ l2 √ n ∑ (ue2,t+1 − ub2,t+1 )2 ≤ sup √ ∑ ( e
u 2,t+1 − b
u 2,t+1 ) ∑
n l2 =⌊ nν ⌋ t= k0 ⌊ nν0 ⌋≤ t≤ n n t= k0 n l =⌊nν ⌋ l2
0 2 0
k 0 + l2 − 1 Z 1
1 2 1
≤ sup √
n
∑ (ue2,t+1 − ub2,t+1 )
λ2 = ν0 λ2
dλ2
⌊ nν0 ⌋≤ t≤ n t= k0
+ o p ( 1)
= o p ( 1) , (A.101)
44
Then the natural unfeasible and feasible estimators of (A.102) are obtained as
!2
T −1 T −1
b2 = 1
φ ∑ ub22,t+1 −
1
∑ ub22,t+1 , (A.103)
n t= k0
n t= k0
and
!2
T −1 T −1
e2 = 1
φ ∑ ue22,t+1 −
1
∑ ue22,t+1 (A.104)
n t= k0
n t= k0
respectively. To demonstrate (e), it is convenient to start from the decomposition of Margaritella and Stauskas
(2024):
! " #!
2 2 2 T −1 2 1 T −1 2 2 2 1 T −1 2 2
e −φ b =
n t∑
ue2,t+1 − ∑ ue2,t+1
n t∑
φ (ue2,t+1 − ub2,t+1 ) + (ue2,t+1 − ub2,t+1 )
= k0
n t=k = k0
0
" #!2
1 T −1 2 2 1 T −1 2 2
n t∑ n t∑
− (ue2,t+1 − ub2,t+1 ) + (ue2,t+1 − ub2,t+1 )
= k0 = k0
!
2 T −1 2 1 T −1
1 T −1 2
∑ ue2,t+1 (ue22,t+1 − ub22,t+1 ) − 2 ∑ ue22,t+1 (ue2,t+1 − ub22,t+1 )
n t∑
=
n t=k n t=k = k0
0 0
! !
T −1 T −1 T −1
1 2 2 1 2 1 1 T −1 2
+ 2
n t∑
( e
u 2,t+1 − ub 2,t+1 )
n ∑ 2,t+1
ue − 2
n ∑ 2,t+1 n ∑ (ue2,t+1 − ub22,t+1)
e
u 2
= k0 t= k0 t= k0 t= k0
!2
1 T −1 2 2
2 1 T −1 2
−
n t∑
ue 2,t+1 − b
u 2,t+1 −
n ∑ (ue2,t+1 − ub22,t+1 )
= k0 t= k0
!
1 T −1 2 1 T −1 2
∑ (ue2,t+1 − ub22,t+1 ) (ue2,t+1 − ub22,t+1 )
n t∑
− 2
n t=k = k0
0
!
2 T −1 2 1 T −1
1 T −1 2
∑ ue2,t+1 (ue22,t+1 − ub22,t+1 ) − 2 ∑ ue22,t+1 (ue2,t+1 − ub22,t+1 )
n t∑
=
n t=k n t=k = k0
0 0
!2
1 T −1 2 2 1 T −1
ue2,t+1 − ub22,t+1 − 3 ∑ (ue22,t+1 − ub22,t+1)
n t∑
− (A.105)
=k
n t=k
0 0
which is consistent if the remainder is negligible. For this purpose, observe that the difference of squares
u2,t+1 (ue2,t+1 − ub2,t+1 ) − (ue2,t+1 − ub2,t+1 )2 so that the fourth term in
can be rewritten as ue22,t+1 − ub22,t+1 = 2e
(A.105) becomes
T −1 T −1 T −1
1 1 1
n ∑ (ue22,t+1 − ub22,t+1) = 2
n ∑ ue2,t+1(ue2,t+1 − ub2,t+1) +
n ∑ (ue2,t+1 − ub2,t+1)2
t= k0 t= k0 t= k0
− 21
= o p (n ) (A.106)
by part (a) and (b) of this lemma. Next, we can show that ue22,t+1 is a consistent estimator of u2t+1 uniformly
in t:
45
1 τ 2 τ 2
≤ 2 sup u2t+1 + 2T −1 sup T 2 + 2 (δet − δ) sup T − 2 zt
k0 ≤ t ≤ T −1 k0 ≤ t ≤ T −1 k0 ≤ t ≤ T −1
−1
= 2 sup u2t+1 + O p ( T ) = O p ( 1) , (A.107)
k0 ≤ t ≤ T −1
τ 2
where T − 2 zt = O p (1) follows from the results of Lemma 3.1 of Magdalinos and Phillips (2009). Now,
this expression implies that the first and second term in (A.105) are
T −1 T −1
1 1
∑ ue22,t+1 (ue22,t+1 − ub22,t+1 ) ≤ 2 sup ue22,t+1 ∑ (ue22,t+1 − ub22,t+1)
n t= k0 k0 ≤ t ≤ T −1 n t= k0
− 21
= o p (n ) (A.108)
and
!
T −1 T −1 T −1
1 1 1
∑ ue22,t+1 ∑ (ue22,t+1 − ub22,t+1) ≤ sup ue22,t+1 ∑ (ue22,t+1 − ub22,t+1)
n t= k0
n t= k0 k0 ≤ t ≤ T −1 n t= k0
− 21
= o p (n ). (A.109)
Hence, it remains to look at the third term. From previous results, we see that
T −1
1 2 1 T −1 2
∑ ue22,t+1 − ub22,t+1 = ∑ u2,t+1 (ue2,t+1 − ub2,t+1 ) − (ue2,t+1 − ub2,t+1 )2
2e
n t= k0
n t=k
0
T −1 T −1
4 4
= ∑ ue22,t+1 (ue22,t+1 − ub22,t+1 ) − ∑ ue2,t+1 (ue2,t+1 − ub2,t+1 )3
n t=k n t= k0
0
T −1
1
(ue2,t+1 − ub2,t+1 )4
n t∑
+ (A.110)
=k 0
Here,
T −1 T −1
1 1
∑ ue22,t+1 (ue22,t+1 − ub22,t+1 ) ≤ sup ue22,t+1 ∑ (ue22,t+1 − ub22,t+1)
n t= k0 k0 ≤ t ≤ T −1 n t= k0
− 21
= o p (n ), (A.111)
while using the result in (A.71) gives
1 T −1
ue2,t+1 (ue2,t+1 − ub2,t+1 )3
n t∑
=k 0
T −1
1 0
=
n ∑ ue2,t+1(ue2,t+1 − ub2,t+1)2 ((Q−N1δbt − δet )′bz0t + (eαt − α)′ e0r,t + α′ e0r,t )
t= k0
T −1
≤ D− 1
D T (Q− 1b e0 T 1 ∑ (ue2,t+1 − ub2,t+1)2 z0t
T sup |ue2,t+1 | sup N δ t − δt ) n T b
k0 ≤ t ≤ T −1 k0 ≤ t ≤ T −1 t= k0
T −1
1 τ 1 τ T1
+T − 2 − 2 sup |ue2,t+1 | sup T 2 + 2 (e
αt − α ) ∑ (ue2,t+1 − ub2,t+1)2 e0r,t
k0 ≤ t ≤ T −1 k0 ≤ t ≤ T −1 nT t= k0
T −1
1 τ T1
+ T − 4 − 2 α0 sup |ue2,t+1 | ∑ (ue2,t+1 − ub2,t+1)2 e0r,t
k0 ≤ t ≤ T −1 nT t= k0
! 21 ! 12
T −1 T −1
− 41 T τ 0 1 1 2
≤ MT sup |ue2,t+1 | sup T D T (Q−
2 1b
N δt − δet ) ∑ (ue2,t+1 − ub2,t+1)4 ∑ z0t
b
n k0 ≤ t ≤ T −1 k0 ≤ t ≤ T −1 T t= k0
T 1+ τ t= k0
46
! 21 ! 12
T −1 T −1
− 12 − τ2 T 1 τ 1 1 2
2+2
+T
n
sup |ue2,t+1 | sup T (eαt − α )
T ∑ (ue2,t+1 − ub2,t+1)4 T ∑ e0r,t
k0 ≤ t ≤ T −1 k0 ≤ t ≤ T −1 t= k0 t= k0
! 21 ! 12
T −1 T −1
1 τ T 0 1 1 2
+T − 4 − 2 α sup |ue2,t+1 | ∑ (ue2,t+1 − ub2,t+1)4 ∑ e0r,t
n k0 ≤ t ≤ T −1 T t= k0
T t= k0
− 21 − 14 τ 1
2−2 − 12 − 12 − τ2 − 12 − 14 − τ2
= op(N T ) + o p (T ) + op(N T ) + op(N T )
− 12 − 14 τ 1
2−2
= op(N T ) + o p (T ). (A.112)
Substituting the expression (A.71) into the quartic term of (A.110) then gives
T −1
1
n ∑ (ue2,t+1 − ub2,t+1)4
t= k0
T −1
1 0
=
n ∑ ((Q−N1 δbt − δet )′ bz0t + (eαt − α)′ e0r,t + α′ e0r,t )4
t= k0
T −1
4 0 ′ 0 4 4 T −1 ′ 0 4 4 T −1 ′ 0 4
∑ (Q− 1b e
n t∑ n t∑
≤ N tδ − δ t ) b
z t + (eα t − α ) e r,t + α er,t
n t= k0 = k0 = k0
T −1
4 0 4 4
≤ ∑ D T (Q− 1b e
N δ t − δt ) D− 1 0
T bzt
n t= k0
T −1
4 1 τ 4 4
+ T −2−2τ ∑ T 2 + 2 (e
α t − α) e0r,t
n t= k0
T −1
4 1 τ 4 4
+ T −1−2τ ∑ T4+2 α e0r,t
n t= k0
T −1
T 4 τ
e0 4 1 4
≤ 4 D−
T
1
sup T 2 D T (Q− 1b
N δ t − δt ) ∑ z0t
b
n k0 ≤ t ≤ T −1 T 1+2τ t= k0
T −1
T 1 τ 4 1 4
+4T −2−2τ sup T 2 + 2 (e
αt − α ) ∑ e0r,t
n k0 ≤ t ≤ T −1 T t= k0
T −1
T 0 4 1 4
+ T −1−2τ α ∑ e0r,t
n T t= k0
T −1
T τ 0 4 1 4
≤ MT −1 sup T 2 D T (Q− 1b e
N δt − δ t ) ∑ z0t
b + O p ( N −2 T −1−2τ )
n k0 ≤ t ≤ T −1 T 2+2τ t= k0
−2 2τ −1
= Op(N ) + O p (T )
= o p ( 1) , (A.114)
1
which is negligible if τ ∈ 0, 2 . Hence,
T −1 2
1
∑ ue22,t+1 − ub22,t+1 = o p ( 1) . (A.115)
n t= k0
47
Therefore, combining all of the results in (A.106)-(A.114) leads to
e2 − φ
φ b2 = o p (1), (A.116)
as desired.
The statistics s f ,j for j = 2, 3, 4 can be power-enhanced, as argued in Pitarakis (2023b). The power adjust-
ment is derived by replacing ue22,t+1 with ŭ2,t+1 = ue22,t+1 − (ue1,t+1 − ue2,t+1 )2 . Clearly, we only have ub22,t+1 ,
therefore, we work with the feasible power-adjustment term. For s fb,2 , for instance, this means that we
need to add
k0 + l20 −1
b 2 = 1 1 √1
△ ∑ (ue1,t+1 − ub2,t+1 )2
b 2 λ02 n
ω t= k0
sPower b 2.
= s fb,2 + △
fb,2
k0 + l20 −1 k0 + l20 −1
1 2 1
√ ∑ (ue1,t+1 − ub2,t+1 ) = √ ∑ (ue1,t+1 − ue2,t+1 + ue2,t+1 − ub2,t+1 )2
n t= k0 n t= k0
k0 + l20 −1 k0 + l20 −1
1 2 2
=√
n
∑ (ue1,t+1 − ue2,t+1 ) + √
n
∑ (ue1,t+1 − ue2,t+1 )(ue2,t+1 − ub2,t+1 )
t= k0 t= k0
k0 + l20 −1
1
+√ ∑ (ue1,t+1 − ub2,t+1 )2
n t= k0
k0 + l20 −1
1
=√ ∑ (ue1,t+1 − ue2,t+1 )2 + o p (1), (A.117)
n t= k0
because the last term is negligible based on Lemma 2. Moreover, by Cauchy-Schwarz inequality,
k0 + l20 −1 k + l 0 −1
1 1 0 2
√
n
∑ (ue1,t+1 − ue2,t+1 )(ue1,t+1 − ub2,t+1 ) ≤ √ ∑ |ue1,t+1 − ue2,t+1||ue2,t+1 − ub2,t+1 |
n t= k0
t= k0
1/2 1/2
k0 + l20 −1 k0 + l20 −1
1 1
≤ √ ∑ (ue1,t+1 − ue2,t+1)2 √n ∑ (ue2,t+1 − ub2,t+1)2
n t= k0 t= k0
= o p ( 1) , (A.118)
due to Lemma 2 again, and the first component coming from the infeasible adjustment term, which is
clearly bounded. The power adjustment terms for s fb,3 and s fb,4 are analogous.
48
7.5.2 Breaks in loadings
We provide a discussion on how we can accommodate structural breaks in factor loadings. Let the load-
ings break at some time point D = ⌊φT ⌋ with φ ∈ (0, 1), such that
′
Λ′i,t ft = I (t < D )Λ1,i ′
ft + I (t ≥ D )Λ2,i ft . (A.119)
where rk(Q) = 2r. This means that our estimated object in the asymptotic analysis is
′ ′ ′ ′ ′ ′ ′
g gt = D N M Λt ft + D N M et = D N M Q gt + D N M et = g0t + e0t ,
b0t = D N M b (A.124)
′ −1 √ ′ − 1′ ′
where g0t = [g′t , 0′(m−2r )×1 ]′ ∈ R m , e0t = [e2r,t Q2r , N (e−2r,t − Q−2r Q2r e2r,t )′ ]′ = [e02r,t , e0−′ 2r,t ]′ ∈ R m ,
√
D N = diag( NI2r , Im−2r ) and
" #
−1 −1
M=
Q2r −Q2r Q−2r = [M , M ] ∈ R m×m . (A.125)
2r −2r
0(m−2r )×2r Im−2r
Note that every loss differential in the asymptotic expansion of the feasible statistic in Pitarakis (2023a)
and Pitarakis (2023b) is an appropriately normalized version of
k 0 + b0
∑ dbt , where dbt ∈ {ue1,t+1 (ue2,t+1 − ub2,t+1 ), ue2,t+1 (ue2,t+1 − ub2,t+1 ), (ue2,t+1 − ub2,t+1 )2 } (A.126)
t= k0
and b0 is a tuning parameter which splits the out-of-sample part of observations, such that b0 T −1 = O(1).
gt instead of bft in the analysis from the
Clearly, b0 ∈ {m0 , l10 , l20 }. Consider that D ∈ (1, k0 ]. Then we use b
beginning and thus the results of Lemma 1 carry over, which makes Lemma 2 hold, as well. Therefore,
the asymptotic results do not change. Next, let D ∈ (k0 , k0 + b0 ]. This means that
k 0 + b0 D −1 k 0 + b0
∑ dbt = ∑ dbt + ∑ dbt , (A.127)
t= k0 t= k0 t= D
where ∑tD=−k10 dbt is based on the sample portion without the break, and therefore it behaves as typical loss
differential analyzed under Lemma 2 (function of bft ), because DT −1 = φ + O( T −1 ). Importantly, ∑k0 +b0 dbt t= D
gt . As long as rk(Q) = 2r, the rates will
behaves in the same way, because its summands are functions of b
49
stay the same, and the final asymptotic result will remain unchanged.
Every variation of the location of b0 will lead to the same conclusions. In the extreme case, let b0 ∈ {l10 , l20 },
where each is strictly smaller than 1 (equivalent to trimming some last out-of-sample observations). Then
if D ∈ [l j , T − 1] for j = 1, 2, the break never appears in our analysis, which implies that Lemma 1 and
Lemma 2 carry over again.
It is important to comment on the difference from Stauskas and Westerlund (2022), where it must be that
D ∈ (1, k0 ], i.e. only in the in-sample period. Take the situation in (A.127). Because of a very different
structure of the statistics that appear in Clark and McCracken (2001), both components are O p (1) instead
of o p (1). In that case, the presence of the break affects the variance of an already highly non-standard
asymptotic distribution. Because the precise location of the break is unknown, simulating the distribution
becomes a notorious challenge. This is not an issue in the current case, because the loss differentials in
the expansions of the feasible statistics will be negligible irrespective of the break location.
1
DGPs
0.9
(1)
(2)
0.8
(3)
(4)
0.7
(5)
(6)
Local power
0.6
(7)
0.5 (8)
(9)
0.4
0.3
0.2
0.1
0
0 0.1 0.2 0.3 0.4 0.5 0.6
Figure 4: Local power evolution of s fb,2 for different DGPs and values of α. Setting: ( N, T ) = (200,600), h =
1, ν0 = 0.8 λ01 = 1, λ02 = 0.65.
50
1
DGPs
0.9 (1)
(2)
0.8 (3)
(4)
0.7 (5)
(6)
Local power
0.6 (7)
(8)
0.5 (9)
0.4
0.3
0.2
0.1
0
0 0.1 0.2 0.3 0.4 0.5 0.6
Figure 6: Local power evolution of s fb,4 for different DGPs and values of α. Setting: ( N, T ) = (200,600), h =
1, ν0 = 0.8 λ01 = 1, λ02 = 0.65.
1
DGPs
0.9
(1)
(2)
0.8
(3)
(4)
0.7
(5)
(6)
Local power
0.6
(7)
0.5 (8)
(9)
0.4
0.3
0.2
0.1
0
0 0.1 0.2 0.3 0.4 0.5 0.6
Figure 5: Local power evolution of s fb,3 different DGPs and values of α. Setting: ( N, T ) = (200,600), h = 1,
ν0 = 0.8 λ01 = 1, λ02 = 0.65.
51
1
0.9
0.8
0.6
0.5
0.4
Persistence
0.3 =0
=0.1
0.2 =0.2
=0.3
0.1 =0.4
=0.49
0
0 0.1 0.2 0.3 0.4 0.5 0.6
Figure 7: Local power evolution of s fb,2 for different values of τ and α. Setting: DGP (2), ( N, T ) = (200,600),
h = 1, ν0 = 0.8 λ01 = 1, λ02 = 0.65.
0.9
0.8
0.7
Local power
0.6
0.5
0.4
Persistence
0.3 =0
=0.1
0.2 =0.2
=0.3
0.1 =0.4
=0.49
0
0 0.1 0.2 0.3 0.4 0.5 0.6
Figure 8: Local power evolution of s fb,3 for different values of τ and α. Setting: DGP (2), ( N, T ) = (200,600),
h = 1, ν0 = 0.8 λ01 = 1, λ02 = 0.65.
52
1
0.9
0.8
0.6
0.5
0.4
Persistence
0.3 =0
=0.1
0.2 =0.2
=0.3
0.1 =0.4
=0.49
0
0 0.1 0.2 0.3 0.4 0.5 0.6
Figure 9: Local power evolution of s fb,4 for different values of τ and α. Setting: DGP (2), ( N, T ) = (200,600),
h = 1, ν0 = 0.8 λ01 = 1, λ02 = 0.65.
In Figures 1, 2 and 3, we report the results on test size and local power across the different DGPs for the
remaining three statistics. Similarly to the MC experiments presented in the main paper, the other three
statistics exhibit excellent power that actually seems to improve upon the results of the first statistics. In-
deed, looking at three figures, these statistics share a very similar performance overall but seem to suffer
moderately in terms of both size and local power under ARCH effects and regression heavy tail settings.
It is however notable that the fourth statistic has the best local power out of all proposed statistics.
We then turn our attention to the effects of higher degrees of persistency on test size and local power.
The results for the remaining three statistics are presented in Figures 4, 5 and 6. It is clear from these
graphical representations that regressor persistency has no implications for test size, but it causes a mild
deterioration in local power for greater values of α which, as discussed for the first statistics, is anticipated
by our theory. As in the above, it seems that the fourth statistics outperforms its peers as well as suffer
from less decay in local power due to nonstationarity.
Finally, we investigate how our tests perform in the case of persistent data and serial correlation in
the idiosyncratic components of the factor model under DGP (3) and different values of τ. Recall in fact
that, under DGP (3), all panel idiosyncratic components ei,t exhibit autocorrelation of order 1 in the range
(0.6, 1), meaning that they have strong persistence in time. Further, as outlined in the proof of Lemma 1
in Section 7.3, the presence of serial correlation in the idiosyncratics reduces the rate of T 11+τ ∑tT=1 zt e′t from
O p ( N −1/2 T −(1+τ )/2 ) in Equation (A.5) to O p ( N −1/2 T −τ/2 ) in Equation (A.1). If τ ∈ (0, 0.5) in addition,
we have to rule out serial correlation in order to have asymptotically valid tests. However, we recog-
nise that it is unrealistic to do so in empirical applications, particularly in the context of macroeconomic
data. Hence, a Monte Carlo experiment under DGP (3) and different values of τ will serve as a robust
benchmark should any potential concerns arise. The results on the local power of the four statistics are
53
1
0.9
0.8
0.6
0.5
0.4
Persistence
0.3 =0
=0.1
0.2 =0.2
=0.3
0.1 =0.4
=0.49
0
0 0.1 0.2 0.3 0.4 0.5 0.6
Figure 10: Local power evolution of s fb,1 for different values of τ and α. Setting: DGP (3), ( N, T ) = (200,600),
h = 1, µ0 = 0.45.
presented in Figures 7 to 11. All things considered, it is very clear that, while the local power worsens
with the persistence as before, the results are still very satisfactory and seem to suggest that our rates in
the case with serial correlation and persistence might possibly be conservative, if anything.
All in all, we conclude based on our MC experiments that the fourth statistics provides the most sat-
isfactory results relative to its peers, but all statistics exhibit excellent power under the local alternative
across all experiments, thus making them realistically applicable and relevant in most empirical settings.
54
Stationary Persistent
Statistic EA Germany France Italy Spain Statistic EA Germany France Italy Spain
GDP GDP
ER 1 2 1 1 1 ER 1 2 1 1 1
# breaks 2 1 2 1 0 # breaks 2 1 2 1 0
break dates 2020:2 2001:3 2009:3 2021:4 - break dates 2020:2 2020:2 2001:3 2009:3 2021:4
2021:4 2010:4 2021:4 2021:4 2010:4
SC -0.002 0.139 0.103 0.095 0.117 SC 0.558 0.580 0.558 0.534 0.511
Abs SC 0.184 0.288 0.309 0.262 0.260 Abs SC 0.559 0.580 0.558 0.535 0.511
XC 0.002 0.002 0.002 0.000 0.000 XC 0.005 0 0.002 0.001 0.007
Abs XC 0.007 0.006 0.002 0.000 0.000 Abs XC 0.036 0 0.014 0.001 0.024
IPMN IPMN
ER 1 2 1 1 1 ER 1 2 1 1 1
# breaks 2 1 2 1 0 # breaks 2 1 2 1 0
break dates 2020:2 2001:3 2009:3 2021:4 - break dates 2020:2 2001:3 2009:3 2021:4 -
2021:4 2010:4 2021:4 2010:4
SC -0.003 0.141 0.102 0.097 0.117 SC 0.555 0.589 0.564 0.532 0.513
Abs SC 0.183 0.287 0.305 0.265 0.261 Abs SC 0.556 0.589 0.564 0.532 0.513
XC 0.002 0.003 0.002 0.000 0.000 XC 0.005 0.000 0.002 0.006 0.007
Abs XC 0.007 0.006 0.002 0.000 0.000 Abs XC 0.040 0.000 0.019 0.031 0.022
HICPNEF HICPNEF
ER 1 2 1 1 1 ER 1 2 1 1 1
# breaks 2 1 2 1 0 # breaks 2 1 2 1 0
break dates 2020:2 2001:3 2009:3 2021:4 - break dates 2020:2 2001:3 2009:3 2021:4 -
2021:4 2010:4 2021:4 2010:4
SC 0.001 0.142 0.111 0.097 0.122 SC 0.555 0.585 0.560 0.534 0.515
Abs SC 0.181 0.286 0.306 0.265 0.256 Abs SC 0.555 0.585 0.560 0.535 0.515
XC 0.002 0.003 0.002 0.000 0.000 XC 0.004 0 0.002 0.005 0.005
Abs XC 0.007 0.006 0.002 0.000 0.000 Abs XC 0.033 0 0.013 0.030 0.021
UNETOT UNETOT
ER 1 2 1 1 1 ER 1 2 1 1 1
# breaks 2 1 2 1 0 # breaks 2 1 2 1 0
break dates 2020:2 2001:3 2009:3 2021:4 - break dates 2020:2 2001:3 2009:3 2021:4 -
2021:4 2010:4 2021:4 2010:4
SC -0.005 0.133 0.107 0.100 0.115 SC 0.555 0.587 0.564 0.535 0.518
Abs SC 0.181 0.286 0.313 0.262 0.258 Abs SC 0.555 0.587 0.564 0.536 0.518
XC 0.002 0.003 0.002 0.000 0.000 XC 0.005 0.000 0.002 0.001 0.006
Abs XC 0.006 0.006 0.002 0.000 0.000 Abs XC 0.035 0.001 0.013 0.002 0.023
IRT3M IRT3M
ER 1 2 1 1 1 ER 1 2 1 1 1
# breaks 2 1 2 1 0 # breaks 2 1 2 1 0
break dates 2020:2 2001:3 2009:3 2021:4 - break dates 2020:2 2001:3 2009:3 2021:4 -
2021:4 2010:4 2021:4 2010:4
SC -0.004 - - - - SC 0.551 - - - -
Abs SC 0.187 - - - - Abs SC 0.552 - - - -
XC 0.002 - - - - XC 0.004 - - - -
Abs XC 0.007 - - - - Abs XC 0.032 - - - -
IRT6M IRT6M
ER 1 2 1 1 1 ER 1 2 1 1 1
# breaks 2 1 2 1 0 # breaks 2 1 2 1 0
break dates 2020:2 2001:3 2009:3 2021:4 - break dates 2020:2 2001:3 2009:3 2021:4 -
2021:4 2010:4 2021:4 2010:4
SC -0.006 - - - - SC 0.552 - - - -
Abs SC 0.185 - - - - Abs SC 0.552 - - - -
XC 0.002 - - - - XC 0.003 - - - -
Abs XC 0.006 - - - - Abs XC 0.030 - - - -
LTIRT LTIRT
ER 1 2 1 1 1 ER 1 2 1 1 1
# breaks 2 1 2 1 0 # breaks 2 1 2 1 0
break dates 2020:2 2001:3 2009:3 2021:4 - break dates 2020:2 2001:3 2009:3 2021:4 -
2021:4 2010:4 2021:4 2010:4
SC 0.002 0.134 0.107 0.096 0.112 SC 0.558 0.621 0.593 0.559 0.539
Abs SC 0.179 0.288 0.307 0.258 0.259 Abs SC 0.559 0.621 0.593 0.559 0.539
XC 0.002 0.002 0.002 0.001 0.000 XC 0.006 0 0.004 0.001 0.000
Abs XC 0.007 0.024 0.002 0.001 0.000 Abs XC 0.038 0 0.024 0.001 0.000
Table 9: ”ER” refers to the estimated number of factors based on the eigenvalue ratio of Ahn and Horenstein (2013). ”# breaks”
and ”break dates” are obtained using the sup-LM test statistics of Chen et al. (2014). ”SC” and ”Abs SC” refer to the average
and average absolute serial correlations of order 1 in the idiosyncratic panel data residuals that are significant at a 5% level
under the Ljung-Box tests. Similarly, ”XC” and ”Abs XC” refer to the average and average absolute pair-wise correlations in the
idiosyncratic panel data residuals using the thresholding procedure of Cai and Liu (2011).
55
Stationary Persistent
Statistic EA Germany France Italy Spain Statistic EA Germany France Italy Spain
WS WS
ER 1 2 1 1 1 ER 1 2 1 1 1
# breaks 2 1 2 1 0 # breaks 2 1 2 1 0
break dates 2020:2 2001:3 2009:3 2021:4 - break dates 2020:2 2001:3 2009:3 2021:4 -
2021:4 2010:4 2021:4 2010:4
SC -0.002 0.134 0.101 0.099 0.118 SC 0.557 0.579 0.566 0.533 0.519
Abs SC 0.183 0.294 0.306 0.266 0.262 Abs SC 0.557 0.579 0.566 0.534 0.519
XC 0.002 0.002 0.002 0.000 0.000 XC 0.005 0.000 0.001 0.007 0.008
Abs XC 0.007 0.006 0.002 0.000 0.000 Abs XC 0.036 0.000 0.015 0.030 0.029
GGLB GGLB
ER 1 2 1 1 1 ER 1 2 1 1 1
# breaks 2 1 2 1 0 # breaks 2 1 2 1 0
break dates 2020:2 2001:3 2009:3 2021:4 - break dates 2020:2 2001:3 2009:3 2021:4 -
2021:4 2010:4 2021:4 2010:4
SC -0.002 0.137 0.109 0.094 0.118 SC 0.555 0.588 0.564 0.504 0.516
Abs SC 0.183 0.291 0.309 0.259 0.262 Abs SC 0.555 0.588 0.564 0.505 0.516
XC 0.002 0.003 0.002 0.000 0.000 XC 0.004 0.000 0.002 0 0.004
Abs XC 0.007 0.006 0.002 0.000 0.000 Abs XC 0.028 0.001 0.014 0 0.017
NFCLB NFCLB
ER 1 2 1 1 1 ER 1 2 1 1 1
# breaks 2 1 2 1 0 # breaks 2 1 2 1 0
break dates 2020:2 2001:3 2009:3 2021:4 - break dates 2020:2 2001:3 2009:3 2021:4 -
2021:4 2010:4 2021:4 2010:4
SC -0.002 0.140 0.108 0.098 0.119 SC 0.551 0.583 0.569 0.532 0.516
Abs SC 0.183 0.289 0.313 0.266 0.262 Abs SC 0.552 0.583 0.569 0.533 0.516
XC 0.002 0.003 0.002 0.000 0.000 XC 0.003 0 0.002 0.001 0.006
Abs XC 0.006 0.006 0.002 0.000 0.000 Abs XC 0.038 0 0.015 0.001 0.022
HHLB HHLB
ER 1 2 1 1 1 ER 1 2 1 1 1
# breaks 2 1 2 1 0 # breaks 2 1 2 1 0
break dates 2020:2 2001:3 2009:3 2021:4 - break dates 2020:2 2001:3 2009:3 2021:4 -
2021:4 2010:4 2021:4 2010:4
SC -0.002 0.138 0.107 0.099 0.122 SC 0.555 0.583 0.572 0.538 0.519
Abs SC 0.184 0.291 0.311 0.266 0.259 Abs SC 0.555 0.583 0.572 0.539 0.519
XC 0.002 0.003 0.001 0.000 0.000 XC 0.005 0.000 0.003 0.001 0.006
Abs XC 0.007 0.006 0.001 0.000 0.000 Abs XC 0.038 0.001 0.029 0.002 0.026
REER42 REER42
ER 1 2 1 1 1 ER 1 2 1 1 1
# breaks 2 1 2 1 0 # breaks 2 1 2 1 0
break dates 2020:2 2001:3 2009:3 2021:4 - break dates 2020:2 2001:3 2009:3 2021:4 -
2021:4 2010:4 2021:4 2010:4
SC -0.004 0.134 0.104 0.093 0.111 SC 0.556 0.601 0.622 0.529 0.524
Abs SC 0.178 0.283 0.309 0.267 0.259 Abs SC 0.556 0.601 0.622 0.529 0.524
XC 0.002 0.002 0.002 0.001 0.000 XC 0.005 0 0.001 0.002 0.001
Abs XC 0.008 0.020 0.002 0.001 0.000 Abs XC 0.038 0 0.006 0.024 0.001
HPRC HPRC
ER 1 2 1 1 1 ER 1 2 1 1 1
# breaks 2 1 2 1 0 # breaks 2 1 2 1 0
break dates 2020:2 2001:3 2009:3 2021:4 - break dates 2020:2 2001:3 2009:3 2021:4 -
2021:4 2010:4 2021:4 2010:4
SC 0.002 0.139 0.109 0.101 0.122 SC 0.549 0.558 0.568 0.523 0.512
Abs SC 0.180 0.289 0.308 0.263 0.257 Abs SC 0.550 0.558 0.568 0.523 0.512
XC 0.002 0.003 0.002 0.000 0.000 XC 0.001 0 0.003 0.002 0.005
Abs XC 0.007 0.007 0.002 0.000 0.000 Abs XC 0.004 0 0.021 0.002 0.019
SHIX SHIX
ER 1 2 1 1 1 ER 1 2 1 1 1
# breaks 2 1 2 1 0 # breaks 2 1 2 1 0
break dates 2020:2 2001:3 2009:3 2021:4 - break dates 2020:2 2001:3 2009:3 2021:4 -
2021:4 2010:4 2021:4 2010:4
SC -0.004 0.134 0.107 0.094 0.109 SC 0.538 0.609 0.602 0.559 0.522
Abs SC 0.178 0.290 0.309 0.267 0.260 Abs SC 0.539 0.609 0.602 0.559 0.522
XC 0.002 0.002 0.002 0.000 0.000 XC 0.006 0.000 0.004 0.005 0.008
Abs XC 0.008 0.006 0.003 0.000 0.000 Abs XC 0.044 0.000 0.027 0.029 0.032
Table 10: ”ER” refers to the estimated number of factors based on the eigenvalue ratio of Ahn and Horenstein (2013). ”#
breaks” and ”break dates” are obtained using the sup-LM test statistics of Chen et al. (2014). ”SC” and ”Abs SC” refer to the
average and average absolute serial correlations of order 1 in the idiosyncratic panel data residuals that are significant at a 5%
level under the Ljung-Box tests. Similarly, ”XC” and ”Abs XC” refer to the average and average absolute pair-wise correlations
in the idiosyncratic panel data residuals using the thresholding procedure of Cai and Liu (2011).
56
1
0.9
0.8
0.7
Local power
0.6
0.5
0.4
Persistence
0.3 =0
=0.1
0.2 =0.2
=0.3
0.1 =0.4
=0.49
0
0 0.1 0.2 0.3 0.4 0.5 0.6
Figure 11: Local power evolution of s fb,2 for different values of τ and α. Setting: DGP (3), ( N, T ) = (200,600),
h = 1, ν0 = 0.8 λ01 = 1, λ02 = 0.65.
0.9
0.8
0.7
Local power
0.6
0.5
0.4
Persistence
0.3 =0
=0.1
0.2 =0.2
=0.3
0.1 =0.4
=0.49
0
0 0.1 0.2 0.3 0.4 0.5 0.6
Figure 12: Local power evolution of s fb,3 for different values of τ and α. Setting: DGP (3), ( N, T ) = (200,600),
h = 1, ν0 = 0.8 λ01 = 1, λ02 = 0.65.
57
1
0.9
0.8
0.6
0.5
0.4
Persistence
0.3 =0
=0.1
0.2 =0.2
=0.3
0.1 =0.4
=0.49
0
0 0.1 0.2 0.3 0.4 0.5 0.6
Figure 13: Local power evolution of s fb,4 for different values of τ and α. Setting: DGP (3), ( N, T ) = (200,600),
h = 1, ν0 = 0.8 λ01 = 1, λ02 = 0.65.
References
Ahn, S. C. and Horenstein, A. R. (2013). Eigenvalue ratio test for the number of factors. Econometrica,
81(3):1203–1227.
Ando, T. and Bai, J. (2016). Panel data models with grouped factor structure under unknown group
membership. Journal of Applied Econometrics, 31(1):163–191.
Bai, J. and Ng, S. (2002). Determining the number of factors in approximate factor models. Econometrica,
70(1):191–221.
Bai, J. and Ng, S. (2006). Confidence intervals for diffusion index forecasts and inference for factor-
augmented regressions. Econometrica, 74(4):1133–1150.
Bai, J. and Ng, S. (2007). Determining the number of primitive shocks in factor models. Journal of Business
& Economic Statistics, 25(1):52–60.
Banerjee, A., Marcellino, M., and Masten, I. (2008). Forecasting macroeconomic variables using diffusion
indexes in short samples with structural change. In Wohar, M. E. and Rapach, D. E., editors, Forecasting
in the Presence of Structural Breaks and Model Uncertainty, page 149–194. Elsevier.
Barigozzi, M., Lissona, C., and Luciani, M. (2024a). Measuring the euro area output gap. FEDS Working
paper.
Barigozzi, M., Lissona, C., and Tonni, L. (2024b). Large datasets for the euro area and its member countries
and the dynamic effects of the common monetary policy. arXiv preprint arXiv:2410.05082.v1.
58
Boivin, J. and Ng, S. (2006). Are more data always better for factor analysis? Journal of Econometrics,
132(1):169–194.
Breitung, J. and Demetrescu, M. (2015). Instrumental variable and variable addition based inference in
predictive regressions. Journal of Econometrics, 187(1):358–375.
Breitung, J. and Eickmeier, S. (2011). Testing for structural breaks in dynamic factor models. Journal of
Econometrics, 163(1):71–84.
Cai, T. and Liu, W. (2011). Adaptive thresholding for sparse covariance matrix estimation. Journal of the
American Statistical Association, 106(494):672–684.
Campbell, J. Y. and Yogo, M. (2006). Efficient tests of stock return predictability. Journal of financial eco-
nomics, 81(1):27–60.
Carriero, A., Clark, T. E., Marcellino, M., and Elmar, M. (2022). Addressing covid-19 outliers in bvars with
stochastic volatility. The Review of Economics and Statistics, 106(5):1403–1417.
Chamberlain, G. (1983). Funds, factors, and diversification in arbitrage pricing models. Econometrica:
Journal of the Econometric Society, pages 1305–1323.
Chen, L., Dolado, J. J., and Gonzalo, J. (2014). Detecting big structural breaks in large factor models.
Journal of Econometrics, 180(1):30–48.
Ciccarelli, M. and Mojon, B. (2010). Global inflation. The Review of Economics and Statistics, 92(3):524–535.
Clark, T. E. and McCracken, M. W. (2001). Tests of equal forecast accuracy and encompassing for nested
models. Journal of econometrics, 105(1):85–110.
Clark, T. E. and McCracken, M. W. (2012). Reality checks and comparisons of nested predictive models.
Journal of Business & Economic Statistics, 30(1):53–66.
Corradi, V. and Swanson, N. R. (2014). Testing for structural stability of factor augmented forecasting
models. Journal of Econometrics, 182(1):100–118.
De Vos, I., Everaert, G., and Sarafidis, V. (2024). A method to evaluate the rank condition for cce estima-
tors. Econometric Reviews, 43(2-4):123–155.
De Vos, I. and Stauskas, O. (2024). Cross-section bootstrap for cce regressions. Journal of Econometrics,
240(1):105648.
Diebold, F. X. (2015). Comparing predictive accuracy, twenty years later: A personal perspective on the
use and abuse of diebold–mariano tests. Journal of Business & Economic Statistics, 33(1):1–1.
Eickmeier, S. and Ziegler, C. (2008). How successful are dynamic factor models at forecasting output and
inflation? a meta-analytic approach. Journal of Forecasting, 27(3):237–265.
Forni, M., Hallin, M., Lippi, M., and Reichlin, L. (2003). Do financial variables help forecasting inflation
and real activity in the euro area? Journal of Monetary Economics, 50(6):1243–1255.
Gonçalves, S., McCracken, M. W., and Perron, B. (2017). Tests of equal accuracy for nested models with
estimated factors. Journal of Econometrics, 198(2):231–252.
Hallin, M. and Liška, R. (2011). Dynamic factors in the presence of blocks. Journal of Econometrics,
163(1):29–41.
59
Hansen, P. R. and Timmermann, A. (2015). Equivalence between out-of-sample forecast comparisons and
wald statistics. Econometrica, 83(6):2485–2505.
Hjalmarsson, E. (2010). Predicting global stock returns. Journal of Financial and Quantitative Analysis,
45(1):49–80.
Karabiyik, H., Reese, S., and Westerlund, J. (2017). On the role of the rank condition in cce estimation of
factor-augmented panel regressions. Journal of Econometrics, 197(1):60–64.
Karabiyik, H. and Westerlund, J. (2021). Forecasting using cross-section average–augmented time series
regressions. The Econometrics Journal, 24(2):315–333.
Lenza, M. and Primiceri, G. E. (2022). How to estimate a vector autoregression after march 2020. Journal
of Applied Econometrics, 37(4):688–699.
Ludvigson, S. C. and Ng, S. (2011). A factor analysis of bond risk premia. In Gilles, D. and Ullah, A.,
editors, Handbook of Empirical Economics and Finance, pages 313–372. Chapman and Hall.
Magdalinos, T. and Phillips, P. C. (2009). Limit theory for cointegrated systems with moderately integrated
and moderately explosive regressors. Econometric Theory, 25(2):482–526.
Margaritella, L. and Stauskas, O. (2024). New tests of equal forecast accuracy for factor-augmented re-
gressions with weaker loadings. arXiv preprint arXiv:2409.20415.
Maroz, D., Stock, J., and Watson, M. (2021). Comovement of economic activity during the covid recession.
Massacci, D. and Kapetanios, G. (2024). Forecasting in factor augmented regressions under structural
change. International Journal of Forecasting, 40(1):62–76.
McCracken, M. and Ng, S. (2020). Fred-qd: A quarterly database for macroeconomic research. Technical
report, National Bureau of Economic Research.
McCracken, M. W. (2007). Asymptotics for out of sample tests of granger causality. Journal of econometrics,
140(2):719–752.
McCracken, M. W. and Ng, S. (2016). Fred-md: A monthly database for macroeconomic research. Journal
of Business & Economic Statistics, 34(4):574–589.
Moench, E., Ng, S., and Potter, S. (2013). Dynamic hierarchical factor models. Review of Economics and
Statistics, 95(5):1811–1817.
Moon, H. R. and Weidner, M. (2015). Linear regression for panel with unknown number of factors as
interactive fixed effects. Econometrica, 83(4):1543–1579.
Peña, D. and Poncela, P. (2006). Nonstationary dynamic factor analysis. Journal of Statistical Planning and
Inference, 136(4):1237–1257.
Pesaran, M. H. (2006). Estimation and inference in large heterogeneous panels with a multifactor error
structure. Econometrica, 74(4):967–1012.
Pitarakis, J.-Y. (2023a). Direct multi-step forecast based comparison of nested models via an encompassing
test. arXiv preprint arXiv:2312.16099.
60
Pitarakis, J.-Y. (2023b). A novel approach to predictive accuracy testing in nested environments. Econo-
metric Theory, pages 1–44.
Stauskas, O. and Westerlund, J. (2022). Tests of equal forecasting accuracy for nested models with esti-
mated cce factors. Journal of Business & Economic Statistics, 40(4):1745–1758.
Stock, J. and Watson, M. (2002). Macroeconomic forecasting using diffusion indexes. Journal of Business &
Economic Statistics, 20(2):147–162.
Stock, J. H. and Watson, M. W. (1999). Forecasting inflation. Journal of monetary economics, 44(2):293–335.
West, K. D. (1996). Asymptotic inference about predictive ability. Econometrica: Journal of the Econometric
Society, pages 1067–1084.
61