Statistical Inference For Score Decomposition
Statistical Inference For Score Decomposition
March 5, 2026
Abstract
arXiv:2603.04275v1 [[Link]] 4 Mar 2026
We introduce inference methods for score decompositions, which partition scoring func-
tions for predictive assessment into three interpretable components: miscalibration, discrim-
ination, and uncertainty. Our estimation and inference relies on a linear recalibration of
the forecasts, which is applicable to general multi-step ahead point forecasts such as means
and quantiles due to its validity for both smooth and non-smooth scoring functions. This
approach ensures desirable finite-sample properties, enables asymptotic inference, and es-
tablishes a direct connection to the classical Mincer-Zarnowitz regression. The resulting
inference framework facilitates tests for equal forecast calibration or discrimination, which
yield three key advantages. They enhance the information content of predictive ability tests
by decomposing scores, deliver higher statistical power in certain scenarios, and formally
connect scoring-function-based evaluation to traditional calibration tests, such as financial
backtests. Applications demonstrate the method’s utility. We find that for survey inflation
forecasts, discrimination abilities can differ significantly even when overall predictive ability
does not. In an application to financial risk models, our tests provide deeper insights into the
calibration and information content of volatility and Value-at-Risk forecasts. By disentan-
gling forecast accuracy from backtest performance, the method exposes critical shortcomings
in current banking regulation.
1 Introduction
Forecasting—whether of macroeconomic conditions, financial market movements, or other un-
certain outcomes—has become an indispensable element of modern economic analysis. From
monetary policy design to investment planning and risk management, economic decisions in-
creasingly rely on the ability to anticipate what lies ahead. The value of these decisions, how-
ever, depends critically on the accuracy and credibility of the underlying forecasts. Therefore,
establishing rigorous standards for evaluating predictions is essential, ensuring that economic
forecasts are not only informative but also empirically grounded and verifiable.
Two main concepts have emerged for evaluating a forecast Xt of an uncertain future real-
valued outcome Yt . First, the statistical theory of forecast comparison ranks competing forecasts
according to their empirical score (or loss) (Savage, 1971; Gneiting, 2011). The appropriate
∗
Faculty of Economics and Business, Goethe University Frankfurt, Campus Westend, Theodor-W.-Adorno-
Platz 4, 60629 Frankfurt am Main, Germany, and Heidelberg Institute for Theoretical Studies (HITS),
dimitriadis@[Link].
†
Institute of Economics, University of Hohenheim, Schloss Hohenheim 1 C, 70593 Stuttgart, Germany,
[Link]@[Link].
1
scoring function depends on the forecast target—e.g., whether the aim is to predict the full
distribution of Yt , its mean, or a particular quantile. For meaningful comparisons, these scoring
functions must be strictly consistent: the ideal forecast should minimize the expected score,
ensuring that lower score values reliably indicate superior predictive performance.
Second, forecasts are evaluated in terms of their alignment with the realized outcomes—a
property referred to as calibration in statistics (Gneiting et al., 2007; Gneiting and Resin, 2023),
forecast optimality, rationality or efficiency in economics (Diebold et al., 1998; Elliott et al.,
2016), and backtesting in finance (Basel Committee, 2019a,b). For example, when forecasting
the mean, the average forecast should coincide with the average realized outcome—both uncon-
ditionally and conditional on the forecast itself, the latter being referred to as auto-calibration.
In economics, this property is most commonly assessed using the test of Mincer and Zarnowitz
(1969, henceforth MZ), which examines whether the forecast Xt equals the conditional expec-
tation of Yt given Xt through a linear regression framework. Extensions to other targets, such
as quantiles, follow naturally by replacing the mean regression with the corresponding quantile
regression (Gaglianone et al., 2011; Guler et al., 2017; Bayer and Dimitriadis, 2022).
These two evaluation paradigms are known to occasionally produce counterintuitive results,
a point illustrated by our following motivational example. In this paper, we develop asymptotic
inference for the components of so-called score decompositions, which provide a transparent
link between the two forecast evaluation regimes described above. Beyond delivering a sharper
theoretical understanding of the respective strengths and limitations of these approaches, the
motivational example underscores their practical relevance by demonstrating how, in applied
settings, the two methods can lead to seemingly contradictory conclusions.
2
Table 1: Mincer and Zarnowitz (1969) test p-values, average scores (QLIKE for variance, and check loss for the
quantile) and its decomposition terms for variance and VaR forecasts. See Sections 1.1 and 5.2 for details.
the same time, the MZ test decisively rejects auto-calibration of the HAR forecasts. In contrast,
the GARCH variance forecasts yield significantly higher scores, yet auto-calibration cannot be
rejected. A similar pattern emerges for the HS VaR forecasts: they record the largest scores but
constitute the only forecasting sequence for which calibration is not rejected.
This example illustrates that good predictive performance and proper auto-calibration do
not necessarily coincide and may diverge sharply. Since both evaluation approaches are widely
applied in practice, developing tools to reconcile and interpret these seemingly contradictory
conclusions is of central importance.
We shed light on these results by examining the score decomposition reported in Table 1.
Score decompositions have a long tradition in meteorology (Murphy, 1973; Dawid, 1986; Murphy
and Winkler, 1987; Bröcker, 2009; Bentzien and Friederichs, 2014) and have recently attracted
renewed interest in statistics (Dimitriadis et al., 2021; Gneiting and Resin, 2023; Gneiting et al.,
2023), yet they have received comparatively little attention in economics. In essence, the aver-
age score Si of forecaster i can be decomposed additively into three non-negative components
capturing miscalibration (MCBi ), discrimination (DSCi ), and a forecast-independent uncertainty
term (UNC):
Intuitively, the MCB component measures the extent of miscalibration relative to the scor-
ing function by quantifying how much the score would improve if the forecasts were perfectly
calibrated—thus capturing a concept closely related to what the MZ test evaluates. Likewise, the
DSC component reflects the forecaster’s ability to distinguish (or “discriminate”) between higher
and lower outcome realizations, conditional on the forecasts being calibrated. The remaining
UNC term does not depend on the forecast itself and therefore serves as a normalization.
The last columns of Table 1 help reconcile the discrepancies between the score-based and
calibration-based assessments. For both forecast types, the HAR model attains the highest
(and thus best) DSC values, which is intuitive given that it incorporates the richest information
set through high-frequency returns. However, it also yields the largest (worst) MCB values,
accounting for its rejection in the MZ test. By contrast, the GARCH model—and even more so
the HS model—exhibits weaker discrimination but achieves substantially better calibration. Yet,
despite the clear interpretive value of these decomposition terms, formal inference methods for
them are not available, limiting their practical usefulness as reflected in Elliott and Timmermann
3
(2016, p. 355), who point out that “am important limitation of these decompositions is that there
are no objective measures for how large or small the terms should be”.
This gap is particularly problematic in settings such as banking regulation under the Basel
framework, where calibration-based backtesting plays a central role. Such procedures risk
creating misguided incentives: the best-performing HAR forecasts may be labeled “invalid,”
whereas HS forecasts—whose poor discriminatory ability makes them slow to respond to emerg-
ing crises—may be deemed acceptable. These tensions highlight the need for both a deeper
understanding of the distinct properties captured by the decomposition and the development of
rigorous statistical inference for its components—needs that this paper directly addresses.
1.2 Contributions
In this paper, we propose augmenting scoring-function-based forecast comparisons with the
corresponding decomposition terms, as illustrated in Table 1, and we are the first to develop
asymptotic inference for these components. To estimate the decomposition components in (1),
we advocate the use of a simple linear regression—either mean or quantile, depending on the
forecast target. This specification is motivated by several considerations: its stability and re-
sistance to overfitting; its good empirical fit in our applications; its close conceptual connection
to the Mincer and Zarnowitz (1969) framework; its attractive non-negativity guarantees for the
resulting decomposition terms; and, crucially, its tractability for developing valid asymptotic
inference procedures. Recent contributions have suggested nonparametric (isotonic) regression
approaches instead (Dimitriadis et al., 2021; Gneiting and Resin, 2023; Arnold et al., 2024).
However, their asymptotic properties remain unclear—particularly in time-series settings—and
they are prone to overfitting, which can introduce biases that further complicate or even under-
mine asymptotic inference.
We derive the joint asymptotic distribution of the estimated MCB and DSC components for
competing forecasts allowing for possibly non-smooth scoring functions, thereby enabling formal
tests of equal miscalibration or equal discrimination for, e.g., mean or quantile forecasts. These
tests extend the classical framework of Diebold and Mariano (1995) by incorporating two layers
of uncertainty: the sampling variation arising from averaging over time and the additional vari-
ation introduced by the linear “recalibration regression” used in constructing the MCB and DSC
components. A key complication is that the limiting distribution depends on whether the popula-
tion values of MCB or DSC are zero or strictly positive. Our testing procedures account for these
cases through a p-value adjustment inspired by a hybrid intersection–union/union–intersection
(IU/UI) testing principles (Berger, 1997).
Our simulations demonstrate that the proposed tests are valid—though potentially conser-
vative due to the underlying IU/UI principles—under their respective null hypotheses for both
mean and quantile forecasting settings. We also show that the tests exhibit strong power prop-
erties, and in particular achieve a substantial power increase relative to the classical DM test in
fair comparisons where differences in average scores stem exclusively from a single component,
either MCB or DSC. This perhaps surprising power gain arises because our tests isolate a single
aspect of forecast performance—either discrimination or miscalibration—whereas score-based
tests must aggregate both components into a single, and consequently noisier, metric.
4
Our first application compares the predictive performance of two survey forecasts for U.S. in-
flation—the Survey of Professional Forecasters (SPF) and the Michigan Survey of Consumers.
While professional forecasters unsurprisingly attain better average scores than consumers, these
differences are typically insignificant (Ehm et al., 2016; Patton, 2020). In contrast, the score
decomposition shows that both groups exhibit similar calibration, whereas our new tests detect
a significant difference in discrimination. This finding is consistent with the power gains demon-
strated in our simulations. Given that professional forecasters draw on a richer information set,
their superior discrimination is a natural consequence.
The second application extends the motivating example of Section 1.1 and clarifies the rela-
tionship between scoring-function-based forecast evaluation and backtests for both, variance and
VaR forecasts. Moreover, we disentangle several widely used VaR backtests by characterizing
them through the conditioning sets they implicitly impose within the framework of conditional
quantile calibration, thereby revealing the precise properties they actually test.
1.3 Literature
Early approaches to score decompositions resembling (1) date back to Sanders (1963), Theil
(1966), and Murphy (1973), and primarily focus on the squared error in the context of probability
forecasts for binary outcomes, which can be interpreted as a mean. DeGroot and Fienberg (1981)
and Bröcker (2009) demonstrated that such decompositions can be derived for any strictly proper
scoring rule, including those for full distributional forecasts. More recently, Gneiting and Resin
(2023) developed a general framework for score decompositions for scoring functions, extending
these concepts to point forecasts. This line of research has spurred a growing body of both
theoretical and applied work; see, for example, Bentzien and Friederichs (2014); Siegert (2017);
Pohle (2020); Fissler et al. (2022); Wüthrich et al. (2025); Brehmer et al. (2024). For a more
detailed historical overview of score decompositions, see Mitchell (2019, Section 1.1).
In practice, the score decompositions rely on estimates of recalibrated forecasts, and the
choice of estimation method remains an active area of research. Recently, isotonic regres-
sion—which estimates, for example, a conditional mean nonparametrically under a monotonicity
constraint—has been widely used for both point and distributional forecasts (Dimitriadis et al.,
2021; Gneiting et al., 2023; Dimitriadis et al., 2024b; Arnold et al., 2024; Allen et al., 2025). A
key reason for its popularity is that it satisfies attractive finite-sample non-negativity conditions,
which facilitate the interpretation of the decomposition terms; these conditions are generally not
satisfied by alternative methods such as binning or kernel regressions (Dimitriadis et al., 2021).
We show that, under mild conditions, these non-negativity properties are also preserved when
recalibration is performed using linear regressions.
As we develop asymptotic inference methods for the estimated decomposition terms on the
right-hand side of (1), our approach is closely related to statistical tests for predictive per-
formance, which focus on the sampling variability of the left-hand side of (1). The seminal
contribution of Diebold and Mariano (1995) laid the foundation for a large econometric litera-
ture on testing for equal forecast performance. Numerous extensions have since been proposed,
including methods for model-based forecasts (West, 1996; Clark and McCracken, 2001), multiple
forecast comparisons (Hansen, 2005; Hansen et al., 2011), conditional predictive ability (Giaco-
5
mini and White, 2006; Li et al., 2022), and anytime-valid procedures (Henzi and Ziegel, 2022;
Choe and Ramdas, 2024). For comprehensive surveys of this literature, see, for example, West
(2006), Clark and McCracken (2013), and Diebold (2015).
In contrast, the assessment of sampling uncertainty for the right-hand side of (1) has received
relatively little attention. Existing contributions, such as Siegert (2014) and Gneiting and Resin
(2023), propose ad-hoc (resampling) methods to quantify the uncertainty for a single forecaster,
but they do not provide tools for comparative inference across multiple forecasters. Our work
addresses this gap by developing a unified framework for inference on the decomposition com-
ponents.
VaR backtests are widely used to evaluate the reliability, i.e., the calibration, of risk models
via their quantile forecasts. Classic procedures include the unconditional coverage test of Kupiec
(1995), which underlies the traffic-light system of Basel Committee (1996, 2019b), as well as the
conditional coverage test of Christoffersen (1998), which additionally assesses the independence
of exceedances. A large body of subsequent work tests slightly differing notions of conditional
calibration (Berkowitz et al., 2011; Gaglianone et al., 2011; Berkowitz et al., 2011; Hoga and
Demetrescu, 2023; Fosten et al., 2024; Wang et al., 2025); see He et al. (2022) for a recent
overview. Our score decompositions—and in particular the application in Section 5.2—show
that well-calibrated forecasts do not necessarily exhibit strong predictive performance. This
underscores the importance of evaluating VaR forecasts through both calibration and discrimi-
nation, rather than relying solely on traditional backtesting metrics; also see Section 3.5.
A growing strand of the macroeconomic literature—including, among many others, Coibion
and Gorodnichenko (2015), Bordalo et al. (2020), Kohlhas and Walther (2021), Angeletos et al.
(2021), Broer and Kohlhas (2024), and Farmer et al. (2024)—examines how forecast errors re-
spond to observable information. This line of work effectively tests conditional calibration, ask-
ing whether forecast errors are predictable given the information available at the time forecasts
are made. Our methodology suggests a natural refinement of these approaches by additionally
assessing forecast discrimination, since calibration captures only one dimension of predictive
performance. This refinement is particularly relevant because forecasts that discriminate well
between outcomes are often found to be miscalibrated, and thus appear to deviate from rational
expectations more frequently.
6
2 Fundamentals of forecast evaluation
After fixing notation in Section 2.1, we discuss relative and absolute forecast evaluation in
Sections 2.2 and 2.3.
for all F ∈ P and all (deterministic) x ∈ A, where EY ∼F [ · ] denotes the expectation with respect
to Y ∼ F ∈ P. A scoring function is said to be strictly consistent if equality in (2) implies
7
x = Γ(F ) such that the optimal forecast Γ(F ) uniquely obtains the lowest expected score. A
functional is said to be elicitable if a strictly consistent scoring function exists.
Holzmann and Eulert (2014, Theorem 1) allows to define consistency of a scoring function
relative to a generic information set Gt ⊆ Ft as
E S Γ(Yt | Gt ), Yt | Gt ≤ E S(X
et , Yt ) | Gt , (3)
for all Yt | Gt ∈ P, and for all Gt -measurable random variables X et . The scoring function is
et = Γ(Yt | Gt ) a.s..
strictly consistent if almost sure equality in (3) implies that X
If the target functional Γ is the conditional mean, under mild regularity conditions, all strictly
consistent scoring functions are characterized by the so-called Bregman class,
parametrized by a strictly convex function ϕ with subgradient ϕ′ (Gneiting, 2011, Theorem 7).
The omnipresent squared error
arises for ϕSE (x) = x2 . For ϕQL (x) = − log(x) we obtain the QLIKE scoring function,
y y
SQL (x, y) = − log − 1, (6)
x x
which is regularly used for the evaluation of variance forecasts in financial econometrics with
positive action and observation domain, A = O = (0, ∞); see e.g., Patton (2011). Under the
standard assumption of a zero mean, the variance coincides with the second moment, thereby
rendering the Bregman class applicable.
If the target functional is the conditional α-quantile, α ∈ (0, 1), all strictly consistent scoring
functions are given by the class of generalized piecewise linear (GPL) loss functions
where the function g is non-decreasing on A = O ⊆ R (Gneiting, 2011, Theorem 9). All GPL
loss functions are non-differentiable at x = y due to the indicator function. The check loss
SCL (x, y) = (1{y ≤ x} − α)(x − y) as the most popular candidate arises for gCL (x) = x.
In practice, scoring functions are employed to compare forecast sequences Xit , i = 1, 2,
SiT = T1 Tt=1 S(Xit , Yt ).
P
t = 1, . . . , T , for the target variable Yt through their average scores b
The sample average serves as an empirical approximation to the expectation appearing in (2)–
(3). Following Diebold and Mariano (1995), a rich body of asymptotic inference methods is
available for constructing tests based on such average score differences.
8
through the theory of forecast calibration. A forecast Xit for Γ(Yt ) is said to be conditionally
calibrated with respect to the information set Git ⊆ Ft if
This property asserts a certain compatibility of the forecast Xit with its forecasting target
Γ(Yt | Git ). A conditionally calibrated forecast w.r.t. the full information set Ft minimizes the
Ft -conditional expectation of the scoring function by (3), such that the econometric literature
also calls it optimal or efficient (Elliott and Timmermann, 2016, Chapter 15). As the forecast
might not have used—or even have access to—all information in Ft , it is, however, often more
appropriate to consider calibration based on some reduced information set Git ⊂ Ft in (8).
An important special case is that of an auto-calibrated forecast, obtained for Git = σ{Xit }, i.e.,
Conditioning on the forecast itself guarantees that the information encoded in the forecast was
indeed used, and it remains feasible even when the evaluator does not have access to the full
information set of the forecaster Iit .
A common tool to assess auto-calibration is to plot x ∈ A against an estimate of Γ(Yt |
Xit = x), where deviations from the diagonal indicate miscalibration. In the meteorological
literature, this is known as a reliability diagram, typically estimated nonparametrically (Murphy
and Winkler, 1992; Pohle, 2020; Dimitriadis et al., 2021). In economic mean forecast evaluation,
this idea goes back to Mincer and Zarnowitz (1969), who regress Yt on Xit and test whether the
intercept is zero and the slope is one—essentially a linear reliability diagram. Gaglianone et al.
(2011); Guler et al. (2017); Bayer and Dimitriadis (2022) cover extensions to other functionals.
A further important tool for calibration assessment are identification functions. Under
smoothness conditions, they correspond to derivatives of the scoring function (Nolde and Ziegel,
2017). Formally, an identification function V : A × O → R is increasing and left-continuous in its
first argument, and it is called strict if EY ∼F V(x, Y ) = 0 if and only if x = Γ(Y ). For example,
V(x, y) = 2(x − y) is the canonical identification function that arises as the derivative of the
squared error loss and for the α-quantile, V(x, y) = 1{y ≤ x} − α is the almost sure derivative of
the check loss. Identification functions allow for a (subject to regularity conditions) equivalent
definition of conditional calibration from (8) as
E V(Xit , Yt ) | Git = 0, a.s., (10)
which specializes to auto-calibration when Git = σ{Xit }. Taking expectations on both sides of
(10) allows to define unconditional calibration of Xit as E V(Xit , Yt ) = 0.
3 Theory
We introduce linear score decompositions in Section 3.1, develop asymptotic inference in Sec-
tion 3.2, discuss testable hypotheses in Section 3.3, and derive low-level inference conditions
for mean and quantile forecasts in Section 3.4. Section 3.5 discusses the connection to VaR
9
backtests. Related decompositions are discussed in Section B.
the Reference forecast r̄ = Γ(Yt ), which is unconditionally calibrated but constant, and the
R
corresponding Reference score S = E S(r̄, Yt ) .
With these ingredients, we define the population score decomposition, consisting of non-
negative components (Gneiting and Resin, 2023):
C R C R
Si = Si − Si − S − Si + |{z}
S . (11)
| {z } | {z }
MCBi DSCi UNC
C
In the canonical case Wit = σ{Xit }, the component MCBi = Si − Si quantifies the increase in
the score’s predictive ability that can be achieved by recalibration, hence providing a measure
of miscalibration such that low MCBi values are desirable. As the recalibrated forecast is a mea-
surable transformation of Xit , it incorporates no additional information; thus, any score gains
stem purely from improved calibration. If MCBi = 0, recalibration yields no score improvement,
indicating that the original forecast was already calibrated.
R C C
The component DSCi = S − Si compares the scores of two recalibrated forecasts: Si based
R
on information in Xit , and S unconditionally without additional information. As this “controls
for calibration”, the remaining score improvement in DSCi reflects the forecast’s information
content, that is, its discrimination ability. Larger values are therefore desirable. The final
R
component UNC = S is forecast-independent and serves as a normalizing constant.
In practice, the components in (11) require estimated versions of both the recalibrated and
reference forecasts. While estimating the reference forecast as the unconditional functional is
straightforward, the estimation of the recalibrated forecast remains debated, with proposals
including binning methods (Murphy and Winkler, 1987), kernel regression (Pohle, 2020), and
more recently isotonic regression (Dimitriadis et al., 2021; Gneiting and Resin, 2023). We instead
propose the use of linear regression, which mitigates biases from nonparametric overfitting that
can hinder valid inference, particularly in settings of close to auto-calibration, aligns with Min-
cer and Zarnowitz (1969), and easily accommodates additional covariates beyond the forecasts
themselves.
10
For this, we consider the recalibration variables Wit = (1, Xit , . . . )⊤ ∈ Rk that generate the
recalibration information Wit = σ{Wit } and impose the following:
Assumption 3.1. S is a strictly consistent scoring function for the functional Γ, the random
variables Yt , Xit and Wit are strictly stationary for all t ∈ N, and Γ(Yt | Wit ) = Wit⊤ θ̄i a.s. for
some true parameter θ̄i ∈ int(Θ) ⊆ Rk , where Θ is convex and compact.
Besides stationarity, Assumption 3.1 imposes linearity of the recalibration function XitC =
Γ(Yt | Wit ) = Wit⊤ θ̄i , which is required for correct specification of the linear “recalibration
regression”. For the finite-sample statements of Theorem 3.2 below, it is crucial that we use
M-estimation for the regression parameters by using the same scoring function that is used for
the decomposition1
T
1X
S Wit⊤ θ, Yt ,
θiT = arg min
b (12)
θ∈Θ T
t=1
where we assume for simplicity that the arg min in (12) is unique. Moreover, the empirical
reference forecast rbT is the unconditional functional Γ of the sample Y1 , . . . , YT . Then, the
empirical counterpart of the score decomposition in (11) is given by
[ iT − DSC
SiT = MCB
b [ T := S
d iT + UNC SCiT − b
biT − b SR
T − S
bC bR
iT + ST , (13)
where
T T T
1X 1X 1X
SiT =
b S(Xit , Yt ), SCiT =
b S(Wit⊤ θbiT , Yt ), SR
b
T = S(b
rT , Yt ).
T T T
t=1 t=1 t=1
The following theorem ensures (strict) positivity of the sampling versions of the decompo-
sition in (13). While similar results for the population decomposition (11) are established in
Gneiting and Resin (2023, Theorem 2.23), it is noteworthy that these results extend to the
sampling versions under linear recalibration, even without imposing correct specification in As-
sumption 3.1.
Theorem 3.2. Assume that the recalibrated forecasts are obtained by (12) using the same strictly
consistent score S as used in (13). Furthermore, let Wit contain a constant and the forecast Xit ,
rT , 0, . . . )⊤ ∈ Θ almost surely. Then, for the decomposition in (13), it holds that:
and let (b
[ iT ≥ 0, which holds with equality if Xit = W ⊤ θbiT a.s. for all t = 1, . . . , T .
(a) MCB it
11
The non-negativity guarantees of Theorem 3.2 are essential, as zero miscalibration and dis-
crimination components are routinely interpreted as perfect calibration and no discrimination,
respectively (Dimitriadis et al., 2021; Gneiting and Resin, 2023; Allen et al., 2025). The proof
of Theorem 3.2 applies to any recalibration method estimated by global score minimization and
which is flexible enough to nest both a constant and the identity line, such as isotonic regres-
sion (Dimitriadis et al., 2021; Gneiting and Resin, 2023). In contrast, kernel regressions (Pohle,
2020), logistic regressions and Beta-CDF models for binary events (Siegert, 2017; Ranjan and
Gneiting, 2010) do not offer such guarantees. Although kernels can represent constants and
the identity line, they are not fitted via global score minimization; Beta-CDFs, being strictly
increasing, cannot produce constants; and logistic regressions generally cannot reproduce the
identity line.
T
−1/2
X ⊤ d
T −1/2 ΩT a1t , c1t , a2t , c2t , et −→ N (0, I5 ),
t=1
where ait := S(Xit , Yt )−E[S(Xit , Yt )], cit := S(Wit⊤ θ̄i , Yt )−E[S(Wit⊤ θ̄i , Yt )], et := S(r̄, Yt )−
E[S(r̄, Yt )] and where
T
!
X ⊤
ΩT = Var T −1/2 a1t , c1t , a2t , c2t , et . (14)
t=1
(iii) The expected score E S(Wit⊤ θ, Yt ) is continuously differentiable in θ for all θ ∈ int(Θ).
T
X
νiT (θ) := T −1/2 S(Wit⊤ θ, Yt ) − E S(Wit⊤ θ, Yt )
(15)
t=1
is stochastically equicontinuous, i.e., for all ε > 0 and η > 0, there exists δ > 0 such that
!
lim sup P sup νiT (θ) − νiT (τ ) ≥ η < ε. (16)
T →∞ {θ,τ ∈Θ: ∥θ−τ ∥≤δ}
Assumption 3.3 provides high-level conditions on the scoring function and the parametric
estimators. These conditions can accommodate a wide range of settings, including various
12
estimation methods, target functionals, and forms of temporal dependence. They are verified for
the mean and quantile cases in Section 3.4 below, and their generality allows for straightforward
extensions to other functionals. In a nutshell, conditions (i) and (ii) can be verified based
on classical dependence and moment conditions on the forecast-realization pairs. Conditions
(iii) and (iv) ensure applicability to non-smooth scoring functions, such as the check loss, by
imposing smoothness on the expected score and controlling empirical deviations via stochastic
equicontinuity (Andrews, 1994; van der Vaart and Wellner, 1996).
Theorem 3.4. Let Assumptions 3.1 and 3.3 hold and assume that MCBi , DSCi > 0 for i = 1, 2.
Then, as T → ∞,
[ 1T
MCB − MCB1 1 −1 0
0 0
√
d
−1/2 DSC1T − DSC1 d 0 −1 0 0 1
T ΞΩT Ξ⊤ −→ N (0, I4 ), with Ξ=
.
[ 2T
MCB − MCB2
0 0 1 −1 0
DSC
d 2T − DSC2 0 0 0 −1 1
The joint asymptotic approximation of Theorem 3.4 is required for the construction of “com-
parative” confidence intervals and hypothesis tests—for instance, tests for equal miscalibration
or discriminatory ability across the two forecast sequences. Notably, the resulting asymptotic
distribution—with covariance matrix ΩT from (14)—is unaffected by the estimation uncertainty
of θbiT and rbT , due to an orthogonality between these estimators and the sampling error from
replacing expected scores with their empirical averages.
Theorem 3.4 requires strict positivity of the population values MCBi , DSCi > 0, hence ex-
cluding the boundary cases MCBi = 0 and DSCi = 0, which necessitate a separate treatment for
which we impose the following high-level assumptions.
(i) For the estimators rbT and θbiT , the following joint asymptotic behavior holds
P −1
! 1 T ′′ (r̄, Y ) T
√ rbT − r̄ T t=1 S t 0 −1/2
X
T =− P
−1 T
S′ (r̄, Yt )Wit + oP (1),
θiT − θ̄i 1 T ′′ ⊤ ⊤
t=1 S (Wit θ̄i , Yt )Wit Wit
b
T t=1
(17)
where T −1/2 Tt=1 S′ (Wit⊤ θ̄i , Yt )Wit satisfies a central limit theorem (CLT) and the terms
P
1 PT ′′ 1 PT ′′ ⊤ ⊤
T t=1 S (r̄, Yt ) and T t=1 S (Wit θ̄i , Yt )Wit Wit are invertible for T large enough and
satisfy laws of large numbers (LLN).
(ii) The scoring function S is three times continuously differentiable in its first argument and
for hT (θ) := T1 Tt=1 S′′′ (Wit⊤ θ, Yt ) ∥Wit ∥3 , we have that
P
As before, the high-level conditions in Assumption 3.5 allow for case-by-case verifications for
specific functionals, scoring functions, and estimators as exemplified in Section 3.4. Condition (i)
13
provides a standard asymptotic linear representation of the estimators, slightly strengthening
Assumption 3.3 (i). Notably, the MCB result below could be derived using only the representa-
tion for θbiT . Condition (ii) applies to smooth loss functions and imposes a uniform LLN type
requirement, although this condition is substantially weakened by the T −1/2 scaling factor.
Theorem 3.6. Suppose that Assumptions 3.1 and 3.5 hold. Then, as T → ∞,
\iT − 1 N ⊤ Υ−1 NiT −→
• if MCBi = 0, then T MCB
P
0; and
2 iT i
d iT − 1 N ⊤ H −1 NiT −→
• if DSCi = 0, then T DSC
P
0,
2 iT i
!!
−1 E[S′′ (r̄, Yt )]−1 0
S′′ (Wit⊤ θ̄i , Yt )Wit Wit⊤ Hi−1
:= E S′′ (r̄, Yt )Wit Wit⊤
where Υi := E , − ,
0 0
:= Var T −1/2 Tt=1 S′ (Wit⊤ θ̄i , Yt )Wit .
P
and NiT ∼ N (0, ΠiT ) with ΠiT
Theorem 3.6 shows for smooth scoring functions that when the population quantities MCBi or
[ iT and T DSC
DSCi attain their boundary value of zero, the non-negative test statistics T MCB d iT
(see Theorem 3.2) converge to generalized χ2 -distributions. Extending Theorem 3.6 to non-
smooth scoring functions via the proof strategy of Theorem 3.4 would necessitate establishing
√
stochastic equicontinuity of T νiT (θ) = Tt=1 S(Wit⊤ θ, Yt )−E S(Wit⊤ θ, Yt ) . To the best of our
P
knowledge, established techniques for such a derivation are currently unavailable in this context.
Recall that the asymptotic variance in the Gaussian case of Theorem 3.4 is unaffected by the
estimation error in rbT and θbiT ; it depends solely on the sampling uncertainty from approximating
expected scores by empirical averages, as captured by ΩT in (14). In contrast, the limiting
distributions in Theorem 3.6 depend exclusively on the estimation effects, as summarized by
Υi , Hi , and ΠiT , which correspond to the population counterparts of the asymptotic expansion
in (17). Technically, this distinction arises because the asymptotic distribution in Theorem 3.4
is driven by first-order Taylor expansion terms, whereas the distributions in Theorem 3.6 are
governed by second-order terms; see the proofs for details.
Feasible inference requires consistent estimators of the population matrices ΩT , ΠiT , Υ−1
i ,
and Hi−1 so that the asymptotic results in Theorems 3.4 and 3.6 remain valid when replacing
population quantities by their estimates. While Υ−1
i and Hi−1 admit consistent estimation via
sample averages, we use HAC estimators for the long-run covariance matrices ΩT and ΠiT to
accommodate serial dependence (Newey and West, 1987; Andrews, 1991). Consistency follows
by standard arguments, see e.g., White (2001, Chapter 6.4) for smooth and Galvao and Yoon
(2024) for non-smooth scoring functions.
14
For example, consider the hypothesis HMCB : MCB1 = MCB2 , which we can rewrite as
HMCB = MCB1 = MCB2
= MCB1 = MCB2 > 0 ∪ MCB1 = MCB2 = 0
= MCB1 = MCB2 > 0 ∪ MCB1 = 0 ∩ MCB2 = 0
+ 0 0
=: HMCB ∪ HMCB 1
∩ HMCB 2
. (18)
+
For the hypothesis HMCB , we employ the selection vector ω = (1, 0, −1, 0)⊤ such that Theo-
rem 3.4 yields
−1/2 √ d
ω ⊤ ΞΩT Ξ⊤ ω
[ 1T − MCB
T MCB [ 2T − MCB1 − MCB2 −→ N (0, 1), (19)
pMCB := max p+ 0 0
MCB , 2 min pMCB1 , pMCB2 . (20)
Similarly, we obtain an asymptotically valid p-value, pDSC , for testing equal discrimination,
HDSC = DSC1 = DSC2 = DSC1 = DSC2 > 0 ∪ DSC1 = 0 ∩ DSC2 = 0 , (21)
by adapting (18) and (20). The Gaussian p-value follows from (19) by setting ω = (0, 1, 0, −1)⊤ ,
while for smooth scoring functions, the corresponding generalized χ2 -probability is computed
analogously using the method of Imhof (1961). For non-smooth scoring functions, we test
DSCi = 0 through the MZ-type regression of the VQR test, but by simply testing the hypothesis
of a zero slope.
A classical DM test for equal expected scores, HS = S1 = S2 , arises as a special case
of Theorem 3.4 for ω = (1, −1, −1, 1)⊤ . Confidence intervals for the individual decomposition
terms as well as their differences can be obtained by inverting the above tests.
15
Assumption 3.7. The following holds for i = 1, 2:
(i) The function ϕ(z) in (4) is five times continuously differentiable with ϕ′′ (z) > 0 for all
z ∈ A = O.
(ii) (Yt , Wit⊤ )t∈N is strong mixing of size −s/(s − 2), s > 2; see White (2001, Definition 3.45).
(iv) The matrix VT := Var T −1/2 Tt=1 ϕ′′ (Wit⊤ θ̄i )Wit (Wit⊤ θ̄i − Yt ) is uniformly positive
P
Assumption 3.7 resembles standard time series conditions for mean regression. Item (i) is
a smoothness and convexity condition on the scoring function. For the SE scoring function,
we often use A = O = R, whereas the QLIKE score requires positive values such as A = O =
(0, ∞). Item (ii) is a classical mixing condition that ensures that suitable LLNs and CLTs apply
and items (iii) and (iv) are standard in M-estimation (Newey and McFadden, 1994, Chapter
3.1). Item (v) collects moment conditions, which simplify considerably for the common choices
ϕSE (z) = z 2 and ϕQL (z) = − log(z) of the SE and QLIKE scores in (5) and (6).
Assumption 3.8. The matrix ΩT from (14) is uniformly positive definite for T large enough.
Positive definiteness of ΩT is stated in the separate Assumption 3.8 as it is required for the
asymptotic normality result in Theorem 3.4, but is violated in the cases MCBi = 0 and DSCi = 0,
and is as such not required for the verification of the conditions of Theorem 3.6. The positive
definiteness of ΩT essentially imposes non-collinearity conditions for the (sum over time of the)
original, recalibrated, and reference scores. Violations arise, for instance, if the recalibrated
score is a constant multiple of the original score for all t ∈ N, in which case MCBi = 0 follows.
Proposition 3.9. Consider a Bregman score from (4) for the decomposition, which is also used
in (12). Then:
• Assumptions 3.1, 3.7, and 3.8 imply Assumption 3.3 such that Theorem 3.4 holds; and
• Assumptions 3.1 and 3.7 imply Assumption 3.5 such that Theorem 3.6 holds.
We continue with the case where Γ is the α-quantile for some α ∈ (0, 1), where all strictly
consistent loss functions are given by the GPL class in (7). For a linear GPL score decomposition,
we employ (generalized) quantile regression by combining the estimator (12) with the loss in (7)
and impose the following conditions.
16
(i) The function g from (7) is twice continuously differentiable with g ′ (z) > 0 for all z ∈ A =
O.
(ii) The sequence (Yt , Wit⊤ )t∈N is strong mixing of size −s/(s − 2) for some s > k + δ with
δ > 0, and where k is the dimensionality of Θ from Assumption 3.1.
(iii) For all t ∈ N, Fit , the conditional distribution of Yt given Wit , belongs to a class of
distributions that possess a positive Lebesgue density fit with 0 < fit (x) ≤ K < ∞,
uniformly on the support of Fit .
(vii) The matrix ΩT from (14) is uniformly positive definite for T large enough.
The conditions of Assumption 3.10 are standard regularity conditions for quantile regression,
but also cover the case of estimation through the more general GPL loss functions. The required
moments of order s > k are required for the derivation of stochastic equicontinuity for strong
mixing processes through Hansen (1996). As we consider low dimensionalities of Θ ⊂ Rk such
as k = 2 in our applications, this is not overly restrictive.
Compared to Assumption 3.7, the conditions in Assumption 3.10 require the conditional dis-
tribution to be absolutely continuous, as specified in item (iii); this is a standard requirement for
quantile regression frameworks. Furthermore, item (vii) is included here because the subsequent
proposition specifically verifies the conditions for the asymptotic normality result established in
Theorem 3.4, but does not treat the case of Theorem 3.6.
Proposition 3.11. Consider a GPL score from (7) for the decomposition, that is also used in
(12). Then, Assumptions 3.1 and 3.10 imply Assumption 3.3 such that Theorem 3.4 holds.
Proposition 3.11 verifies the high-level conditions of Theorem 3.4 through standard conditions
from (time series) quantile regression. Notice that Theorem 3.6 only applies to smooth loss
functions, and hence, the cases MCBi = 0 and DSCi = 0 require a different treatment for the
case of quantile forecasts. For the case MCBi = 0, we employ the quantile-specific MZ-test
of Gaglianone et al. (2011), which we adapt for the case of DSCi = 0 by simply testing for a
zero slope in the MZ-regression. The resulting p-values are accordingly used in the combination
method in (20).
17
Table 2: Summary of prominent VaR backtests, including the underlying hypothesis, methodology and relation
to conditional calibration in (10) through Git = σ(Git ). We use the shorthand Vit−l := V(Xit−l , Yt−l ) for l ≥ 0.
such as by Kupiec (1995) and the traffic light system of Basel Committee (1996, 2019b) typically
focus on unconditional counts for the “violations” or “hits”, where realized losses exceed the VaR
predictions as captured through the quantile identification functions, Vit = V(Xit , Yt ) = 1{Yt ≤
Xit } − α, from (10). These unconditional tests, however, remain insensitive to the clustering of
violations—a phenomenon particularly critical during financial crises, when consecutive days of
unpredicted losses can lead to rapid capital depletion.
Following the seminal work of Christoffersen (1998), a plethora of conditional backtests has
been proposed over the past decades. See e.g., Campbell (2007) and Nieto and Ruiz (2016)
for comprehensive reviews on VaR backtesting. Most of these methods test hypotheses closely
related to the conditional calibration property in (10) by utilizing varying information sets
Git = σ(Git ). Table 2 provides a non-exhaustive list of prominent VaR backtests, illustrating the
intrinsic connection between their choice of covariates (or instruments) Git and our recalibration
variables Wit . Notably, the equivalence between (8) and (10) permits specifications based on
either quantile regressions for Yt , or mean regressions for the respective identification function.
As a canonical case of assessing auto-calibration, the quantile-specific adaptation of the MZ
backtest by Gaglianone et al. (2011) utilizes the same recalibration technique as our linear score
decomposition when Wit⊤ = (1, Xit ). By matching the recalibration variables Wit with the
covariates Git , both the score decompositions and the underlying inference can be tailored to
the specific notion of conditional calibration underlying the other backtests presented in Table 2.
4 Simulations
We now assess the finite-sample performance of the proposed tests from Section 3.3. Section 4.1
describes the simulation setup and Section 4.2 reports the results.
18
iid
with εY,t ∼ N (0, 1). The predictor variables in Vt follow independent autoregressive AR(1)
processes Kt = βKt−1 + εK,t , Lt = βLt−1 + εL,t , Mt = βMt−1 + εM,t , Nt = βNt−1 + εN,t , with
iid
fixed β = 0.25 and εK,t , εL,t , εM,t , εN,t ∼ N (0, 1), which are all mutually independent.
Forecaster X1t has exclusive access to Kt and forecaster X2t to Lt , while both forecasters
share information on Mt and Nt as formalized by I1t = σ {Kt , Mt , Nt } and I2t = σ {Lt , Mt , Nt }.
Notice that our setup requires both “joint” predictors Mt and Nt in order to construct non-
collinear forecasts with equal MCB and/or DSC values.
Based on the distinct information sets I1t and I2t , the forecasters issue mean forecasts
and δ0 , ξ0 ∈ R. The imposed zeros in δ and ξ reflect the absence of the corresponding predictor
variables in the information sets of the respective forecaster. The coefficients (δ0 , δ ⊤ ) and (ξ0 , ξ⊤ )
and their relation to γ allows to generate forecasts with different calibration and discrimination
properties, as formalized in the following proposition for the squared error score from (5).
Proposition 4.1. For realizations Yt and forecasts X1t and X2t following (22)–(24) with δ ̸= 0
and ξ ̸= 0, and given the recalibration information Wit = σ{Wit }, with Wit = (1, Xit )⊤ , i = 1, 2,
the population squared error score decomposition of the forecasts X1t is
(δ ⊤ δ − δ ⊤ γ)2 (δ ⊤ γ)2
MCB1 = δ02 + ς · , DSC1 = ς · , UNC = ς · γ ⊤ γ + 1, (25)
δ⊤δ δ⊤δ
1
where ς = 1−β 2
is the variance of the zero-mean AR(1) processes Kt , Lt , Mt , Nt . For MCB2 and
DSC2 interchange δ02 and δ with ξ02 and ξ, respectively.
The expressions in Proposition 4.1 illustrate that for non-constant forecasts, MCB1 = 0 holds
if and only if δ ⊤ δ = δ ⊤ γ and δ0 = 0. Hence, perfect calibration, MCB1 = 0, can be achieved
even if the forecaster omits relevant variables. In fact, MCB1 = 0 implies that if a predictor
(say Mt ) is used, it must be used with exactly the right parameter value (δM = γM ). This
also implies that the use (e.g., δM ̸= 0) of irrelevant predictor variables (e.g., γM = 0) causes
miscalibrated forecasts.
The DSC1 component is unaffected by the intercept δ0 that does not carry any information.
DSC1 is positive whenever δ ⊤ γ ̸= 0, that is, whenever X1t reacts to the signal in Yt , even
(δ ⊤ γ)2
incorrectly such as in the wrong direction. Moreover, by Cauchy-Schwarz, δ⊤ δ
≤ γ ⊤ γ, such
that the maximal value of DSC1 is ς γ ⊤ γ, which is attained if and only if the parameters δ are
(δ ⊤ γ)2 c2 (γ ⊤ γ)2
chosen proportionally to γ, i.e., δ = cγ for some c ∈ R as then, δ⊤ δ
= c2 γ ⊤ γ
= γ ⊤ γ. Hence,
maximal discrimination can arise even for miscalibrated forecasts. If only some relevant variables
are observed, the best achievable discrimination is obtained when the available predictors are
used correctly. Analogous conclusions hold for the second forecaster X2t by replacing δ with ξ.
The closed-form expressions in Proposition 4.1 enable diverse specifications of MCBi and
DSCi , allowing us to study the size and power of our proposed tests across different settings by
appropriately choosing the parameters γ, (δ0 , δ ⊤ ), and (ξ0 , ξ⊤ ). In detail, Table 3 specifies twelve
different parameterizations that result in illustrative settings for MCBi and DSCi as given in the
19
Table 3: Twelve mean forecasting specifications of the parameters γ, δ0 , δ, ξ0 , ξ in (22)–(24) based on a flexible
k ∈ {0, 0.05, . . . , 0.45, 0.5}. All setups guarantee non-collinear forecasts and generate different relations of the
MCB and DSC values as given in the numbered rows and the last column termed “Implications”.
8k2
1) MCB1 (k) = MCB2 (k) = 15
1 1 √k + 1 k 1 k 1
0 4 4 0 0 0 0 2 4 0 0 0 2 + 4 2 + 4 0 0 < DSC1 < DSC2
1 1 1 k 1 k 1 k 1 k 1
4 4 4 0 0 2 + 4 0 2 + 4 0 0 0 2 + 4 2 + 4 0 0 < DSC1 = DSC2
8k2
2) MCB1 = 0 and MCB2 (k) = 15
1 1 1 k 1 k 1
0 4 4 0 0 0 0 4 0 0 0 2 + 4 2 + 4 0 0 < DSC1 < DSC2
1 1 1 1 1 k 1 k 1
4 4 4 0 0 4 0 4 0 0 0 2 + 4 2 + 4 0 0 < DSC1 = DSC2
16k2
3) MCB1 (k) = 2 MCB2 (k) = 15
1 1
0 4 4 0 0 0 0 k + 14 0 0 0 k
2 + 1
4
k
2 + 1
4 0 0 < DSC1 < DSC2
1 1 1 √k 1 √k
4 4 4 0 0 2
+ 4 0 + 14
2
0 0 0 k
2 + 1
4
k
2 + 1
4 0 0 < DSC1 = DSC2
8k2
4) DSC1 (k) = DSC2 (k) = 15
k 1 k 1 k 1 k 1
0 0 0 k 0 2 + 4√ 2
0 0 2 + 4√ 2
0 0 2 + 4 0 2 + 4 0 < MCB1 < MCB2
k 1 k 1 k 1 k 1
0 0 0 k 0 2 + 4 0 0 2 + 4 0 0 2 + 4 0 2 + 4 0 < MCB1 = MCB2
8k2
5) DSC1 = 0 and DSC2 (k) = 15
1 k 1 k 1
0 0 0 k 0 0 0 4 0 0 0 2 + 4 0 2 + 4 0 < MCB1 < MCB2
1 1 k 1 k 1
0 0 0 k 0 4 0 4 0 0 0 2 + 4 0 2 + 4 0 < MCB1 = MCB2
16k2
6) DSC1 (k) = 2 DSC2 (k) = 15
1 k 1 k 1
0 0 0 k 0 0 0 0 k+ 4 0 0 2 + 4 0 2 + 4 0 < MCB1 < MCB2
0 0 0 k √1 0 0 0 k+ 1
0 0 k
+ 1
0 k
+ 1
0 < MCB1 = MCB2
15 4 2 4 2 4
20
Table 4: Four quantile forecasting specifications of the parameters γ, δ0 , δ, ξ0 , ξ in (22) and (26) based on a flexible
k ∈ {0, 0.05, . . . , 0.95, 1}. All setups guarantee non-collinear forecasts and generate different relations of the MCB
and DSC values as given in the numbered rows and the last column termed “Implications”. The values of ξ0 (k, α)
are chosen numerically such that MCB1 (k) = MCB2 (k) holds; see Appendix C and Figure C.1 for details.
Realization α
Forecaster X1t α
Forecaster X2t
γK γL γM γN δ0 δK δL δM δN ξ0 ξK ξL ξM ξN Implications
facet labels of Figure 1. To analyze test power, we continuously misspecify the respective null
hypotheses through a parameter k > 0 that affects γ, (δ0 , δ ⊤ ), and (ξ0 , ξ⊤ ). The specifications
8k2
of Table 3 are chosen to guarantee that either MCB2 (k) or DSC2 (k) depend on k and equal 15
for all k ≥ 0. Figure D.1 in Appendix D plots the population values MCBi and DSCi , i = 1, 2
as functions of k for all twelve setups.
8k2
For example, we generate the scenario MCB1 = 0 and MCB2 (k) = 15 while maintaining
0 < DSC1 < DSC2 independent of k by setting γ⊤ = (0, 1 1
4 , 4 , 0) as well as (δ0 , δ ⊤ ) = (0, 0, 0, 14 , 0)
and (ξ0 , ξ⊤ ) = (0, 0, 41 + k2 , 14 + k2 , 0). Here, the realizations are only driven by Lt and Mt with
equal magnitudes, γL = γM . The first forecaster X1t does not have access to Lt , but she
correctly uses Mt through δM = γM such that her forecasts are perfectly calibrated, MCB1 = 0.
1
The sole access to Mt however limits her discrimination to DSC1 = 15 . By contrast, the second
forecaster X2t uses both relevant variables Lt and Mt , and hence achieves a higher discrimination
2
DSC2 = 15 > DSC1 for any k ≥ 0. However, for values k > 0, she uses the information incorrectly
8k2
with ξ ̸= γ such that MCB2 (k) = 15 increases with k.
For the evaluation of our tests for quantile forecasts, we use the same DGP in (22). As
the information in Vt only affects the conditional mean of Yt in (22), we generate the quantile
forecasts at level α ∈ (0, 1) as
α
X1t = Vt⊤ δ + δ0 + Φ−1 (α), and α
X2t = Vt⊤ ξ + ξ0 + Φ−1 (α), (26)
using the inverse of the standard normal CDF, denoted by Φ−1 (·).
We deliberately use the pure location process (22) for quantile forecasts, allowing us to ex-
plicitly control the population miscalibration and discrimination values for the check loss in
Proposition C.1, albeit these closed-form expressions become much more complicated as for the
squared error. Due to this additional complication, we only consider four informative param-
21
a)
MCB1(k ) = MCB2(k ) = 8k 2 15 MCB1 = 0 and MCB2(k ) = 8k 2 15 MCB1(k ) = 2MCB2(k ) = 16k 2 15
1.0
0.4
0.2
0.0
1.0
5
0.
0.
0.
0.
0.
0.
0.
0.
0.
0.
0.
0.
0.
0.
0.
0.
0.
0.
k
b)
DSC1(k ) = DSC2(k ) = 8k 2 15 DSC1 = 0 and DSC2(k ) = 8k 2 15 DSC1(k ) = 2DSC2(k ) = 16k 2 15
1.0
0.4
0.2
0.0
1.0
5
0.
0.
0.
0.
0.
0.
0.
0.
0.
0.
0.
0.
0.
0.
0.
0.
0.
0.
k
Figure 1: Empirical rejection rates for mean forecasts and the squared error score for the proposed tests of equal
miscalibration (in blue) in the upper panel a), equal discrimination (in orange) in the lower panel b), and of the
DM test of equal predictive performance (in gray) in the lower plots rows of both panels. We generate the data
according to (22)–(24) with the twelve parameterizations of Table 3, which depend on the parameter k ≥ 0 that
is displayed on the x-axes. We use T = 500 and a nominal level of 10%.
eterizations as given in Table 4. For these, Appendix C provides the closed-form expressions
for MCBi and DSCi as explicit functions of k, α and the parameters of the DGP. Notably, for
the last parametrization, (approximate) equality of the miscalibration values is achieved by a
numerical choice for the values of ξ0 = ξ0 (k, α); see Figure C.1 and the text of Appendix C.
22
of the test for equal discrimination HDSC : DSC1 = DSC2 in orange color in the lower panel. For
comparison, we also show rejection rates of the DM test of equal predictive power HS : S1 = S2
with gray and dotted lines. We only show these rejection rates when the underlying population
score difference coincides with the respective difference in MCB or DSC values. E.g., in the lower
row of panel a), we ensure DSC1 = DSC2 such that S1 − S2 = MCB1 − MCB2 and the tests of
the hypotheses HS and HMCB are comparable.
Throughout the four plots in the left column, as well as for k = 0 in all twelve plots, the data
is generated under the respective null hypotheses of equal miscalibration in panel a), and equal
discrimination in panel b). Here, we find correct or conservative rejection rates below the nominal
level of 10% for our tests, which is to be expected given the Bonferroni-type corrections from
Section 3.3. In the left panel, our tests approach the correct size for larger values of k as e.g., for
the MCB-test, miscalibration of both forecasts increases with k such that 2 min{p0MCB1 , p0MCB2 }
becomes much smaller than p+
MCB , which operates under its null hypothesis and hence generates
asymptotically exact size control.
The middle and right plot columns show that for increasing k, power increases naturally in
all subplots. This illustrates the robustness of our testing procedure to zero or positive values
of the respective population MCBi and DSCi values. Importantly, in the cases where score
differences coincide with MCB (DSC) differences in the upper (lower) panel, the component
tests’ power increases substantially with respect to the DM test. This can be explained as in
these cases, the DM test’s (estimated) variance is affected by fluctuations of both, (estimated)
MCB and DSC values whereas the MCB test is unaffected by noise in the DSC component, and
vice versa. Hence, we can expect power gains by testing on the respective components, as will
become evident in our applications in Section 5.
Figure 2 displays rejection rates for quantile forecasts at levels α ∈ {0.01, 0.05, 0.1, 0.25, 0.5}
and within the four most informative parameterizations that allow for comparison with the DM
test. These four scenarios correspond to the lower left and middle plots in panels a) and b) from
Figure 1. The results for median forecasts with α = 0.5 are comparable to the mean forecasting
results from Figure 1: The tests are conservative under the null, develop power with k > 0 and
despite their conservative size2 , our tests dominate the DM-test in terms of power. For more
extreme quantile levels, power naturally declines. Importantly, and especially for the DSC test,
the dominance of the DM test is even more pronounced for extreme quantile levels.
2
In unreported simulations, we find that the elevated type I error of the DSC test in the third column for
α = 0.01 does not increase with k and declines as the sample size grows. This pattern suggests that the distortion
is attributable to finite-sample effects.
23
MCB1(k ) = MCB2(k ) MCB1(k ) = 0 and MCB2(k ) MCB1 = MCB2 > 0 MCB1 ≈ MCB2 > 0
DSC1 = DSC2 > 0 DSC1 = DSC2 > 0 DSC1(k ) = DSC2(k ) DSC1 = 0 and DSC2(k )
1.00
0.75
α = 0.01
0.50
0.25
0.00
1.00
0.75
α = 0.05
0.50
0.25
0.00
empirical rejection rate
1.00
0.75
α = 0.1
0.50
0.25
0.00
1.00
0.75
α = 0.25
0.50
0.25
0.00
1.00
0.75
α = 0.5
0.50
0.25
0.00
00
25
50
75
00
00
25
50
75
00
00
25
50
75
00
00
25
50
75
00
0.
0.
0.
0.
1.
0.
0.
0.
0.
1.
0.
0.
0.
0.
1.
0.
0.
0.
0.
1.
k
proposed test on MCB difference proposed test on DSC difference DM test
Figure 2: Empirical rejection rates for α-quantile forecasts and the check loss for the proposed tests of equal
miscalibration (in blue) in the left two columns, equal discrimination (in orange) in the right two columns, and
of the DM test of equal predictive performance (in gray) in all subplots. We generate the data according to (22)
and (26) with the four parameterizations of Table 4, which depend on the parameter k ≥ 0 that is displayed on
the x-axes. We use T = 500 and a nominal level of 10%.
24
5 Applications
Here, we apply the proposed tests for equal miscalibration and discrimination in two central fore-
casting domains: inflation rates in Section 5.1 and financial risks in Section 5.2. Replication code
and details on the data are available at [Link] replications.
8
Inflation Rate (percent)
-2
1982Q4 1988Q1 1993Q3 1999Q1 2004Q2 2009Q4 2015Q2 2020Q4
Figure 3: Quarterly year-over-year CPI inflation rate together with the SPF and Michigan forecasts. For details,
see the text in Section 5.1.
25
0.8
3
46
1.
0.6
SPF
DSC
0.4
3
00 p0MCB = 0.00
2.
SPF
p0DSC = 0.00
0.2
Michigan
Michigan ΔMCB = 0.09
0.0 ΔDSC = -0.45** p0DSC = 0.22
Figure 4: Inflation forecast evaluation results for the SPF and Michigan forecasts using the squared error scoring
function. The left “MCB–DSC” plots show the average score as iso-lines together with its decomposition into
miscalibration and discrimination components for the four competing forecasts. The gray boxes in the right panels
display the p-values of tests for zero miscalibration and discrimination, respectively. The white box displays the
average score, miscalibration and discrimination differences, together with 0 to 3 stars indicating significance at
the 10%, 5%, and 1% levels.
(2020) and construct a corresponding SPF measure by averaging the one- to four-quarter-ahead
consensus forecasts (where the consensus is defined as the median across respondents). Averaging
is appropriate because SPF forecasts are stated in terms of annualized quarter-over-quarter
inflation rates. As the evaluation target Yt , we use the four-quarter logarithmic growth rate
of the consumer price index (CPI), based on the 2023:Q3 vintage, over the year following the
forecast date.
The left panel of Figure 4 reports the average squared error score together with its decom-
position into miscalibration and discrimination components for the two competing forecasts.
Calibration appears broadly similar across the two surveys. In contrast, the SPF forecasts dis-
play a markedly higher degree of discrimination—a difference that is not visible in the purely
scoring-function-based analyses of Ehm et al. (2016) and Patton (2020). The right panel cor-
roborates their main finding: the overall difference in predictive performance is statistically
insignificant, with a p-value of 0.284.
Consistent with the visual impression from the left panel, the difference in miscalibration
is not statistically significant. However, the SPF forecasts exhibit significantly stronger dis-
crimination, with a p-value of 0.037. For the Michigan survey forecasts, we cannot even reject
the null hypothesis of no discrimination. These findings suggest that professional forecasters’
predictions contain more information about future inflation outcomes, which is plausible given
their field-specific expertise and experience.
Figure D.3 in Appendix D reports diagnostic checks for the linear recalibration regression and
reveals no evidence of misspecification, supporting the use of linear recalibration for inference
on score decompositions.
The macroeconomic literature following Coibion and Gorodnichenko (2015), as discussed
26
in the introduction, examines the sensitivity of forecast errors to additional covariates. These
covariates can be integrated into our framework through the recalibration vector Wit . Conse-
quently, their tests for rational expectations focus on a specific dimension of forecast calibration.
However, such evaluations fail to account for the underlying forecast discrimination—a compo-
nent that our application to survey forecasts reveals to be fundamentally distinct from calibration
performance.
From a methodological perspective, this application underscores the utility of our proposed
inference framework. Beyond partitioning forecast performance into interpretable economic
components, our tests can exhibit superior statistical power compared to the DM test in these
settings. Furthermore, the framework is applicable in general time-series environments, including
multi-step-ahead forecasting horizons.
27
We further use two models that use high-frequency information through the RV estimator.
The third model is the Heterogeneous Autoregressive (HAR) model proposed by Corsi (2009),
20
1/2 1/2 iid
X
RVt = ϕ0 + wl RVt−l +εt , with εt ∼ N (0, σ 2 ),
l=1
where the MIDAS weights wl are parameterized through the exponential Almon polynomial.
We construct VaR forecasts for these four models by using the α-quantile of the imposed
residual distribution multiplied by the square root of the respective variance forecast. For
quantile forecasts, we additionally use the classical Historical Simulation (HS) method, that
uses the empirical α-quantile of the preceding 250 trading days as a quantile forecast.
For the first four models, we estimate the parameters by maximum likelihood using the
imposed residual distributions, based on a fixed estimation window from January 1, 2000, to
December 31, 2007, corresponding to approximately the first third of the sample. The subsequent
period from January 1, 2008 to January 1, 2022 serves as evaluation period, yielding T = 3617
one-step ahead forecasts and corresponding realizations. Figure D.2 in Appendix D shows the
variance and VaR forecasts together with their evaluation targets.
Figure 5 presents the evaluation of variance forecasts under the QLIKE and SE scoring func-
tions. The left panels, following Dimitriadis et al. (2024b), plot the estimated miscalibration and
discrimination of the four forecasting methods along the axes. Gray iso-lines represent average
score levels, with observations closer to the upper-left corner indicating superior performance.
We find that the RV-based forecasting methods exhibit markedly stronger discrimination, which
translates into better overall predictive accuracy as reflected in the average scores.
The right panel of Figure 5 reports the corresponding test results: Gray-shaded boxes dis-
play p-values for tests of zero miscalibration and zero discrimination. While all models show
significantly positive discrimination, the null of correct calibration is rejected more frequently
for the RV-based models than for the two GARCH-type specifications.
The remaining off-diagonal panels present pairwise differences in average scores (∆b
S), mis-
[ and discrimination (∆DSC),
calibration (∆MCB), d with significance indicated by 0 to 3 stars
corresponding to the 10%, 5%, and 1% levels. Focusing on the lower column, which compares
the GARCH model to the three alternatives, most differences are highly significant. In particu-
lar, the inferior performance of the GARCH model is primarily driven by its significantly weaker
discrimination, consistent with the informational advantage of RV-based models that exploit
high-frequency data.
Turning to the VaR evaluation results, Table 5 reports the p-values of seven standard VaR
backtests—described in Section 3.5 and Table 2—along with the unconditional hit frequency
in the final column. Strikingly, particularly at the 1% level, the simple HS model appears
28
a) QLIKE Score
HAR
MIDAS 1.75 p0DSC = 0.00
HAR
ΔS = -0.02*** p0MCB = 0.00
MIDAS
6.5
ΔMCB = -0.00
ΔDSC = 0.02*** p0DSC = 0.00
DSC
GJR
ΔS = 0.48*** ΔS = 0.51*** p0MCB = 0.00
GJR
b) SE Score
p0MCB = 0.05
HAR
MIDAS
7.99
p0DSC = 0.00
11
ΔMCB = 0.00
10
ΔS = -0.24 ΔS = 0.09 p0MCB = 0.75
GJR
GARCH
ΔMCB = -0.24 ΔMCB = -0.24 ΔMCB = -0.02
ΔDSC = -1.88*** ΔDSC = -2.21*** ΔDSC = -1.89** p0DSC = 0.00
Figure 5: Variance forecast evaluation results for the E-mini futures using a) the QLIKE and b) the squared error
scoring functions. The left “MCB–DSC” plots show the average score as iso-lines together with its decomposition
into miscalibration and discrimination components for the four competing forecasts. The gray boxes in the right
panels display the p-values of tests for zero miscalibration and discrimination, respectively. The white boxes
display the average score, miscalibration and discrimination differences, together with 0 to 3 stars indicating
significance at the 10%, 5%, and 1% levels.
29
Table 5: Backtesting results for 1% and 5% VaR forecasts. The first panel displays the p-values of the backtests
summarized in Table 2. The column “Basel” also displays the colors of their traffic light system. The last column
reports the relative frequencies of hits, i.e., how often the realizations exceed the VaR forecasts.
p-values backtests
Model UC Basel CC NZ VQR DQ DQX hit freq.
HS 0.137 0.024 0.149 0.124 0.321 0.002 0.004 1.3%
GARCH 0.001 0.000 0.002 0.000 0.006 0.023 0.016 1.8%
1%-VaR GJR 0.000 0.000 0.000 0.000 0.000 0.002 0.001 2.0%
HAR 0.000 0.000 0.000 0.000 0.003 0.000 0.000 2.3%
MIDAS 0.000 0.000 0.000 0.000 0.000 0.000 0.000 2.3%
to be the best calibrated, whereas all competing models are consistently rejected across the
backtests. Although evidence at the 5% level and for conditional calibration backtests is weak
or inconclusive, the 1% level remains the most relevant for risk management (Basel Committee,
2019b, p. 81).
Score decompositions with their associated inference help explain the seemingly superior
performance of the HS model. Figure 6 presents these decompositions, analogous to Figure 5,
using the check loss function for 1% and 5% VaR forecasts.
At the 1% level, the MCB–DSC plot in Figure 6 shows that, although the HS model exhibits
the best calibration, it has by far the weakest discrimination. This results in clearly inferior
predictive performance according to the check loss, as further confirmed by the strongly sig-
nificant results in the right panel. These findings resolve the seemingly contradictory results
highlighted in the motivational example from Section 1.1 and underscore the risk of evaluat-
ing forecasts solely on calibration. As such, they point to a potentially serious shortcoming in
current banking regulation.
At the 5% level, the HS model shows markedly poorer conditional calibration, as reflected
both in the backtest results in Table 5 and in the score decomposition in Figure 6. The relative
ordering of the other four models, however, is similar to that observed at the 1% level, with
generally stronger significance, which can be attributed to lower noise at less extreme probability
levels.
A general pattern evident in the MCB–DSC plot at the 1% level—and, to a lesser extent,
at the 5% level and for the variance forecasts in Figure 5—is that models incorporating more
information can convert it into higher discrimination. However, this often comes at the cost
of increased miscalibration. Notably, achieving auto-calibration conditional on Wit = (1, Xit )
is relatively easy when forecasts Xit are nearly constant, but much more difficult for highly
discriminating forecasts that vary substantially over time. This observation calls into question
the common practice of evaluating risk measure forecasts primarily through MZ-type backtests
(Gaglianone et al., 2011; Guler et al., 2017; Bayer and Dimitriadis, 2022), which focus on auto-
30
a) VaR 1%
HAR
HAR p0DSC = 0.00
GJR
GARCH ΔS = -0.03 p0MCB = 0.00
MIDAS
1.5
ΔMCB = 0.00
ΔDSC = 0.03 p0DSC = 0.00
GJR
b) VaR 5%
4 p0MCB = 0.52
HAR
MIDAS 12.47
p0DSC = 0.00
HAR
ΔS = -0.01 p0MCB = 0.57
MIDAS
GJR
Figure 6: VaR forecast evaluation results for the E-mini futures at quantile levels a) 1% and b) 5% for the check
loss function and its decomposition terms. For details, see the caption of Figure 5.
31
calibration.
Figures D.4–D.5 in Appendix D present diagnostic checks for the linear recalibration re-
gressions underlying the score decompositions of the variance and VaR forecasts. The pointwise
confidence bands largely cover the null specification, providing no evidence against correct model
specification. This supports the use of linear recalibration for score decompositions in these ap-
plications, combining robustness and adequate fit with the ability to conduct inference, which
is essential in the settings considered.
6 Conclusion
This paper has established a comprehensive inferential framework for score decompositions,
enabling a more granular assessment of predictive performance through the lenses of miscali-
bration, discrimination, and uncertainty. By utilizing a linear recalibration technique, we have
shown that it is possible to derive robust inference for a wide range of point forecasts—including
those governed by non-smooth loss functions—while maintaining attractive finite sample non-
negativity conditions and a formal bridge to the classical Mincer and Zarnowitz (1969) regression.
Our results demonstrate that this decomposition approach not only enriches the informational
content of traditional predictive ability tests but also offers tangible gains in statistical power.
Through our empirical applications, we have highlighted how these methods can reveal signifi-
cant latent differences in forecast discrimination and, crucially, identify systemic deficiencies in
financial risk backtesting that current banking regulations fail to capture.
The assumption of linear recalibration is guaranteed to hold under the null hypotheses of
zero discrimination or perfect calibration, though it may be violated in more general settings.
Nevertheless, this linear specification provides essential stability and allows for the inclusion
of further recalibration variables, allowing for direct generalizations to e.g., cross-calibration
(Strähl and Ziegel, 2017), and enabling direct links to VaR backtests and the literature on
macroeconomic survey forecasts following Coibion and Gorodnichenko (2015). The linearity
assumption further guards against the overfitting that often plagues nonparametric techniques
in this context. Such overfitting typically manifests as a positive bias in the non-negative MCB
and DSC components—a risk that is particularly acute in the small-to-medium sample sizes
characteristic of our applications. Notably, we find no empirical evidence against the linearity
assumption in our data, which supports the robustness of our approach.
Nonetheless, particularly in data-rich environments, semiparametric extensions that estimate
the recalibration curve nonparametrically—using kernel, spline, or isotonic regression—represent
promising avenues for future research. Such developments could build on the seminal work of
Newey (1994) regarding two-step semiparametric inference and more recent advances in isotonic
plug-in estimation, such as Xu (2022).
We treat the possibly model-based forecasts as the finalized output of a specific sequence of
modeling choices, including parameter estimation. Consequently, we formulate our hypotheses
based on the realized quality of these forecasts rather than on the properties of the underlying
models themselves. In contrast, the framework of West (1996) assesses whether two forecasting
models based on their pseudo-true parameters would exhibit equal predictive accuracy, hence
requiring that the test statistic’s standardization explicitly accounts for the asymptotic noise
32
introduced by parameter estimation. This extension can be integrated into our framework
by redefining the population MCB and DSC components in a model-based environment and
replacing the central limit theorem in Assumption 3.3 (ii) with the asymptotic approximations
of West (1996), or with the numerous refinements and generalizations proposed in subsequent
work.
Generalizations of our framework to other elicitable functionals, such as expectiles, are
straightforward. However, functionals that are merely jointly or multi-objective elicitable fol-
lowing Fissler and Ziegel (2016) and Fissler and Hoga (2024)—such as the variance (when the
mean is unknown), the Expected Shortfall, or the systemic risk measures Co-Value-at-Risk and
Marginal Expected Shortfall—require additional care. In these cases, the underlying recali-
bration regressions necessitate the estimation of conditional models for the associated nuisance
components, such as the mean or the VaR (Patton et al., 2019; Bayer and Dimitriadis, 2022;
Dimitriadis and Hoga, 2026). The propagation of estimation risk from these nuisance quantities
into the final score decomposition represents a non-trivial extension of the current inferential
framework.
In conclusion, the contributions of this paper extend beyond their immediate applications
in econometrics and finance. The principles of score decomposition and linear recalibration
constitute a versatile toolkit for any domain where predictive accuracy is paramount, from
meteorology to machine learning. The framework proposed herein represents a fundamental
shift in perspective: from merely asking which forecast is better to diagnosing why it is better.
As the complexity of our models and the richness of our data continue to grow, this diagnostic
approach will be essential for robust and reliable decision-making.
Acknowledgments
Timo Dimitriadis acknowledges funding by the German Research Foundation (DFG) through
the projects 502572912 and 568876076. We gratefully acknowledge the Hohenheim Datalab
(DALAHO) for providing access to Refinitiv TickHistory. We thank all seminar and conference
participants who provided feedback on earlier versions of this paper. In particular, we are grate-
ful to Sam Allen, Tillman Gneiting, Alexander Jordan, Robert Jung, Fabian Krüger, Andrew
Patton, Winfried Pohlmeier, Johannes Resin, Karsten Schweikert, and Johanna Ziegel for their
valuable comments and suggestions.
33
Appendix
The appendix contains all proofs in Appendix A, a discussion of related score decompositions
in Appendix B, details on the quantile simulation DGP in Appendix C, and additional figures
for the simulations and applications in Appendix D.
A Proofs
Proof of Theorem 3.2. For item (a), notice that by the definition of the M-estimator in (12),
we have that
T T
1X 1X
C
S Wit⊤ θbiT , Yt
MCBiT = SiT − SiT =
[ b b S Xit , Yt −
T T
t=1 t=1
T T
1 X 1X
S Wit⊤ θ, Yt ≥ 0,
= S Xit , Yt − min
T θ∈Θ T
t=1 t=1
since Wit is assumed to contain an intercept and the forecast Xit . If Xit = Wit⊤ θbiT almost
[ iT = 0 is immediate.
surely, MCB
For item (b), let Xit ̸= W ⊤ θbiT for some t ∈ {1, . . . , T } with positive probability. Then, we
it
get that
T T T
1X 1X ⊤b 1X
S Wit⊤ θ, Yt ,
S Xit , Yt > S Wit θiT , Yt = min
T T θ∈Θ T
t=1 t=1 t=1
with positive probability, where the inequality holds as the latter minimum is assumed to be
unique and (0, 1, 0, . . . )⊤ ∈ Θ, for which W ⊤ (0, 1, 0, . . . )⊤ = Xit . Thus, MCB
it
[ iT > 0 with
positive probability follows.
For item (c), first notice that the empirical distribution of the sample Y1 , . . . , YT is con-
tained in P as the convex combination of Dirac measures at Y1 , . . . , YT , such that the strict
consistency of S implies that the unconditional functional rbT coincides with the M-estimator
arg min T1 Tt=1 S r, Yt . Hence, we have that
P
r∈A
T T
1X 1X
R C
S Wit⊤ θbiT , Yt
DSCiT = SiT − SiT =
d b b S rbT , Yt −
T T
t=1 t=1
T T
1 X 1X
S Wit⊤ θ, Yt ≥ 0,
= min S r, Yt − min
r∈A T θ∈Θ T
t=1 t=1
as Wit is assumed to contain an intercept, which nests the first minimization over all r ∈ A. If
rbT = W ⊤ θbiT almost surely, DSC
it
d iT = 0 is immediate.
For item (d), assume that rbT ̸= Wit⊤ θbiT for some t ∈ {1, . . . , T } with positive probability.
34
Then,
T T T T
1X 1X 1X 1X
S Wit⊤ θbiT , Yt = min S Wit⊤ θ, Yt ,
min S r, Yt = S rbT , Yt >
r∈A T T T θ∈Θ T
t=1 t=1 t=1 t=1
with positive probability, where the strict inequality holds as the latter minimum is assumed
to be unique and (brT , 0, . . . )⊤ ∈ Θ, for which W ⊤ (b
rT , 0, . . . )⊤ = rbT . Thus, DSC
it
d iT > 0 with
positive probability follows, which concludes the proof.
where
T T
1X 1X
AiT := ait := S(Xit , Yt ) − E[S(Xit , Yt )]
T T
t=1 t=1
T T
1 X 1 X
BiT := bitT := S(Wit⊤ θbiT , Yt ) − S(Wit⊤ θ̄i , Yt )
T T
t=1 t=1
T T
1 X 1 X
CiT := cit := S(Wit⊤ θ̄i , Yt ) − E[S(Wit⊤ θ̄i , Yt )]
T T
t=1 t=1
T T
1 X 1 X
DT := dtT := S(r̂T , Yt ) − S(r̄, Yt )
T T
t=1 t=1
T T
1 X 1 X
ET := et := S(r̄, Yt ) − E[S(r̄, Yt )].
T T
t=1 t=1
Here, AiT , CiT and ET capture the difference between sample and population versions of the
scores of the original forecasts, the (population) recalibrated and the (population) reference
forecasts, respectively. The terms BiT and DT capture the estimation effect of the recalibrated
and reference forecast, respectively.
√ √
We start to show that the estimation effects captured in T BiT (and similarly, in T DT )
vanish in probability. For this, we define for all θ ∈ Θ,
T T
b iT (θ) := 1 1X
X
S(Wit⊤ θ, Yt ) E S(Wit⊤ θ, Yt ) ,
Q and QiT (θ) := (28)
T T
t=1 t=1
√
such that νiT (θ) = b iT (θ) − QiT (θ) for νiT (θ) defined in (15).
T Q
Then, we employ a mean-value theorem in the third equality below to get for some (possibly
35
random) θeiT that is on the line between θbiT and θ̄i that
√ √ √
T BiT = TQb iT (θbiT ) − T Q
b iT (θ̄i )
√
= T QiT (θbiT ) − QiT (θ̄i ) + νiT (θbiT ) − νiT (θ̄i )
√ (29)
= ∇θ Q (θeiT ) · T θbiT − θ̄i + νiT (θbiT ) − νiT (θ̄i )
iT
We now provide more details on the derivation of the order terms in the last line of (29). For
the last summand, we use the stochastic equicontinuity imposed in Assumption 3.3 (iv): For
any η > 0 and ε > 0, there exists a δ > 0 such that
lim sup P νiT (θiT ) − νiT (θ̄i ) > η
b
T →∞
≤ lim sup P νiT (θbiT ) − νiT (θ̄i ) > η, ∥θbiT − θ̄i ∥ ≤ δ + lim sup P ∥θbiT − θ̄i ∥ > δ
T →∞ T →∞
!
≤ lim sup P sup νiT (θ) − νiT (θ̄i ) > η + o(1)
T →∞ {θ∈Θ: ∥θ−θ̄i ∥≤δ}
< ε,
= o(1),
where the final equality requires that we choose δ < ε/C. Hence, we get that
the conditional functional Γ(Yt | Wit ) = Wit⊤ θ̄i by Assumption 3.1 together with (3) and the
√
tower property. An equivalent reasoning shows that T DT = oP (1) (by simply setting Wit = 1
in (28) and the following.)
36
√ √
Hence, using T BiT = oP (1) and T DT = oP (1), we can write
[ 1T
MCB − MCB1 1 −1 0 0 0
√ DSC1T
d T
= 0 −1 0 0 1 T −1/2
− DSC1 X ⊤
T a1t , c1t , a2t , c2t , et + oP (1),
[ 2T
MCB − MCB2 0 0 1 −1 0
t=1
DSC
d 2T − DSC2 0 0 0 −1 1
and as
T
−1/2 −1/2
X ⊤ d
T ΩT a1t , c1t , a2t , c2t , et −→ N (0, I5 )
t=1
Hence, as Wit⊤ θ̄i = Γ(Yt | Wit ) a.s. by assumption, and the sigma-fields σ{Xit } ⊆ σ{Wit }
are nested, invoking Holzmann and Eulert (2014, Corollary 2) shows that the equality in (31)
implies that Xit = Wit⊤ θ̄i a.s..
Then, by using the definitions from (28) and arguments as in (29), we get that
XT
S Xit , Yt − S Wit⊤ θbiT , Yt
T MCBiT − MCBi =
\
t=1
T
X
S Wit⊤ θbiT , Yt − S Wit⊤ θ̄i , Yt
=−
t=1
XT
S′ (Wit⊤ θ̄i , Yt ) Wit⊤ θbiT − Wit⊤ θ̄i
=−
t=1 (32)
T
1 X ′′ 2
− S (Wit⊤ θ̄i , Yt ) Wit⊤ θbiT − Wit⊤ θ̄i
2
t=1
T
1X 3
− S′′′ (Wit⊤ θeiT , Yt ) Wit⊤ θbiT − Wit⊤ θ̄i
6
t=1
(1) (2) (3)
=: −T BiT − T BiT − T BiT ,
for some random variable θeiT between θbiT and θ̄i by a second-order Taylor expansion and a
mean-value theorem.
(3) (1) (2)
In the following, we show that T BiT = oP (1), whereas T BiT = OP (1) and T BiT = OP (1)
[ iT − MCBi .
contribute to the asymptotic distribution of T MCB
37
By Assumption 3.5 (i), it holds that
T
!−1 T
!
√ 1 X ′′ X
S (Wit⊤ θ̄i , Yt )Wit Wit⊤ T −1/2 S′ (Wit⊤ θ̄i , Yt )Wit
T θbiT − θ̄i = − + oP (1).
T
t=1 t=1
(33)
(1)
Hence, for BiT , by using (33), we get that
T √
(1)
X
−1/2
S′ (Wit⊤ θ̄i , Yt )Wit⊤ ×
−T BiT = −T T θbiT − θ̄i
t=1
T
!⊤ T
!−1
X 1 X ′′
= T −1/2 S′ (Wit⊤ θ̄i , Yt )Wit × S (Wit⊤ θ̄i , Yt )Wit Wit⊤
T
t=1 t=1
T
!
X
× T −1/2 S′ (Wit⊤ θ̄i , Yt )Wit + oP (1).
t=1
(2)
Similarly, for BiT , using (33) twice, we get that
T
(2) 1 X ′′ 2
−T BiT =− S (Wit⊤ θ̄i , Yt ) Wit⊤ θbiT − Wit⊤ θ̄i
2
t=1
T
!
1 √ b ⊤ X
′′
√
(Wit⊤ θ̄i , Yt )Wit Wit⊤
=− T (θiT − θ̄i ) × S × T (θbiT − θ̄i )
2
t=1
T
!⊤ T
!−1
1 −1/2
X
′ 1 X ′′
=− T S (Wit⊤ θ̄i , Yt )Wit × S (Wit⊤ θ̄i , Yt )Wit Wit⊤
2 T
t=1 t=1
T
!
X
× T −1/2 S′ (Wit⊤ θ̄i , Yt )Wit + oP (1).
t=1
(3)
Finally, for the absolute value of T BiT , we get that
T
(3) 1 X ′′′ 3
T BiT = S (Wit⊤ θeiT , Yt ) Wit⊤ θbiT − Wit⊤ θ̄i
6
t=1
T
1X 3
≤ S′′′ (Wit⊤ θeiT , Yt ) × Wit⊤ θbiT − Wit⊤ θ̄i
6
t=1
T √
1 −1/2 1 X ′′′
≤ T S (Wit⊤ θeiT , Yt ) ∥Wit ∥3 × ∥ T (θbiT − θ̄i )∥3
6 T
t=1
where we now provide details for the first oP (1) term. For hT (θ) :=
38
PT
1
T t=1 S′′′ (Wit⊤ θ, Yt ) ∥Wit ∥3 , we have that for any ε > 0 small enough
P T −1/2 hT (θeiT ) > ε
−1/2 −1/2
≤P T hT (θiT ) − hT (θ̄i ) > ε/2 + P T
e hT (θ̄i ) > ε/2
!
≤ P T −1/2 sup hT (θ) − hT (θ̄i ) > ε/2 + P T −1/2 hT (θ̄i ) > ε/2 + P ∥θeiT − θ̄i ∥ > δ
∥θ−θ̄i ∥≤δ
= o(1),
(34)
where the first two terms above are o(1) by Assumption 3.5 (ii) and the third is o(1) by the
consistency of θeiT .
Combining the previous three terms, we get that
T
!⊤ T
!−1
1 −1/2
X
′ 1 X ′′
(Wit⊤ θ̄i , Yt )Wit S (Wit⊤ θ̄i , Yt )Wit Wit⊤
[ iT − MCBi
T MCB = T S ×
2 T
t=1 t=1
T
!
X
× T −1/2 S′ (Wit⊤ θ̄i , Yt )Wit + oP (1).
t=1
[ iT −
Then, given a Gaussian random variable NiT ∼ N (0, ΠiT ), the quantity of interest T MCB
MCBi is approximated in distribution by
[ iT − MCBi − 1 NiT
⊤ −1
T MCB Υi NiT = oP (1),
2
Hence, as Wit⊤ θ̄i = Γ(Yt | Wit ) a.s. by assumption and r̄ = Γ(Yt ), and as the sigma-fields
σ{∅} ⊆ σ{Wit } are nested, we obtain from Holzmann and Eulert (2014, Corollary 2) that
r̄ = Wit⊤ θ̄i a.s. for all t ∈ N. Hence, we have that
T T
X X
d iT − DSCi =
T DSC rT , Yt ) − S(r̄, Yt ) +
S(b S(Wit⊤ θ̄i , Yt ) − S(Wit⊤ θbiT , Yt )
t=1 t=1
= T DT − T BiT .
A third order Taylor expansion of T DT gives that for some reT on the line between rbT and r̄,
T
X
T DT = rT , Yt ) − S(r̄, Yt )
S(b
t=1
39
T T T
X 1X 2 1 X 3
= S′ (r̄, Yt ) rbT − r̄ + S′′ (r̄, Yt ) rbT − r̄ + S′′′ (e
rT , Yt ) rbT − r̄
2 6
t=1 t=1 t=1
(1) (2) (3)
=: T DT + T DT + T DT .
(1)
!
(1) (1)
T DT
T DT − BiT = 1 −1 × (1)
T BiT
T
!
X S′ (r̄, Yt )(b
rT − r̄)
= 1 −1 ×
t=1
S′ (r̄, Yt )W ⊤ (θbiT − θ̄i )
it
T
! !
−1/2
X S′ (r̄, Y t) 0 √ rbT − r̄
= 1 −1 × T × T
t=1
0 S′ (r̄, Yt )Wit⊤ θbiT − θ̄i
T
!
X S′ (r̄, Yt ) 0
= 1 −1 × T −1/2
t=1
0 S′ (r̄, Yt )Wit⊤
P −1
1 T ′′ (r̄, Y ) T
T t=1 S t + o P (1) 0 X
× P −1 × T −1/2 S′ (r̄, Yt )Wit
1 T ′′ ⊤ ⊤
T t=1 S (Wit θ̄i , Yt )Wit Wit + oP (1) t=1
T
X
= T −1/2 S′ (r̄, Yt ) −S′ (r̄, Yt )Wit⊤
t=1
P −1
1 T ′′ (r̄, Y ) T
T t=1 S t 0 −1/2
X
× P
−1 × T
S′ (r̄, Yt )Wit + oP (1)
1 T ′′ ⊤
T t=1 S (r̄, Yt )Wit Wit t=1
T
!
X
= T −1/2 S′ (r̄, Yt )Wit⊤ (36)
t=1
T
!−1 P −1
1 T ′′
1 X S (r̄, Yt ) 0
× S′′ (r̄, Yt )Wit Wit⊤ − T t=1
T 0 0
t=1
T
!
X
−1/2
× T S′ (r̄, Yt )Wit⊤ + oP (1).
t=1
In the last step, we have used the fact that for any (column vector) a ∈ Rk and matrix B ∈ Rk×k
with first/upper-left entries a1 , B11 ∈ R, respectively, it holds that
! !!
B11 0 ⊤ B11 0
a1 −a⊤ × × a = −a × B− × a, (37)
B 0 0
where the 0 entries have the appropriate dimensions. This reformulation allows to express the
term above as a quadratic form.
40
For the second terms in (35), similar derivations yield that
(2) (2)
T DT − BiT
(2)
!
T DT
= 1 −1 × (2)
T BiT
2 !
1 X S′′ (r̄, Yt ) rbT − r̄
= 1 −1 × 2
2
t=1
S′′ (r̄, Yt ) Wit⊤ θbiT − Wit⊤ θ̄i
√ P √
1 T ′′ (r̄, Y ) ×
1 T (brT − r̄) × T t=1 S t T (b
rT − r̄)
= 1 −1 × √ ⊤ P √
2 T θbiT − θ̄i × T1 Tt=1 S′′ (r̄, Yt )Wit Wit⊤ × T θbiT − θ̄i
√ !
1 T rbT − r̄ 0
= 1 −1 × √ ⊤
2 0 T θbiT − θ̄i
1 PT ′′ (r̄, Y )
! √ !
S t 0 T r T − r̄
× T t=1 × √
b
1 PT ′′ ⊤
0 T t=1 S (r̄, Yt )Wit Wit T θbiT − θ̄i
√ !⊤ 1 PT ′′ (r̄, Y )
! √ !
1 T rbT − r̄ t=1 S t 0 T r T − r̄
= × √ × T × √
b
− T1 Tt=1 S′′ (r̄, Yt )Wit Wit⊤
2
P
T θbiT − θ̄i 0 T θbiT − θ̄i
P −1 ⊤
T 1 T ′′ (r̄, Y )
1 X T t=1 S t 0
= T −1/2 S′ (r̄, Yt )Wit⊤ × P −1
2 1 T ′′ ⊤
t=1 T t=1 S (r̄, Yt )Wit Wit
!
1 P T ′′ (r̄, Y )
S t 0
× T t=1
− T1 Tt=1 S′′ (r̄, Yt )Wit Wit⊤
P
0
P −1
1 T ′′ (r̄, Y ) T
T t=1 S t 0 −1/2
X
× P
−1 × T
S′ (r̄, Yt )Wit + oP (1)
1 T ′′ ⊤
T t=1 S (r̄, Yt )Wit Wit t=1
T
!
X
= − T −1/2 S′ (r̄, Yt )Wit⊤
t=1
T
!−1 P −1
1 T ′′
1 X S (r̄, Yt ) 0
× S′′ (r̄, Yt )Wit Wit⊤ − T t=1
T 0 0
t=1
T
!
X
−1/2
× T S′ (r̄, Yt )Wit⊤ + oP (1).
t=1
In the penultimate equality, we have plugged in the (joint) asymptotic expansion of the esti-
mators from (17). In the last equality above, we have used the fact that for a (column) vector
b = (b1 0)⊤ ∈ Rk+1 with first entry b1 ∈ R, b1 ̸= 0 and an invertible matrix B ∈ Rk×k , we have
that
! ! ! !
b−1 0 b b⊤ b1 0
b B × 1
× = b−1
1 b −Ik × = b−1
1 bb
⊤
−B = − B.
0 −B −1 B B 0 0
41
For the third term in (35), we get that
PT
′′′ r , Y )(b 3
(3)
!
DT 1 t=1 S (e T t rT − r̄)
(3) = PT
BiT 6 S′′′ (W ⊤ θ eiT , Yt )(W ⊤ θbiT − W ⊤ θ̄i )3
t=1 it it it
PT
′′′ (e 3
1 t=1 |S r T , Yt )| · |br t − r̄|
≤ PT ′′′ ⊤θ
3
6 S (W eiT , Yt ) · W ⊤ θbiT − W ⊤ θ̄i
t=1 it it it
( T
!) √ 3 !
1 −1/2 1 X |S′′′ (e rT , Yt )| 0 T ∥b
rt − r̄∥
≤ T √ 3
6 T
t=1
0 |S′′′ (θeiT , Yt )| · ∥Wit⊤ ∥3 T θbiT − θ̄i
!
OP (1)
= oP (1) ×
OP (1)
= oP (1), (38)
where the proof of the oP (1)-term follows analogously to (34). Notice for this that as Wit
contains a constant by assumption, the procedure in (34) together with Assumption 3.5 (ii) also
implies that T −1/2 T1 Tt=1 |S′′′ (e
P
rT , Yt )| = oP (1).
Hence, from the above three derivations together with (35), we obtain
T
!
1 −1/2
X
′
d iT − DSCi
T DSC = T S (r̄, Yt )Wit⊤
2
t=1
T
!−1 P −1
1 T ′′
1 X S (r̄, Yt ) 0
× S′′ (r̄, Yt )Wit Wit⊤ − T t=1
T 0 0
t=1
T
!
X
−1/2
× T S′ (r̄, Yt )Wit + oP (1),
t=1
such that the claim of the theorem follows by applying a CLT to the outer terms and a law of
large numbers to the inner one.
Proof of Proposition 3.9. Notice that ϕ(Yt ) and ϕ(Wit⊤ θ) are measurable functions of α-
mixing processes of size −s/(s − 2), s > 2, and thus itself mixing of the same size by White
(2001, Theorem 3.49) together with Assumption 3.7(ii). Next, consider the objective functions
T
1X
QiT (θ) :=
b S(Wit⊤ θ, Yt ) and
T
t=1
T
" #
h i 1 X h i
QiT (θ) := E Q b iT (θ) = E S(Wit θ, Yt ) = E S(Wit⊤ θ, Yt ) .
⊤
T
t=1
42
The first and second derivatives of the objective functions with respect to θ are
T
1 X ′′
∇θ Q
b iT (θ) := ϕ (Wit⊤ θ)Wit (Wit⊤ θ − Yt ),
T
t=1
h i h i
∇θ QiT (θ) := E ∇θ Q b iT (θ) = E ϕ′′ (Wit⊤ θ)Wit (Wit⊤ θ − Yt )
T
1 X ′′′
∇θθ QiT (θ) :=
b ϕ (Wit⊤ θ)Wit Wit⊤ (Wit⊤ θ − Yt ) + ϕ′′ (Wit⊤ θ)Wit Wit⊤ ,
T
t=1
h i h i
∇θθ QiT (θ) := E ∇θθ Q b iT (θ) = E ϕ′′′ (Wit⊤ θ)Wit Wit⊤ (Wit⊤ θ − Yt ) + ϕ′′ (Wit⊤ θ)Wit Wit⊤ .
d2 (Wit , Yt ) = sup ϕ′′′ (Wit⊤ θ)Wit Wit⊤ (Wit⊤ θ − Yt ) + ϕ′′ (Wit⊤ θ)Wit Wit⊤ ,
θ∈Θ
a.s.
with equality if and only if Wit⊤ θ̄i = Wit⊤ θ. From this we get the equivalences
a.s. a.s.
Wit⊤ θ̄i = Wit⊤ θ ⇐⇒ Wit⊤ θ̄i − θ = 0
⊤
2 ⊤ a.s.
⇐⇒ Wit θ̄i − θ = θ̄i − θ Wit Wit⊤ θ̄i − θ = 0,
which implies
⊤ h i
E Wit Wit⊤ θ̄i − θ = 0.
θ̄i − θ (40)
As per Assumption 3.7(iii), the matrix E Wit Wit⊤ is positive definite, the equality in (40) can
hold only if θ̄i −θ = 0. Hence, θ = θ̄i is the unique minimizer of QiT (θ), which verifies condition
43
(i) of Newey and McFadden (1994, Theorem 2.1).
Condition (ii) of Newey and McFadden (1994, Theorem 2.1), i.e., compactness of Θ, holds
by imposing Assumption 3.1 and condition (iii) requires continuity of the objective function
QiT (θ) which is fulfilled as per Assumption 3.7(i).
b iT (θ) − QiT (θ) =
Condition (iv) of Newey and McFadden (1994, Theorem 2.1) requires supθ∈Θ Q
oP (1) which follows from Davidson (1994, Theorem 21.9) given that
b iT (θ) − QiT (θ) = oP (1) for all θ ∈ Θ, and
(a) Q
(b) Q b iT (·) − QiT (·) is stochastically equicontinuous.
We first prove the point-wise law of large numbers for the objective function in (a). From
White (2001, Theorem 3.49), it follows that S(Wit⊤ θ, Yt ) is stationary and α-mixing of size
−s/(s − 2), s > 2, by Assumptions 3.1 and 3.7(ii). Further, as of White (2001, Proposition
3.44), S(Wit⊤ θ, Yt ) is ergodic because it is a measurable function of (Yt , Wit⊤ ). Since
the ergodicity Theorem 3.34 in White (2001) is verified and the pointwise law of large numbers
in (a) follows.
Concerning condition (b) of Davidson (1994, Theorem 21.9), we show that
T
1X ⊤
S Wit θ, Yt − S Wit⊤ θ,
e Yt ≤ BT θ − θe for all θ, θe ∈ Θ,
T
t=1
By the mean value theorem, which can be applied due to convexity of Θ as per Assumption 3.1,
together with triangle inequality, we get for some mean value θ ∗ on the line joining θe and θ,
that
T T T
1X ⊤ 1X
e Yt = 1
X
S Wit θ, Yt − S Wit⊤ θ, ϕ′′ (Wit⊤ θ ∗ )Wit⊤ (Yt − Wit⊤ θ ∗ ) θ − θe
T T T
t=1 t=1 t=1
T
1 X
≤ ϕ′′ (Wit⊤ θ ∗ ) · ∥Wit ∥ · Wit⊤ θ ∗ − Yt · θ − θe
T
t=1
T
1 X
≤ ϕ′′ (Wit⊤ θ ∗ ) · ∥Wit ∥ · (∥Wit ∥∆θ + |Yt |) · θ − θe
T
t=1
≤ BT θ − θe ,
(43)
44
where
T
1X
BT := sup ϕ′′ (Wit⊤ θ) ∥Wit ∥ · (∥Wit ∥∆θ + |Yt |) . (44)
T θ∈Θ t=1
where the existence of the involved expectations is implied by Assumption 3.7(v). It follows
that Qb iT (·) − QiT (·) is stochastically equicontinuous. Hence, we can apply Theorem 2.1 of
P
Newey and McFadden (1994) to conclude that the M-estimator is consistent, i.e. θbiT −→ θ̄i .
Secondly, we check conditions (i)–(v) of Newey and McFadden (1994, Theorem 3.1). Since
θ̄i ∈ int(Θ) by Assumption 3.1, condition (i) holds. By Assumption 3.7(i), the objective function
Qb iT (θ) is twice continuously differentiable, ensuring that condition (ii) is fulfilled. To verify
√
condition (iii), we have to show that a central limit theorem holds for T ∇θ Q b iT (θ̄i ). By the
Cramér–Wold device, this reduces to show univariate convergence, i.e.,
T
−1/2 d
X
T −1/2 λ⊤ VT ϕ′′ Wit⊤ θ̄i Wit Yt − Wit⊤ θ̄i −→ N (0, 1), T → ∞, (46)
t=1
for any λ ∈ Rk satisfying λ⊤ λ = 1. We show this by verifying the required conditions of the
CLT in White (2001, Theorem 5.20). Note that by White (2001,
n Theorem 3.49) and Assumption o
−1/2
3.7(ii), the sequences ϕ Wit θ̄i Wit Yt − Wit θ̄i and λ VT ϕ′′ Wit⊤ θ̄i Wit Yt − Wit⊤ θ̄i
′′ ⊤
⊤
⊤
are α-mixing of size −s/(s − 2), s > 2. Since θ = θ̄i uniquely minimizes the population objective
(see the discussion around (40)) the first-order condition,
T
" #
1 X ′′
∇θ QiT (θ̄i ) = E ϕ (Wit⊤ θ̄i )Wit (Wit⊤ θ̄i − Yt )
T
t=1
(47)
T
1 X h i h i
= E ϕ′′ (Wit⊤ θ̄i )Wit Wit⊤ θ̄i − E ϕ′′ (Wit⊤ θ̄i )Wit Yt = 0,
T
t=1
holds by the law of iterated expectations, E ϕ′′ (Wit⊤ θ̄i )Wit Yt = E ϕ′′ (Wit⊤ θ̄i )Wit E [Yt | Wit ] =
E ϕ′′ (Wit⊤ θ̄ih)Wit Wit⊤ θ̄i , together with Assumptioni 3.1. This first-order condition further im-
−1/2
plies that E λ⊤ VT ϕ′′ Wit⊤ θ̄i Wit Yt − Wit⊤ θ̄i = 0. By sub-multiplicativity of norms and
is finite for s > 2, ∀t. Therefore, given Assumption 3.7(iv), the univariate convergence result in
(46) follows. Further, by the Cramér–Wold device, the conditions of the univariate central limit
45
theorem for mixing series hold for all linear combinations, leading to the desired result
T √
−1/2 −1/2 d
X
T −1/2 VT ϕ′′ Wit⊤ θ̄i Wit Yt − Wit⊤ θ̄i = VT T ∇θ Q
b iT (θ̄i ) −→ N (0, Ik ). (48)
t=1
For the fourth condition of Newey and McFadden (1994, Theorem 3.1), we use Newey and
McFadden (1994, Lemma 2.4) to show uniform convergence in probability for the Hessian, i.e.
P
sup ∇θθ Q
b iT (θ) − ∇θθ QiT (θ) −→ 0. (49)
θ∈Θ
Since we consider stationary random variables, a compact Θ by Assumption 3.1, and the objec-
tive function is continuous for all θ ∈ Θ by Assumption 3.7(i), we are left to show that there
exists a dominating function d3 (Wit , Yt ) ≥ ∇θθ Q
b iT (θ) for all θ ∈ Θ, with E [d3 (Wit , Yt )] < ∞,
which holds for
d3 (Wit , Yt ) := sup ϕ′′′ Wit⊤ θ · ∥Wit ∥2 · Wit⊤ θ − Yt + sup ϕ′′ Wit⊤ θ · ∥Wit ∥2 , (50)
θ∈Θ θ∈Θ
whose expectation is finite by Assumption 3.7(v). Hence, the desired convergence result in (49)
follows.
Lastly, condition (v) of Newey and McFadden (1994, Theorem 3.1) requires a non-singular
expected Hessian, ∇θθ QiT (θ̄i ) = E ϕ′′′ (Wit⊤ θ̄i )Wit Wit⊤ (Wit⊤ θ̄i − Yt ) + ϕ′′ (Wit⊤ θ̄i )Wit Wit⊤ .
Notice that by Assumption 3.1 and the law of iterated expectations, we have that
E ϕ′′′ (Wit⊤ θ̄i )Wit Wit⊤ (Wit⊤ θ̄i − Yt ) = E ϕ′′′ (Wit⊤ θ̄i )Wit Wit⊤ Wit⊤ θ̄i − E[Yt | Wit ] = 0, and
Assumption 3.7(i) assures ϕ′′ (Wit⊤ θ̄i ) > 0. This, together with the positive definiteness of
E Wit Wit⊤ , allows us to apply Dimitriadis and Bayer (2019, Lemma B.2), which establishes
that the matrix E ϕ′′ (Wit⊤ θ̄i )Wit Wit⊤ is positive definite as well. Thus, ∇θθ QiT (θ̄i ) is non-
singular.
We conclude the verification of Assumption 3.3(i) by noting that the above results generalize
to the case where the reference forecast is considered, which corresponds to the setting Wit = 1.
We continue to verify the high-level Assumption 3.3(ii), which invokes a CLT for the vector
of scores T −1/2 Tt=1 a1t , c1t , a2t , c2t , et . For this, we use the Cramér-Wold Device under which
P
T
−1/2
X ⊤ d
T −1/2 λ⊤ ΩT a1t , c1t , a2t , c2t , et −→ N (0, 1), (52)
t=1
46
theorem for mixing variables in White (2001, Theorem 5.20). To do so, note that ait , cit , et ,
i = 1, 2, are defined as difference in scores between the sample hand population values. Hence
⊤ i
these terms are centered around zero by construction and thus E λ⊤ a1t , c1t , a2t , c2t , et = 0.
It further holds that
−1/2 ⊤ s −1/2 s s
E λ⊤ ΩT ≤ λ⊤ ΩT
a1t , c1t , a2t , c2t , et E a1t , c1t , a2t , c2t , et < ∞, (53)
which validates the required moment conditions for White (2001, Theorem 5.20).
−1/2 s
Since λ⊤ ΩT is finite by Assumption 3.8, we particularly analyze the vector
(a1t , c1t , a2t , c2t , et ) and note that by linearity of expectations
⊤ s
E a1t , c1t , a2t , c2t , et = E|a1t |s + E|c1t |s + E|a2t |s + E|c2t |s + E|et |s < ∞, (54)
s
provided that each E|ait |s , E|cit |s , E|et |s < ∞. To verify this, we detail the bound for E|cit |s and
show that it is finite by applying the cs -inequality:
s
E|cit |s = E S(Wit⊤ θ, Yt ) − E[S(Wit⊤ θ, Yt )]
h i h i s
≤ 2s−1 E |S(Wit⊤ θ, Yt )|s + E S(Wit⊤ θ, Yt ) < ∞. (55)
Finiteness of the sum in (55) follows from the moment conditions in Assumptions 3.7(v) by
again applying the cs -inequality, i.e.,
s h si h s i
E S(Wit⊤ θ, Yt ) ≤ 32s−1 E [|ϕ(Yt )|s ] + E ϕ(Wit⊤ θ) + E ϕ′ (Wit⊤ θ) Yt − Wit⊤ θ < ∞.
The same reasoning applies to E|ait |s and E|et |s . Since all norms on a finite-dimensional space
are equivalent, there exists a constant K > 0 such that, for the vector (a1t , c1t , a2t , c2t , et ), the
⊤
Euclidean norm is at most K times the Ls -norm, i.e., we can write a1t , c1t , a2t , c2t , et ≤
⊤
K a1t , c1t , a2t , c2t , et . Raising this expression to the s-th power, taking expectations, and
s
using (54), we get that
⊤ s ⊤ s h i
E a1t , c1t , a2t , c2t , et ≤ K sE a1t , c1t , a2t , c2t , et ≤ K s E|a1t |s + . . . + E|et |s < ∞,
s
because K s is a constant. Hence, the moment condition of White (2001, Theorem 5.20) is
fulfilled and the univariate convergence result in (52) follows, given Assumption 3.8. Further,
by the Cramér–Wold device, the conditions of the univariate central limit theorem for mixing
series hold for all linear combinations, leading to the desired result,
T
−1/2
X ⊤ d
T −1/2 ΩT a1t , c1t , a2t , c2t , et −→ N (0, I5 ). (56)
t=1
For the verification of the high-level Assumption 3.3(iii), we notice that by Assumption
3.7(i), ϕ(Wit⊤ θ) is continuously differentiable in θ, implying that S(Wit⊤ θ, Yt ) is also continuously
47
differentiable in θ with derivative
∥ϕ′′ (Wit⊤ θ)Wit (Wit⊤ θ − Yt )∥ ≤ sup ∥ϕ′′ (Wit⊤ θ)Wit (Wit⊤ θ − Yt )∥,
θ∈Θ
whose expectation is finite by Assumption 3.7(v). Hence, differentiability carries over to the
expectation by the Leibniz integral rule. Moreover, under the same assumptions and reasoning,
both the first- and second-order derivatives of the expected score with respect to θ remain
continuously differentiable.
For the verification of high-level Assumption 3.3(iv), we can write
T
1 X
∇θ νiT (θ) = √ Sit , with Sit := Wit S′ (Wit⊤ θ, Yt ) − E[Wit S′ (Wit⊤ θ, Yt )],
T t=1
with
such that Sit is a process centered around zero. Stationarity in Assumption 3.1 implies that
T 2 T 2
1 X 1 X
E √ Sit = E Sit
T t=1 T
t=1
T T
1 XX h i
= E Sit⊤ Siu (57)
T
t=1 u=1
T
1X h ⊤ i 2 X
= E Sit Sit + E[Sit⊤ Siu ],
T T
t=1 1≤t<u≤T
where the second equality follows from the definition ∥Sit ∥2 = Sit⊤ Sit which is then split into
a diagonal (t = u) and an off–diagonal (t ̸= u) part in the final step. Concerning the diagonal
part, Assumption 3.1 implies E ∥Sit ∥2 = E ∥S1t ∥2 , which is finite by the moment bounds
T T
1X h ⊤ i 1X h i h i
sup E Sit Sit = sup E ∥Sit ∥2 = E ∥Si1 ∥2 < ∞. (58)
T ∈N T t=1 T ∈N T t=1
Concerning the off-diagonal part, observe that the summand depends only on the lag ℓ = u−t.
Using stationarity, as assumed in Assumption 3.1, we can rewrite the expression as
T −1 ∞
2 X 2 X X h i
E[Sit⊤ Siu ] = ⊤
(T − ℓ)E[Si1 Si,1+ℓ ] ≤ 2 ⊤
E Si1 Si,1+ℓ ,
T T
1≤t<u≤T ℓ=1 ℓ=1
where the inequality follows by taking absolute values and noting the simple fact that (T −ℓ)/T <
48
1 for all ℓ = 1, . . . , T − 1.
We now show that ∞
P ⊤
ℓ=1 E Si1 Si,1+ℓ < ∞ through the (component-wise) covariance in-
equality for a α-mixing sequence given in Corollary 14.3 of Davidson (1994). Assumption 3.7(ii)
assures Sit to be α-mixing of size −s/(s − 2) for some s > 2, and Assumption 3.7(v) give
E [∥Sit ∥s ] < ∞ for the same s. In Davidson’s notation, we chose p = r = s > 2, which always
(j)
satisfies the required condition r ≥ p/(p − 1). We denote Sit to be the j-th, j = 1, . . . , k,
component of Sit , to get by the aforementioned corollary that, for every lag ℓ ≥ 1,
k k k
⊤ X (j) (j) X (j) (j)
X (j) 2
≤ 6α(ℓ) 1−2/s
E Si1 Si,1+ℓ = E Si1 Si,1+ℓ ≤ Cov Si1 , Si,1+ℓ E ∥Si1 ∥s ,
j=1 j=1 j=1
(60)
where we factor out 6α(ℓ) 1−2/s in the last step since the decay of the mixing coefficients applies
(j)
uniformly to each component Sit .
To see that the mixing coefficients are sufficiently fast decaying, recall that
per definition
s
− s−2 −ϵ
(White, 2001, Definition 3.45) there exists some ϵ > 0, such that α(ℓ) = O ℓ . From
s
this, we get α(ℓ)1−2/s = O ℓ−( s−2 +ϵ)(1−2/s) = O ℓ−[1+ϵ(1−2/s)] and given that s > 2 by
Assumption 3.7(ii), we obtain 1 + ϵ(1 − 2/s) > 1. Hence, the sum ∞ −[1+ϵ(1−2/s)] converges,
P
ℓ=1 ℓ
P∞
which implies that ℓ=1 α(ℓ)1−2/s < ∞. Therefore,
∞ k ∞
X ⊤ X (j) 2 X
α(ℓ) 1−2/s < ∞,
E Si1 Si,1+ℓ ≤ 6 E ∥Si1 ∥s (61)
ℓ=1 j=1 ℓ=1
P∞ ⊤
which does not involve T and it immediately follows that also supT ∈N ℓ=1 E Si1 Si,1+ℓ < ∞.
Using (57) together with (58) and (61) gives
T 2
1 X h i
sup E √ Sit = sup E ∥∇θ νiT (θ)∥2 < ∞,
T ∈N T t=1 T ∈N
which, according to Davidson (1994, Theorem 21.10) and the discussion below that theorem,
implies that the empirical process νiT (θ) = T −1/2 Tt=1 S(Wit⊤ θ, Yt )−E S(Wit⊤ θ, Yt ) is stochas-
P
tically equicontinuous.
Hence, we have shown that Assumption 3.3 holds such that Theorem 3.4 applies under the
given conditions. We now continue to verify the high-level conditions in Assumption 3.5.
49
For the verification of the high-level Assumption 3.5(ii), recall that
T
1 X ′′′ ⊤
hT (θ) = S Wit θ, Yt ∥Wit ∥3
T
t=1
T
1 X
= −ϕ′′′ (Wit⊤ θ) + ϕ′′′′ (Wit⊤ θ)(Wit θ − Yt ) · ∥Wit ∥3 (62)
T
t=1
T
1 X
=: gt (θ),
T
t=1
Concerning the first part of the high level Assumption 3.5(ii), we use that hT (θ) is Lipschitz-
continuous with (random) Lipschitz constant L = supθ∈Θ ∥h′T (θ)∥ as ϕ is assumed to be five
times differentiable with bounded derivatives; see Assumption 3.7(v). Therefore, we get that
To establish that T −1/2 Lδ = oP (1), it suffices to verify a uniform law of large numbers for
∥h′T (θ)∥. By taking the supremum of ∥gt′ (θ)∥, we obtain the dominating function
d4 Wit⊤ θ, Yt := sup ϕ′′′′′ Wit⊤ θ · Wit⊤ θ − Yt · ∥Wit ∥4 ,
θ∈Θ
Assumption 3.1 the data is stationary and the parameter space Θ is compact. As the function
∥gt′ (θ)∥ is continuous for all θ ∈ Θ by Assumption 3.7(i), we can apply the uniform law of large
numbers in Newey and McFadden (1994, Lemma 2.4) to get that
T
1X ′
∥gt (θ)∥ − E ∥gt′ (θ)∥ −→ 0.
P
sup (64)
θ∈Θ T t=1
50
inequality to gt (θ̄i ) which, together with Assumption 3.7(v) results in
h i h i
E [gt (θ)] ≤ E ϕ′′′ Wit θ̄i ∥Wit ∥3 + E ϕ′′′′ Wit θ̄i · Wit θ̄i − Yt · ∥Wit ∥3 < ∞.
(65)
Given (65), Theorem 3.34 in White (2001) implies the point-wise law of large numbers and we
have that hT (θ̄i ) = E gt (θ̄i ) + oP (1). Hence, T −1/2 hT (θ̄i ) = oP (1), as T → ∞.
For the verification of the high-level Assumption 3.5(i), we not that the true parameter
θ̄i lies in int(Θ) such that consistency ensures that θbiT is in the interior of the parameter
space with probability approaching one as T → ∞. This implies that the first-order condition
0 = ∇θ Q
b iT (θbiT ) holds with probability approaching one as T → ∞.
Based on this, we apply the mean-value theorem to get that
0 = ∇θ Q
b iT θ̄i + ∇θθ Q
b iT θeiT θbiT − θ̄i , (66)
where θeiT is a mean value on the line joining θbiT and θ̄i . Note, by the mean value theorem, it
holds that θeiT − θ̄i ≤ θbiT − θ̄i . Thus consistency of the estimator and continuous mapping
h i−1 h i−1
implies ∇θθ Q b iT θeiT = ∇θθ Q b iT θ̄i + oP (1). Therefore, we get that
√ √
h
i−1
T θiT − θ̄i = − ∇θθ QiT θ̄i
b b + oP (1) T ∇θ Q
b iT θ̄i
(67)
h i−1 √
= − ∇θθ Q
b iT θ̄i T ∇θ Q
b iT θ̄i + oP (1).
√
Since the central limit theorem applies to T ∇θ Q
b iT (θ̄i ), as shown in (48), and the law of large
numbers for the Hessian (established in (49)) ensures convergence, it follows that
−1/2
√
h
i−1 h i−1
d
∇θθ Q
b iT θ̄i VT ∇θθ Q
b iT θ̄i T θbiT − θ̄i −→ N (0, Ik ).
Recall that Equation (67) establishes the asymptotic behavior of an individual es-
timator and note that under Assumptions 3.1, 3.7(ii), and 3.7(v), we get that
1 PT ′′′ ⊤ ⊤ ⊤ P
T t=1 ϕ (Wit θ̄i )Wit Wit (Wit θ̄i −Yt ) −→ 0, which allows us to simplify the empirical Hessian
at the true parameter, i.e.,
T
b iT (θ̄i ) = 1
X
∇θθ Q ϕ′′′ (Wit⊤ θ̄i )Wit Wit⊤ (Wit⊤ θ̄i − Yt ) + ϕ′′ (Wit⊤ θ̄i )Wit Wit⊤
T
t=1
T
1 X
= ϕ′′ (Wit⊤ θ̄i )Wit Wit⊤ + oP (1).
T
t=1
Similarly the asymptotic behavior of the unconditional expectation r̄ follows for Wit = 1, i.e.,
1 PT ′′′ P
T t=1 ϕ (r̄)(r̄ − Yt ) −→ 0. Based on this, we stack the estimators emerging from (67) to get
51
that
P −1
! 1 T ′′ (r̄) + o (1)
√ rbT − r̄ T t=1 ϕ P 0
T = − P −1 ×
θiT − θ̄i 1 T ′′ ⊤ ⊤
t=1 ϕ (Wit θ̄i )Wit Wit + oP (1)
b
T (68)
T
X
T −1/2 ϕ′′ (Wit⊤ θ̄i )Wit (Wit⊤ θ̄i − Yt ) + oP (1).
t=1
The law of large numbers established in (49) ensures convergence of the first term in (68) and
invertibility follows from Assumption 3.7(iii) for T large enough. Asymptotic normality of the
second term in (68) follows from the central limit theorem in (48). This verifies Assumption
3.5(i) and hence concludes this proof.
Proof of Proposition 3.11. We check the conditions of Assumption 3.3, for which we use
T
b iT (θ) := 1
X
Q S(Wit⊤ θ, Yt ) and
T
t=1
T
" #
h i 1 X h i
QiT (θ) := E Q b iT (θ) = E S(Wit⊤ θ, Yt ) = E S(Wit⊤ θ, Yt ) ,
T
t=1
with
While the scoring function S itself is non-smooth (at Yt = Wit⊤ θ), we show that its expectation
is twice (as required below) continuously differentiable everywhere, which implies condition (iii)
of Assumption 3.3. For this, we use the law of iterated expectations and write
h i
QiT (θ) = E E (1{Yt ≤ Wit⊤ θ} − α)(g(Wit⊤ θ) − g(Yt )) | Wit
"Z ⊤ #
h i Wit θ
= E Fit (Wit⊤ θ) − α g(Wit⊤ θ) − E
g(y)fit (y)dy + αE[g(Yt )]. (69)
−∞
not depend on θ. For thehR middle term, as gi and fit are continuous, the fundamental theorem
⊤θ
Wit
of calculus implies ∇θ E −∞ g(y)fit (y)dy = E Wit g(Wit⊤ θ)fit (Wit⊤ θ) . Hence, we obtain
differentiability such that condition (iii) of Assumption 3.3 is satisfied and get that
h i
∇θ QiT (θ) = E Wit g ′ (Wit⊤ θ) Fit (Wit⊤ θ) − α . (70)
52
For the second derivative, which will be used below, we similarly get that
h n oi
∇θθ QiT (θ) = E (Wit Wit⊤ ) g ′′ (Wit⊤ θ) Fit (Wit⊤ θ) − α + g ′ (Wit⊤ θ)fit (Wit⊤ θ) ,
We next show condition (iv) of Assumption 3.3, i.e., stochastic equicontinuity of the empirical
process νiT (θ) from (15). For this, we first show the following lemma:
Lemma A.1. Given that the function g is Lipschitz with constant Lg , the GPL loss (7) is
Lipschitz with constant Lg max(α, 1 − α).
Proof of Lemma A.1. We start to write the function S from (7) as S(x, y) = ρ(g(x), g(y)) with
α(v − u), u ≤ v,
ρ(u, v) = (72)
(1 − α)(u − v), u > v,
which is a piecewise linear function with slopes α and 1 − α, respectively. Hence, we get that for
all u1 , u2 , v ∈ R that ρ(u1 , v) − ρ(u2 , v) ≤ max(α, 1 − α) u1 − u2 and as a direct consequence,
Lemma A.2. Assume that the function g from (7) is Lipschitz continuous, (Yt , Wit⊤ ) is α-
s
mixing of size −s/(s − 2) for some s > k + δ for some δ > 0, for which E Wit < ∞ and
s
E g(Wit⊤ θ) − g(Yt ) < ∞. Then, the empirical process νiT (θ) from (15) based on the GPL loss
S from (7) is stochastically equicontinuous; see (16).
Proof of Lemma A.2. Using Lemma A.1, we get for all θ, τ ∈ Θ that
is Lipschitz continuous. Hence, we employ Hansen (1996, Theorem 3) for mixing arrays, for
which we verify his Assumption 4. For this, we choose s > q = k̃ = k + δ for some small δ > 0,
where k is the dimensionality of Θ, and the function space F := S(Wit⊤ θ, Yt ) : θ ∈ Θ .
First, by the mixing condition from Assumption 3.10 with mixing coefficients αm , we have
P∞ (s−2)/s
that m=1 αm < ∞. As s− k̃s
k̃
< s−2
s for any s > 2 and k̃ > 2, we also have that
P∞ (s−k̃)/(k̃s)
m=1 αm < ∞, such that the first condition Hansen (1996, Assumption 4) holds.
53
The second condition of Hansen (1996, Assumption 4) holds as we impose stationarity and
s
E g(Wit⊤ θ) − g(Yt ) < ∞ such that for any S(Wit⊤ θ, Yt ) ∈ F ,
( T
)1/2
1 X s 2/s s 1/s
lim sup E S(Wit⊤ θ, Yt ) = E S(Wit⊤ θ, Yt ) < ∞.
n→∞ T
t=1
s
as E Wit < ∞ by assumption.
Hence, we can employ Hansen (1996, Theorem 3) to conclude that their Condition 1 holds
q 1/q
with the Lq -norm ρq (·) = E · . To see that the formulation of stochastic equicontinuity
in Hansen (1996, Condition 1) implies (16), notice that
n o n o
f1 , f2 ∈ F : ρq (f1 − f2 ) < δ = θ, τ ∈ Θ : ρq S(Wit⊤ θ, Yt ) − S(Wit⊤ τ , Yt ) < δ
n o
q
= θ, τ ∈ Θ : E S(Wit⊤ θ, Yt ) − S(Wit⊤ τ , Yt ) < δ q
δq
q
⊇ θ, τ ∈ Θ : θ − τ < q ,
Lg max(α, 1 − α)q E∥Wit ∥q
such that the supremum over the the left-hand side (in Hansen (1996)) is larger or equal than
the supremum over the right-hand side (in (16)). Furthermore, Markov’s inequality can be
used to transfer the Lq -statement in Hansen (1996) to the probability statement in (16), which
concludes this proof.
We continue to show consistency of θbiT through Newey and McFadden (1994, Theorem 2.1),
which is required for the verification of Assumption 3.3 (i). Notice that Θ is compact by
assumption and that the population objective QiT (θ) is continuous on Θ as shown above. As
the GPL scoring function S from (7) is strictly consistent for the α-quantile, we employ Holzmann
and Eulert (2014, Theorem 1) to conclude that
h i h i
E S(Wit⊤ θ̄i , Yt ) ≤ E S(Wit⊤ θ, Yt ) , (73)
with equality if and only if Wit⊤ θ = Wit⊤ θ̄i , which is equivalent to θ = θ̄i as shown around (40).
Thus, the objective function θ 7→ QiT (θ) = E S(Wit⊤ θ, Yt ) is uniquely minimized by θ̄i .
It remains to show the ULLN, supθ∈Θ Q b iT (θ) − QiT (θ) = oP (1) through Davidson (1994,
b iT (θ) − QiT (θ) = oP (1) for all θ ∈ Θ, follows by
Theorem 21.9). For this, the pointwise LLN, Q
White (2001, Theorem 3.34) as S(Wit⊤ θ, Yt ) is ergodic by White (2001, Proposition 3.44) and as
E S(Wit⊤ θ, Yt ) ≤ E g(Wit⊤ θ) + E g(Yt ) < ∞ by assumption. Furthermore, the second condi-
tion of Davidson (1994, Theorem 21.9), that Q b iT (θ)−QiT (θ) is stochastically equicontinuous,
54
follows immediately as
T
X
νiT (θ) = T −1/2 S(Wit⊤ θ, Yt ) − E S(Wit⊤ θ, Yt ) = T 1/2 Q
b iT (θ) − QiT (θ)
t=1
√ T
X
b iT (θ̄i ) = T −1/2 Wit g ′ (Wit⊤ θ̄i ) 1{Yt ≤ Wit⊤ θ̄i } − α ,
TD (74)
t=1
are mixing of the same size as the underlying random variables. Furthermore, for s > 2, we have
h si h si h i
E g ′ (Wit⊤ θ̄i )λ⊤ Wit 1{Yt ≤ Wit⊤ θ̄i } − α ≤ E g ′ (Wit⊤ θ̄i )λ⊤ Wit ≤ E ∥g ′ (Wit⊤ θ̄i )Wit ∥s ,
T
!
X
λ⊤ VT λ = Var T −1/2 g ′ (Wit⊤ θ̄i ) 1{Yt ≤ Wit⊤ θ̄i } − α λ⊤ Wit
t=1
is uniformly bounded away from zero by assumption, White (2001, Theorem 5.20) together with
the Cramér-Wold device yields that
√ −1/2 d
T ViT b iT (θ̄i ) −→
D N (0, Ik ). (75)
Finally, item (v) of Newey and McFadden (1994, Theorem 7.1) is implied by stochastic
equicontinuity of the process νiT shown in Lemma A.2; see the discussion on page 2187 of
Newey and McFadden (1994).
Hence, we can conclude that
55
√ √
which particularly implies T θbiT − θ̄i = OP (1). Invoking Wit = 1, this also implies T rbT −
r̄ = OP (1) such that condition (i) of Assumption 3.3 holds.
The high-level Assumption 3.3(ii) holds by the same arguments following (52), applied to
the GPL loss function and using the respective moment conditions from Assumption 3.10.
Proof of Proposition 4.1. Consider the squared error scoring function and notice that
h i
MCBi = E S(Xit , Yt ) − S(E[Yt | Wit ], Yt )
h i
= E (Xit − Yt )2 − (E[Yt | Wit ] − Yt )2
(76)
= E[Xit2 ] − E[E[Yt | Wit ]2 ] − 2E[Xit Yt ] + 2E[Yt E[Yt | Wit ]]
= E[Xit2 ] + E[E[Yt | Wit ]2 ] − 2E[Xit Yt ],
and
h i h i
DSCi = E S(E[Yt ], Yt ) − E S(E[Yt | Wit ], Yt )
h i
= E (E[Yt ] − Yt )2 − (E[Yt | Wit ] − Yt )2
h i h i (77)
= −E E[Yt | Wit ]2 + 2E Yt E[Yt | Wit ]
h i
= E E[Yt | Wit ]2 ,
where the final equality follows from the law of iterated expectations.
1
Recall that ς = Var(Mt ) = Var(Kt ) = Var(Lt ) = Var(Nt ) = . As the processes
1 − β2
Mt , Kt , Lt , Nt are independent, it also follows that E[Mt Kt ] = E[Mt Lt ] = E[Mt Nt ] = E[Kt Lt ] =
E[Kt Nt ] = E[Lt Nt ] = 0. Hence, for forecaster i = 1, we have E[X1t ] = δ0 and
2
E[X1t ] = δ02 + (δM
2 2
+ δK 2
+ δN )ς = δ02 + ς · δ ⊤ δ, and (78)
E[X1t Yt ] = (δK γK + γM δM + γN δN )ς = ς · δ ⊤ γ. (79)
2
E[X2t ] = ξ02 + (ξM
2
+ ξL2 + ξN
2
)ς = ξ02 + ς · ξ⊤ ξ, and (80)
E[X2t Yt ] = (ξM γM + γL ξL + ξN γN )ς = ς · ξ ⊤ γ. (81)
For the term E E[Yt | Wit ]2 , with Wit = (1, Xit )⊤ , we use E[Yt | Wit ] = E[Yt | Xit ]. Given
with
56
Using the expressions (78)–(81), we get
" 2 #
δ⊤γ (δ ⊤ γ)2
E E[Yt | W1t ]2 = E 2
X1t =ς· , (82)
δ⊤δ δ⊤δ
" 2 #
ξ⊤ γ (ξ ⊤ γ)2
E E[Yt | W2t ]2 = E 2
X2t =ς· . (83)
ξ⊤ ξ ξ⊤ ξ
Hence, the expression for MCB1 from (25) follows by substituting (78), (79) and (82) into (76)
and by using the identity a + b2 /a − 2b = (a − b)2 /a with a = δ ⊤ δ and b = δ ⊤ γ. The formula
for forecaster X2t follows analogously.
For the uncertainty component under squared error loss UNC = E (E[Yt ] − Yt )2 = E[Yt2 ]
since E[Yt ] = 0. Since εY,t ∼ N (0, 1) and Var(εY,t ) = E[ε2Y,t ] = 1, the uncertainty component is
given by
UNC = E[Yt2 ]
2
= γK E[Kt2 ] + γL2 E[L2t ] + γM
2
E[Mt2 ] + γN
2
E[Nt2 ] + E[ε2Y,t ]
2
= (γK + γL2 + γM
2 2
+ γN )ς + 1
= ς · γ ⊤ γ + 1,
resolution, and uncertainty. This representation is prevalent in the literature and it is often
applied to the Brier (1950) score, which describes the squared error in the context of probability
forecasts for binary outcomes; see e.g., Elliott and Timmermann (2016) and Pohle (2020). As
such, our paper augments the original Murphy (1973) decomposition with inference under linear
recalibration.
Second, the squared error score decomposition of Mincer and Zarnowitz (1969, Equation 5a)
reads as
In terms of our notation, and the refinement of unconditional and conditional miscalibration
terms of Gneiting and Resin (2023), the components in (84) are Si = uMCBi + cMCBi +
UNC + DSCi , where uMCBi and cMCBi are denoted as unconditional and conditional mis-
calibration components, which are non-negative and satisfy MCBi = uMCBi + cMCBi given that
57
they are obtained from
C
uMCBi = Si − E[S(Xit + c, Yt )] and cMCBi = E [S(Xit + c, Yt )] − Si ,
with constant c chosen such that Xit + c is unconditionally calibrated, i.e., E[Xit + c] = E[Yt ].
In this sense, the score decomposition in Mincer and Zarnowitz (1969) closely mirrors the
structure considered here. Although the residual component in (84) conflates uncertainty and
discrimination, this is unproblematic for comparative evaluation, as the uncertainty term is
forecast-invariant and cancels out in pairwise comparisons. Accordingly, the residual component
can be viewed as a re-normalized version of discrimination. Mincer and Zarnowitz (1969) however
place little emphasis on this residual component, referring to it merely as a residual term without
further interpretation or diagnostic use.
In summary, the aforementioned decompositions can be viewed as special cases of the general
score decomposition in (11). The asymptotic theory developed in this paper is sufficiently general
to be readily applicable to these alternative representations. For example, in the case of the
Mincer and Zarnowitz (1969) decomposition, inference for the combined term UNC + DSCi
follows directly from our results, and inference for uMCBi and cMCBi only requires one additional
unconditional estimation step for the constant c.
Proposition C.1. Consider the DGP in (22) and (26). For α ∈ (0, 1), define zα = Φ−1 (α).
Using Wit = σ{Wit } with Wit = (1, Xitα )⊤ and the check loss Sα (x, y) = (1{x ≥ y} − α)(x − y),
we get
MCB1 = m1 Φ(κ1 ) + α − 1 + s1 ϕ(κ1 ) − σ1|X ϕ(zα ), (85)
DSC1 = ϕ(zα ) σY − σ1|X , (86)
2 α (δ ⊤ γ)2
σ1|X := Var(Yt | X1t ) = σY2 − ς
δ⊤δ
m1 := E[Yt − Xitα ] = −(δ0 + zα ),
s21 := Var(Yt − Xitα ) = 1 + ς γ ⊤ γ + δ ⊤ δ − 2δ ⊤ γ ,
m1
κ1 := .
s1
(δ ⊤ γ)2 (δ ⊤ γ)2
It further holds that DSC1 is increasing with δ⊤ δ
and MCB1 is increasing in s1 and in δ⊤ δ
.
α , equivalent statements hold with (δ , δ) replaced by (ξ , ξ).
For the second forecaster X2t 0 0
For the four parameterizations given in Table 4, we continue to simplify the closed-form
expressions for MCBi and DSCi from Proposition C.1 as functions of k, α, and ς such that
58
the MCB and DSC relations from Table 4 can be verified. Notice that in the following, the
expressions for κ(k) vary for the different DGPs denoted 1), 2), 4) and 5).
First, for the check loss decomposition for DGP 1) in Table 4, we have
q
MCB1 (k) = MCB2 (k) = −zα Φ −z α
s(k) + α − 1 + s(k) ϕ −zα
s(k) − 1+ ς
16 ϕ(zα ),
q q
3ς ς
DSC1 = DSC2 = ϕ(zα ) 1 + 16 − 1 + 16 , (87)
r
k2
s(k) = Var(Yt − Xitα ) = 1+ς 2 + 1
16 .
Here, s(k) is increasing in k, such that MCB1 (k) = MCB2 (k) are increasing in k. MCB1 (0) =
MCB2 (0) can be obtained by simplifying the above formula for k = 0. Furthermore, DSC1 =
DSC2 > 0 does not depend on k.
Second, for the check loss decomposition for DGP 2) in Table 4, we have
MCB1 = 0,
q
MCB2 (k) = −zα Φ s−z α
2 (k)
+ α − 1 + s 2 (k) ϕ −zα
s2 (k) − 1+ ς
16 ϕ(zα ),
q q (88)
3ς ς
DSC1 = DSC2 = ϕ(zα ) 1 + 16 − 1 + 16 ,
q
α ς ςk2
s2 (k) = Var(Yt − X2t ) = 1+ 16 + 2 .
As before, s2 (k) is increasing in k, such that MCB2 (k) is increasing in k, and DSC1 = DSC2 > 0
are independent of k. MCB2 (0) can again be obtained by simplifying the above formula for
k = 0.
Third, for the check loss decomposition for DGP 4) in Table 4, we have
q 2
MCB1 (k) = MCB2 (k) = −zα Φ −z α
s(k) + α − 1 + s(k) ϕ −zα
s(k) −
ςk
2 + 1 ϕ(zα ),
p q
2
DSC1 (k) = DSC2 (k) = ϕ(zα ) ςk 2 + 1 − ςk2 + 1 , (89)
q
2
s(k) = Var(Yt − Xitα ) = 1 + ςk2 + 8ς .
Here, DSC1 (k) = DSC2 (k) are increasing in k (as k 2 grows faster than k 2 /2) and we have
DSC1 (0) = DSC2 (0) = 0. Furthermore, MCB1 (k) = MCB2 (k) > 0 holds by (89).
59
α = 0.01 α = 0.05 α = 0.1 α = 0.25 α = 0.5
0.8 0.70 0.65 0.56 0.50
0.7 0.65
0.60 0.54 0.48
ξ0
0.60
0.6 0.55 0.52
0.55 0.46
25
50
75
00
00
25
50
75
00
00
25
50
75
00
00
25
50
75
00
00
25
50
75
00
0.
0.
0.
0.
1.
0.
0.
0.
0.
1.
0.
0.
0.
0.
1.
0.
0.
0.
0.
1.
0.
0.
0.
0.
1.
k
Figure C.1: Values for ξ0 (k, α) used for the quantile DGP 5) in Table 4 for k ∈ [0, 1] and α ∈
{0.01, 0.05, 0.1, 0.25, 0.5}, such that MCB1 (k) = MCB2 (k) holds as accurately as possible for all k and α for
the quantities from (90).
Fourth, for the check loss decomposition for DGP 5) in Table 4 with ξ0 = ξ0 (k, α), we have
DSC1 = 0,
p q
ςk 2
DSC2 (k) = ϕ(zα ) ςk 2 + 1 − 2 +1 ,
p
MCB1 (k) = −(δ0 + zα ) Φ −(δs10(k)
+zα )
+ α − 1 + s1 (k) ϕ −(δs10(k)
+zα )
− ςk 2 + 1 ϕ(zα ),
q 2 (90)
−(ξ0 +zα ) −(ξ0 +zα )
MCB2 (k) = −(ξ0 + zα ) Φ s2 (k) + α − 1 + s2 (k) ϕ s2 (k) − ςk2 + 1 ϕ(zα ),
q
α
s1 (k) := Var(Yt − X1t ) = 1 + ςk 2 + 8ς ,
q
α 2
s2 (k) := Var(Yt − X2t ) = 1 + ςk2 + 8ς .
As before, DSC2 (k) is increasing in k (as k 2 grows faster than k 2 /2) and we have DSC2 (0) = 0.
Notice for this that the choice of ξ0 does not affect the discrimination of the second forecast at
all.
Equality of the MCBi , i = 1, 2 terms in (90) is difficult to obtain analytically. Hence, we
choose ξ0 = ξ0 (k, α) depending on α ∈ {0.01, 0.05, 0.1, 0.25, 0.5} and a sequence of values
of k ∈ [0, 1] numerically such that MCB1 (k, α) = MCB2 (k, α) holds as accurately as possible.
Figure C.1 displays these numeric solutions for ξ0 = ξ0 (k, α) as functions of k and α.
Figure C.2 displays the expressions of MCBi and DSCi , i = 1, 2 from (87)–(90) for the
considered values of k and α, which confirms the properties of our simulation setup discussed
above. In particular, Figure C.2 confirms that the numeric choices of ξ0 (k, α) yield the desired
results.
Proof of Proposition C.1. We prove the claims for i = 1; the case i = 2 follows by replacing
(δ0 , δ) with (ξ0 , ξ).
First, under (22), the vector Vt has mean zero and covariance matrix ςI4 , hence γ ⊤ Vt ∼
N (0, ς γ ⊤ γ) and
Yt = γ ⊤ Vt + εY,t ∼ N (0, σY2 ), σY2 = ς γ ⊤ γ + 1.
α = V⊤ δ + δ + z we have
Moreover, with X1t t 0 α
α
Var(X1t ) = ς δ ⊤ δ, and α
Cov(Yt , X1t ) = ς δ ⊤ γ.
60
MCB1(k ) = MCB2(k ) MCB1(k ) = 0 and MCB2(k ) MCB1 = MCB2 > 0 MCB1 ≈ MCB2 > 0
DSC1 = DSC2 > 0 DSC1 = DSC2 > 0 DSC1(k ) = DSC2(k ) DSC1 = 0 and DSC2(k )
0.0125 0.0125 0.06
α = 0.01
0.0075 0.0075 0.0050
0.0050 0.0050
0.02
0.0025
0.0025 0.0025
α = 0.05
0.02 0.02
0.010 0.050
0.04 0.04
α = 0.1
0.03 0.03 0.02 0.06
0.02 0.02
0.01 0.03
0.01 0.01
α = 0.25
0.04
0.04 0.04 0.050
0.02
0.02 0.02 0.025
0.075 0.075
0.06 0.06
α = 0.5
25
50
75
00
00
25
50
75
00
00
25
50
75
00
00
25
50
75
00
0.
0.
0.
0.
1.
0.
0.
0.
0.
1.
0.
0.
0.
0.
1.
0.
0.
0.
0.
1.
k
MCB1 DSC1 MCB2 DSC2
Figure C.2: Population values of MCBi and DSCi , i = 1, 2 from (87)–(90) for a range of values for k, our five
choices of α, and the parameter values ξ0 (k, α) displayed in Figure C.1.
61
α ) is jointly Gaussian, Y | X α is Gaussian with constant conditional variance
Since (Yt , X1t t 1t
2
α )2
Cov(Yt , X1t (δ ⊤ γ)2
σ1|X = Var(Yt ) − α = σY2 − ς ⊤ . (91)
Var(X1t ) δ δ
We continue by deriving the formula for UNC. Since the α-quantile functional is strictly
consistent for the check loss, the expected loss E[Sα (x, Yt )] is uniquely minimized at the true
conditional quantile Qα (Yt ) = σY zα = σY Φ−1 (α).
To obtain an explicit formula for the check loss at the true quantile, let Yt = σY Zt with
Zt ∼ N (0, 1). Then,
E[Sα (σY zα , Yt )] = E (α − 1{Zt ≤ zα })(σY Zt − σY zα )
= σY E (α − 1{Zt ≤ zα })(Zt − zα )
= σY αE[Zt − zα ] − E (Zt − zα )1{Zt ≤ zα } .
Since E[Zt ] = 0, the first term becomes −αzα . For the second term,
Z zα
E (Zt − zα )1{Zt ≤ zα } = (z − zα )ϕ(z) dz = −ϕ(zα ) − zα α,
−∞
which yields
E[Sα (σY zα , Yt )] = σY − αzα + ϕ(zα ) + zα α = σY ϕ(zα ). (92)
We continue to derive the close-form expression for the discrimination component. Because
α ) is jointly Gaussian, the conditional distribution Y | X α is normal with conditional
(Yt , X1t t 1t
2
variance σ1|X α ) from (91) that does not depend on X α . Hence,
= Var(Yt | X1t 1t
α α 2 α α
Yt | X1t ∼ N E[Yt | X1t ], σ1|X and Qα (Yt | X1t ) = E[Yt | X1t ] + σ1|X zα .
By the check loss identity for Gaussian distribution from (92), we obtain
α α
E Sα Qα (Yt | X1t ), Yt | X1t = σ1|X ϕ(zα ) a.s.,
Consequently,
α
DSC1 = UNC − E Sα Qα (Yt | X1t ), Yt = σY ϕ(zα ) − σ1|X ϕ(zα ),
α α α
MCB1 = E[Sα (X1t , Yt )] − E[Sα (Qα (Yt | X1t ), Yt )] = E[Sα (X1t , Yt )] − σ1|X ϕ(zα ),
α , Y )] in the following.
for which we derive a closed form for E[Sα (X1t t
62
α . Then E is Gaussian with mean
For this, define the forecast error E1t := Yt − X1t 1t
α
m1 = E[E1t ] = E[Yt ] − E[X1t ] = −(δ0 + zα ),
and variance
Hence,
E[Sα (x, U )] = α E[E] − E E 1{E < 0} .
and therefore
E[Sα (x, U )] = m Φ(m/s) + α − 1 + s ϕ(m/s).
α and U = Y and using κ = m /s gives
Applying this with x = X1t t 1 1 1
α
E[Sα (X1t , Yt )] = m1 Φ(κ1 ) + α − 1 + s1 ϕ(κ1 ).
(δ ⊤ γ)2
We continue to establish monotonicity of MCB1 and DSC1 with respect to κ1 and R1 = δ⊤ δ
.
For DSC1 , it is immediate that σ1|X is decreasing in R1 , such that DSC1 is increasing in R1 . The
same argument applies to MCB1 being increasing in R1 as only the latter part in (85) depends
on σ1|X .
From (85), only the first two terms of MCB1 depend on s1 . Define
m1
g(s1 ) := m1 Φ(κ1 ) + α − 1 + s1 ϕ(κ1 ), κ1 = .
s1
dκ1
Differentiating with respect to s1 and using ds1 = −κ1 /s1 , we obtain
where we used ϕ′ (x) = −xϕ(x) and κ1 = m1 /s1 . Since the remaining term −σ1|X ϕ(zα ) does not
depend on s1 , MCB1 is strictly increasing in s1 .
63
a)
MCB1(k ) = MCB2(k ) = 8k 2 15 MCB1 = 0 and MCB2(k ) = 8k 2 15 MCB1(k ) = 2MCB2(k ) = 16k 2 15
0.1
Population value
0.0
0.1
0.0
0
5
0.
0.
0.
0.
0.
0.
0.
0.
0.
0.
0.
0.
0.
0.
0.
0.
0.
0.
k
b)
DSC1(k ) = DSC2(k ) = 8k 2 15 DSC1 = 0 and DSC2(k ) = 8k 2 15 DSC1(k ) = 2DSC2(k ) = 16k 2 15
0.1
Population value
0.0
0.1
0.0
0
5
0.
0.
0.
0.
0.
0.
0.
0.
0.
0.
0.
0.
0.
0.
0.
0.
0.
0.
k
MCB1 DSC1 MCB2 DSC2
Figure D.1: Population values of MCBi and DSCi , i = 1, 2 for the mean forecasts from the DGP in (22)–(24) for
the twelve parametrizations from Table 3, following the close-form expressions of Proposition 4.1.
D Additional Figures
Figure D.1 plots the population values for MCBi and DSCi , i = 1, 2 for the mean forecasts for
the twelve DGPs from Table 3.
Figure D.2 plots the time series of the variance and VaR forecasts together with their corre-
sponding evaluation targets for the application of Section 5.2.
Figures D.3–D.5 show model diagnostics for the linear recalibration regressions that are used
for the score decompositions for the forecasts of the applications of Sections 5.1–5.2. For the
mean forecasts in Figures D.3–D.4, we plot the model residuals Yt −W ⊤ θbiT on the y-axis against it
the forecasts Xit (as the covariates of the underlying recalibration regressions) on the x-axis. The
blue curve and its pointwise 95% confidence band are obtained from a local polynomial mean
regression, implemented via the loess function in the statistical software R, with automatically
selected tuning parameters. The black-dashed zero-line being inside these confidence bands for
the majority of forecast values indicates no evidence against correct specification of the linear
recalibration regressions.
64
Variance VaR
8
5
6
0
HAR
4
-5
2
-10
0
8
5
6
0
MIDAS
4
-5
2
-10
0
8
5
6
GARCH
0
value
4
-5
2
-10
0
8
5
6
0
GJR
4
-5
2
-10
0
HS
-5
-10
2019 2020 2021 2022 2019 2020 2021 2022
Time
Variance 1% VaR 5% VaR
Figure D.2: Daily E-mini returns (right) and Realized Variances (left) in gray, together with the corresponding
variance forecasts in blue (left) and VaR forecast at levels 1% in orange and 5% in purple (right) for the forecasting
models considered in Section 5.2 and the time period after the year 2019.
65
Michigan Forecast SPF Forecast
5.0
2.5
Residual
0.0
-2.5
-5.0
2 4 6 2 4 6 8
Forecast
Figure D.3: Model diagnostics for the linear recalibration regression for the inflation rate forecasts in Section 5.1.
We plot the model residuals against the forecasts (as the covariates of the underlying recalibration regression)
together with a nonparametric polynomial regression and its 95% confidence bands. The zero-line being mostly
contained in the confidence bands implies no evidence against correct model specification. For further details, see
the text of Appendix D.
2.5
Residual
0.0
-2.5
-5.0
0.1 1.0 10.0 0.1 1.0 10.0 0.3 1.0 3.0 10.0 30.0 0.3 1.0 3.0 10.0 30.0
Forecast
2.5
Residual
0.0
-2.5
-5.0
0.1 1.0 10.0 0.1 1.0 10.0 0.3 1.0 3.0 10.0 30.0 0.3 1.0 3.0 10.0 30.0
Forecast
Figure D.4: Model diagnostics for the linear recalibration regression for the variance forecasts in Section 5.2. We
plot the model residuals against the forecasts (as the covariates of the underlying recalibration regression) on a
log-transformed x-axis together with a nonparametric polynomial regression and its 95% confidence bands. The
zero-line being mostly contained in the confidence bands implies no evidence against correct model specification.
For further details, see the text of Appendix D.
66
a) Quantile Mincer-Zarnowitz Regression for 1% VaR
HAR MIDAS GARCH GJR HS
1.00
Generalized Residual
0.75
0.50
0.25
0.00
-16 -12 -8 -4 0 -12 -8 -4 0 -10 -5 -10 -5 -7.5 -5.0 -2.5
Forecast
0.75
0.50
0.25
0.00
Figure D.5: Model diagnostics for the linear recalibration regression for the VaR forecasts in Section 5.2. We plot
the generalized model residuals against the forecasts (as the covariates of the underlying recalibration regression)
together with a nonparametric polynomial regression and its 95% confidence bands. The zero-line being mostly
contained in the confidence bands implies no evidence against correct model specification. For further details, see
the text of Appendix D.
In Figure D.5, we follow Pohle et al. (2024) and Dimitriadis and Hoga (2026, Figures 4 and
6) and plot the generalized quantile residuals, i.e., the identification functions, V Wit⊤ θbiT , Yt ) =
1{Yt ≤ W ⊤ θbiT } − α, against the quantile forecasts Xit together with a corresponding nonpara-
it
metric mean regression and their pointwise 95% confidence bands. The zero-line being mostly
contained in the confidence bands again does not provide evidence against correct specification
of the linear recalibration regressions.
References
Allen, S., Burnello, J., and Ziegel, J. (2025). Assessing the conditional calibration of interval forecasts
using decompositions of the interval score. Preprint. [Link]
Andrews, D. (1994). Empirical process methods in econometrics. In Engle, R. and McFadden, D., editors,
Handbook of Econometrics, volume 4, chapter 37, pages 2247–2294. Elsevier, Amsterdam.
Andrews, D. W. K. (1991). Heteroskedasticity and autocorrelation consistent covariance matrix estima-
tion. Econometrica, 59(3):817–858.
Angeletos, G.-M., Huo, Z., and Sastry, K. A. (2021). Imperfect macroeconomic expectations: Evidence
and theory. NBER Macroeconomics Annual, 35:1–86.
Arnold, S., Walz, E.-M., Ziegel, J., and Gneiting, T. (2024). Decompositions of the mean continuous
ranked probability score. Electronic Journal of Statistics, 18(2):4992–5044.
Basel Committee (1996). Overview of the amendment to the capital accord to incorporate market risks.
Technical report, Bank for International Settlements. [Link]
Basel Committee (2019a). CRE: Calculation of RWA for credit risk. Technical report, Bank for In-
ternational Settlements. [Link]
20191215&published=20191215&export=pdf.
67
Basel Committee (2019b). Minimum capital requirements for market risk. Technical report, Bank for
International Settlements. [Link]
Bayer, S. and Dimitriadis, T. (2022). Regression-based Expected Shortfall backtesting. Journal of
Financial Econometrics, 20(3):437–471.
Bentzien, S. and Friederichs, P. (2014). Decomposition and graphical portrayal of the quantile score.
Quarterly Journal of the Royal Meteorological Society, 140(683):1924–1934.
Berger, R. L. (1997). Likelihood ratio tests and intersection-union tests. In Panchapakesan, S. and
Balakrishnan, N., editors, Advances in Statistical Decision Theory and Applications, pages 225–237.
Birkhäuser, Boston, MA.
Berkowitz, J., Christoffersen, P., and Pelletier, D. (2011). Evaluating Value-at-Risk models with desk-level
data. Management Science, 57(12):2213–2227.
Bollerslev, T. (1986). Generalized autoregressive conditional heteroskedasticity. Journal of Econometrics,
31(3):307–327.
Bollerslev, T., Hood, B., Huss, J., and Pedersen, L. H. (2018). Risk everywhere: Modeling and managing
volatility. The Review of Financial Studies, 31(7):2729–2773.
Bordalo, P., Gennaioli, N., Ma, Y., and Shleifer, A. (2020). Overreaction in macroeconomic expectations.
American Economic Review, 110(9):2748–82.
Brehmer, J. R., Kraus, K., Gneiting, T., Herrmann, M., and Marzocchi, W. (2024). Enhancing the
statistical evaluation of earthquake forecasts—an application to Italy. Seismological Research Letters,
96(3):1966–1988.
Brier, G. W. (1950). Verification of forecasts expressed in terms of probability. Monthly Weather Review,
78(1):1–3.
Bröcker, J. (2009). Reliability, sufficiency, and the decomposition of proper scores. Quarterly Journal of
the Royal Meteorological Society, 135(643):1512–1519.
Broer, T. and Kohlhas, A. N. (2024). Forecaster (mis-)behavior. Review of Economics and Statistics,
106(5):1334–1351.
Campbell, S. D. (2007). A review of backtesting and backtesting procedures. Journal of Risk, 9(2):1–17.
Choe, Y. J. and Ramdas, A. (2024). Comparing sequential forecasters. Operations Research, 72(4):1368–
1387.
Christoffersen, P. (1998). Evaluating interval forecasts. International Economic Review, 39(4):841–862.
Clark, T. and McCracken, M. (2013). Advances in forecast evaluation. In Elliott, G. and Timmermann, A.,
editors, Handbook of Economic Forecasting, volume 2 of Handbook of Economic Forecasting, chapter 20,
pages 1107–1201. Elsevier, Amsterdam.
Clark, T. E. and McCracken, M. W. (2001). Tests of equal forecast accuracy and encompassing for nested
models. Journal of Econometrics, 105(1):85–110.
Coibion, O. and Gorodnichenko, Y. (2015). Information rigidity and the expectations formation process:
A simple framework and new facts. American Economic Review, 105(8):2644–2678.
Corsi, F. (2009). A simple approximate long-memory model of Realized Volatility. Journal of Financial
Econometrics, 7(2):174–196.
Davidson, J. (1994). Stochastic Limit Theory: An Introduction for Econometricians. Advanced Texts in
Econometrics. Oxford Univerity Press, Oxford.
Dawid, A. P. (1986). Probability forecasting. In Kotz, S., Johnson, N. L., and Read., C. B., editors,
Encyclopedia of Statistical Sciences, volume 7, page 210–218.
DeGroot, M. H. and Fienberg, S. E. (1981). Assessing probability assessors: Calibration and refinement.
Technical report.
Diebold, F. X. (2015). Comparing predictive accuracy, twenty years later: A personal perspective on the
use and abuse of Diebold–Mariano tests. Journal of Business & Economic Statistics, 33(1):1–9.
Diebold, F. X., Gunther, T. A., and Tay, A. S. (1998). Evaluating density forecasts with applications to
financial risk management. International Economic Review, 39(4):863–883.
Diebold, F. X. and Mariano, R. S. (1995). Comparing predictive accuracy. Journal of Business &
Economic Statistics, 20(1):134–144.
Dimitriadis, T. and Bayer, S. (2019). A joint quantile and Expected Shortfall regression framework.
68
Electronic Journal of Statistics, 13(1):1823–1871.
Dimitriadis, T., Fissler, T., and Ziegel, J. (2024a). Characterizing M-estimators. Biometrika, 111(1):339–
346.
Dimitriadis, T., Gneiting, T., and Jordan, A. I. (2021). Stable reliability diagrams for probabilistic
classifiers. Proceedings of the National Academy of Sciences, 118(8):1–10.
Dimitriadis, T., Gneiting, T., Jordan, A. I., and Vogel, P. (2024b). Evaluating probabilistic classifiers:
The triptych. International Journal of Forecasting, 40(3):1101–1122.
Dimitriadis, T. and Hoga, Y. (2025). Dynamic CoVaR modeling and estimation. Journal of Business &
Economic Statistics. [Link]
Dimitriadis, T. and Hoga, Y. (2026). Regressions under adverse conditions. Journal of Business &
Economic Statistics, 44(1):227–241.
Duchesne, P. and Lafaye De Micheaux, P. (2010). Computing the distribution of quadratic forms: Further
comparisons between the Liu-Tang-Zhang approximation and exact methods. Computational Statistics
and Data Analysis, 54:858–862.
Ehm, W., Gneiting, T., Jordan, A., and Krüger, F. (2016). Of quantiles and expectiles: Consistent scoring
functions, choquet representations and forecast rankings. Journal of the Royal Statistical Society Series
B: Statistical Methodology, 78(3):505–562.
Elliott, G., Ghanem, D., and Krüger, F. (2016). Forecasting conditional probabilities of binary outcomes
under misspecification. Review of Economics and Statistics, 98(4):742–755.
Elliott, G. and Timmermann, A. (2016). Economic Forecasting. Princeton University Press.
Engle, R. F. and Manganelli, S. (2004). CAViaR: Conditional autoregressive value at risk by regression
quantiles. Journal of Business & Economic Statistics, 22(4):367–381.
Farmer, L. E., Nakamura, E., and Steinsson, J. (2024). Learning about the long run. Journal of Political
Economy, 132(10):3334–3377.
Faust, J. and Wright, J. H. (2013). Forecasting inflation. In Elliott, G. and Timmermann, A., editors,
Handbook of Economic Forecasting, volume 2, chapter 1, pages 2–56. Elsevier, Amsterdam.
Fissler, T. and Hoga, Y. (2024). Backtesting systemic risk forecasts using multi-objective elicitability.
Journal of Business & Economic Statistics, 42(2):485–498.
Fissler, T., Lorentzen, C., and Mayer, M. (2022). Model comparison and calibration assessment:
User guide for consistent scoring functions in machine learning and actuarial practice. Preprint.
[Link]
Fissler, T. and Ziegel, J. F. (2016). Higher order elicitability and Osband’s principle. The Annals of
Statistics, 44(4):1680 – 1707.
Fosten, J., Gutknecht, D., and Pohle, M.-O. (2024). Testing quantile forecast optimality. Journal of
Business & Economic Statistics, 42(4):1367–1378.
Gaglianone, W. P., Lima, L. R., Linton, O., and Smith, D. R. (2011). Evaluating Value-at-Risk models
via quantile regression. Journal of Business & Economic Statistics, 29(1):150–160.
Galvao, A. F. and Yoon, J. (2024). HAC covariance matrix estimation in quantile regression. Journal of
the American Statistical Association, 119(547):2305–2316.
Ghysels, E., Kvedaras, V., and Zemlys, V. (2016). Mixed frequency data sampling regression models:
The R package midasr. Journal of Statistical Software, 72(4):1–35.
Ghysels, E., Santa-Clara, P., and Valkanov, R. (2006). Predicting volatility: Getting the most out of
return data sampled at different frequencies. Journal of Econometrics, 131(1):59–95.
Giacomini, R. and White, H. (2006). Tests of conditional predictive ability. Econometrica, 74(6):1545–
1578.
Glosten, L. R., Jagannathan, R., and Runkle, D. E. (1993). On the relation between the expected value
and the volatility of the nominal excess return on stocks. The Journal of Finance, 48(5):1779–1801.
Gneiting, T. (2011). Making and evaluating point forecasts. Journal of the American Statistical Associ-
ation, 106(494):746–762.
Gneiting, T., Balabdaoui, F., and Raftery, A. E. (2007). Probabilistic forecasts, calibration and sharpness.
Journal of the Royal Statistical Society Series B: Statistical Methodology, 69(2):243–268.
Gneiting, T. and Resin, J. (2023). Regression diagnostics meets forecast evaluation: Conditional calibra-
69
tion, reliability diagrams, and coefficient of determination. Electronic Journal of Statistics, 17(2):3226–
3286.
Gneiting, T., Wolffram, D., Resin, J., Kraus, K., Bracher, J., Dimitriadis, T., Hagenmeyer, V., Jordan,
A. I., Lerch, S., Phipps, K., and Schienle, M. (2023). Model diagnostics and forecast evaluation for
quantiles. Annual Review of Statistics and Its Application, 10(1):597–621.
Guler, K., Ng, P. T., and Xiao, Z. (2017). Mincer-Zarnowitz quantile and expectile regressions for forecast
evaluations under aysmmetric loss functions. Journal of Forecasting, 36(6):651–679.
Hansen, B. E. (1996). Stochastic equicontinuity for unbounded dependent heterogeneous arrays. Econo-
metric Theory, 12(2):347–359.
Hansen, P. R. (2005). A test for superior predictive ability. Journal of Business & Economic Statistics,
23(4):365–380.
Hansen, P. R., Lunde, A., and Nason, J. M. (2011). The model confidence set. Econometrica, 79(2):453–
497.
He, X. D., Kou, S., and Peng, X. (2022). Risk measures: Robustness, elicitability, and backtesting.
Annual Review of Statistics and Its Application, 9:141–166.
Henzi, A. and Ziegel, J. F. (2022). Valid sequential inference on probability forecast performance.
Biometrika, 109(3):647–663.
Hoga, Y. and Demetrescu, M. (2023). Monitoring Value-at-Risk and Expected Shortfall forecasts. Man-
agement Science, 69(5):2954–2971.
Hoga, Y. and Dimitriadis, T. (2023). On testing equal conditional predictive ability under measurement
error. Journal of Business & Economic Statistics, 41(2):364–376.
Holzmann, H. and Eulert, M. (2014). The role of the information set for forecasting—with applications
to risk management. The Annals of Applied Statistics, 8(1):595 – 621.
Imhof, J. P. (1961). Computing the distribution of quadratic forms in normal variables. Biometrika,
48(3/4):419–426.
Kohlhas, A. N. and Walther, A. (2021). Asymmetric attention. American Economic Review, 111(9):2879–
2925.
Kuester, K., Mittnik, S., and Paolella, M. S. (2005). Value-at-Risk prediction: A comparison of alternative
strategies. Journal of Financial Econometrics, 4(1):53–89.
Kupiec, P. H. (1995). Techniques for verifying the accuracy of risk measurement models. The Journal of
Derivatives, 3(2):73–84.
Li, J., Liao, Z., and Quaedvlieg, R. (2022). Conditional superior predictive ability. The Review of
Economic Studies, 89(2):843–875.
Liu, L. Y., Patton, A. J., and Sheppard, K. (2015). Does anything beat 5-minute RV? A comparison of
realized measures across multiple asset classes. Journal of Econometrics, 187(1):293–311.
Mincer, J. and Zarnowitz, V. (1969). The evaluation of economic forecasts. In Mincer, J., editor, Economic
Forecasts and Expectations: Analysis of Forecasting Behavior and Performance, pages 3–46. National
Bureau of Economic Research, New York.
Mitchell, K. (2019). Score decompositions in forecast verification. PhD thesis, University of Exeter,
Exeter, United Kingdom.
Mühlemann, A. (2021). The role of loss functions in regression problems. PhD thesis, University of Bern,
Bern, Switzerland.
Murphy, A. H. (1973). A new vector partition of the probability score. Journal of Applied Meteorology
and Climatology, 12(4):595–600.
Murphy, A. H. and Winkler, R. L. (1987). A general framework for forecast verification. Monthly Weather
Review, 115(7):1330 – 1338.
Murphy, A. H. and Winkler, R. L. (1992). Diagnostic verification of probability forecasts. International
Journal of Forecasting, 7(4):435–455.
Newey, W. K. (1994). The asymptotic variance of semiparametric estimators. Econometrica, pages
1349–1382.
Newey, W. K. and McFadden, D. (1994). Large sample estimation and hypothesis testing. In Engle,
R. and McFadden, D., editors, Handbook of Econometrics, volume 4, chapter 36, pages 2111–2245.
70
Elsevier, Amsterdam.
Newey, W. K. and West, K. D. (1987). A simple, positive semi-definite, heteroskedasticity and autocor-
relation consistent covariance matrix. Econometrica, 55(3):703–708.
Nieto, M. R. and Ruiz, E. (2016). Frontiers in VaR forecasting and backtesting. International Journal
of Forecasting, 32(2):475–501.
Nolde, N. and Ziegel, J. F. (2017). Elicitability and backtesting: Perspectives for banking regulation.
The Annals of Applied Statistics, 11(4):1833–1874.
Patton, A. J. (2011). Volatility forecast comparison using imperfect volatility proxies. Journal of Econo-
metrics, 160(1):246–256.
Patton, A. J. (2020). Comparing possibly misspecified forecasts. Journal of Business & Economic
Statistics, 38(4):796–809.
Patton, A. J., Ziegel, J. F., and Chen, R. (2019). Dynamic semiparametric models for Expected Shortfall
(and Value-at-Risk). Journal of Econometrics, 211(2):388–413.
Pohle, M.-O. (2020). The Murphy decomposition and the calibration-resolution principle: A new per-
spective on forecast evaluation. Preprint. [Link]
Pohle, M.-O., Demetrescu, M., Dimitriadis, T., and Wermuth, J.-L. (2024). Generalised residual diag-
nostics. Mimeo.
R Core Team (2026). R: A language and environment for statistical computing. R Foundation for
Statistical Computing, Vienna, Austria.
Ranjan, R. and Gneiting, T. (2010). Combining probability forecasts. Journal of the Royal Statistical
Society Series B: Statistical Methodology, 72(1):71–91.
Sanders, F. (1963). On subjective probability forecasting. Journal of Applied Meteorology and Climatol-
ogy, 2(2):191–201.
Savage, L. J. (1971). Elicitation of personal probabilities and expectations. Journal of the American
Statistical Association, 66(336):783–801.
Siegert, S. (2014). Variance estimation for Brier score decomposition. Quarterly Journal of the Royal
Meteorological Society, 140(682):1771–1777.
Siegert, S. (2017). Simplifying and generalising Murphy’s Brier score decomposition. Quarterly Journal
of the Royal Meteorological Society, 143(703):1178–1183.
Strähl, C. and Ziegel, J. (2017). Cross-calibration of probabilistic forecasts. Electronic Journal of Statis-
tics, 11(1):608–639.
Theil, H. (1966). Applied economic forecasting. North-Holland, Amsterdam.
van der Vaart, A. W. and Wellner, J. A. (1996). Weak convergence and empirical processes: With
applications to statistics. Springer Science & Business Media.
Wang, Q., Wang, R., and Ziegel, J. (2025). E-backtesting. Management Science.
[Link]
West, K. D. (1996). Asymptotic inference about predictive ability. Econometrica, 64(5):1067–1084.
West, K. D. (2006). Forecast evaluation. Handbook of Economic Forecasting, 1:99–134.
White, H. (2001). Asymptotic Theory for Econometricians. Academic Press, San Diego, 2nd edition.
Wüthrich, M. V., Richman, R., Avanzi, B., Lindholm, M., Maggi, M., Mayer, M., Schelldorfer, J., and
Scognamiglio, S. (2025). AI tools for actuaries. Preprint. [Link]
Xu, M. (2022). Semiparametric estimation with plug-in isotonic estimator. Preprint.
[Link] n/view?usp=sharing.
Zeileis, A. (2004). Econometric computing with HC and HAC covariance matrix estimators. Journal of
Statistical Software, 11(10):1–17.
71