0% found this document useful (0 votes)
2 views87 pages

Forecast Evaluation and Combination Techniques

This document discusses forecast evaluation and combination methods in economic forecasting, focusing on key concepts such as unbiasedness, efficiency, and predictive accuracy. It outlines various tests for evaluating forecast performance and compares different forecasting models. The document also emphasizes the potential benefits of combining forecasts to improve accuracy.

Uploaded by

saidul
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
2 views87 pages

Forecast Evaluation and Combination Techniques

This document discusses forecast evaluation and combination methods in economic forecasting, focusing on key concepts such as unbiasedness, efficiency, and predictive accuracy. It outlines various tests for evaluating forecast performance and compares different forecasting models. The document also emphasizes the potential benefits of combining forecasts to improve accuracy.

Uploaded by

saidul
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 1 / 87

Applied Economic Forecasting using Time Series


Methods

Eric Ghysels and Massimiliano Marcellino

Companion Slides - Chapter 4 Forecast Evaluation and


Combination

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 2 / 87


Forecast Evaluation and Combination

Introduction
Unbiasedness and efficiency
Evaluation of fixed event forecasts
Tests of predictive accuracy
Forecast comparison tests
The combination of forecasts
Forecast encompassing
Evaluation and combination of density forecasts
Examples using simulated data
Empirical examples
Concluding remarks

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 3 / 87


Introduction

Given different competing forecasts, in this chapter we try to answer


the following questions:
(i) How “good,” in some sense, is a particular set of forecasts?
(ii) Is one set of forecasts better than another one?
(iii) Is it possible to get a better forecast as a combination of various
forecasts?

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 4 / 87


Introduction

To address these questions:


(i) we define some key properties a good forecast should have and
discuss how to test them.
(ii) we introduce some basic statistics to assess whether one forecast is
equivalent or better than another with respect to a given criterion.
(iii) we discuss how to combine the forecasts and why the resulting
pooled forecast can be expected to perform well.

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 5 / 87


Forecast Evaluation and Combination

Introduction
Unbiasedness and efficiency
Evaluation of fixed event forecasts
Tests of predictive accuracy
Forecast comparison tests
The combination of forecasts
Forecast encompassing
Evaluation and combination of density forecasts
Examples using simulated data
Empirical examples
Concluding remarks

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 6 / 87


Unbiasedness and efficiency

The optimal forecast under a MSFE loss function is the conditional


expectation of the variable, so that it should be Et (yt+h ) = ŷt+h|t ,
where Et (·) is the conditional expectation given information at time t.
Unbiasedness means that the expected value of the forecast error
should be equal to zero, implying that on average the forecast
should be correct.

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 7 / 87


Unbiasedness and efficiency

Efficiency is related to the efficient use of the available information.

So that the optimal h-steps ahead forecast error should be at most


correlated of order h − 1, and uncorrelated with available information
at the time the forecast is made.

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 8 / 87


Unbiasedness and efficiency
Inefficient forecasts can still be unbiased, and biased forecasts can be
efficient.
If yt is a random walk, namely,
yt = yt−1 + εt ,
and we use as a forecast
eyt+h|t = yt−g ,
then eyt+h|t is unbiased but not efficient.
If yt is a random walk with drift, namely,
yt = a + yt−1 + εt ,
and we use as a forecast
eyt+h|t = yt ,
then eyt+h|t is biased but efficient.
c Eric Ghysels & Massimiliano Marcellino 2018 Edition 9 / 87
Unbiasedness and efficiency

In order to test whether a forecast is unbiased, let us consider the


regression

yi+h = α + βŷi+h|i + εi+h , i = T, ..., T + H − h, h<H (1a)

where h is the forecast horizon and T + 1, ..., T + H the evaluation


sample. The forecasts ŷi+h|i are recursively updated in periods i = T,
. . . , T + H - h.
Sufficient condition:
α = 0, β = 1 (1b)
Necessary condition:

α = (1 − β)E(ŷi+h|i ). (1c)

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 10 / 87


Unbiasedness and efficiency

The sufficient condition (α = 0, β = 1) can be tested by a “robust”


F−test, where the fact that εi+h is in general autocorrelated (at least)
of order h − 1 is taken into account in the derivation of the HAC
variance.

The necessary condition (formula 1c) is equivalent to τ = 0 in the


regression:
ei+h = yi+h − ŷi+h|i = τ + εi+h ,
which can be tested with a robust version of the t−test.

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 11 / 87


Unbiasedness and efficiency
Note that (α = 0, β = 1) also implies that the forecast and forecast
errors are uncorrelated. (formula 1a) can be rewritten as

ei+h = α + (β − 1)ŷi+h|i + εi+h ,

so that

E(ŷi+h|i ei+h ) = αE(ŷi+h|i ) + (β − 1)E(ŷi+h|i ) + E(ŷi+h|i εi+h ) = 0.


| {z }
=0

Moreover, when (α = 0, β = 1), we have that

var(yi+h ) = var(ŷi+h|i ) + var(ei+h ),

implying that the volatility of the variable should be larger than that of
the (optimal) forecast, the more so the larger the variance of the
forecast error.
c Eric Ghysels & Massimiliano Marcellino 2018 Edition 12 / 87
Unbiasedness and efficiency

The coefficient of determination (R2 ) from formula 1a can also be


used as an indicator of the forecast quality, with good forecasts
associated with high R2 .

However, one should also consider that persistent variables are


easier to forecast than volatile variables, given that their past is a
useful leading indicator.

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 13 / 87


Unbiasedness and efficiency

To test for weak efficiency: ei+h is correlated, across time, at most


of order h - 1, so that no lagged information beyond h − 1 can explain
the forecast errors.

We can fit a moving average model of order h − 1, MA(h − 1), to the


h-steps ahead forecast error and testing that the resulting residuals
are white noise.

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 14 / 87


Unbiasedness and efficiency

To test for strong efficiency: No indicators available when the


forecasts were formulated can improve h−step forecast and
therefore explain the h−step ahead forecast error.

We can test if γ = 0 in the regression

ei+h = γ 0 zi + εi+h ,

where zi is a vector of potentially relevant variables for explaining the


forecast errors. ei+h are realizations of a set of h-steps ahead
forecast errors, namely, ei+h = yi+h - ŷi+h|i .

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 15 / 87


Forecast Evaluation and Combination

Introduction
Unbiasedness and efficiency
Evaluation of fixed event forecasts
Tests of predictive accuracy
Forecast comparison tests
The combination of forecasts
Forecast encompassing
Evaluation and combination of density forecasts
Examples using simulated data
Empirical examples
Concluding remarks

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 16 / 87


Evaluation of fixed event forecasts

So far we have considered forecasts for period T + h made in period


T where T progressively increases, namely, {ŷi+h|i } for i = T, . . . ,
T + H − h.

As an alternative, we can consider {ŷτ |τ −h }, h = 1, 2, . . . i.e.,


forecasts for a fixed target value (yτ ) made at different time periods
that become closer and closer to τ. The {ŷτ |τ −h } are known as fixed
event forecasts. For example, for an AR(1) process, it is

ŷτ |τ −h = ρh yτ −h .

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 17 / 87


Evaluation of fixed event forecasts
The properties of fixed events forecasts were studied by, e.g.,
Clements(1997). Let us decompose the forecast error as:

eτ |τ −h = yτ − ŷτ |τ −h = vτ |τ −h+1 + vτ |τ −h+2 + ... + vτ |τ , (2)

where

vτ |J = ŷτ |J − ŷτ |J−1 , ŷτ |τ = yτ , J = τ − h + 1, . . . , τ. (3)

For the AR(1) example, it is

vτ |J = ρτ −J εJ,

and
eτ |τ −h = ρh−1 ετ −h+1 + ... + ετ .
Unbiasedness requires that

E(eτ |τ −h ) = 0, ∀ τ − h.

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 18 / 87


Evaluation of fixed event forecasts

For weak efficiency, the following should hold:

E(eτ |τ −h |vτ |τ −h , ..., vτ |1 ) = 0, ∀ τ − h,

This condition is equivalent to

E(vτ |τ −h |vτ |τ −h−1 , ..., vτ |1 ) = 0, ∀ τ −h

These conditions also imply that

ŷτ |J − ỹτ |J−1 = ρ(τ −J) εJ , (4)

This property can be assessed by testing if forecast revisions are


white noise.

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 19 / 87


Evaluation of fixed event forecasts

Strong efficiency can be defined as the lack of explanatory power


for vτ |J of variables z included in the information set for period J.
This property can be assessed by regressing vτ |J on zJ and testing
for the non-significance of zJ .

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 20 / 87


Forecast Evaluation and Combination

Introduction
Unbiasedness and efficiency
Evaluation of fixed event forecasts
Tests of predictive accuracy
Forecast comparison tests
The combination of forecasts
Forecast encompassing
Evaluation and combination of density forecasts
Examples using simulated data
Empirical examples
Concluding remarks

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 21 / 87


Tests of predictive accuracy

Tests of predictive accuracy compare an estimate of the forecast


error variance obtained from the past residuals with the actual
MSFE of the forecasts.
Hence, they provide a measure of how well the model performs in
the future relative to the past.
We will focus on Wald-type tests.

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 22 / 87


Tests of predictive accuracy

if yt admits the Wold MA(∞) representation:

yt = ψ(L)εt ,

then the h − step ahead minimum MSFE predictor is



X
ŷT+h|i = ψJ εT+h−J ,
J=h

with associated forecast error


h−1
X
eT+h = ψJ εT+h−J ,
J=0

where ψ0 = 1.

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 23 / 87


Tests of predictive accuracy

We can group the errors in forecasting (yT+1 , ..., yT+h ) conditional on


period T in
eh = ψεh , (5)
where eh = (eT+1 , ..., eT+h )0 , εh = (εT+1 , ..., εT+h )0 and
 
1 0 ... ... 0 0
 ψ1 1 ... . . . 0 0
 
 ψ2 ψ1 . . . . . . 0 0
 
ψ=  ... ..
.
..
.
.. .. 
. .
.
 
 .. .. .. 
 . . . 1 0
ψh−1 ψh−2 . . . . . . ψ1 1

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 24 / 87


Tests of predictive accuracy

If we define
Φh = E(eh e0h ) = ψE(εh ε0h )ψ 0 = σε2 ψψ 0 ,
and if the appropriate model over [1, ..., T] remains valid over the
forecast horizon and ε ∼ N, then

Q = e0h Φ−1 2
h eh ∼ χ (h),

where Φ−1 −2 −1 0 −1
h = σε (ψ ) ψ .

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 25 / 87


Tests of predictive accuracy

Writing the autoregressive (AR) approximation of the Wold


representation (see again the next chapter for details) as

ϕ(L)yt = εt , ϕ(L) = ψ(L)−1 ,

we also have
εh = ϕeh , ϕ = ψ −1 ,
Φ−1 −2 0
h = σε ϕ ϕ.

We can therefore rewrite the statistic Q as


h
e0h ϕ0 ϕeh ε0h εh 1 X 2
Q= = = εT+J .
σε2 σε2 σε2
J=1

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 26 / 87


Tests of predictive accuracy

An operational version of the test is therefore:


h
1 X 2
Q̂ = 2 eT+J|T+J−1 ∼ F(h, T − p),
σ̂ε
J=1

where p is the number of parameters used in the model and


eT+J|T+J−1 indicates for clarity the one-step ahead forecast error.

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 27 / 87


Forecast Evaluation and Combination

Introduction
Unbiasedness and efficiency
Evaluation of fixed event forecasts
Tests of predictive accuracy
Forecast comparison tests
The combination of forecasts
Forecast encompassing
Evaluation and combination of density forecasts
Examples using simulated data
Empirical examples
Concluding remarks

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 28 / 87


Forecast comparison tests

The most common approach to compare alternative predictions is to


rank them according to the associated loss function, typically the
MSFE or MAFE. However, these comparisons are deterministic.
We will now consider two tests for the hypothesis that two forecasts
are equivalent, in the sense that the associated loss difference is not
statistically different from zero:
(i) Morgan - Granger - Newbold test
(ii) Diebold - Mariano test

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 29 / 87


Morgan - Granger - Newbold test

The test requires the forecast errors to be zero mean, normally


distributed, and uncorrelated.
If we indicate by e1 and e2 the forecast errors from the competing
models, the test is based on the auxiliary variables:

u1,T+J = e1,T+J − e2,T+J , u2,T+J = e1,T+J + e2,T+J . (6)

It is
E(u1 u2 ) = MSFE1 − MSFE2 ,
so that the hypothesis of interest is whether u1 and u2 are correlated
or not.

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 30 / 87


Morgan - Granger - Newbold test

The proposed Morgan - Granger - Newbold test statistic is


r
p ∼ tH−1 ,
(H − 1)−1 (1 − r2 )

where (1) tH−1 a Student t distribution with H − 1 degrees of


freedom, (2) H is the length of the evaluation sample and
PH
u1,T+i u2,T+i
r = qP 1 .
H 2 PH 2
u
1 1,T+i u
1 2,T+i

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 31 / 87


Diebold - Mariano test

It relaxes the requirements on the forecast errors and can deal with
the comparison of general loss functions.
Let us define the Diebold-Mariano test statistic as
PH
1/2 j=1 dj /H d
DM = H = H 1/2 , (7)
σd σd
where
dj = g(e1j ) − g(e2j ),
g is the loss function of interest, e.g., the quadratic loss g(e) = e2 or
the absolute loss g(e) = |e|, e1 and e2 are the errors from the two
competing forecasts, and σd2 is the variance of d.

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 32 / 87


Diebold - Mariano test

In order to take into account the serial correlation of the forecast


errors, σd2 can be estimated as

h−1
! H
X X
−1
bd2
σ = γ0 + 2 γi with γk = H (dt − d)(dt−k − d),
i=1 t=k+1

where h is the forecast horizon, so that for h = 1 there is no


correlation and the standard formula for variance estimation can be
used.
Under the null hypothesis that E(d) = 0, the statistic DM has an
asymptotic standard normal distribution.

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 33 / 87


Diebold - Mariano test

Harvey, Leybourne and Newbold (1998) suggested a modified


version of the DM statistic,
1/2
H + 1 − 2h + H −1 h(h − 1)

HLN = DM,
HH

to be compared with critical values from the Student t distribution


with H − 1 degrees of freedom, in order to improve the finite sample
properties of the DM test.

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 34 / 87


Diebold - Mariano test

When the models underlying the forecasts under comparison are


nested, for example an AR(1) and an AR(2), then the asymptotic
distribution of the DM test becomes non-standard and a functional of
Brownian motions.
A simple solution in this case is the use of rolling rather than
recursive estimation.

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 35 / 87


Forecast Evaluation and Combination

Introduction
Unbiasedness and efficiency
Evaluation of fixed event forecasts
Tests of predictive accuracy
Forecast comparison tests
The combination of forecasts
Forecast encompassing
Evaluation and combination of density forecasts
Examples using simulated data
Empirical examples
Concluding remarks

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 36 / 87


The combination of forecasts
Let us assume that two forecasts ŷ1 and ŷ2 are available for the
same target y, with associated forecast errors e1 and e2 . We want to
construct the combined (linear) forecast

ŷc = αŷ1 + (1 − α)ŷ2 , (8)

where the weights can be chosen in order to minimize the MSFE of


ŷc , From formula 8 we have

ec = y − ŷc = αe1 + (1 − α)e2 (9)

so that

MSFEc = α2 MSFE1 + (1 − α)2 MSFE2


(10)
+ 2α(1 − α)ϕ(MSFE1 MSFE2 )1/2

where ϕ is the correlation coefficient between e1 and e2 .


c Eric Ghysels & Massimiliano Marcellino 2018 Edition 37 / 87
The combination of forecasts

The optimal pooling weights, the minimizers of formula10,

MSFE2 − ϕ(MSFE1 MSFE2 )1/2


α∗ = ,
MSFE1 + MSFE2 − 2ϕ(MSFE1 MSFE2 )1/2

which yields

MSFE1 MSFE2 (1 − ϕ2 )
MSFEc∗ =
MSFE1 + MSFE2 − 2ϕ(MSFE1 MSFE2 )1/2

and
MSFEc∗ ≤ min(MSFE1 , MSFE2 ),
where equality holds if either ϕ2 = MSFE1 /MSFE2 (i.e., e2 = e1 + u) or
ϕ2 = MSFE2 /MSFE1 (i.e., e1 = e2 + u), which implies that ŷ1 or ŷ2 is
the optimal forecasts.

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 38 / 87


The combination of forecasts

If the forecast errors are uncorrelated (ϕ = 0), α∗ only depends on


the relative size of MSFE1 , and MSFE2 , which are commonly used
weights in empirical applications even with correlated errors.
In practice α is not known and must be estimated. An easier way to
obtain an estimate of α is to run, over the evaluation sample, the
regression
y = αŷ1 + (1 − α)ŷ2 + e, (11)
or
e2 = α(ŷ1 − ŷ2 ) + e. (12)

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 39 / 87


The combination of forecasts
In general, a lower MSFEc can be obtained by running the
unrestricted regression

y = α0 + α1 ŷ1 + α2 ŷ2 + u, (13)

with combined forecast

ỹc = α0 + α̂1 ŷ1 + α̂2 ŷ2 .

But the residuals from formula 13, i.e., b u = y − ỹc , will be in general
serially correlated when the restriction α̂1 + αˆ2 = 1 is not imposed.
Indeed,
X2 X2
u = −α̂0 + (1 −
b α̂i )y + α̂i ei .
i=1 i=1

Hence, a proper estimation method (such as GLS) should be


adopted.
c Eric Ghysels & Massimiliano Marcellino 2018 Edition 40 / 87
The combination of forecasts

In the presence of a rather large set of alternative forecasts, from a


practical point view a combined forecast obtained by simply
averaging all the alternative available forecasts tends to work well.

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 41 / 87


Forecast Evaluation and Combination

Introduction
Unbiasedness and efficiency
Evaluation of fixed event forecasts
Tests of predictive accuracy
Forecast comparison tests
The combination of forecasts
Forecast encompassing
Evaluation and combination of density forecasts
Examples using simulated data
Empirical examples
Concluding remarks

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 42 / 87


Forecast encompassing

One model encompasses another with respect to a certain property


if from the first model it is possible to deduce the property of interest
in the second model.

Forecast encompassing concerns whether the one-step forecast


of one model can explain the forecast errors made by another.

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 43 / 87


Forecast encompassing

From an operational point of view, we can use the regression


e2 = α(ŷ1 − ŷ2 ) + e and test for α = 0. If α 6= 0 the difference between
ŷ1 and ŷ2 can partly explain e2 , and therefore the second model
cannot forecast encompass the first one.

Similarly, if β 6= 0 in the regression

e1 = β(ŷ2 − ŷ1 ) + v, (14)

the first model cannot forecast encompass the second one.

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 44 / 87


Forecast encompassing

A second test can be based on the regression

e1 = δŷ2 + ϕ, (15)

and it requires δ = 0 for the second model not to forecast


encompass the first one.
A third alternative is a test for α1 = 1 , α2 = 0,

y = α0 + α1 ŷ1 + α2 ŷ2 + u.

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 45 / 87


Forecast encompassing

Even if the procedures for forecast combination and encompassing


are similar, the suggestions from the two methods are different.
The former simply indicates to combine the competing forecasts, the
latter to respecify the models that produced the forecasts, because
both of them are somewhat misspecified.

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 46 / 87


Forecast Evaluation and Combination

Introduction
Unbiasedness and efficiency
Evaluation of fixed event forecasts
Tests of predictive accuracy
Forecast comparison tests
The combination of forecasts
Forecast encompassing
Evaluation and combination of density forecasts
Examples using simulated data
Empirical examples
Concluding remarks

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 47 / 87


Evaluation and combination of density forecasts

Let us indicate the density forecast by fT+h|T , given information up


to T (and XT+h ) with horizon h, and its cumulative distribution
function (CDF) by FT+h|T . Similarly, we indicate the true density of
the target variable by gT+h|T and its CDF by GT+h|T .
In the case of the linear regression model considered in Chapter 1,
under the assumption of normal errors, we have seen that the
(optimal) density forecast (fT+h|T ) is

yT+h ∼ N(ŷT+h , V(eT+h )),

where ŷT+h = XT+h β̂T and V(eT+h ) denotes the variance of the
forecast error. The true density (gT+h|T ) is instead

yT+h ∼ N(XT+h β, σε2 ).

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 48 / 87


Evaluation of density forecasts

It is convenient to introduce the Probability Integral Transformation


(PIT), defined as
PITt (x) ≡ Ft+h|t (x), (16)
for any forecast x.
It can be shown that if Ft+h|t = Gt+h|t for all t, then the PITt s are
independent U[0, 1] variables, where U denotes the uniform
distribution.
Therefore, to assess the quality of density forecasts we can check
whether their associated PITs are independent and uniformly
distributed.

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 49 / 87


Evaluation of density forecasts

Uniformity (typically defined as probabilistic calibration) can be


evaluated qualitatively, by plotting the histogram of the PITt s for the
available evaluation sample.
For a more formal assessment of probabilistic calibration, let us
consider the inverse normal transformation:

zt = Φ−1 (PITt ), (17)

where Φ is the CDF of a standard normal variable. If PITt is


iid iid
∼ U(0, 1) then zt is ∼ N(0, 1).

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 50 / 87


Evaluation of density forecasts

It is more convenient to assess probabilistic calibration using zt s


rather than PITt s.
Mitchell and Walls (2011) provide a list of tests for uniformity and
normality.
If the zt s are indeed normally distributed and therefore
independence and lack of correlation are equivalent, we can use any
of the tests for no correlation in the errors described in Chapter 2.

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 51 / 87


Comparison of density forecasts

It is convenient to introduce the logarithmic score, defined as

log Sj (x) = log fj,t+h|t (x), (18)

where j indicates the alternative densities.

If one of the densities under comparison coincides with gt+h|t (the


true density), then the expected value of the differences in the
logarithmic scores coincides with the Kullback-Leibler Information
Criterion (KLIC):

KLICj,t = Eg [log gt+h|t (x) − log fj,t+h|t (x)] = E[dj,t (x)].

We can interpret dj,t as a density forecast error, so that the KLIC is a


kind of “mean density error”.

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 52 / 87


Comparison of density forecasts

To compare two densities, fj and fk , we can then use:

4Lt = logSj (xt ) − log Sk (xt ). (19)

To assess whether statistically the two densities are different


(basically, to construct the counterpart of the Diebold-Mariano test in
a density context), we can use:

P
4Lt
T( /[Link].) → N(0, 1),
T

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 53 / 87


Combination of density forecasts

Following, e.g.,Wallis(2005), starting from n forecast densities fj ,


j = 1, ..., n, the combined density forecast is
n
X
fc = wj fj ,
j=1
Pn
where wj ≥ 0, j = 1, ..., n, and j=1 wj = 1. The combined density fc
is therefore a finite mixture distribution.
Defining values for the weights wj , j = 1, ..., n, is not easy. A simple
solution that often works well in practice , is to set wj = 1/n, j = 1, . . . ,
n. Alternatively, the weights could be chosen optimally, to maximize
a certain objective function or minimize the KLIC with respect to the
true unknown density.

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 54 / 87


Forecast Evaluation and Combination

Introduction
Unbiasedness and efficiency
Evaluation of fixed event forecasts
Tests of predictive accuracy
Forecast comparison tests
The combination of forecasts
Forecast encompassing
Evaluation and combination of density forecasts
Examples using simulated data
Empirical examples
Concluding remarks

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 55 / 87


Examples using simulated data

In the previous three chapters we have seen forecasts produced by


linear regressions, the same models augmented with dummy
variables and also models with lagged variables,

yt = α1 + α2 xt + εt
yt = α1 + β1 Dt + α2 xt + β2 Dt xt + εt
yt = α1 + β1 Dt + α2 xt + β2 Dt xt + γ1y yt−1 + γ1x xt−1 + εt

Naturally, model (mis-)specification affects the forecast accuracy. In


this chapter we will compare the performance of all three
specifications.

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 56 / 87


Examples using simulated data

Table 1 represents simple forecast evaluation statistics for all the


models considered. The RMSFE and MAFE are the smallest for the
dynamic regression model.
Forecasting Model RMSFE RMSFE MAFE MAFE
recursive recursive

Linear Regression 9.536 9.490 7.995 7.950


Dummy variable model 10.130 9.642 8.314 7.938
Dynamic model 1.108 0.965 0.865 0.760
Table 1: Forecasts evaluation: RMSFE and MAFE

The next step is to compare individual forecasts. In section 5 we


discussed two main tests: the Diebold-Mariano (DM) and
Morgan-Granger-Newbold (MGN) tests. We start with the DM test.

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 57 / 87


Diebold - Mariano(DM) test

The null hypothesis of the DM test is that the average loss


differential between the forecasts of compared models is equal to
zero. The DM test results for all three models are described in Table
2.
Comparisons M1 vs M2 M1 vs M3 M2 vs M3
DM test statistics -1.113 8.582 11.445
P-value 0.132 0.000 0.000
Table 2: Diebold-Mariano test for equality of MSFE

This result indicates that at 5% significance level, there is no


difference between forecasts of Model and 2, but the dynamic
models provides statistically significant better forecasts.

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 58 / 87


Examples using simulated data

The null hypothesis of the MGN test is that the MSFEs associated
with two forecasts are equal. The test results appear in Table 3 and
are in line with the DM test findings.
Comparisons M1 vs M2 M1 vs M3 M2 vs M3
MGN test statistics -1.100 42.932 65.137
P-value 0.273 0.000 0.000
Table 3: Morgan-Granger-Newbold test

We also compare recursive forecasts using DM and MGN tests.


The results are omitted as they convey the same message.

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 59 / 87


Unbiasedness tests

Let us reconsider the formula 1a :

yi+h = α + βŷi+h|i + εi+h , i = T, ..., T + H − h, h<H

Unbiasedness test can be implemented with a t−test (we do not


need a robust one with h = 1) in the following regression:

et+1 = yt+1 − ŷt+1|t = τ + εt+1 .

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 60 / 87


Unbiasedness tests

The results for Models 1 and 2 appear in respectively Tables 4 and 5.


Coefficient Std. Error t-Statistic Prob.

τ -1.744 0.953 -1.830 0.069

R-squared 0.000 Mean dep var -1.744


Adjusted R-squared 0.000 [Link] var 9.609
[Link] regression 9.609 Akaike IC 7.368
Sum squared resid 18375.640 Schwarz IC 7.385
Log likelihood -735.834 Hannan-Quinn 7.375
DW stat 0.790
Table 4: Unbiasedness and weak efficiency (via DW) tests for Model 1

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 61 / 87


Unbiasedness tests

Coefficient Std. Error t-Statistic Prob.

τ -3.223 0.810 -3.979 0.000

R-squared 0.000 Mean dep var -3.223


Adjusted R-squared 0.000 [Link] var 9.625
[Link] regression 9.625 Akaike IC 7.372
Sum squared resid 18436.060 Schwarz IC 7.388
Log likelihood -736.162 Hannan-Quinn 7.378
DW stat 1.369
Table 5: Unbiasedness and weak efficiency (via DW) tests for Model 2

For M1 we accept the null τ = 0 at the 5 % level, while for M2 we


clearly reject the null. This means that forecasts for M1 appear to be
unbiased, while the opposite is true for M2.

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 62 / 87


Weak efficiency tests

For weak efficiency test, we can look at the same regression output
as the DW statistic tells us whether there is order-one
autocorrelation in the εt , which amounts to testing the
autocorrelation of et .

There is evidence for serial correlation in the forecast errors for both
models.

In general, with h > 1 we need to fit a moving average model of


order h − 1, MA(h − 1), to the h-steps ahead forecast error and test
that the resulting residuals are white noise.

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 63 / 87


Forecast Evaluation and Combination

Introduction
Unbiasedness and efficiency
Evaluation of fixed event forecasts
Tests of predictive accuracy
Forecast comparison tests
The combination of forecasts
Forecast encompassing
Evaluation and combination of density forecasts
Examples using simulated data
Empirical examples
Concluding remarks

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 64 / 87


Forecasting Euro area GDP growth

Using again the Euro area data series, we consider the Models 1
through 4, estimated over the period 1996Q1 to 2006Q4; and using
2007Q1 to 2013Q2 as the forecast evaluation period. Recall that
Models 1 through 4 and the ARDL Model 2 contain different
elements in their information sets:

yt = α + Xt β + εt

Model 1: Xt = (iprt , sut , prt , srt )


Model 2: Xt = (iprt , sut , srt )
Model 3: Xt = (iprt , sut )
Model 4: Xt = (iprt , prt , srt )
ARDL Model 2: : Xt = (yt−1, iprt , sut−1, srt−1 )

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 65 / 87


Forecasting Euro area GDP growth
To proceed we need to re-estimate the ARDL Model 2, presented in
Chapter 3, over the period 1996Q1 to 2006Q4 using this estimation
sample, with the results appearing in Table 6.
Variable Coefficient Std. Error t-Statistic Prob.

C 0.333 0.083 4.037 0.000


Y(-1) 0.197 0.139 1.413 0.166
IPR 0.201 0.069 2.907 0.006
SU(-1) -0.002 0.018 -0.140 0.889
SR(-1) 0.009 0.006 1.435 0.159

R-squared 0.513 Mean dep var 0.564


Adjusted R-squared 0.463 S.D. dep var 0.355
S.E. of regression 0.260 Akaike IC 0.251
Sum squared resid 2.638 Schwarz IC 0.454
Log likelihood -0.518 Hannan-Quinn 0.326
F-statistic 10.256 DW stat 2.220
Prob(F-statistic) 0.000
Table 6: Estimation output: ARDL Model 2

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 66 / 87


Forecasting Euro area GDP growth

ARDL Model 2 has a better in sample than Models 1 - 3 in terms of


2
AIC, and in most cases, R2 and R . By the same measures, it seems
to be not as good as Model 4.

However, that a good in-sample fit does not necessarily imply a


good out-of-sample performance. To compare the forecasts
produced by these models, we need to look at forecast sample
evaluation statistics and also to perform forecast comparison tests.

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 67 / 87


Forecasting Euro area GDP growth

We first produce one-step ahead static forecasts for the period


2007Q1 to 2013Q2 and compute the RMSFE and the MAFE of the
forecasts produced by each of the models, as reported in Table 7.
M1 M2 M3 M4 ARDL M2
RMSFE 0.383 0.378 0.394 0.415 0.353
MAFE 0.327 0.331 0.344 0.364 0.310
Table 7: Forecast evaluation measures

One-step ahead forecasts produced by the ARDL Model 2


outperforms the other forecasts as indicated by both the values of
RMSFE and MAFE.

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 68 / 87


Forecasting Euro area GDP growth

Unbiasedness can be tested by regressing each forecast error on


an intercept and then test its significance. It turns out that it is
significant for all the models, indicating the presence of a bias.

Weak efficiency is not rejected for any of the models, however, as


the one-step ahead forecasts errors are all serially uncorrelated.

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 69 / 87


Forecasting Euro area GDP growth

Assuming the loss function is quadratic, we first compare the


one-step ahead static forecasts using the use of the
Diebold-Mariano test. The test statistics are reported in Table 8.
vs M1 vs M2 vs M3 vs M4
Test Stat. 0.721 0.737 1.331 1.535
p-value 0.235 0.230 0.091 0.062
Table 8: DM tests ARDL Model 2 against Models 1 through 4

The null hypothesis is that the average loss differential, d, is equal to


zero, i.e., H0 : d = 0. The test statistic has an asymptotic standard
a
normal distribution, i.e., DM ∼ N (0, 1) .

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 70 / 87


Forecasting Euro area GDP growth

If we consider a significance level of 5%, the null hypothesis that the


loss differential between the forecasts from ARDL Model 2 and
those of the other models being zero cannot be rejected.

However, at 10% significance level, the test results suggest the loss
differential between the forecasts from ARDL Model 2 and those of
Model 3 and Model 4.

Moreover, the reported test statistics here are all positive. This is an
indication that the loss associated with Model 1, 2, 3, and 4 is larger
than that of ARDL Model 2.

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 71 / 87


Forecasting Euro area GDP growth

The computed test statistics and the corresponding p-values of


Morgan-Granger-Newbold (MGN test) are reported in Table 9.
vs M1 vs M2 vs M3 vs M4
Test Stat. 1.756 0.735 0.900 1.651
p-value 0.090 0.468 0.375 0.110
Table 9: MGN tests ARDL Model 2 against Models 1 through 4

The null hypothesis is that the mean of the loss differential is zero,
meaning the variances of the two forecast errors are [Link] test
statistic follows a Student t distribution with H − 1 degree of freedom
(in this case H = 26).

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 72 / 87


Forecasting Euro area GDP growth

At 5% significance level, the null hypothesis cannot be rejected for


all four cases.

Note however, that the null hypothesis can be rejected for the first
case at the 10% significance level.

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 73 / 87


Forecasting Euro area GDP growth

We now look into whether there is any change in the results if we


consider recursive forecasts. Tables 10 and 11 present the test
results.
vs M1 vs M2 vs M3 vs M4
Test stat. 1.716 0.737 1.331 1.535
p-value 0.043 0.230 0.091 0.062
Table 10: DM tests recursive forecasts ARDL Model 2 vs Models 1 - 4

vs M1 vs M2 vs M3 vs M4
Test stat. 1.898 1.281 1.208 1.693
p-value 0.068 0.211 0.237 0.102
Table 11: MGN tests recursive forecasts ARDL Model 2 vs Models 1 - 4

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 74 / 87


Forecasting Euro area GDP growth

Looking at the Diebold-Mariano test statistics, the null hypothesis


that the loss differential between the one-step ahead recursive
forecasts from ARDL Model 2 and those from Models 2 through 4
being zero cannot be rejected at the 5% significance level.

Whereas the null hypothesis that the one-step ahead recursive


forecasts from ARDL Model 2 are no difference from those of Model
1 can be rejected at 5% significance level.

The Morgan-Granger-Newbold test results confirm this finding at the


10 % level.

At the 10 % level we can also reject the DM test null for the ARDL
Model 2 against Models 3 and 4.

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 75 / 87


Forecasting Euro area GDP growth

We can construct pooled forecasts by combining the 5 available


predictors, using equal weights for simplicity and as the sample size
is rather short.

It turns out that the MSFE of the combined forecasts is lower than
that of the forecasts associated with Models 1 through 4, but larger
than that of the ARDL Model 2.

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 76 / 87


Forecasting US GDP growth

Using again the US data series employed in the previous chapters,


we consider Models 1 through 4 and ARDL Models , estimated over
the period 1985Q1 to 2006Q4; and using 2007Q1 to 2013Q4 as the
forecast evaluation period.
Model 1: Xt = (iprt , sut , prt , srt )
Model 2: Xt = (iprt , sut , srt )
Model 3: Xt = (iprt , sut )
Model 4: Xt = (iprt , prt , srt )
ARDL Model 2: : Xt = (yt−1, iprt , sut−1, srt−1 )

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 77 / 87


Forecasting US GDP growth

We compute the RMSFE and the MAFE of the forecasts produced


by these models, as reported in Table 12.
M1 M2 M3 M4 ARDL M2
RMSFE 0.537 0.531 0.539 0.537 0.567
MAFE 0.402 0.399 0.403 0.418 0.438
Table 12: Simple forecast evaluation statistics

It is clear that one-step ahead forecasts produced by the ARDL


Model 2 is outperformed by the other forecasts

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 78 / 87


Forecasting US GDP growth

We first compare the one-step ahead static forecasts from ARDL


Model 2 with those of Model 1, 2, 3 and 4, using Diebold - Mariano
(DM) test and Morgan - Granger - Newbold (MGN) test.
Table 13 reports the test statistics and the corresponding p-values of
DM test.
vs M1 vs M2 vs M3 vs M4
Test stat. -1.516 -1.748 -1.375 -2.279
p-value 0.064 0.040 0.084 0.011
Table 13: DM test on one-step ahead forecasts from ARDL model 2 against one-step ahead
forecasts from Model 1 through 4.

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 79 / 87


Forecasting US GDP growth

At 5% significance level, the null hypothesis that the loss differential


between the forecasts from ARDL Model 2 and those of the other
models being zero can be rejected for Models 2 and 4.
Moreover, at 10% significance level, the test results suggest the loss
differential between the forecasts from ARDL Model 2 and all
models is non-zero.
Note, however, that the reported test statistics here are all negative.
This is an indication that the loss associated with Model 1, 2, 3 and 4
is smaller than that of ARDL Model 2.

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 80 / 87


Forecasting US GDP growth

The computed test statistics and the corresponding p-values of


MGN test are reported in Table 14.
vs M1 vs M2 vs M3 vs M4
Test stat. -1.109 -1.308 -0.946 -2.007
p-value 0.276 0.201 0.352 0.054
Table 14: MGN test on one-step ahead forecasts from ARDL model 2 against one-step ahead
forecasts from Model 1 through 4

The null hypothesis of the MGN test is that the mean of the loss
differential is zero, meaning the variances of the two forecast errors
are equal. The null hypothesis cannot be rejected for first 3 cases,
and can be rejected in the last case at the 10% significance level.

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 81 / 87


Forecasting US GDP growth

We now look into whether there is any change in the results if we


consider recursive forecasts. Tables 15 and 16 present the test
results.
vs M1 vs M2 vs M3 vs M4
Test stat. -1.473 -1.748 -1.375 -2.279
p-value 0.070 0.040 0.084 0.011
Table 15: DM test on one-step ahead recursive forecasts from ARDL Model 2 against those
from Model 1 through 4

vs M1 vs M2 vs M3 vs M4
Test stat. -1.258 -1.243 -0.735 -1.641
p-value 0.219 0.224 0.468 0.112
Table 16: MGN test on one-step ahead recursive forecasts from ARDL Model 2 against those
from Model 1 through 4

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 82 / 87


Forecasting US GDP growth

Looking at the Diebold-Mariano test statistics, the null hypothesis


that the loss differential between the one-step ahead recursive
forecasts from ARDL Model 2 and those from Model 2, 3, and 4
being zero can be rejected at 10% significance level.

However, the Morgan-Granger-Newbold test results show that the


null hypothesis of the mean of the loss differential being zero cannot
be rejected at 10% significance level for all 4 cases.

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 83 / 87


Default risk

Let us now revisit the empirical default risk models.


ARDL Model 1: OASt = α + β1 OASt−1 + β2 VIXt−1 + εt
ARDL Model 2: OASt = α + β1 OASt−1 + β2 SENTt−1 + εt
ARDL Model 3: OASt = α + β1 OASt−1 + β2 PMIt−1 + εt
ARDL Model 4: OASt = α + β1 OASt−1 + β2 SP500t−1 + εt

We found that the in-sample Model 4 featured the best fit, whereas
Model 3 had the best out-of-sample RMSFE.

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 84 / 87


Default risk
We use the Morgan-Granger-Newbold test for pairwise
comparisons of the one-step ahead forecasts.
Model 1 Model 2 t-stat p-val Superior

VIX SENT 0.505 0.617 2


VIX PMI 0.575 0.569 2
VIX SP500 -1.042 0.305 1
SENT PMI 0.662 0.513 2
SENT SP500 -1.410 0.168 1
PMI SP500 -1.463 0.153 1
Table 17: MGN tests default risk models

The results appear in Table 17. The results indicate that the
out-of-sample differences across the different models is not
statistically significant. The DM statistics also confirm the same
finding.
c Eric Ghysels & Massimiliano Marcellino 2018 Edition 85 / 87
Forecast Evaluation and Combination

Introduction
Unbiasedness and efficiency
Evaluation of fixed event forecasts
Tests of predictive accuracy
Forecast comparison tests
The combination of forecasts
Forecast encompassing
Evaluation and combination of density forecasts
Examples using simulated data
Empirical examples
Concluding remarks

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 86 / 87


Concluding remarks

The main focus in this book is forecasting using econometric


models. It is important to stress, however, that the methods reviewed
in this chapter developed for the purpose of evaluating forecasts
apply far beyond the realm of econometric models. Indeed, the
forecasts could simply be, say, analyst forecasts at least not
explicitly related to any specific model. Hence, the reach of the
methods discussed in this chapter is wide.

c Eric Ghysels & Massimiliano Marcellino 2018 Edition 87 / 87

You might also like