Economic Forecasting: Methods & Evaluation
Economic Forecasting: Methods & Evaluation
February 7, 2007
Abstract
Forecasts guide decisions in all areas of economics and finance and their value can only be
understood in relation to, and in the context of, such decisions. We discuss the central role
of the loss function in helping determine the forecaster’s objectives. Decision theory provides
a framework for both the construction and evaluation of forecasts. This framework allows an
understanding of the challenges that arise from the explosion in the sheer volume of predic-
tor variables under consideration and the forecaster’s ability to entertain an endless array of
forecasting models and time-varying specifications, none of which may coincide with the ‘true’
model. We show this along with reviewing methods for comparing the forecasting performance
of pairs of models or evaluating the ability of the best of many models to beat a benchmark
specification.
1 Introduction
Forecasting problems are ubiquitous in all areas of economics and finance where agents’ decisions
depend on the uncertain future value of one or more variables of interest. When a household decides
how much labor to supply or how much to save for a rainy day, this presumes an ability to forecast
a stream of future wages and returns on savings. Similarly, firms’ choice of when to invest, how
much to invest and how to finance it (the capital structure decision) depends on their forecasts
of future cash flows from potential investments, future stock prices and interest rates. Indeed, all
present value calculations, and hence the vast majority of questions in asset pricing, have embedded
in them forecasts of future cash flows generated by uncertain payoff streams. In public finance,
decisions on whether to go ahead with large infrastructure projects such as the construction of a
new bridge or a tunnel require projecting traffic flows and income streams over the project’s lifetime
which may well be several decades.
Recent research has seen a virtual revolution in how economists compute, apply and evalu-
ate forecasts. This research has occurred as a result of extensive developments in information
technology that have opened access to thousands of new potential predictor variables (including
tick-by-tick trading data, disaggregate survey forecasts and real-time macroeconomic data) and a
∗
We thank the editor, Roger Gordon, three anonymous referees and Lutz Kilian, Michael McCracken, Barbara
Rossi and Norm Swanson for providing detailed comments on the paper. We also thank Gray Calhoun for excellent
research assistance.
1
wealth of new techniques that facilitate search over and estimation of the parameters of increasingly
complicated forecasting models. Questions such as which particular predictor variables to include,
which functional form to use for the forecasting model and how to weight old versus more recent
data have become an essential part of forecast construction and evaluation.
Economic forecasting is unique in that forecasters are forced to ‘show their hand’ in real time
as they generate their forecasts. Future outcomes of most predicted variables are observed within
a reasonable period of time, so a direct sense of how well a forecasting model performed can be
gained. If forecasting performance is poor, this will become clear to the forecaster once data on
realizations of the predicted variable is revealed. This real time feedback may in turn lead to a
change in the forecasting model itself, thus posing unique challenges to the process of evaluating how
fast the forecaster is learning over time. This is in stark contrast to many econometric problems.
For example, evaluation of an estimate of the effect of schooling on wages may take generations.
In many economic problems we do not obtain an objective confirmation of how good the original
estimate is since we do not have reference data for evaluating the economic prediction.
Often the result of the feedback from forecasts has been disheartening, both to econometricians
trying to utilize data as efficiently as possible and to economists whose theories result in predictions
that appear unable to explain as much of the variation in the data as they had hoped. To take one
example, a seemingly simple task such as estimating the weights on different models in least squares
forecast combination regressions is commonly outperformed on real data by using a simple equal-
weighted average of forecasts (Clemen (1989)). An infamous result of Meese and Rogoff (1983)
shows that despite a great deal of theoretical work on exchange rates−and even with the benefit of
using future data suggested by theory as relevant−the random walk ‘no change’ prediction cannot
be beaten. This result has to a great extent held up for exchange rate forecasts (Kilian (1999)).
While the performance of a forecasting model often can be observed fairly quickly, only limited
economic conclusions can be drawn from the model’s historical track record. Forecasting models
are best viewed as greatly simplified approximations of a far more complicated reality and need
not reflect causal relations between economic variables. Indeed, simple mechanical forecasting
schemes−such as the random walk−are often found to perform well empirically although they do
not provide new economic insights into the underlying variable (Clements and Hendry (2002)).
Conversely, models aimed at uncovering true unconditional relationships in the data need not be
well suited for forecasting purposes.
In a unified framework this paper provides an understanding of the properties, construction
and evaluation of economic forecasts. Our objective is to help explain differences among the many
approaches used by various researchers and understand the breadth of results reported in the
empirical forecasting literature.
Our coverage emphasizes the importance of integrating economic forecasts (including model
specification, variable selection and parameter estimation) in a decision theoretical framework.
This is, in our view, the defining characteristic of economic forecasts. From this perspective,
forecasts do not have any intrinsic value and are only useful in so far as they help improve economic
decisions. What constitutes a good forecast depends on how costly various prediction errors are
to the forecaster and hence reflects both the forecaster’s preferences and the manner in which
forecasts are mapped into economic decisions. Economic forecasting is not an exercise in modelling
the data disjoint from the purpose of the forecast provision. Section 2 illustrates these points
2
initially through two examples from economics and finance.
We next provide a formal statement of the forecasting problem. Section 3 reviews both classical
and Bayesian approaches to forecasting and introduces the individual components of the economic
forecasting problem such as the forecaster’s prediction model, the underlying information set as
well as the forecaster’s loss function. Although often only treated implicitly, the loss function is
essential to all forecasting problems and so we devote Section 4 to a deeper discussion of various
types of loss functions and the restrictions and assumptions they embody.
Using the decision theoretic framework set out in Section 3, Section 5-7 review several issues
that arise in the practical construction of economic forecasts. Each of these topics has been active
areas of research in recent years. An over-arching problem in economic forecasting is the myriad
of data that a forecaster could potentially employ. Estimating models with a large number of
parameters relative to the sample size undermines one of the central methods of econometrics–
OLS justified through properties such as asymptotic efficiency of the parameter estimates–and
opens the possibility that other estimation techniques are better suited to the task of constructing
forecasts. In concert with differences over loss functions this provides a partial explanation of the
myriad of estimation methods seen in practice.
Section 5 addresses the choice of functional form of the forecasting model. Lack of guidance from
economic theory is often an issue and so the functional form is commonly chosen on grounds such
as empirical ‘fit’ or an ability to capture certain episodes in the historical data sample. We review
several methods aimed at approximating unknown functional forms in a parsimonious, yet flexible
manner−a task made essential by the short samples available in most forecasting applications.
A related problem is related to how the forecasting model and the underlying data evolve over
time. When forecasting models are viewed as simple approximations to a complex and evolving
reality that changes due to shifts in legislation, institutions and technology−or even wars and
natural catastrophes−it is to be expected that the ‘true’ but unknown data generating process
changes over time. In the forecasting literature this has been captured through various approaches
that deal with model- and parameter instability. Since all estimation techniques essentially average
over past data to obtain a forecasting model, this raises the problem of exactly how to choose the
data sample and how to weight ‘old’ versus ‘new’ data. Other approaches attempt to directly model
breaks in the model parameters in order to increase the effective data sample. These are covered
in Section 6.
An important part of the analysis of economic forecasts is to assess how good they are. Until
recently, forecasts were largely evaluated without the use of standard errors that account for pa-
rameter estimation error and model specification search. It is well understood that data mining
programs that search over many models tend to overfit and hence inflate estimates of forecasting
performance by using the same data for estimation and evaluation purposes. However, little or
no account is typically made for such model search that precede the analysis. Standard practice
for dealing with data mining has been to hold back some data and check whether the forecasting
model still performed well in future out-of-sample periods. Again, often average losses are com-
pared without any regard to pre-testing biases. Recent work has resulted in methods that account
for sampling error in various forecasting situations. These are reviewed in section 7.
Economic theory rarely identifies a single forecasting model that works well in practice and
leaves open many degrees of freedom in forecast construction. The resulting plethora of economic
3
forecasting models has given rise to procedures for forecast comparisons that can handle even
situations with a very large set of models. An alternative to evaluating particular models and
attempting to select a single dominant model is to average over various forecasting methods. Both
forecast comparison and forecast combination are reviewed in Section 8. Section 9 provides an
empirical analysis of forecasts of inflation and stock returns. Finally, Section 10 concludes.
4
conditional forecasts are usually computed in the context of a structural model for the economy.
Since current and future interest rates are affected by the central bank’s decisions, the central
bank’s forecasting problem cannot be separated from its decisions regarding current and future
interest rates.
Central bankers come with certain subjective views about how the economy operates which they
may wish to impose on their forecasting model−a theme that naturally leads to Bayesian forecasting
methods or other methods that can trade off theoretical and empirical coherence. Should the central
bank use a simple vector autoregression (VAR) fitted to historical data or maybe as a way to capture
expectations as was done at least at some point by the FRB/US model used at the Board (Brayton
and Tinsley (1996)) and thus use a model tailored to fit historical features of the data? Should
it use a more theoretically coherent dynamic stochastic general equilibrium (DSGE) model? Or,
should it use some combination of the two? If the central bank adjusts forecasts from a formal
model using judgemental information, an additional issue arises, namely how much weight to assign
to the data versus the judgemental forecast. Implicitly or explicitly, such weights will reflect the
bank’s prior beliefs.
Model instability or ‘breaks’ are likely to be empirically relevant for central bankers trying to
forecast inflation. In fact, inflation appears to be among the least stable macroeconomic variables
exactly because it depends on monetary policy regimes, macroeconomic shocks and other factors.
Stock and Watson (1999b) report evidence of instability in the parameters of the Phillips curve.
Quite frequently forecasters find themselves in situations that differ in important regards from
the historical sample used to estimate their forecasting models. Pagan (2003) refers to the difficulties
and uncertainties the Bank of England faced in their forecasts following the events of 11 September,
2001. Indeed, an important part of maintaining a good forecasting model is to monitor and evaluate
its performance both historically and in real time. Because past forecast errors have often been
found to have predictive power over future errors, monitoring for serial correlation in forecast errors
potentially offers a simple way to improve upon a forecast. More generally, if the process generating
the predicted variable is subject to change, it is conceivable that a forecasting model that performed
well historically may have failed to do so in the more recent past.
5
weights. Due to estimation error, often the raw estimates are shrunk towards their values implied
by a simple benchmark model such as the CAPM (Ledoit and Wolf (2003)). Alternatively, the
investor’s choice variables−the portfolio weights−can be restricted through short sale restrictions
and maximum holding limits (Jagannathan and Ma (2003)).
If the mean and variance of returns are allowed to depend on time-varying state variables,
the question immediately arises which state variables to select among interest rates (levels and
spreads), macroeconomic activity variables, technical variables such as price momentum or rever-
sals, valuation measures such as the price-earnings or book-to-market ratios or the dividend yield
etc. Asset pricing theory provides little guidance to the exact identity of the relevant state variables;
this raises several questions such as how to avoid over-fitting the forecasting model−a risk always
encountered when multiple prediction models are considered−and how to assess the forecasting
models’ performance against a benchmark strategy such as simply holding the market portfolio.
Another problem that is more unique to forecasting models for financial returns is that any
predictability patterns that do not capture time-varying risk premia must, if markets are efficient,
be non-stationary because their discovery should lead to their self-destruction once investors act to
take advantage of such predictability. For example, there is evidence suggesting that popular models
for predicting stock returns based on the dividend yield ceased to be successful at some point during
the nineties, perhaps because of changes in firms’ dividend payout and share repurchase practices
or perhaps because investors incorporated earlier evidence of predictability. Only if a model’s
forecasting performance is tracked carefully through time can this sort of evidence be uncovered.
These examples indicate the complexity of many of the issues involved in economic forecasting.
To further understand these points, we next provide a formal statement of the objectives underlying
the calculation of actual forecasts.
6
forecasting model nonparametrically. However, this ignores the short data samples and the large
dimension of the set of potential predictor variables in most empirical forecasting problems. In
practice a flexible parametric forecasting model is often the best one can hope to achieve.
A third element concerns which type of information to report for the outcome of interest. We
could report a single number (point estimate), a range estimate or perhaps an estimate of the
full probability distribution of all possible values. Most of the theory of frequentist forecasting has
been directed towards point forecasting. A more recent literature has examined interval forecasts or
forecasts of the conditional distribution of the variable of interest rather than a summary statistic.
From a Bayesian perspective similar issues arise, although it is natural in this approach to provide
the full predictive distribution.
3.1 Notation
Throughout the analysis we let Y be the random variable that generates the value to be forecast.
To begin with, we restrict attention to point forecasts, f , which are functions of the available data
at the time the forecast is made. Hence, if we collect all relevant information at the time of the
forecast into the outcome z of the random variable Z, then the forecast is f (z). Discovering which
variables are informative from a forecasting standpoint is important in practice. We examine this
in greater depth later on but for now simply think of this information as being incorporated into
some random variable that generates our data. Exactly how z maps into the forecast f (z) depends
on a set of unknown parameters, θ, that typically have to be estimated from data. We emphasize
this by writing f (z, θ).
The loss function is a function L(f, Y, Z) that maps the data, Z, outcome, Y , and forecast,
f , to the real number line, i.e. for any set of values for these random variables the loss function
returns a single number. The loss function describes in relative terms how bad any forecast might
be given the outcome and possibly other observed data that accounts for any state dependence in
the loss. The loss function and its properties are examined in greater detail in Section 4.
7
chosen, f ). That this is a function of θ will become clear below. A sensible rule has low risk and
minimizes expected loss or equivalently maximizes expected utility.1
Assuming the existence of a density for both Y given Z and for Z (denoted pY (y|z, θ) and
pZ (z|θ), respectively), we can write the risk as
Z Z
R(θ, f ) = L(f (z, θ), y, z)pY (y|z, θ)pZ (z|θ)dydz. (2)
z y
It is this risk that forecasting methods – methods for choosing f (z, θ) – attempt to control and
minimize.
Forecasts are generally viewed as ‘poor’ if they are far from the observed realization of the
outcome variable. However, as is clear from (2), point forecasts aim to estimate not the realization
of the outcome but rather a function of its distribution. The inner integral in the risk function (2)
is EY [L(f (z, θ), Y, Z)|Z] which removes Y from the expression leaving the loss function relating
the forecast to θ and the realization of the data z. For example in the case of mean squared loss
this is the variance of Y given Z plus the squared difference between the forecast and the mean of
Y given Z. Both the conditional mean and variance of Y are functions of θ and z.
As noted in section 2.1, we are often interested in conditional forecasts, i.e. forecasts of Y
conditional on a specific path taken by another random variable, W . In the above analysis and
in what follows, the results can be extended to this case by replacing pY (y|z, θ) and pZ (z|θ) with
the distributions conditional on the outcome of W being set to w, i.e. pY (y|z, w, θ) and pZ (z|w, θ).
While this extension may seem trivial conceptually, it can be difficult to implement in practice.
For example consider a VAR in interest rates and inflation, where we want to predict inflation one
period ahead conditional on the value of interest rates one period ahead. The density of future
inflation given future interest rates as well as current and past values of both variables is relatively
simple to write down or estimate and is the density of a rotated VAR. However, the density of past
values of both variables conditional on future interest rates presents some difficulties in practice,
see, e.g. Waggoner and Zha (1999) for a Bayesian example.
where L0 (evaluated using the data) is often called the generalized forecast error. Assuming squared
1
This representation of the problem limits further choices of the loss function and requires assumptions on the
underlying random variables to ensure that the risk exists.
8
loss in the difference between the forecast and outcome, we have
Z
L(f (z, θ), y, z)py (y|z, θ)dy = E[Y − f (Z, θ)|Z = z, θ]2 . (4)
y
This is minimized by choosing f (z, θ) = E[Y |Z = z, θ], i.e. the conditional mean as a function of
the data, Z, the forecasting model and its parameters, θ.
In practice the parameters θ are almost always unknown and so the second step in the classical
approach involves selecting a ‘plug-in’ estimator for θ. The resulting estimator θ̂(z) is a function
of the data z and hence the forecasting rule f (z, θ̂(z)) is only a function of the observable data.
For example under squared loss and z = {yt , xt }Tt=1 , when the conditional mean of YT +1 is θ 0 xT
we might use OLS estimates from a regression of yt+1 on xt over the available sample as the plug-in
P P
estimator for θ. The forecast is then f (z, θ̂(z)) = ( Tt=2 x0t−1 yt )0 ( Tt=2 x0t−1 xt−1 )−1 xT . Alternative
plug-in estimators are discussed in detail below.
In choosing between plug-in estimators, one approach is to examine the risk functions for the
various methods, R(θ, f ). These are functions of both the method f and the parameters θ. Typically
no risk function dominates uniformly over all θ, i.e. some are better for some values of θ but work
less well for other values. The classical forecaster could then choose a method that minimizes worst
case risk or could alternatively consider a weighting function over θ, choosing the best method
for that particular weighting. Denoting the weighting function by π(θ), one would choose the
R
method that minimizes R(θ, f )π(θ)dθ, i.e. the risk averaged over all models that are thought to
be important.
9
3.5 Relating the Methods
The complete class theorem tells us that if the classical method does not correspond to a Bayesian
procedure for some prior, then it is inadmissible. Under the same problem setting (i.e. for identical
loss function and densities), one could therefore find a Bayesian method with equal or smaller risk
than the classical procedure for all possible values of θ. Conversely, if there is equivalence between
the two methods, then the classical approach cannot be beaten. Admissibility is only interesting for
weights π(θ) relevant to the forecasting problem, but will of course be preferred given such weights.
The first problem that can arise in the classical setting is that the ‘plug in’ method of con-
structing f (z, θ) using the estimate θ̂(z) is ad hoc. Often forecasters choose estimators that yield
nice properties of the parameters themselves. For example estimators that are consistent, asymp-
totically normal and asymptotically efficient for θ may be employed. However, because the goal of
forecasting is not to estimate θ, but to construct the forecast, f (z, θ), such methods may not yield
good forecasting rules even for reasonable weighting functions π(θ).
Some practical considerations cloud the picture. In practice, differences between the risks of
optimal and ad hoc methods need not be large enough to justify differences in computational costs.
Moreover, such comparisons require that the model be correctly specified which is against both
the spirit and practice of modern forecasting. Still, the Bayesian approach offers a construction
method that is guaranteed to be admissible for the specified model. Even if the true model is not
necessarily the one used to construct the forecasts, provided that the forecasting model is close to
the true model we are assured to use a method that works well for this possible true model.
It follows from this discussion that it is difficult to find optimal solutions even for very simple
forecasting problems. Furthermore, even for simple models such as linear regression models OLS
may not be the best approach (this is discussed at length in section 5.1). Hence for various combi-
nations of distributions of the data and values of the parameters of the model there is leeway for
alternative methods to dominate. This lack of a single dominant approach explains much of the
interest in different forecasting approaches seen in the last two decades.
Closed form solutions are not always available, so numerical integration over the density forecasts
is often required to evaluate the risk.
Forecasters with different loss functions will generally construct different optimal forecasts even
though the density for the data is the same for each of them. For example, suppose that the cost
of a forecast error, e = y − f , is (1 − α)|y − f | for y < f and α|y − f | for y ≥ f . Then Koenker
and Bassett (1978) show that (in the absence of z) the optimal forecast is f = F −1 (α) where F is
the cumulative distribution function of a continuous outcome variable Y and F −1 is the so-called
10
quantile function. The higher the relative cost of positive forecast errors (higher α), the larger the
optimal forecast and hence the smaller the probability of observing costly positive forecast errors.
Two forecasters with different loss functions in this family (different values of α) will want
different quantiles of the distribution, so an agency that merely reports a single number could
never give them both the optimal forecast. It would be sufficient to provide the entire distribution
(density forecast) because this has all the quantile information and hence works for any piece-wise
linear loss function. Of course, under MSE loss, all forecasters will agree that only the mean of the
predictive density is required.
Under the Classical approach, a plug-in estimate of θ is generally used to construct the predictive
density. The Bayesian equivalent to this approach is to provide the predictive density by removing θ
through integration over the prior distribution π(θ) rather than through estimation. The Bayesian
chooses f (z) to minimize
Z µZ ½Z ¾ ¶
r(π, a) = L(f (z), y, z)pY (y|z, θ)dy pZ (z|θ)π(θ)dθ dz
z θ y
Z µZ ½Z ¾ ¶
= m(z) L(f (z), y, z)pY (y|z, θ)dy π(θ|z)dθ dz
z θ y
Z µZ ½ Z ¾ ¶
= m(z) L(f (z), y, z) pY (y|z, θ)π(θ|z)dθ dy dz
z y θ
Z µZ ¶
= m(z) {L(f (z), y, z)pY (y|z)} dy dz. (7)
z y
R
Now pY (y|z, θ)π(θ|z)dθ = pY (y|z) is the predictive density obtained by integrating over θ using
π(θ|z) as weights. Conditioning on information known to the forecaster, z, the optimal forecast
minimizes the bracketed integral.
4 Loss Functions
Short of the special (and uninteresting) case with perfect foresight it will not be possible to find
a forecasting method that always sets f (z, θ) equal to the outcome y. A formal method of trading
off potential forecast errors of different signs and magnitudes is therefore required. This is the role
of the loss function which describes in relative terms how costly any forecast is given the outcome
and possibly other observed data. In mathematical terms, the loss function L(f (Z, θ), Y, Z) maps
the data, outcome and forecast to the real number line, i.e. for any set of values of these random
variables the loss function returns a single number.
Forecasters thus must pay attention to how errors will affect their results, which means con-
structing a mathematical representation of potential losses. A natural foundation for a loss function
is a utility function that involves both the outcome and the forecast. For a given indirect utility
function U (f (Z, θ), Y, Z), we can set the loss L(f, Y, Z) = −U (f, Y, Z) in order to see how one
might elicit the loss function (Granger and Machina (2006), Skouras (2001)).
11
4.1 Determinants of the Loss Function
Economic insights about the forecasting problem should be used to guide the choice of loss function
in any given situation. In particular, the choice of loss should address issues such as (i) the
relative cost of over- and underpredicting the outcome variable, i.e. the issue of symmetric versus
asymmetric loss; (ii) how economic decisions are influenced by the forecast, which may involve
strategic considerations; and (iii) which variables affect the forecaster’s loss, i.e. forecast error
alone or perhaps the level of the predicted variable matters as well.
On the first point, symmetry versus asymmetric loss, most empirical work in forecasting assumes
mean squared error (MSE) loss which of course implies symmetric loss. Apart from the fact that
using MSE loss represents ‘conventional practice’, this choice is likely to reflect difficulties in putting
numbers on the relative cost of over- and underpredictions. Construction of a loss function requires
a deep understanding of the forecaster’s objectives and this may not always be easily accomplished.
Still, the implicit choice of MSE loss by the majority of studies in the forecasting literature
seems difficult to justify on economic grounds. As noted by Granger and Newbold (1986, p. 125),
“.. an assumption of symmetry about the conditional mean ... is likely to be an easy one to accept
... an assumption of symmetry for the cost function is much less acceptable.”
Papers that consider the properties of optimal forecasts under asymmetric loss from a theoretical
perspective include Granger (1969, 1999), Varian (1974), Zellner (1986), Weiss (1996), Christof-
fersen and Diebold (1997), Batchelor and Peel (1998), Granger and Pesaran (2000), Pesaran and
Skouras (2002) and Patton and Timmermann (2006, 2007).
Many economic considerations can help in deriving the loss function and determining the extent
of any asymmetry. Consider a firm involved in forecasting the sales of a new product. Overpre-
dicting sales leads to inventory and insurance costs and ties up capital. It may also give rise to
discounts needed to sell the remaining surplus. Such costs are mostly known or can at least be
estimated with a fair degree of precision. Contrast this with the cost of underpredicting sales which
leads to stock-out costs, loss of goodwill and reputation and lost current and future sales. Such
costs are less tangible and can be difficult to quantify. Nonetheless, this must be attempted in
order to construct forecasts that properly trade off the costs of over- and underpredictions.
For money managers, asymmetric loss may be linked to loss aversion or concerns related to
liquidity, bankruptcy or regulatory constraints (Lopez and Walter (2001)). Under the Basel II
accord, banks are required to forecast their Value at Risk which is a measure of how much they
expect to lose with a certain probability such as 1%. Capital provisions are affected by this forecast:
Overpredicting the Value at Risk ties up more capital than necessary, while underpredicting it could
lead to regulatory penalties and the need for increased future capital provisions.
Empirical applications of asymmetric loss include exchange rate forecasting (Ito (1990), West,
Edison and Cho (1993)), Budget forecasts (Artis and Marcellino (2001)) and the Federal Reserve’s
Greenbook forecasts (Capistran (2005)).
12
4.1.2 Use of Forecasts
Turning to the second point, i.e. the use of the forecast, interesting issues in constructing the
loss function arise when the forecast is itself best viewed as a signal in a strategic game that
explicitly accounts for the forecast provider’s incentives. The papers by Ehrbeck and Waldmann
(1996), Ottaviani and Sorensen (2006), Scharfstein and Stein (1990) and Truman (1994) suggest
more complicated loss functions grounded on game theoretical models. Forecasters are assumed
to differ by their ability to forecast. The chief objective of the forecaster is to influence clients’
assessment of their ability. Such objectives are common for business analysts or analysts employed
by financial services firms such as investment banks or brokerages whose fees are directly linked to
clients’ assessment of analysts’ forecasting ability.
An interesting example comes from financial analysts’ earnings forecasts which are commonly
found to be upward biased (e.g., Hong and Kubik (2003) and Lim (2001)). By reporting a rosier
(i.e. upwards biased) picture of a firm’s earnings prospects, analysts may get favored by the
firm’s management and get access to more precise and timely information. Too strong a bias will
compromise the precision of the analysts’ forecast and will be detrimental to the position of the
analysts in the regular rankings that are important to their career prospects, particularly for “buy-
side” analysts. Ultimately, forecasts must trade off bias against precision. In general we would not
expect the cost of over- and under-predicting earnings to be identical and so biases are likely to
persist.
13
usually normalized so L(y, y, z) = 0 for all y and z. For this to be a unique minimum we have
L(f (z, θ), y, z) > 0 for all f 6= y.
Restrictions on the form of the loss function are also needed to make sense of the ideas of
minimizing risk, in particular we require that the expected loss exists. The existence of expected
loss depends both on the loss function and on the conditional distribution of the outcome variable.
Recall that expected loss is
Z
EY [L(f (z, θ), Y, z] = L(f (z, θ), y, z)pY (y|z, θ)dy. (8)
Hence issues with the existence of expected loss revolve around how large the loss becomes for tail
behavior of the predicted variable.
Symmetry of the loss function is the constraint that, for all d,
This is a particularly tractable loss function since there are no unknown parameters and the optimal
forecast is simply the conditional mean of Y : f ∗ (Z, θ) = E[Y |Z, θ]. Hence under MSE loss the
classical ‘plug in’ approach to forecasting simply involves estimating the conditional mean of Y .
This relates naturally to regression analysis and the greater part of econometric theory. Under the
R
Bayesian approach the optimal forecast is the mean of the predictive density, f (z) = yPY (y|z)dy.
Mean absolute error (MAE) loss is also very common:
For all continuous distributions pY (y|z, θ) the optimal forecast is the conditional median of Y.
These two loss functions are nested in the general family of loss functions considered by Elliott,
Komunjer and Timmermann (2005)
L(f (Z, θ), Y, Z; p, α) ≡ [α + (1 − 2α) · 1(Y − f (Z, θ) < 0)] · |Y − f (Z, θ)|p . (12)
For p = 1 this gives the lin-lin (piece-wise linear) loss function which nests MAE loss when α = 0.5.
For p = 2 the asymmetric quadratic loss function, which nests MSE loss when α = 0.5, is obtained.
Optimal forecasts from the lin-lin loss function are conditional quantiles, while those from the
asymmetric quadratic loss function are expectiles (Newey and Powell (1987)).
Varian (1974), Zellner (1986) and Christoffersen and Diebold (1997) studied linex loss,
L(f (Z, θ), Y, Z) = exp(b(Y − f (Z, θ))) − b(Y − f (Z, θ)) − 1, (13)
where b > 0 is a parameter that controls the degree of asymmetry. If b > 0, large underpredictions
(f < y) are costlier than overpredictions of the same magnitude, with the relative cost increasing
14
as the magnitude of the forecast error rises. Conversely, for b < 0, large overpredictions are costlier
than equally large underpredictions.
Direction-of-change based loss is another class that has generated considerable interest in recent
work. The simplest of these takes the form
(
0 if sign(Y ) = sign(f )
L(f (Z, θ), Y, Z) = . (14)
1 otherwise
Sign loss functions are popular in finance since they are closely related to market timing in financial
markets and also are linked to volatility forecasting (Christoffersen and Diebold (2006)). To see
this, we modify the objective function so that it reflects the magnitude of the outcome variable. For
example, consider the decisions of a ‘market timer’ whose utility is linear in the payoff, U (y, δ(z)) =
δy, where y is the return on the market portfolio (possibly in excess of a risk-free rate) and the
action rule, δ(z), is to invest one unit in the market portfolio if this has a positive expected return
and otherwise be short one unit, i.e.
(
1 if f > 0
δ= . (15)
−1 if f ≤ 0
and so depend on both the signs of y and f as well as the magnitude of y. Notice that this objective
does not adhere to our definition of a loss function. Intuitively this is because large forecast errors
for forecasts with the correct sign lead to smaller loss than small forecast errors for forecasts with
the wrong sign.4
15
value of p, α is the single parameter that controls the degree of asymmetry among the loss functions
in (12). When p = 2, α/(1 − α) measures the relative cost of positive and negative forecast errors
of the same magnitude. For example, a value of α = 0.4 suggests that positive errors are two-thirds
as costly as negative errors of the same magnitude.
The estimate of α thus provides economic information about the degree of asymmetry required
to justify the observed sequence of forecasts. Sometimes such estimates can be rejected on economic
grounds. Suppose, for example, that an estimate α = 0.1 is required to justify rationality of the
observed forecasts. This suggests that it is almost ten times costlier to underpredict than to
overpredict the variable of interest. This may be deemed implausible on economic grounds and so
asymmetric loss is unlikely to be the explanation for the observed behavior of the forecast error.
How the forecaster maps predictions into actions may also be helpful in explaining properties of
observed forecasts. Leitch and Tanner (1991) studied forecasts of T-bill futures contracts and found
that professional forecasters reported predictions with higher mean squared error than those from
simple time-series models. This is puzzling since the time-series models presumably incorporate far
less information than the professional forecasts. When measured by their ability to correctly forecast
the direction of future interest rate movements−a metric related to the forecasters’ ability to make
money−the professional forecasts did better than the time-series models. A natural conclusion to
draw from this is that the professional forecasters’ objectives are poorly approximated by the mean
squared error (MSE) loss function and are closer to a directional or ‘sign’ loss function. This would
make sense if the investor’s decision rule is to go long if an asset’s payoff is predicted to be positive
and otherwise go short.5
16
to allow the functional form or the parameters of the forecasting model to change over time.
Each of these issues highlights a theme in recent work on economic forecasting which we review
below. First, however, we review the workhorse in the forecasting literature, namely the linear
forecasting model and its extension to vector autoregressions (VARs). Challenges faced in using
VARs for forecasting foreshadow issues that arise for the general problem.
17
residuals (mean zero and variance σ 2 I with σ 2 known) and prior β ∼ N (β 0 , Ω), the posterior
distribution for the regression parameters is normal with mean (σ −2 Z̃ 0 Z̃ +Ω−1 )−1 (σ −2 Z̃ 0 y +Ω−1 β 0 )
and variance (σ −2 Z̃ 0 Z̃ + Ω−1 )−1 . The plug-in forecast simply uses the mean of the posterior, which
takes the form of a shrinkage estimator. Setting Ω = σ 2 k−1 I the estimator becomes (Z̃ 0 Z̃+kI)−1 Z̃ 0 y
which is in the form of a Ridge estimator. Under MSE loss, normally distributed β̂ and any general
prior distribution π(β), it can be shown that the Bayes rule forecast in a linear regression takes the
form of a correction to the OLS estimator.
To employ these methods in practice requires specifying the prior (and typically the results
are extended beyond the known variance case). Litterman (1980, 1986) and Doan, Litterman
and Sims (1984) suggested “Minnesota” priors on the parameters of a VAR which more heavily
weight the parameter configuration towards a model where variables follow individual random
walks. More distant lags are shrunk towards zero more heavily. This type of prior can be helpful
in obtaining better forecasts of macroeconomic outcomes. Robertson and Tallman (1999) examine
the forecasting performance of flat prior VARs (i.e. the usual OLS estimates) versus Bayesian
methods based on more informative priors. They find that extensions to the Litterman priors
revolving around long run properties provides forecasting gains for a number of macroeconomic
variables (GDP growth, unemployment, Fed funds rate and CPI inflation). Kadiyala and Karlsson
(1993, 1997) examine more extensive priors than the Litterman approach (allowing for dependence
between the equations) and find examples of improvements over the Minnesota prior.
A promising recent literature uses dynamic stochastic general equilibrium (DSGE) models to
constrain VARs in a Bayesian setting. In this vein, Del Negro et al (2006) cast DSGE models as
(reduced-form) VARs that include an error correction term. Theoretical restrictions from firms’
and households’ optimizing behavior subject to their intertemporal budget constraints, along with
assumptions about government expenditures imply a set of cross-equation restrictions on the pa-
rameters of the VAR. Del Negro et al (2006) relate these constraints to the priors on the model
parameters; ignoring the theory corresponds to diffuse priors while informative priors pull the pa-
rameters towards the theoretical constraints. In a simulation study these authors find evidence
that, for a range of macroeconomic variables, the DSGE-based VAR produces better out-of-sample
forecasting performance than the standard unconstrained VAR.
A variety of estimation and variable selection methods have been suggested as plug-in estimators
to attain a better trade-off between the bias and variance components of the forecast MSE. Most
OLS
of the alternative plug-in estimators are modifications of the OLS estimator, β̂ i , and fall in the
s
general class of shrinkage estimators, β̂ i , of the form
s OLS
β̂ i = (1 − γ i )β̂ i + γ i β̃ i 0 < γ i < 1, i = 1, ..., k. (18)
where β̃ i is the shrinkage target. The shrinkage weight, γ i , generally depends on the data (James
and Stein (1961)). It is common practice to set β̃ i = 0, in which case the OLS estimators are
shrunk towards zero. The expected gain from this approach under MSE is to reduce the variance
in the bias-variance trade-off that arises from the plug-in estimator for the regression coefficients.
For empirical evidence on shrinkage methods, see Zellner and Hong (1989).
18
Estimators differ in how they specify the shrinkage weights γ i and shrinkage target β̃ i . These
include bagging and subset selection methods (explained in detail in section 5.2.1 below), as well
as Stein regression and many Bayes and empirical Bayes methods. Other examples of shrinkage
methods that have been less employed for forecasting include ridge regression (which is a special case
of (18)), as well as the lasso (Tibshirani (1996)) and the (non negative) garrote (Breiman (1995)).
The latter two methods are somewhat more complicated in estimation and are essentially penalized
least squares estimation methods. For the lasso the data are first normalized (see Miller (2002)
for a textbook treatment) and then the sum of squared errors from the regression is minimized
subject to the constraint that the sum of the absolute value of the regression coefficients is less
than a chosen value. For the garrote, first least squares estimates from a regression of yt+1 on xt
are obtained, then values ci are chosen to minimize
T −1
à !2
X X OLS
yt+1 − ci β̂ i xit
t=1 i
subject to the constraints that each ci ≥ 0 and that their sum falls below a chosen value.
It is common to use standard least-squares estimation techniques for model parameters without
reference to the loss function. An example is using linear regression and examining if the forecasts
capture turning points in the data. Despite such practice, recent studies have suggested that
forecasting performance can often be significantly improved by using the same loss function in the
estimation and evaluation stages. This has been found both in simulation experiments (Elliott and
Timmermann (2004)) and in empirical studies (Christoffersen and Jacobs (2004)).
To understand such gains, note that, for loss functions other than MSE, optimal forecasts
generally require use of the loss function either directly in estimation or through the specification
of the model that is estimated. For example in the case of lin-lin loss, the optimal forecast is a
conditional quantile and so quantile methods can be employed to estimate the parameters of the
forecasting model. When this is specified up to a set of unknown parameters−i.e. the function
f (.) of the forecasting model, f (z, θ), is known but θ is not−then the loss function can be used to
estimate the unknown parameters through M estimation, i.e.
T
X −1
θ̂ = arg min(T − 1)−1 L(f (zt , θ), yt+1 , zt ). (19)
t=1
When the loss function is differentiable extremum estimators based on the first order moment
conditions can also be used. Approximate methods have been suggested by Weiss (1996). In
the context of forecast combination, Elliott and Timmermann (2004) propose methods of moments
estimators and approximate methods for a variety of loss functions including lin-lin and asymmetric
quadratic loss.
19
readily available from governments and other organizations. The virtual explosion in the number of
potential predictor variables is exacerbated by the fact that the dynamic structure of the forecast-
ing model is typically unknown. Adding an extra variable therefore increases the model dimension
not only by a single parameter but by many parameters to account for the dynamic effects of this
variable on the outcome. At the same time, the length of the available data is often relatively
short because the frequency of time series observations and the period over which data have been
constructed combine to limit the sample size. Model instability (which will be discussed below)
may further limit the useful length of the data.
Short samples, whether due to limited data availability or model instabilities, along with large
sets of predictor variables mean that forecasters always face a trade-off in terms of how complicated
the model can be versus how well its parameters can be estimated. Methods that place constraints
on the number of parameters may produce better forecasts even if the population model for the
data does not share these constraints. For example, small and very parsimonious models have
been found to work well in many empirical studies. Reductions in loss resulting from adding more
variables are often more than offset by the resulting increase in parameter estimation error. For
MSE loss, additional variables reduce bias, but estimation error increases the variance, leading to a
bias-variance trade-off. Methods other than omitting relevant variables and using OLS are available
to exploit this trade-off. This has led to the proposal of a wide range of estimation techniques in
the forecasting literature which we next turn to.
20
Inoue and Kilian (2006) discuss conditions under which a variety of tools for model selection−e.g.,
ranking by recursive MSE, rolling MSE or model selection by the AIC or BIC−will identify the
model with the lowest true out of sample MSE among a finite set of forecasting models. They find
that selection by AIC and ranking by recursive MSE yield inconsistent results and have a positive
probability of choosing a model which does not have the best forecasting performance while the
SIC is consistent for nested models. Of course it should be borne in mind that consistency is not
the most important criterion to satisfy here and how a method handles bias-variance trade-offs may
be more important to forecasting performance.
An alternative approach is to evaluate each variable on its own or in smaller groups – a method
advocated by, e.g., Hendry and Krolzig (2004). This step-wise procedure employs t-tests to remove
individual variables with statistically insignificant coefficient estimates. While computationally
highly attractive, this approach gives rise to problems of its own: Insignificant estimates can arise
not only because the true parameters are small but also because of large sampling error. Moreover,
if one re-estimates the model after removal of some parameters and again examines statistical
significance, the method becomes path dependent. This matters because classical pre-tests need
not result in consistent estimates of the model. Finally, the ‘all or nothing’ approach of either using
the OLS estimate or omitting the regressor may be too restrictive (i.e. restricting γ i at zero or
one).
Methods that attempt to exploit the bias-variance trade-off without the ‘all or nothing’ approach
have thus been suggested. The estimation method closest to pre-testing is bagging (bootstrap
averaging) proposed by Breiman (1996). In this method the forecast model is bootstrapped, re-
estimated using tests for significance to omit variables, and the forecast from the bootstrapped
model is computed. A final forecast is then computed as the average of the forecast over the
bootstrapped models (see Inoue and Kilian (2005) for a more complete description of the method).
Since variables are unlikely to be omitted for all bootstrap replications, the estimator can be viewed
as a smoothed version of the ‘all or nothing’ approach. Bagging still sets β̃ i = 0 but γ i is now equal
to the average over this indicator function across the bootstrap replications. In this way bagging
changes from a hard threshold (zero or one) to a soft one (some number between these values).
An alternative to excluding variables from a large dimensional set of predictors is to extract com-
mon features from the data and then use these as the basis for the forecasting model. Indeed,
the dominant classical approach for dealing with large dimensional data is to extract a set of com-
mon factors of much lower dimensionality than the original variables to summarize an otherwise
overwhelming amount of information. Suppose Z contains N economic variables whose common
dynamics can be represented through the factors
Zt = Λ(L)ψ t + et , (20)
where et is a vector of idiosyncratic shocks, ψ t is a vector of common factors and Λ(L) is a matrix
of lag polynomials representing dynamic effects. For low dimensional systems (small N ), dynamic
factor models can be estimated through the Kalman filter. When N is large, Stock and Watson
(2002) propose a principal components approach to obtain the common factors as the solution to
21
a simple least squares problem. An alternative approach proposed by Forni et al (2000, 2003) is to
extract principal components from the frequency domain using spectral methods.
While the construction of a set of common factors resolves the question of how to aggregate an
otherwise far too large dimensional state vector, use of these techniques in forecasting also raises
new issues. There is a risk that the factor extraction serves as a ‘black box’ approach void of any
economic interpretation. This risk arises in situations where the factors are not clearly identifiable
with underlying blocks of economic variables, although in many situations such blocks can be used
to good avail for interpretation purposes (Ludvigsson and Ng (2005, 2007)). The aim is to ensure
that the first few factors can be interpreted in a way that links them to a particular subset of
variables.
In an empirical analysis, Stock and Watson (2005) find that methods that include the first
few (and most significant) principal components in addition to own-variable autoregressive dy-
namics generally work best and that a few principal components are responsible for most of the
improvement in forecasting performance.
Assuming that E[εt+1 |zt ] = 0, this nests the linear model when g(.) = 0.
Extending the set of forecasting models under consideration to nonlinear specifications substan-
tially expands the model set. Let the set of models M represent the combinations of parameters
and functional forms for the models under consideration. When the true model M0 ∈ M, the
model is said to be correctly specified, otherwise it is misspecified. A misspecified model may yield
forecasts that are difficult to beat, even if the coefficients are not meaningful (Clements and Hendry
(1998)). For example a linear forecasting model estimated by OLS results in the linear model that
minimizes Kullback Leibler distance between the estimated ‘approximate’ model and the unknown
nonlinear model. That said, however, when forecasters have to ‘sell’ their forecasts to decision
makers, the misspecified model may be difficult to put a story to.
The forecaster’s problem is to choose the best available model M ∈ M. Because this involves
a search over functional forms, to be able to actually perform this search the forecaster needs to
restrict the problem further. As with the specification of the variables to be included in the regres-
sion model, economic theory is often not particularly precise on the exact form of the nonlinearity
to be expected.
22
Tests for the functional form g(.) are complicated because θ2 or a subset of this vector is not
identified under the null hypothesis of linearity. As a result, standard methods such as the general-
ized likelihood ratio test lose optimality properties and are no longer approximated asymptotically
by a chi-squared distribution. A large literature has arisen to deal with these issues, see Terasvirta
(2006).
The danger of using a misspecified forecasting model is a real problem given the lack of theo-
retical underpinning of the choice of M. Moreover, often M is a very small subset of the possible
model specifications. Furthermore, for many of the tests of the null of linearity for these models,
rejection does not necessarily imply that a particular nonlinear model chosen is implied. Hence it is
important to understand the estimation and forecasting properties of these models under misspec-
ification. Such misspecification and the approximation properties of nonlinear models may explain
why they sometimes generate extreme forecasts.
The literature has pursued two broad themes: (i) the formalization of specific nonlinear models
that generalize the linear model in an intuitive way; or (ii) the use of more global approximation
procedures that seek to approximate the unknown nonlinear function. Which approach should be
adopted depends on how much is known about the type of nonlinearity to be expected in a given
situation. We next examine each approach in turn.
23
with weights on the individual states that are determined from the updated state probabilities. We
cover these models in more detail in Section 6.2 on breaks.
The large literature on autoregressive conditional heteroskedasticity (ARCH) in asset returns
(reviewed in the context of forecasting by Andersen et al (2006)) is another example of nonlinear
dynamics that could be important to portfolio managers concerned with predicting returns and
managing the risk of their assets. These models imply that the conditional variance of asset returns
is persistent and hence partially predictable, particularly at short horizons. Such predictability of
the conditional variance can be captured by extending (21) to
Nonlinear forecasting models can adapt more quickly to changes in the underlying time-series
dynamics and avoid smoothing the data as much as linear models. This same feature means that
nonlinear models can also be highly sensitive to the sort of outliers found in many economic and
financial time series and may be more prone to overfitting than linear models. Furthermore, the
parameters capturing nonlinear dynamics are often associated with a few episodes such as the
change in the dynamics of US interest rates during the ‘monetarist experiment’ from 1979-1982.
As a consequence, these parameters can be very imprecisely estimated for the typical sample sizes
available to macroeconomic forecasters and so these models often produce quite poor out-of-sample
MSE performance.6 On the other hand, nonlinear forecasting models may perform quite well in
certain states (e.g. recessions or periods of financial crises) and so can be used either in conjunction
with other models that generate more smooth forecast (see Section 8.4 on combination) or for non-
convex loss functions such as the sign function (14) that put smaller weight on outliers.
Often the simplification gj (zt , θj ) = g̃j (zt )θ j is employed to make the forecasting model linear in
the parameters (or at least more so, since additional parameters can be hidden inside g̃(.)). The
idea is to choose the functions g(.) carefully enough and the number of them, J, large enough to
approximate a wide variety of possible nonlinear functions, see e.g. Swanson and White (1995).
There are a large number of theoretically well motivated choices for the basis functions gj (zt , θj ).
Most popular in the economic forecasting literature are artificial neural network models, where
gj (zt , θj ) = θj1 (1 + exp(−zt0 θj2 ))−1 and various methods are employed for choosing or estimating
6
Indeed, some simulation studies find that even when the nonlinear model is correctly specified, it often produces
less precise forecasts than a simple misspecified linear approximation due to the greater uncertainty about the
nonlinear model’s parameters (Psaradakis and Spagnolo (2005)).
24
θj2 . Other basis functions include Fourier series, polynomials, piecewise polynomials and splines.
Methods such as wavelets, ridgelets, and the Gallant (1981) flexible fourier form also belong to
this set. Both theoretically and in practice, different methods work well against different classes of
functions.
Practical problems arise for these methods both in terms of estimation and forecasting. First,
unless g̃(.) is fully specified, estimation requires nonlinear optimization. Variations have arisen
in attempts to find simple methods to specify the functional form, so as to leave the remaining
estimation linear in the parameters and hence estimable by OLS. This is standard for example in
the application of neural net models.
Second, the order J must be chosen. Since an infinite number of possible terms could be included
and the in-sample fit is improved by choosing J as large as possible, overfitting is likely unless the
number of included terms is somehow restricted. Overfitted models tend to produce very good
results for the data used to estimate the models, but very poor forecasts on fresh (out-of-sample)
data. To deal with these issues, methods such as information criteria and cross validation are
used to select among this class of models. Overfitting remains the Achilles heel of these methods,
however.
Finally, as with parametric nonlinear models, the risk of generating relatively extreme forecasts
remains when forecasting from sample points where the data is relatively sparse. Many practitioners
use ‘insanity’ filters, replacing these forecasts with a smoothed value when the forecast is too far
from the outcome or its mean. Even so, the track record of these models in forecasting has been
mixed. Nonlinearities do seem to be present in many macroeconomic series, but the data samples
for these variables tend to be relatively short, thus hampering the precise estimation of nonlinear
forecasting models.
For financial returns the signal-to-noise ratio−i.e. the fraction of predictable variation in asset
returns−tends to be very low. Often this means that the observed nonlinearities are poorly identi-
fied, imprecisely estimated and so the risk of overfitting is very high. As a result, the deterioration
in out-of-sample forecasting performance is likely to be very high when compared against the pre-
dictive performance during the training sample used to estimate the parameters of the model (see,
e.g. Racine (2001)). Consequently there is little evidence that such forecasting models dominate
simple linear specifications, at least under MSE loss.
Overall, the difficulties that arise when forecasting with nonlinear models revolve around the
question whether a fitted nonlinear model provides a good approximation to the true nonlinear
model. In the case of the ad-hoc specifications, rejection of a linear model in favor of a particular
nonlinear specification does not necessarily indicate that the latter will produce good forecasts.
Rejections of typical tests for nonlinearity tend to indicate a range of possible models rather than
a particular model. In the case of model approximation, model specification search tends to result
in models that overfit the data, again causing problems for forecasting.
25
situations the forecaster encounters a multi-period forecasting problem.
Computing multi-period forecasts is simple if the predictors are weakly exogenous but issues
arise in practice. To illustrate this point, when forecasting from a univariate first-order autoregres-
sive model yt = φyt−1 + εt , forward iteration gives
h−1
X
h
yt+h = φ yt + φi εt+h−i . (24)
i=0
Assuming that E[εt+j |yt ] = 0 for j > 1 and that both the model and its parameters are known,
the optimal forecast under MSE loss is simply φh yt .
When the parameters are unknown, however, the problem becomes far more complicated. A
h
simple solution would be to use the plug-in OLS estimate, φ̂ yt . However, this is clearly only a
h
solution of convenience: φ̂ is generally not unbiased and thus, φ̂ will not be unbiased for φh either.
h
Even if φ̂ were unbiased, in general φ̂ would not inherit this property.
An obvious alternative to iterating forward on a single-period model is to tailor the forecasting
model directly to the forecast horizon. This is more in spirit with viewing forecasting models as
misspecified simplifications of the underlying data generating process and entails a model of the
form
yt+h = g(zt , θ) + εt+h . (25)
The chief problem is now the overlap in the forecast errors that will generally exhibit behavior
similar to that of a moving average process of order h − 1. For example, if h = 2, the forecast
errors will be serially correlated according to an MA(1) process even if the true forecasting model
is used and its parameters are known. Such serial dependence can be handled through a number
of procedures that account for autocorrelation in the forecast errors.
Which approach is best−the direct or the iterated−is an empirical matter since it involves
trading off estimation efficiency against robustness to model misspecification. It is also not clear
how iterating multiple periods ahead on a misspecified model will affect the quality of the forecast.
For example, the initial value of the conditioning information (zt ) could well matter in this situation.
Even when the models are correctly specified, there is a trade-off between the cumulative effect on
the forecast of using plug-in parameter estimates (which is avoided in the direct approach), versus
the greater efficiency of the iterated approach that comes from estimating the forecasting model
on data measured at a higher frequency.
Marcellino, Stock and Watson (2006) address these points empirically using a data set of 170
US monthly macroeconomic time series. They find that the iterated approach generates the lowest
MSE-values, particularly if long lags of the variables are included in the forecasting models and if
the forecast horizon is long. This suggests that reducing parameter estimation error can be more
important than concerns related to model misspecification, an issue that cannot be decided ex-ante
on theoretical grounds alone (Schorfheide (2005)).
Special problems may arise when forecasting multiple steps ahead with nonlinear models, where
numerical methods are typically required due to the nonlinear form. This problem stems directly
from taking the ad hoc model to be the true model, which is or course a doubtful assumption. To
illustrate this, suppose that
yt+1 = g(yt ; θ) + εt+1 . (26)
26
Iterating forward to the two-period horizon, we have
Hence the function g (presumed to be nonlinear) needs to be invoked as many times as the length
of the forecast horizon. Moreover, the entire distribution of ε becomes crucial even under MSE loss
where interest is limited to forecasting the conditional mean. For example, if g is quadratic, the
variance of ε matters to forecasting the conditional mean two or more periods ahead.
6 Model Instability
Economic institutions, tax rules and political regimes change over time and the economy evolves in
response to technological and macroeconomic shocks such as the oil price changes in the seventies.
One of the stylized facts of empirical macroeconomics is the ‘Great Moderation’, i.e. the lower
volatility of many macroeconomic series after the mid-eighties. Events such as these make it
plausible that the underlying data generating process changes over time. In many forecasting
situations economic theory is silent on the exact form of these instabilities and instead offers
general guidelines that can help determining the source (and possibly timing) of instabilities such
as changes in monetary policy (due, e.g., to a change in the Federal Reserve chairman) or changes
in economic institutions.
In the construction of the forecasting problem in Section 3, the specification of the likelihood
for the data pY (y|z) did not require that the relationship between the data remains stable over
time, or that the underlying data itself is stable through time, although a model for the process
that generates instability is required. A difficulty that arises in the presence of changes in the data
generating process is the existence of a multitude of models that can capture potential instabilities.
Unit root models are popular, although the root could be near one rather than exactly equal to
one. Fractionally integrated models allow similar behavior at low frequencies. Breaks in regression
parameters can also mimic this type of behavior. Beyond this we could allow the root to be
stochastic and near one.
Model instability introduces at least three problems for forecasters. First, it complicates speci-
fication of the likelihood for the data. From a Bayesian perspective this can make it more difficult
(at least analytically) to determine a closed form forecasting rule, depending on the form of the
nonstationarity. Second, since the parameterization of the nonstationarity results in a larger di-
mension of θ, estimation is also affected. Finally, nonstationary data makes averaging over the
past to obtain plug-in estimates more difficult. In classical estimation this can be a large problem.
Further complications arise through nonstandard properties of the estimators that frequently arise
in these models.
27
of these. Clements and Hendry (2006) also stress instability as a key determinant of forecasting
performance.
Most work on forecasting models with unstable parameters has considered linear specifications
of the form
¡ ¢
Yt+1 = β t − β̄ Z1t + γZ2t + ut+1 , (28)
where the coefficients β t on Z1t are changing over time while the remaining coefficients are constant.
The first problem that arises in the construction of a forecasting model of this type is that
there are many ways in which β t can be nonconstant. We could parameterize β t as a stochastic
process (either mean reverting or not) or as a step function that changes at random times by
random amounts. Examples of such models include the popular unobserved components model
(where Yt+1 = β t + ut+1 and β t follows a random walk process) and extensions of this to the entire
vector β t (as in the models of West and Harrison (1997)).
If the variation in β t is not permanent in the sense that it can be characterized by a mean
reverting process the linear specification that omits the breaks in β t is essentially a heteroskedastic
model and least squares estimation of the parameters will not be too misleading (White (2001)).
When the breaks are permanent, the coefficients of the linear model lose meaning and become
similar to sample averages of a random walk, changing with time and not related directly to any
parameter of the model. In either case, knowing the true model will enable better forecasts.
Forecasters must decide whether or not to (a) use only part of the data available, assuming that
the retained data is sufficiently stationary that it will provide a good approximation to a model
with constant coefficients, or (b) attempt to model the breaking process.
To illustrate these approaches, suppose β t is constant apart from a single break of unknown
size δ at an unknown date τ ,
(
βZ1t + γZ2t + ut+1 t<τ
Yt+1 = . (29)
(β + δ)Z1t + γZ2t + ut+1 t ≥ τ
An example of the first approach would be to try and estimate τ̂ and base the estimates of the
forecasting model solely on data after the break. Alternatively, a forecaster might consider con-
structing estimates for both the break date τ̂ and the size of the break δ̂ in order to construct a
forecast from the full data incorporating the break into the model (and possibly attempt to forecast
future breaks). This would be an example of the second approach.
Unfortunately, while tests for nonconstant parameters are quite good at detecting breaking
behavior of this nature, they are not capable of distinguishing the particular type of nonstationarity
beyond the distinction of ‘permanent’ deviations versus the mean reverting deviations mentioned
above. Nearly all popular tests have no power against temporary deviations of β t from its mean.
Conversely, nearly all tests have similar power against a host of possible processes for β t when it
does depart permanently from any value. This is true for models with few breaks, many breaks, or
breaks every period.7
7
Stock and Watson (1998) show that tests for a single break have power against random walk breaks. Elliott
and Mueller (2006) consider a wide class of breaking processes and show that optimal tests for each of the breaking
processes have equivalent asymptotic power against all of the other breaking processes in a wide class.
28
The implication for forecasting is that once one has found evidence of breaks in the parameters
of some variables, there is still a great deal of uncertainty as to the nature of the breaking process.
Because it will be difficult to pin down the appropriate model, parameterizing and estimating the
breaking process will generally be quite difficult.
Even if it were known that the forecasting model has a single break point, estimates of the break
date are often not particularly useful in practice. Tests for a break will often reject stability even
though the break size is too small to permit precise estimation of exactly when the break occurred.
One can still proceed with the first approach and try to estimate the window of data to use for the
forecast. However, some account for the uncertainty surrounding the timing of the break is likely
to be an important part of a successful forecasting strategy in the presence of breaks.
For the alternative of estimating the full model including both the break date and the size
of the break, Elliott (2005) shows that estimation of the break size and its location results in
very poor forecasts relative to knowing these parameters. Instead a method of averaging over all
possible break dates with weights that depend on sample estimates of the probability that each
date is the true break date is suggested, with substantial gains over least squares estimates of
these two parameters. In an empirical application to exchange rate forecasting, Rossi (2006) found
evidence of widespread instabilities and showed that Elliott’s (2005) method works well in practice
for forecasting a range of currencies.
Similarly, Pesaran and Timmermann (2005a, 2006) find that the estimation window matters sig-
nificantly to the out-of-sample forecasting performance of simple time-series models in the presence
of breaks. Given the considerable uncertainty surrounding the time and the size of the break, they
consider approaches that average forecasts generated under different estimation windows. They
also derive analytical results for the normal model under MSE loss and demonstrate that the gains
in forecast accuracy from using pre-break data increases when breaks are small and occur late in
the sample.
When there is more than one break, things become even more difficult. Bai (1997) and Bai and
Perron (1998) suggest an approach to determine the number of breaks through repeated tests on the
data. This approach has been applied to forecast stock returns by Paye and Timmermann (2006)
and Rapach and Wohar (2006). Their results suggest the presence of multiple breaks in standard
forecasting models for stock returns and reveal wide variation in the extent of predictability in
stock returns across break segments. Paye and Timmermann (2006) also find that the break dates
are difficult to pin down, vary greatly across different model specifications and do not seem to be
common across international markets. This makes the task of forecasting stock returns particularly
difficult since the question of ‘how much historical data to use’ and how to weight new versus old
data is both very important in practice and difficult to come up with a satisfactory solution to.
Structural breaks in parameters can also cause a forecasting model’s performance to deviate
significantly and erratically from the outcome expected on the basis of its in-sample fit. Giaco-
mini and Rossi (2006) refer to these situations as “model breakdowns” and develop a method to
empirically detect them.
The two most common types of instability found in macroeconomic and financial data are breaks
in the model parameters and unit root or long memory behavior of the data. We next discuss each
of these.
29
6.2 Modeling the Break Process
The presence of historical breaks in a time-series model requires that the possibility of future breaks
be considered. This means that the process generating breaks must itself be modeled. In this regard,
the forecasting problem is unique compared with the problem of detecting and dating past breaks.
Approaches that do not model the break process itself and treat breaks as deterministic (such as
Bai and Perron (1998)) are not directly applicable to forecasting.
This is not a problem for the time-varying parameter specifications which directly posit a model
for how the parameters evolve in future periods. A popular approach is to parameterize β t as a
random walk and use the Kalman Filter to estimate the path for β t and produce a forecast (Harvey
(2006) covers the classical approach while West and Harrison (1997) cover the Bayesian approach).
The simplest example arises when Zt = 1, β t is a random walk and both the innovations to β t
and Yt are normally distributed, so the forecast of YT +1 is β̂ T . Then
where φt depends on the variances of the two error terms. For a given choice of initial value, this
recursion can be used to generate forecasts in real time. In the limit φt can be approximated by a
constant, which for a given value of φ yields the exponentially weighted average (IMA(1,1)) model
t−1
(1 − φ) X s
β̂ t = φ Yt−s . (31)
(1 − φt ) s=0
This approach is equivalent to the discounted least squares model which puts a decreasing weight
on data further back in time. These methods can readily be extended to the general model with
time-varying predictor variables.
Examples of empirical application of these models are plentiful and include tracking of the skills
of mutual fund managers (Mamaysky, Spiegel and Zhang (2006)) and prediction of variables such
as GDP growth, inflation and electricity demand (Harvey and Koopman (1993)).8
Another example is the recurring stochastic breaks models embodied in the Markov switching
approach of Hamilton (1989). This assumes that changes to the parameters of the model are
driven by a latent state variable, S, that follows a first-order Markov chain. For example, the linear
forecasting model could be modified as follows:
Provided that no state is absorbing, this model gives rise to recurring shifts in the parameters.
Standard practice seems to be not to conduct much testing to identify the number of regimes, k,
and many papers simply assume the presence of two states.
8
Risk Metrics use this method to track conditional volatility in financial markets and typically sets φ close to one.
30
Again these models have been applied extensively in empirical analysis. Garcia and Perron
(1996) and Ang and Bekaert (2002) use regime-switching models to capture the dynamics in US in-
terest rates, while Perez-Quiros and Timmermann (2000) use these models to predict stock returns.
Some papers find evidence that letting the state transition probabilities depend on forward-looking
variables such as the leading indicator helps improve forecasting performance.
Perhaps surprisingly, relatively little work has been undertaking on merging the linear dynamic
factor VARs with multivariate nonlinear specifications such as Markov switching models as proposed
by Diebold and Rudebusch (1996). Some empirical findings have indicated the potential of this
type of model (see, e.g., Chauvet (1998) for an application to factor modeling and Guidolin and
Timmermann (2006, 2007) in the context of multivariate regime switching models applied to forecast
stock and bond returns and interest rates). Moreover, Markov Chain Monte Carlo methods which
are useful for estimating these models are now widely available (Kim and Nelson (1998)), so the
application of these types of models to real time forecasting is less of a challenge than previously.
By accounting for model uncertainty, standard Bayesian methods are directly applicable for
estimating the parameters of forecasting models even when these are time-varying. Furthermore,
since the procedure is conditional on Z the fact that risk averages across sample information is
incorporated through the prior i.e. by the weighting of the relevant parameters.
As an example of a Bayesian analysis, Pesaran, Pettenuzzo and Timmermann (2006) propose
a hidden Markov chain approach to forecast time series subject to multiple structural breaks.
They assume a hierarchical prior setting in which breaks are viewed as shifts in an underlying
Bernoulli process. The parameters of each regime are realizations of draws from a stationary meta
distribution. Information about this distribution gets updated recursively through time as new
breaks occur. This approach provides a way to forecast the frequency and size of future breaks.
Their empirical findings for US interest rates suggest that accounting for breaks in out-of-sample
forecasts can be important, particularly at long forecast horizons.9 Koop and Potter (2004) also
develop Bayesian methods for forecasting under breaks.
31
but increases risk for more distant alternatives, with the effect dependent on the forecast horizon.
Sample size is also important since estimates become more precise with more data swinging the
balance in favor of less constrained estimation methods.
The most prominent alternative to OLS estimation is the pretest estimator which sets the
estimate equal to one if the pretest fails to reject a unit root (since this is the null being tested)
and otherwise selects the OLS estimate. Diebold and Kilian (2000) examine this method, which
increases the gain from always imposing a unit root at the cost of doing worse on average than the
OLS estimator when the coefficient is further away from a unit root.
Intuition is more complicated in multivariate models because of the higher dimensionality of
the problem which results in a much broader set of trade-offs for the effect of parameter estimation
on risk. Transitory dynamics can also affect both the magnitude and in some cases the sign of the
results. Hence general results−whether analytical or through simulation−are difficult to arrive at.
The issue of unit roots or near nonstationarity also arises for the common linear forecasting
model ŷt = β̂ 0 + β̂ 1 zt−1 when the regressor, zt−1 , has a trend of unknown form. There are many
examples of such models being applied. Forecasts of stock returns by highly persistent variables
such as the dividend-price ratio or earnings-price ratio fit this situation, as does inflation forecasting
using interest rate levels, or forecasting changes in exchange rates with the forward premium. While
methods have been proposed for hypothesis testing in these models, there is not much evaluation
of the effect on forecasting.
As in the unit root case, when the innovations to the regressor are correlated with the residuals of
the forecasting equation, risk becomes a nonconstant function of the nuisance parameters describing
the form of the persistence in the data such as the degree of persistence of z and the covariance
between innovations to y and z. While most theoretical work has focussed on testing β 1 = 0,
little attention has been paid to designing good forecasting procedures or examining the trade-offs
between possible forecasting methods.
7 Forecast Evaluation
As noted in the introduction, one of the major differences between standard econometric problems
and the forecasting problem is that the researcher receives feedback on how well their forecast
actually performed. Thus, when a central bank forecasts next-year output growth or inflation, the
following year it is able to see how far off the forecasts were. Evaluating forecasting procedures in
light of this new information generates a dynamic process through which a number of important
issues arise.
Forecast evaluation usually comprises two separate, but related, tasks, namely (i) providing
summary statistics for measuring the precision of past forecasts; and (ii) testing optimality proper-
ties of the forecasts by means of a variety of diagnostics. The latter involves checking whether the
conditions implied by an optimal forecast hold in a particular sample. If the loss function is known
up to a finite set of unknown parameters, this is a straightforward process. From the forecaster’s
first order condition (3) the generalized forecast errors, L0 (f, y, z) or L0 for short, should themselves
be unpredictable, i.e. follow a martingale difference sequence given all current information used to
construct the forecasts.
32
The nature of this orthogonality condition will of course depend both on the shape of the
forecaster’s loss function and on the presumed data generating process underlying future values
of Y used to calculate the conditional expectation EY [L0 |Z]. For example, under MSE loss the
optimal forecast is, as we have seen, the conditional expectation of Y given all current information,
Z, and the generalized forecast error is simply proportional to the forecast error. Forecast errors
should therefore have zero mean, be serially uncorrelated and be unpredictable given all current
information. These properties are particular to the MSE loss function and need not hold in general
(Patton and Timmermann (2007)).
33
a covariance matrix that depends on the randomness of the out-of-sample observations and has
additional terms reflecting the variation that arises through the forecasts’ dependence on estimated
parameters.
The results of West (1996) show that it is appropriate only in special cases to use the standard
asymptotic variance covariance matrix that ignores randomness in the estimated parameters of
the forecasting model. One situation is when the same loss function is employed for estimating
the parameters θ and evaluating the forecast provided that the data is covariance stationary.13
In this case orthogonalities between the out of sample errors and the estimated model deliver the
asymptotic equivalence. The most interesting case is the linear forecasting model used to minimize
MSE loss for which standard errors can be computed as usual from the sequence of realized losses.
Alternatively, if the estimation sample is large relative to the sample over which the forecasts are
evaluated, then the additional variation due to estimating θ will be small since parameter estimates
will be close to their true values and hence estimation is negligible asymptotically.
Some issues limit direct application of these results, however. When the hold-out sample either
remains a fixed or a negligible proportion of the full sample, the coefficients of the forecasting models
converge to their pseudo-true values. If two or more of the models are asymptotically equivalent
(for example if one nests another) then asymptotically the forecasts will be perfect correlated and
the asymptotic approximation to the covariance matrix of the risks during the hold out sample is
singular.
When choosing the measure in which to report forecasting results, it should be borne in mind
that a forecast that may be good according to one measure (MSE), may not be good in terms of
another measure, e.g. correctly predicted signs. To see this, consider the following simple example
from Satchell and Timmermann (1995):
y = f + ε,
where y is the outcome, f is the forecast, and ε is the forecast error which has standard deviation
σ. If f and ε are independent, the probability of predicting the sign of y is a decreasing function
of the mean squared prediction error, σ 2 . Conversely, if f and ε are dependent, in general no such
relationship between the mean squared prediction error and the probability of predicting the sign
of y will hold. To see this, consider the 2 × 2 case
ε\f -1 1
−σ 1 p11 p12
σ2 p21 p22
34
Now choose 0 < δ < pij such that
p̃11 = p11 + 2∆
p̃12 = p12 − ∆
p̃21 = p21 − ∆
p̃22 = p22
We have thus increased the probability of correctly predicting the sign yet simultaneously increased
the MSE. The general message from this simple example is that forecasting models with low MSE
need not also be the ones with a high proportion of correctly predicted signs which is what may be
most important in some applications.
Moreover, realizations of L0 (f (zt , θ), yt+1 , zt ) should also be uncorrelated with any information
available at time t. Hence it is common to test the condition that
R+P
X
−1
P L0 (f (zt , θ), yt+1 , zt )vt = 0, (35)
t=R+1
where vt is any function of {zs }ts=1 . Such a test can be conducted by regressing L0 on vt and testing
that the OLS coefficients are zero. A particular function of vt that is often employed is the forecast
itself, which is a function of zt and hence is a possible choice for vt .
Under MSE loss, L0 (f (zt , θ), yt+1 , zt ) ∝ yt+1 − f (zt , θ) = et+1 and thus is proportional to the
forecast error. Hence (34) simply tests if the forecast errors have zero mean, and (35) tests that
forecast errors are uncorrelated with any information available at the time that the forecast is made.
For these reasons (34) is known as an unbiasedness test and (35) is known as an orthogonality test.
The most popular form of these tests is the Mincer-Zarnowitz (1969) regression
where ut+1 is an error satisfying E[ut+1 |zt ] = 0. Unbiasedness can now be tested through the joint
constraint that β c = 0 and β = 1.
Tests such as (34) and (35) examine whether or not the information in zt has been used efficiently
in the construction of the forecast. This is an important issue because a rejection of the test would
35
suggest that improved forecasts are possible given the available data. It is also important from
the perspective of testing rationality when the forecasts f (zt , θ) are constructed by agents that are
expected to be acting rationally and zt is data that would have been available to those agents when
they constructed their forecasts.
To examine these tests from an econometric perspective, recall that f (zt , θ)−and possibly also
the instrument vt −is constructed using parameter estimates based on data up to time t. When
evaluating the sampling distribution for the regression estimates in the unbiasedness or orthogo-
nality tests (34) and (35), we must therefore consider the sampling variability that arises through
the fact that the variables in the regression are constructed. West and McCracken (1998) provide
results for these regressions covering a number of methods for constructing the forecasts and vt .
Under assumptions similar to those in West (1996) they show that the coefficients in the regression
tests are asymptotically normal, although the variance covariance matrix may need to be adjusted
to allow for the additional variation arising from sampling variation in the underlying parameter
estimates.
An additional practical concern involves the specification of vt as a function of zt . Often there
are numerous candidate variables in zt . This, combined with the possibility that we could use any
functional form of zt as an instrument, means that the list of candidates is practically unlimited.
Any test of orthogonality has power only in the direction of the included instrument, vt . For
example, in forecasting inflation with vt set to past interest rates, the test would be capable of
picking up any additional explanatory power in interest rates but not for other variables. The
same is true for getting the functional form correct. Avoiding the first problem – picking the
wrong zt to include – is difficult. For the second problem, Corradi and Swanson (2002) suggest a
nonparametric method for estimating a general function of the included elements of zt .
For loss functions other than mean squared loss, L0 (.) is no longer equivalent to the forecast
errors. Hence it is possible that forecast errors are not mean zero and that past information may
well be correlated with forecast errors even when the forecast is constructed optimally. Indeed, it
is clear from (34) and (35) that the tests rely on the use of the correct loss function. Keane and
Runkle (1990, p. 719) write “If forecasters have differential costs of over- and underprediction, it
could be rational for them to produce biased forecasts. If we were to find that forecasts are biased, it
could still be claimed that forecasters were rational if it could be shown that they had such differential
costs.”
Rationality tests may thus reject, not because the forecaster is using information inefficiently
but because the loss function has not been correctly specified. This is an important issue since the
loss function is generally unknown even though it is invariably assumed to be of the MSE type.
Elliott, Komunjer and Timmermann (2005) examine a class of asymmetric quadratic loss functions
L(et+1 ; α) ≡ [α + (1 − 2α)1I(et+1 < 0)] |et+1 |2 , (37)
where α (0 < α < 1) is the asymmetry parameter. This loss function reduces to MSE when α = 0.5.
Regressing forecast errors on vt (as would be appropriate for MSE loss) results in coefficients on vt
that converge to the true coefficient plus an extra term (1 − 2α)E[vt vt0 ]−1 E[vt |et+1 |]. If vt contains a
constant term (which is usually the case) then E[vt |et+1 |] is always nonzero and orthogonality tests
based on MSE loss will reject asymptotically as a result of using a misspecified loss function.14
14
For an interesting nonparametric evaluation of forecasts, see Campbell and Ghysels (1995).
36
In general, future values of L0 (.) should not themselves be predictable given any variables in
the forecaster’s current information set. A joint test of forecast efficiency (rationality) can thus
readily be conducted within the context of a given family of loss functions which yields L0 (.) as a
function of a finite set of unknown parameters. If the test is rejected, either the forecaster did not
use information efficiently or the family of loss functions was incorrectly specified.
In situations where the loss function is not known up to a small set of shape parameters, it
is possible to use tests that trade off assumptions about the underlying data generating process
against much weaker assumptions on the loss function (such as homogeneity properties). Patton
and Timmermann (2006) show that when loss is only required to be a homogenous function of
the forecast error, while the data generating process can have dynamics in the first- and second
conditional moments (thus covering a large range of nonlinear-in mean specifications, ARCH models
etc.), a simple quantile regression test can be used to test forecast optimality.
When several forecasts at multiple horizons are simultaneously available (as in the case with
many survey forecasts), this offers significant advantages in terms of constructing tests for forecast
efficiency that do not depend on knowing which information was available to the forecaster. As-
suming that the forecaster makes efficient use of all historical information, under MSE loss and
a stationary data generating process we have that M SEhL > M SEhS , where hL > hS are long
and short forecast horizons respectively.15 To see this, suppose the bias is zero at all horizons and
that the variance of the optimal two-period forecast is smaller than that of the optimal one-period
forecast. In this case the variance of last period’s two-step-ahead forecast must be smaller than
the variance of the current one-period forecast, contradicting the assumption that the current one-
period forecast was optimal in the first place. Hence, under appropriate stationarity assumptions,
expected loss must be non-decreasing in the length of the forecast horizon.
37
The real-time nature of economic forecasting affects all stages of the forecasting process: Mod-
els must be formulated, selected and estimated in real time. Evidence of model break-down or
misspecification must also be examined in real time. It is not clear, for example, what one can
conclude from full-sample evidence of forecast inefficiency. Unless the inefficiency was detectable
at an earlier stage of the sample, using information that was available historically, it cannot be
established that the forecaster acted irrationally. For example, under MSE loss, the forecast errors
should be mean-zero conditional only on the available information (including the forecasting model)
at the point the forecast was formulated and not conditional on full-sample information.16
Real-time considerations even pertain to the data “vintage” that was available at a given point
in time and could have been used to formulate and evaluate a forecasting model. Croushore (2006)
and Croushore and Stark (2003) make it clear that key macroeconomic data such as GDP growth
are subject to important revisions, partly due to regular updates from preliminary to secondary and
later data releases, partly due to changes in the methodology used to measure a particular variable.
These revisions can lead not only to changes in the estimated parameters but can also affect the
dynamic lag structure or functional form of the forecasting model and hence change conclusions
regarding predictive relationships (Amato and Swanson (2001)). Data revisions are even more
important for composite series such as the index of leading indicators whose composition may
change due to past failures in forecasting (Diebold and Rudebusch (1991)).
These points emphasize that it is important to use the original data vintages when simulating the
real-time out-of-sample forecasting process and evaluating the precision of the resulting forecasts.
38
For the second case, consider a binary outcome. The density for a binary outcome is equivalent
to an estimate of a probability of the positive outcome, which in turn is simply a parameter
estimate. Hence for this special case parameter estimation and density estimation are equivalent.
In the case of a misspecified parametric density, Elliott and Lieli (2006) show that estimation that
takes the loss function into account can provide a better estimate of the probability of a positive
outcome for those loss functions. The estimator depends on the loss function. However each case
yields a different estimate of the density. This is generally true for parametric estimation. Because
parametric density estimation is the estimation of the parameters of the density, and different loss
functions suggest different estimation techniques for the parameters, estimating the density without
paying attention to the loss function and ultimate use of the forecast density involves estimation
trade-offs (either implicit or explicit) that favor some users at the expense of others.
Although density forecasts are still not commonly reported, a literature has emerged on how
such forecasts should be evaluated. A basic tool used to this end is the so-called probability inte-
gral transform. This is simply the inverse of the cumulative density function, F −1 , implied by a
particular parametric forecasting model. When applied to the actual realization of the predicted
variable (y), F −1 (y) should be drawn from a uniform distribution and be independently and iden-
tically distributed over time. This argument ignores the effect of using estimated parameters of
course, but this type of test has regained popularity following the study by Diebold, Gunther and
Tay (1998). Corradi and Swanson (2006a,b) cover these methods and provide a comprehensive
summary of current tests in this area.
R
Bayesian methods provide the predictive density fY (y|z) = f (y|z, θ)π(θ|z)dθ. When the out-
come is realized, it can be compared to the density the model suggests it should be a draw from.
A natural statistic to compute is the p-value of this outcome, y, i.e. P (y < Y p ) where Y p is the
random variable with density fY (y). If this p-value is extreme it might bring the quality of the fore-
casting model into question. This evaluation method is used by, for example, Pesaran, Pettenuzzi
and Timmermann (2006) to assess the quality of forecasts of interest rates from various models.
Despite issues with estimation, there is one major advantage of the provision of a density fore-
cast, especially when the decision maker and the forecaster are different. Density forecasts convey
the uncertainty in the decision making environment, in perhaps a better way than expressions such
as MSE do to decision makers. For an interesting example, see Whiteman (1996) who recounts his
experience with providing density forecasts to Iowa state officials.
39
to a class of particular (parametric) specifications. Forecasting methods are a broader concept and
comprise rules used to select a particular forecasting model at a given point in time as well as
the approach used to estimate the forecasting model’s parameters−e.g. rolling versus expanding
windows.
40
equally precise in the sense that they have identical risk:
where θ 1 , .., θn are the pseudo-true parameters under models 1,.., n. This is the null hypothesis that
West (1996) and many subsequent studies consider. This is akin to viewing forecasting performance
as a specification test for the underlying models.
In contrast, Giacomini and White (2006) consider the comparison of forecasting methods which
comprise not only the prediction model but also the estimation method and length of the estimation
sample. The null they study is quite different from that in (38). Their null hypothesis is concerned
with testing that the accuracy of all models is identical. This involves taking expectations over
Y, Z and the parameter estimates θ̂1 , .., θ̂n which are random variables.
The Giacomini-White analysis shifts the focus away from comparisons based on average perfor-
mance towards the conditional expectation of differences in performance across forecasting methods.
One advantage of this approach is that it directly accounts for the effect of parameter uncertainty
by expressing the null in terms of estimated parameters and estimation windows. Provided that
estimation uncertainty does not vanish asymptotically, nested models can thus be compared under
this approach. In practice this means that the Giacomini-White approach is most relevant to the
comparison of forecasting models estimated using rolling windows.
Turning to the second point, in an important paper that spurred many of the subsequent
studies in the literature, Diebold and Mariano (1995) suggest using the standard t−statistic for
testing equivalence in forecasting performance for pairwise model comparisons (n = 2) by taking
the difference of the estimated losses and testing if the resulting time series has zero mean. For
scaling they suggest a robust estimator of the variance, and suggest comparing this t-statistic to
the standard normal distribution. Special cases of the West (1996) results are able to justify use
of standard variance estimators for the Diebold and Mariano test.17 Their approach does not,
however, account for parameter estimation errors.
West (1996) is the first paper to account for the effect of parameter estimation error on forecast
comparisons when forecasts are updated recursively through time. Clark and West (2004) provide
an interesting illustration of how important parameter estimation can be in the comparison of a
benchmark model with few or none parameters (e.g. the prevailing mean) versus a more heavily
parameterized alternative model that may include time-varying predictor variables. Even when the
larger model is true, because it involves estimation of more parameters and hence is more subject
to parameter estimation error, we would expect this model to perform worse in finite samples
than the simpler (biased) model, unless the predictive power of the extra regressor(s) is sufficiently
large. Clark and West propose a test that accounts for this problem by correcting for parameter
estimation error.
Turning to the third and final point, a limitation of the methods developed in Diebold and
Mariano (1995) and West (1996) is that they only apply when the forecasting models are non-
nested. Clark and McCracken (2001) extend this work and develop tests for comparing nested
models in the presence of parameter estimation uncertainty. This is the case most commonly faced
by empirical researchers and thus is an important step forward in this area.
17
The most prominent special case is when the expected value of the derivative of the loss function with respect
to θ is zero evaluated at the true θ.
41
Whether the nested or non-nested case applies to a given forecast comparison can be surprisingly
difficult to determine. In practice it is often forecasting methods as opposed to forecasting models
that are being compared, whereas most theory is developed for comparing forecasting models. A
given forecasting method, when applied recursively through time, may select different forecasting
models at different points in time. This means that the models selected by two different forecasting
methods sometimes could be nested while at other times could be non-nested. To our knowledge,
no test exists at the present time that handles this complication, making forecast comparisons a
tricky exercise in practice.
where R(f b (z), θ̂b ) is the benchmark performance. The alternative hypothesis is that the best of
the forecast methods outperforms the benchmark, i.e. that it has lower risk:
Since the distribution of the risk differentials is asymptotically normal under the assumptions of the
method (based on the results of West (1996)) this amounts to constructing a test for the maximum
of a set of joint normals with unknown covariance matrix. White solves this problem by employing
a bootstrap procedure to the estimates of risk. This bypasses the need to compute the unknown
variance covariance matrix and directly estimates the p-value for the test.
Hansen (2005) shows that when poor models are added to the set of candidate models, such
that the benchmark is better than other models, the asymptotic normal result fails. He suggests a
procedure where underperforming models are first removed in an initial step.
Controlling for data snooping can be important empirically. In the context of forecasting models
for daily stock market returns based on technical trading rules, Sullivan, Timmermann and White
(1999) find that data snooping can account for what otherwise appears to be strong evidence of
return predictability.
42
8.3 Forecast Encompassing
Chong and Hendry (1986) introduced the idea of forecast encompassing, which can be applied when
choosing between forecasting models. Under MSE loss the idea is similar to the orthogonality
regressions (35) although the additional information vt is no longer a subset of zt but instead
consists of forecasts or forecast errors from other forecasting methods. The idea is simple: If
other forecasts have information relevant for the predicted variable that is not contained in the
original forecast, then such forecasts will enter the orthogonality regression with a nonzero weight.
This would mean that the original forecast did not include all relevant information. Conversely,
if orthogonality holds, then the first forecast is said to encompass the other forecasts because it
incorporates all the relevant information that the other forecasts have.
Romer and Romer (2000) provide an interesting comparison of the Federal Reserve Green Book
inflation forecasts with private sector forecasts using encompassing regressions. Assuming MSE
loss, they find evidence that the Fed inflation forecasts encompass the private sector forecasts. This
conclusion is questioned by Capistran (2005) who finds evidence of significant biases of opposite
sign in the Fed’s forecasts during the pre- and post-Volker periods. Averaged over the full sample
the bias is small, but this conceals evidence of a tendency to underpredict inflation in the pre-Volker
sample followed by subsequent overpredictions.
To test if a particular forecast (null model) encompasses a set of alternative forecasts, a regres-
sion of the forecast error from the null model on the difference between the other forecast errors
and that of the null model (e∗t ) can be undertaken using a t− or an F −test (see, e.g., Clements
and Hendry (1998, p. 265))
Alternative forms of this test have been suggested by Harvey et. al. (1998) and Clark and Mc-
Cracken (2001) when two forecasts are being compared. To handle the problem that forecast errors
depend on the estimated parameter, θ, the results of West and McCracken (1998) can be used.
When the models are nested, the singularity of the joint distribution of the forecast errors is again
a problem. Clark and McCracken (2001) show that in these cases the asymptotic distribution
of a rescaled statistic can be approximated with a function of Brownian motions and hence the
distribution is nonstandard in this case.
43
testing for the presence of real time predictability under the conditions facing actual forecasters
in finite samples, then the use of a hold-out sample may make sense. In the latter case, the bias-
variance trade-off may benefit small, misspecified models even though these models do not have
good population properties.
Under MSE loss, the problem of comparing forecast models reduces to the question of which
forecast procedure is closest to the conditional expectation. This can be tested in the full sample
without problems of nesting of the models so long as the data are sufficiently stationary and not
too dependent, although such tests are of course subject to the earlier mentioned caveats.18
The desire to test out-of-sample forecasting performance is closely related to the uncertainty
about the underlying data generating process and also a concern for the effect of any pre-testing
that might have occurred in constructing the forecasting models. Inoue and Kilian (2004) have
questioned the practice of using a hold-out sample altogether arguing that it does not protect
against data mining because the information available in ‘pseudo real time’ experiments is the
same as that available to someone with access to the full sample.
44
different forecasts diversify against modeling risk−which depends on the correlation in forecast
errors across models−and how much weight to assign to the various forecasts.
A direct answer to the question of how to obtain a set of combination weights is provided by
Bates and Granger (1969) who suggest simply regressing the predicted variable y on the individual
forecasts fi (z, θ) along with a constant
n
X
y = β0 + β i fi (z, θ) + ε. (42)
i=1
When the individual forecasts are believed to be unbiased, it is common to omit the intercept
term and restrict the slope coefficients to sum to one in which case they can be interpreted as
forecast combination weights. This approach assumes MSE loss but has been generalized to other
loss functions and method of moment type estimators (Elliott and Timmermann (2004)).
A comparison of the combination approach to encompassing explains why combining may be
expected to be a more reasonable approach than selecting a single forecast, unless of course the
true model is known to be included in the set of models under consideration and can be identified
in practice. A single forecast only gets selected when the combination puts full weight on one of
the forecasting methods while the rest are given zero weights. This is precisely the case where the
forecast with a weight of one encompasses the other forecasts. However, this is a special case of the
general concept of forecast combination, and so might be expected to be less commonly supported
empirically than more evenly distributed weights.
In practice, although empirical evidence suggests that forecast combinations tend to outperform
forecasts from a single model, strategies designed to obtain optimal combination weights are often
outperformed by simple measures such as averaging the raw forecasts (i.e. giving all forecasts
equal weights) or a trimmed set of these. If the models use roughly the same data sources and
empirical techniques so differences in the performance across forecasting models are too small to be
easily rejected by the data, they will tend to have similar error variances and covariances. In this
situation, giving each forecast identical weights can be relatively efficient. Palm and Zellner (1992)
suggest other reasons– e.g. instability of the covariance between forecast errors or estimation error
in the combination weights.
Which combination methodology is best may well depend on the state of the economy because
the speed with which different forecasts incorporate shifts in the economy could vary. For example,
when the economy is running at a normal pace, time-series models may provide the most accurate
forecast because they make efficient use of historical information. However, these models may be
slower at capturing or predicting turning points−such as the emergence of a recession−compared
with seasoned professional forecasters with access to a much larger information set. This idea is
consistent with findings reported in Elliott and Timmermann (2005) who use a regime switching
approach to track variations in the forecasting performance of time-series and survey forecasts of
six key macroeconomic variables and form a combined forecast.
Bayesian approaches to forecast combination are becoming increasingly popular in empirical
studies. Bayesian Model Averaging has been proposed by, inter alia, Leamer (1978) and Raftery
et al (1997). Under this approach, the predictive density can be computed by averaging over a set
45
of models, Mi , i = 1, ..., n :
n
X
pY (y|z) = p (Mi |z) pY (y|Mi , z) . (43)
i=1
Here p (Mi |z) is the posterior probability of model Mi obtained from the model priors π (Mi ), the
priors for the unknown parameters of each model, π (θi |Mi ), and the likelihood of the models under
consideration. pY (y|Mi , z) is the predictive density of y under the ith model, Mi , given z obtained
after uncertainty about the parameters θi has been integrated out using the posterior. Unlike the
weights used in the classical least-squares combination literature, these weights do not account
for correlations between forecasts and the weights are always confined to the zero-unity interval.
Palm and Zellner (1992) develop a general Bayesian framework for combinations where individual
forecasts can be biased and the covariance matrix of the forecast errors may be unknown. More
details are provided in Timmermann (2006) and Geweke and Whiteman (2006).
9 Empirical Application
To illustrate many of the issues discussed above we consider the predictability of US inflation
and stock returns. For inflation we use log first differences of the CPI while stock returns are
captured by the value-weighted portfolio of US stocks tracked by the Center for Research in Security
Prices (CRSP). Both series are measured at the monthly frequency and the sample period is
1959:1-2003:12. To initialize our parameter estimates we use data from 1959:1 - 1969:12. We then
generate out-of-sample forecasts from 1970:01 to 2003:12. Parameter estimates are either updated
recursively, expanding the estimation window by one observation each month, or by means of a
10-year rolling window. Only data up to the previous month is therefore used to estimate the
model parameters and generate forecasts for the current month. This is commonly referred to as a
pseudo out-of-sample forecasting exercise.
We consider twelve forecasting approaches. The first is an autoregressive (AR) model
k
X
yt+1 = β 0 + β j yt+1−j + εt+1 , (44)
j=1
where k is selected to minimize the BIC with a maximum of 18 lags and εt+1 here and in subsequent
models is regarded as white noise. The second model is a factor augmented AR model, using up
to five common factors:
k
X q
X
yt+1 = β 0 + β j yt+1−j + γ j fj,t + εt+1 , (45)
j=1 j=1
where fj,t is the jth factor and k and q are again selected to minimize the BIC (with k ≤ 18 and
q ≤ 5). Factors are obtained using the principal components approach of Stock and Watson (2002)
to a cross-section of 131 macroeconomic time series which begin in 1960. The factors are extracted
in (simulated) real time using either a recursive or a rolling 10-year estimation window.
46
The third and fourth models are Bayesian VARs (BVARs) fitted to the variable of interest
(inflation or stock returns) and the five factors:
k
X
zt+1 = β 0 + β j zt+1−j + εt+1 . (46)
j=1
Here zt = (yt , f1t , . . . , f5t )0 and we include the most recent six months lags, i.e. k = 6. Following
Litterman, own-lag terms at lag j have a prior variance of 0.04/j 2 , while off-diagonal lags have
a prior variance of 0.0004/j 2 . Both a random walk prior and a white noise prior are considered.
Under the random walk prior, the autoregressive parameters are shrunk towards unity, while under
the white noise prior they are shrunk towards zero. Clearly the random walk prior is reasonable for
the inflation example while the white noise prior is more reasonable for stock returns. We report
both for each example to show the effect of the differences in prior choice.
Turning to the non-linear specifications, we consider two logistic STAR models of the form
where
η t = (1, yt )0
(
1/(1 + exp(γ 0 + γ 1 yt−3 ))
dt = .
1/(1 + exp(γ 0 + γ 1 (yt − yt−6 )))
with two hidden units in the first layer (n1 = 2) and one hidden layer in the second layer (n2 = 1).
For both neural net models, g is the logistic function and η t = (1, yt , yt−1 , yt−2 ). Estimation uses
search methods since αj enters nonlinearly.
We also consider more traditional time-series forecasting methods such as exponential smoothing
where the forecast ft is generated by the recursion
47
where f1 = 0, f2 = y2 and λ2 = (y2 − y1 ). Here α and α and β, respectively, are determined so as
to minimize the sum of squared forecast errors in real time.
We finally consider a forecast combination approach that simply uses the equal-weighted average
in addition to a very different approach that, at each point in time, selects the forecasting model
with the best track record up to the present time and then uses this to generate a forecast for the
following period.
In all cases, we apply the following ‘insanity filter’ which constrains outlier forecasts: If the
predicted change in the underlying variable is greater than any of the historical changes up to a
given point in time, the forecast is replaced with a ‘no change’ forecast.
Results in the form of out-of-sample, annualized root mean squared forecast errors (computed
by multiplying the monthly RMSE values by the square root of 12) are presented in Table 1.
First consider the results under recursive parameter estimation. For inflation, the best model is
the average forecast followed by exponential smoothing, the previous best model, the simple and
factor-augmented AR models and the two-layer neural net model. Slightly worse forecasts are
generated by the one-layer neural net and the BVARs, while the STAR models generate somewhat
worse performances.
Overall, these results indicate that there is not much to differentiate between a cluster of the
best forecasting models. This point is reinforced by the plots of predicted values from three of the
models shown in Figure 1. Inflation forecasts from seemingly very different approaches are quite
similar and dominated by a persistent common component.
Turning to the stock returns and focusing again on the results under recursive estimation, Table
1 shows that the best overall performance is delivered by the combined forecast and the simple
and factor-augmented AR models. Once again the BVAR models perform rather poorly as do
the STAR models and single layer neural nets. For stock returns which are not dominated by a
strongly persistent component, there is more to differentiate between the time-series of forecasts
as shown in Figure 2. Overall, however, while a few approaches perform quite poorly it is difficult
to distinguish with statistical precision between the forecasting performance among a cluster of
reasonable forecasting models.
Forecast precision tends to deteriorate significantly for the BVAR and double exponential
smoothing forecasts under the 10-year rolling estimation window. This happens both for infla-
tion and stock returns. This deterioration is likely due to the larger estimation error associated
with using a shorter estimation window, although one should not forget that there is a trade-off in
the form of faster adaptability as witnessed by the improved forecasting performance observed for
the first STAR model’s inflation forecasts.
Evaluation of the forecasts from these models is complicated because some of them are pair-wise
nested (for example, the STAR and neural net models nest the AR models), while others are not
(e.g. the factor-augmented AR models are not nested by the neural net models). As a diagnostic
test, we simply compare the MSE performance of pairs of models using the Giacomini-White (2006)
approach. Our results assume a rolling estimation window corresponding to 10 years of monthly
observations, i.e. 120 data points and therefore reflects the RMSE values reported in columns two
and four in Table 1.
Results from these pairwise comparisons are reported in Table 2. While the BVAR and double
exponential smoothing inflation forecasts are soundly rejected against those produced by the better
48
models, for most of the other comparisons these tests do not have sufficient power to choose one
model over another. The results are somewhat different in the case of the stock returns which,
unlike the inflation series, do not contain a large persistent component and hence are more difficult
to predict. There is little evidence to distinguish between the simple and factor-augmented AR
models, the exponential smoothing, average and previous best forecasts of stock returns. Con-
versely, the BVARs, double exponential smoothing, STAR and neural net forecasts are generally
rejected against the first group of forecasts. Parsimony seems to be key to successfully predict stock
returns, particularly when a relatively short rolling estimation window of 120 observations is used.
Once again it is clear that although a few approaches perform very poorly and can be rejected
out of hand, it is difficult to systematically differentiate between many of the other approaches.
The discussion has so far assumed MSE loss. Theory suggests that the form of the loss function
alters the optimal functional form of the forecasting model. To illustrate this, we next generated
forecasts under lin-lin loss, setting p = 1 and α equal to 0.35, 0.50 or 0.65 in equation (12),
and considering either the AR model or the factor-augmented AR model.19 All forecasts were
generated using a recursive estimation window. Results from this analysis are presented in Table 3.
For inflation the average value of the lin-lin loss function under the simple AR model is generally
significantly below the values produced under the factor-augmented AR specification. This holds
irrespective of which quantile is being considered. In contrast, for stock returns the two models
produce almost identical out-of-sample forecasting performance.
These empirical results support many of the themes of our theoretical analysis. First, forecasts
from seemingly very different approaches (e.g. linear versus nonlinear models) often produce very
similar results - witness the similar RMSE performance of the neural nets and the autoregressive
models. In part these similarities arise because we truncate the forecasts from the nonlinear models
when these are too far away from the historical sample data.20 In other cases nonlinear models
can generate poor forecasts due to their sensitivity to outliers and their imprecisely estimated
parameters. This last point is illustrated through the performance of the STAR models which
generally was quite poor.
Secondly, it is difficult to outperform simple approaches such as a parsimonious autoregressive
model. Simple forecasting approaches tend to generate relatively smooth and stable forecasts
without being subject to too much parameter estimation error.
Third, and as an extension of the previous point, it appears that in many cases there are only
marginal gains (in terms of out-of-sample RMSE performance) over and above projection on past
values of the series themselves from considering the additional information that can be extracted
from large data sets. For persistent variables such as inflation, a linear autoregressive component
is clearly the single most important predictive component, while for stock returns it is difficult to
come up with predictor variables with significant predictive value.
Fourth, our results support the finding that forecast combination offers an attractive approach
for many economic and financial variables. The average forecast produced the best or second
19
These forecasts were generated using quantile regression, see Koenker and Bassett (1978). This already presents
a nonlinear optimization problem so we only consider linear quantile specifications in our analysis.
20
When extreme forecasts are not truncated, the RMSE for the stock return forecasts rises to 50 and 32 under
the one- and two-layer neural net models, respectively. These values are three times and twice as large as the values
reported in Table 1.
49
best performance among all approaches for both inflation and stock returns. Thus, while forecast
combination does not always generate the single best performance, it usually beats most alternatives
unless some extremely poor models have been left in the mix of models that get combined.
Fifth, the loss function clearly matters in practice. We saw that under MSE loss, the purely
autoregressive and factor-augmented autoregressive models produced essentially indistinguishable
forecasting performance. In contrast, under lin-lin loss, the simple autoregressive forecasts were
better for the inflation series, although they were nearly identical in the case of the stock returns.
Finally, model instability and/or sensitivity of forecasting performance to the sample period is
clearly an issue. Table 1 compares the MSE performance under a recursive estimation approach
which uses an expanding estimation window against that of a rolling 10-year window which can
better accommodate shifts in the underlying data generating process. In many cases, the choice
of estimation window makes a sizeable difference. If estimation error was the predominant effect,
we would expect the ten-year rolling forecasts uniformly to be worse than the forecasts based on
the expanding estimation window. This is exactly what we find for stock returns where there is
no evidence that shortening the estimation window leads to improvements in any of the models.
For inflation, however, we see that for half of the models the out-of-sample forecasting performance
either improves or stays the same as a result of going from the expanding to the rolling estimation
window. Indeed, in the case of the first STAR model, the latter approach produced substantially
better forecasts.
10 Conclusion
The menu of forecasting methodologies available to the applied economist has expanded vastly over
the last few decades. No single approach is currently dominant and choice of forecasting method is
often dictated by the situation at hand such as the forecast user’s particular needs, data availability
and expertise in experimenting with different classes of models and estimation methods. Economic
forecasts are often only one piece of information used in conjunction with a decision maker’s prior
beliefs and other sources of information. Moreover, such forecasts are often used as a way to assign
different weights on various possible scenarios. Purely statistical approaches based on complicated
‘black box’ approaches have with few exceptions so far failed to generate much attention among
economists and are not used to the extent one might otherwise have expected.
Although the situation is still evolving, recent research in the forecasting literature has sup-
ported some broad conclusions:
• Careful attention to the forecaster’s objectives is important not only in the forecast evaluation
stage but also in the estimation and model selection stages. For example, if the forecaster’s
loss function suggests that overpredictions are more costly than underpredictions (or vice
versa) and a particular quantile of the forecast distribution best summarizes the economic
objectives of the forecasting exercise, then quantile rather than least squares estimation should
be used;
• Models of economic and financial time series are often found to be unstable through time
and so forecasting models are best viewed as “approximations” or tracking devices. As a
50
consequence one should not expect that the same forecasting model will continue to dominate
in different historical samples;
• Choice of the sample period used to estimate the parameters of the forecasting model is
therefore important. Using the longest possible data sample or a simple rolling window is not
necessarily the best approach if more precise information about the cause of model instability
is available (e.g. institutional shifts, changes in tax policy or legislation, large technology or
supply shocks). Since the nature and form of model instability may often not be very clear,
more research is required to design robust forecasting approaches that detect and incorporate
model instability in a variety of situations;
• Forecast combination has often been found to offer an attractive alternative to the approach of
seeking to identify a single best forecasting model. In part this stems from the fact that com-
bination allows forecasters to hedge against model uncertainty and shifts in models’ (relative)
forecasting performance;
• Overfitting is an overriding concern in forecasting because of the short time series often
encountered and the difficulty in getting independent data samples that can be used to cross-
validate the forecasting models. This problem is exacerbated for financial time series where
the signal-to-noise ratio tends to be very low. Parameter estimation error also is the likely
reason why including additional economic variables in a forecasting model, which may seem
justified ex ante, often fails to lead to the expected improvement in terms of out-of-sample
forecasting performance;
• It is often difficult to distinguish with much statistical precision between the forecasts gener-
ated by seemingly very different forecasting methods. When large differences in forecasting
performance occur, this often has to do with the tendency of nonlinear forecasting models
to generate outliers in the forecast error distribution due to their sensitivity to the partic-
ular sample used for parameter estimation. How such outliers are dealt with then becomes
important in practice;
• Guidance from economic theory is important at several stages of the forecasting process.
Besides assisting in the choice of the forecaster’s objective function, economic theory can
be helpful in selecting categories of variables to be considered as potential predictors and in
imposing long-run restrictions which may reduce parameter estimation error. Econometric
methods can then be used for variable selection among the predictors deemed potentially
relevant from a theoretical perspective (often a large set), for specification of the short-run
dynamics and for determination of the functional form of the forecasting model.
References
[1] Amato, J.D. and N.R. Swanson, 2001, The Real Time Predictive Content of Money for
Output. Journal of Monetary Economics 48, 3-24.
51
[2] Andersen, T.G., T. Bollerslev, P.F. Christoffersen and F.X. Diebold, 2006, Volatility and
Correlation Forecasting. Pages 777-877 in G. Elliott, C.W.J. Granger and A. Timmermann
(eds.) Handbook of Economic Forecasting. Amsterdam: North Holland.
[3] Ang, A. and G. Bekaert, 2002, Regime Switches in Interest Rates, Journal of Business and
Economic Statistics, 20, 163-182.
[4] Artis, M. and M. Marcellino, 2001, Fiscal Forecasting: The Track Record of the IMF, OECD
and EC. Econometrics Journal 4, S20-S36.
[5] Bai, J., 1997, Estimation of a Change Point in Multiple Regression Models. Review of Eco-
nomics and Statistics 79, 551-563.
[6] Bai, J. and P. Perron, 1998, Estimating and Testing Linear Models with Multiple Structural
Changes. Econometrica 66, 47-78.
[7] Bates, J.M. and C.W.J. Granger, 1969, The Combination of Forecasts. Operations Research
Quarterly 20, 451-468.
[8] Batchelor R. and P. Dua, 1991, Blue Chip Rationality Tests. Journal of Money, Credit and
Banking 23, 692-705.
[9] Batchelor, R. and D.A. Peel, 1998, Rationality Testing under Asymmetric Loss. Economics
Letters 61, 49-54.
[10] Box, G. and G. Jenkins, 1970, Time Series Analysis: Forecasting and Control. Holden-Day,
San Francisco.
[11] Brayton, F. and P. Tinsley, 1996, A Guide to FRB/US.A Macroeconomic Model of the United
States. Federal Reserve Board working paper 1996-42.
[12] Breiman, 1995, Better Subset Regression Using the Nonnegative Garrote. Technometrics 37,
373-384.
[14] Brown, B.Y. and S. Maital, 1981, What do Economists Know? An Empirical Study of
Experts’ Expectations. Econometrica 49, 491-504.
[15] Campbell, B. and E. Ghysels, 1995, Federal Budget Projections: A Nonparametric Assess-
ment of Bias and Efficiency. Review of Economics and Statistics, 17-31.
[16] Capistran, C., 2005, Bias in Federal Reserve Inflation Forecasts: Is the Federal Reserve
Irrational or Just Cautious. Mimeo, Banco de Mexico.
[17] Chauvet, M., 1998, An Econometric Characterization of Business Cycle Dynamics with Factor
Structure and Regime Switches. International Economic Review 39, 969-96.
[18] Chong, Y.Y. and D.F. Hendry, 1986, Econometric evaluation of linear macro-economic mod-
els, Review of Economic Studies 53:671-690.
52
[19] Christoffersen, P.F. and F.X. Diebold, 1997, Optimal Prediction under Asymmetric Loss.
Econometric Theory 13, 808-817.
[20] Christoffersen, P.F. and F.X. Diebold, 2006, Financial Asset Returns, DIrection-of-Change
Forecasting, and Volatility Dynamics. Management Science 52, 1273-1287.
[21] Christoffersen, P.F. and K. Jacobs, 2004, The Importance of the Loss Function in Option
Valuation. Journal of Financial Economics, 72, 291-318.
[22] Clark, T.E. and M.W. McCracken, 2001, Tests of Equal Forecast Accuracy and Encompassing
for Nested Models. Journal of Econometrics 105, 85-110.
[23] Clark, T.E. and M.W. McCracken, 2007, Tests of Equal Predictive Ability with Real Time
Data. Mimeo, Federal Reserve Board.
[24] Clark, T.E. and K.D. West. 2004. Using Out-of-Sample Mean Squared Prediction Errors
to Test the Martingale Difference Hypothesis. Working Paper 04-03, Kansas City Federal
Reserve, Kansas City, USA.
[25] Clemen, R.T., 1989, Combining Forecasts: A Review and Annotated Bibliography. Interna-
tional Journal of Forecasting 5, 559-581.
[26] Clements, M.P. and D.F. Hendry, 1998, Forecasting Economic Time Series, Cambridge Uni-
versity Press.
[27] Clements, M.P. and D.F. Hendry, 2006, Forecasting with Breaks in Data Processes, in C.W.J.
Granger, G. Elliott and A. Timmermann (eds.) Handbook of Economic Forecasting, 605-657,
Amsterdam, North-Holland.
[28] Clements, M.P. and D.F. Hendry, 2002, Modelling Methodology and Forecast Failure. Econo-
metrics Journal 5, 319-344.
[29] Corradi, V. and N.R. Swanson, 2002, A Consistent Test for Out of Sample Nonlinear Predic-
tive Ability. Journal of Econometrics 110, 353-381.
[30] Corradi, V. and N.R. Swanson, 2006a, Predictive Density Evaluation, in C.W.J. Granger, G.
Elliott and A. Timmermann (eds.) Handbook of Economic Forecasting, 197-286, Amsterdam,
North-Holland.
[31] Corradi, V. and N.R. Swanson, 2006b, Predictive Density and Conditional Confidence Interval
Accuracy Tests. Journal of Econometrics 135, 187-228.
[32] Corradi, V. and N.R. Swanson, 2007, Nonparametric Bootstrap Procedures for Predictive
Inference Based on Recursive Estimation Schemes. International Economic Review.
[33] Croushore, D., 2006, Forecasting with Real-Time Macroeconomic Data. Pages 961-982 in G.
Elliott, C. Granger and A. Timmermann (eds.) Handbook of Economic Forecasting. North-
Holland: Amsterdam.
53
[34] Croushore, D. and T. Stark, 2003, A Real-time Data Set for Macroeconomists: Does the
Data Vintage Matter? Review of Economics and Statistics 85, 605-617.
[35] Davies, A. and K. Lahiri, 1995, A New Framework for Analyzing three-dimensional Panel
Data. Journal of Econometrics 68, 205-227.
[36] Del Negro, M., F. Schorfheide, F. Smets and R. Wouters, 2006, On the Fit of New-Keynesian
Models. Forthcoming in Econometric Reviews.
[37] Diebold, F., Gunther, T., and A., Tay, 1998, Evaluating Density Forecasts, International
Economic Review 39, 863-883.
[38] Diebold, F. X. and L. Kilian, 2000, Unit-Root Tests are Useful for Selecting Forecasting
Models, Journal of Business and Statistics, 18, 265-273.
[39] Diebold, F.X. and R. Mariano, 1995, Comparing Predictive Accuracy. Journal of Business
and Economic Statistics 13, 253-65.
[40] Diebold, F.X. and G.D. Rudebusch, 1991, Forecasting Output with the Composite Leading
Index: A Real-Time Analysis, Journal of the American Statistical Association, 86, 603-610.
[41] Diebold, F.X. and G.D. Rudebusch, 1996, Measuring Business Cycles: A Modern Perspective.
Review of Economics and Statistics 78, 67-77.
[42] Doan, T., R. Litterman and C. Sims, 1984, Forecasting and Conditional Projection using
Realistic Prior Distributions", Econometric Reviews, 3, 1-144 (with discussion).
[43] Ehrbeck, T. and Waldmann, R., 1996, Why are Professional Forecasts Biased? Agency versus
Behavioral Explanations, Quarterly Journal of Economics 111, 21-40.
[44] Elliott, G., 2005, Forecasting in the presence of a break. Mimeo, UCSD.
[45] Elliott, G., I. Komunjer and A. Timmermann, 2005, Estimating Loss Function Parameters.
Review of Economic Studies 72, 1107-1125.
[46] Elliott, G., I. Komunjer and A. Timmermann, 2006, Biases In Macroeconomic Forecasts:
Irrationality or Asymmetric Loss? Mimeo UCSD.
[47] Elliott, G. and R. Lieli, 2006, Predicting Binary Outcomes, manuscript, UCSD.
[48] Elliott, G. and U. Mueller, 2006, "Efficient Tests for General Persistent Time Variation in
Regression Coefficients”, Review of Economic Studies, 73, 907-940.
[49] Elliott, G. and A. Timmermann, 2004, Optimal Forecast Combinations Under General Loss
Functions and Forecast Error Distributions. Journal of Econometrics 122, 47-79.
[50] Elliott, G. and A. Timmermann. 2005. Optimal Forecast Combination Weights Under Regime
Switching. International Economic Review 46, 1081-1102.
54
[51] Engle, R.F. and C.W.J. Granger, 1987, Co-integration and Error Correction: Representation,
Estimation and Testing. Econometrica 55, 251-276.
[52] Figlewski, S. and P. Wachtel, 1981, The Formation of Inflationary Expectations. Review of
Economics and Statistics 63, 1-10.
[53] Forni, M., M. Hallin, M. Lippi, and L. Reichlin, 2000. The Generalized Factor Model: Iden-
tification and Estimation. Review of Economics and Statistics 82, 540-554.
[54] Forni, M., M. Hallin, M. Lippi, and L. Reichlin, 2003, The Generalized Dynamic Factor
Model: Forecasting and One Sided Estimation, CEPR working paper 3432.
[55] Franses, P.H. and D. van Dijk, 2005, The forecasting performance of various models for season-
ality and nonlinearity for quarterly industrial production, International Journal of Forecasting
21, 87-102.
[56] Gallant, R. 1981, On the Bias in Flexible Functional Forms and an Essentially Unbiased
Form: The Fourier Flexible Form. Journal of Econometrics 15, 211-245.
[57] Garcia, R. and P. Perron, 1996, An Analysis of the Real Interest Rate under Regime Shifts.
Review of Economics and Statistics 78(1), 111-125.
[58] Geweke, J., 2005, Contemporary Bayesian Econometrics and Statistics. Wiley: New York.
[59] Geweke, J. and C. Whiteman, 2006, Bayesian Forecasting. Pages 3-80 in G. Elliott, C.W.J.
Granger and A. Timmermann (eds.) Handbook of Economic Forecasting. North-Holland:
Amsterdam.
[60] Giacomini, R. and B. Rossi, 2006, Detecting and Predicting Forecast Breakdowns. Duke
University Working Paper.
[61] Giacomini, R., and H., White, 2006, Tests of Conditional Predictive Ability. Econometrica
74, 6, 1545-1578.
[62] Granger, C.W.J., 1966, The Typical Spectral Shape of an Economic Variable. Econometrica
34, 179-192.
[63] Granger, C.W.J., 1969, Prediction with a generalized cost function, OR, 20, 199-207.
[64] Granger, C.W.J., 1999, Outline of Forecast Theory Using Generalized Cost Functions. Span-
ish Economic Review 1, 161-173.
[65] Granger, C.W.J. and M. Machina, 2006, Forecasting and Decision Theory. Pages 81-98 in G.
Elliott, C.W.J. Granger and A. Timmermann (eds.), Handbook of Economic Forecasting.
[66] Granger, C.W.J. and P. Newbold, 1986, Forecasting Economic Time Series, 2nd Edition.
Academic Press, New York.
[67] Granger, C.W.J. and M.H. Pesaran, 2000, Economic and Statistical Measures of Forecast
Accuracy. Journal of Forecasting 19, 537-560.
55
[68] Guidolin, M. and A. Timmermann, 2006, An Econometric Model of Nonlinear Dynamics in
the Joint Distribution of Stock and Bond Returns. Journal of Applied Econometrics 21, 1-22
[70] Hamilton, J.D., 1989, A New Approach to the Economic Analysis of Nonstationary Time
Series and the Business Cycle. Econometrica 57, 357-384.
[71] Hansen, P.R., 2005, A Test for Superior Predictive Ability. Journal of Business and Economic
Statistics 23, 365-380.
[72] Harvey, A.C. and S.J. Koopman, 1993, Forecasting Hourly Electricity Demand Using Time-
Varying Splines. Journal of American Statistical Association 88, 1228-1236.
[73] Harvey, A.C., 2006, Forecasting with Unobserved Components Time Series Models. Pages
327-412 in G. Elliott, C.W.J. Granger and A. Timmermann.
[74] Harvey, D.I., S.J. Leybourne and P. Newbold, 1998, Tests for forecast encompassing, Journal
of Business and Economic Statistics 16, 254-59.
[75] Hendry, D.F. and H-M. Krolzig, 2004, Automatic Model Selection: A New Instrument for
Social Science. Electoral Studies, 23, 525—544.
[76] Hong, H. and J. D. Kubik, 2003, Analyzing the Analysts: Career Concerns and Biased
Earnings Forecasts. Journal of Finance 58,1, 313-351.
[77] Inoue, A. and L. Kilian. 2004, In-sample or out-of-sample Tests of Predictability: Which one
should we use? Econometric Reviews 23(4), 371-402.
[78] Inoue, A. and L. Kilian. 2005, How Useful is Bagging in Forecasting Economic Time Series? A
Case Study of U.S. CPI Inflation. Forthcoming in Journal of American Statistical Association.
[79] Inoue, A. and L. Kilian. 2006, On the Selection of Forecasting Models, Journal of Economet-
rics 130(2), 273-306.
[80] Ito, T., 1990, Foreign Exchange Rate Expectations: Micro Survey Data, American Economic
Review, 80, 434-449.
[81] Jagannathan, R. and T. Ma, 2003, Risk Reduction in Large Portfolios: Why imposing the
wrong constraints helps. Journal of Finance 58, 1651-1684
[82] James, W. and Stein, C., 1961, Estimation with Quadratic Loss. Proceedings of the Fourth
Berkeley Symposium on Mathematical Statistics and Probability, Vol. 1 Berkeley, CA: Uni-
versity of California Press, 361-379.
[83] Kadiyala, K.R. and S. Karlsson, 1993, Forecasting with Generalized Bayesian Vector Autore-
gressions. Journal of Forecasting 12, 365-378.
56
[84] Kadiyala, K.R. and S. Karlsson, 1997, Numerical Methods for Estimation and Inference in
Bayesian VAR-Models", Journal of Applied Econometrics, 12, 99-132.
[85] Keane, M.P. and D.E. Runkle, 1990, Testing the Rationality of Price Forecasts: New Evidence
from Panel Data. American Economic Review 80, 714-735.
[86] Kilian, L., 1999, Exchange Rates and Monetary Fundamentals: What Do We Learn from
Long- Horizon Regressions?, Journal of Applied Econometrics 14, 491-510.
[87] Kilian, L. and S. Manganelli, 2006, The Central Banker as a Risk Manager: Estimating the
Federal Reserve’s Preferences under Greenspan. Mimeo, University of Michigan and ECB.
[88] Kim, C.-J. and Nelson, C.R., 1998, State Space Models with Regime Switching: Classical and
Gibbs Sampling Approaches with Applications. Cambridge, Mass.: MIT Press.
[89] Koenker, R.W. and G.W. Bassett, 1978, Regression Quantiles, Econometrica 46, 33-50.
[90] Koop, G. and S. Potter, 2004, Forecasting and Estimating Multiple Change-point Models
with an Unknown Number of Change-points. Forthcoming, Review of Economic Studies.
[92] Ledoit, O. and M. Wolf, 2003, Improved Estimation of the Covariance Matrix of Stock Returns
with an Application to Portfolio Selection. Journal of Empirical Finance 10, 603-621.
[93] Leitch, G. and J.E. Tanner, 1991, Economic Forecast Evaluation: Profits Versus the Conven-
tional Error Measures, American Economic Review 81, 580-90.
[94] Lim, T., 2001, Rationality and Analysts’ Forecast Bias. Journal of Finance 56-1, 369-385.
[95] Litterman, R.B., 1980, A Bayesian Procedure for Forecasting with Vector Autoregressions.
Working Paper, Massachusetts Institute of Technology.
[96] Litterman, R.B., 1986, Forecasting with Bayesian Autoregressions – Five Years of Experi-
ence, Journal of Business and Economic Statistics, 4, 25-38.
[97] Lopez, J.A. and C.A. Walter, 2001, Evaluating Covariance Matrix Forecasts in a Value-at-
Risk Framework. Journal of Risk 3, 69-91.
[98] Ludvigsson, S.C. and S. Ng, 2005, Macro Factors in Bond Risk Premia. Mimeo, New York
and Michigan University.
[99] Ludvigsson, S.C. and S. Ng, 2007, The Empirical Risk-Return Relation: A Factor Analysis
Approach. Journal of Financial Economics 83, 171-222.
[100] Makridakis, S. and M. Hibon, 2000, The M3-Competition: Results, Conclusions and Impli-
cations. International Journal of Forecasting 16 451-476.
[101] Mamaysky, H., M. Spiegel and H. Zhang, 2006, Improved Forecasting of Mutual Fund Alphas
and Betas. Mimeo, Yale University.
57
[102] Marcellino, M., 2004, Forecast pooling for short time series of macroeconomic variables,
Oxford Bulletin of Economic and Statistics 66:91-112.
[103] Marcellino, M., J.H. Stock and M.W. Watson, 2006, A comparison of Direct and Iterated
Multistep AR Methods for Forecasting Macroeconomic Time Series. Journal of Econometrics
135, 499-526.
[104] Meese, R.A. and K. Rogoff, 1983, Empirical exchange rate models of the seventies : Do they
fit out of sample? Journal of International Economics 14, 3-24.
[105] Miller, A. 2002, Subset Selection in Regression, Second Edition, Chapman and Hall, Boca
Raton, Florida.
[106] Mincer, J. and V. Zarnowitz, 1969, The Evaluation of Economic Forecasts. In J. Mincer, ed.,
Economic Forecasts and Expectations. National Bureau of Economic Research, New York.
[107] Mishkin, F.S., 1981, Are Markets Forecasts Rational? American Economic Review 71, 295-
306.
[108] Nelson, C. and C. Plosser, 1982, Trends and Random Walks in Macroeconomic Time Series:
Some Evidence and Implications. Journal of Monetary Economics 10, 139-162.
[109] Newey, W. and J. Powell, 1987, Asymmetric Least Squares Estimation and Testing. Econo-
metrica 55, 819-847.
[110] Ottaviani, M. and P.N. Sorensen, 2006, The Strategy of Professional Forecasting. Journal of
Financial Economics 81, 441-466.
[111] Pagan, A., 2003, Report on Modelling and Forecasting at the Bank of England. Bank of
England.
[112] Palm, F. C. and A. Zellner, 1992, To combine or not to combine? Issues of combining
forecasts, Journal of Forecasting 11, 687-701.
[113] Patton, A. and A. Timmermann, 2006, Testing Forecast Optimality Under Unknown Loss.
Forthcoming in Journal of American Statistical Association.
[114] Patton, A. and A. Timmermann, 2007, Properties of Optimal Forecasts under Asymmetric
Loss and Nonlinearity. Journal of Econometrics 140, 884-918.
[115] Paye, B. and A. Timmermann, 2006, Instability of Return Prediction Models. Journal of
Empirical Finance 13 (3), 274-315.
[116] Perez-Quiros, G. and A. Timmermann, 2000, Firm Size and Cyclical Variations in Stock
Returns. Journal of Finance, 1229-1262.
[117] Pesaran, M.H., D. Pettenuzzo and A. Timmermann, 2006, Forecasting Time Series Subject
to Multiple Structural Breaks. Review of Economic Studies 73, 1057-1084.
58
[118] Pesaran, M.H. and S. Skouras, 2002, Decision-based Methods for Forecast Evaluation. In
Clements, M.P. and D. F. Hendry (Eds.), A Companion to Economic Forecasting. Blackwell,
Oxford.
[119] Pesaran, M.H. and A. Timmermann, 2005a, Small Sample Properties of Forecasts from Au-
toregressive Models under Structural Breaks. Journal of Econometrics 129, 183-217.
[120] Pesaran, M.H. and A. Timmermann, 2005b, Real time Econometrics. Econometric Theory
11, 212-231
[121] Pesaran, M.H. and A. Timmermann, 2006, Selection of Estimation Window in the Presence
of Breaks. Forthcoming in Journal of Econometrics.
[122] Pesaran, M.H. and M. Weale, 2006, Survey Expectations. Pages 715-776 in the Handbook
of Economic Forecasting, G. Elliott, C.W.J. Granger, and [Link] (eds.), North-
Holland.
[124] Racine, J., 2001, On the Nonlinear Predictability of Stock Returns using Financial and Eco-
nomic Variables. Journal of Business and Economic Statistics 19, 380-382.
[125] Raftery, A.E., D. Madigan and J.A. Hoeting, 1997, Bayesian model averaging for linear
regression models, Journal of the American Statistical Association 92, 179-191.
[126] Rapach, D. and M. Wohar, 2006, Structural Breaks and Predictive Regression Models of
Aggregate US Stock Returns. Journal of Financial Econometrics 4(2), 238-274.
[127] Robertson, J. and E. Tallman, 1999, Vector Autoregressions: Forecasting and Reality, Federal
Reserve Bank of Atlanta Economic Review, First Quarter.
[128] Romer, C.D. and D.H. Romer, 2000, Federal Reserve Information and the Behavior of Interest
Rates. American Economic Review 90(3), 429-457.
[129] Rossi, B., 2006, Are Exchange Rates Really Random Walks? Some Evidence Robust to
Parameter Instability. Macroeconomic Dynamics 10, 20-38.
[130] Satchell, S. and A. Timmermann, 1995, An Assessment of the Economic Value of Nonlinear
Foreign Exchange Rate Forecasts. Journal of Forecasting 14(6), 477-498.
[131] Scharfstein, D. and J. Stein, 1990, Herd Behavior and Investment. American Economic Review
80, 464-479.
[132] Schorfheide, F., 2005, VAR Forecasting under Misspecification. Journal of Econometrics 128,
2005, 99-136
[133] Siliverstovs, B., T. Engsted, and N. Haldrup, 2004, Long-run Forecasting in Multi-
Cointegrated Systems. Journal of Forecasting 23, 315-335.
59
[134] Sims, C.A., 1980, Macroeconomics and Reality. Econometrica 48, 1-48.
[135] Sims, C.A., 2002, The Role of Models and Probabilities in the Monetary Policy Process.
Mimeo, Princeton University.
[137] Stock, J.H. and M.W. Watson, 1996, Evidence on structural instability in macroeconomic
time series relations. Journal of Business and Economic Statistics 14, 11-30.
[138] Stock, J.H. and M.W. Watson, 1998, Median Unbiased Estimation of Coefficient Variance
in a Time Varying Parameter Model, Journal of the American Statistical Association, 93,
349-358.
[139] Stock, J.H. and M.W. Watson, 1999a, A Comparison of Linear and Nonlinear Models for
Forecasting Macroeconomic Time Series. In R. Engle and H. White (eds.), Cointegration,
Causality and Forecasting: A Festschrift in Honour of Clive W.J. Granger. Oxford University
Press.
[140] Stock, J.H. and M.W. Watson, 1999b, Forecasting Inflation. Journal of Monetary Economics
44, 293-335.
[141] Stock, James H., and Mark W. Watson, 2002, Macroeconomic Forecasting Using Diffusion
Indexes. Journal of Business and Economic Statistics 20:147-162.
[142] Stock, J.H. and M.W. Watson. 2005, An Empirical Comparison of Methods for Forecasting
Using Many Predictors. Mimeo, Harvard and Princeton University.
[143] Sullivan, R., A. Timmermann and H. White, 1999, Data-Snooping, Technical Trading Rules
and the Bootstrap. Journal of Finance 54, 1647-1692.
[144] Svensson, L.E.O., 1997, Inflation Forecast Targeting: Implementing and Monitoring Inflation
Targets. European Economic Review 41, 1111-1146.
[145] Swanson, N. and H. White, 1995, A Model Selection Approach to Assessing the Information in
the Term Structure using Linear Models and Artificial Neural Networks. Journal of Business
and Economic Statistics 13, 265-276.
[146] Tay, A.S. and K.F. Wallis, 2000, Density Forecasting: A Survey. Journal of Forecasting 19,
235-254.
[147] Taylor, M.P. and Sarno, L., 2002, Purchasing Power Parity and the Real Exchange Rate.
International Monetary Fund Staff Papers 49, 65-105.
[148] Terasvirta, T., 2006. Forecasting Economic Variables with Nonlinear Models. Pages 423-458
in G. Elliott, C.W.J. Granger, A. Timmermann, eds. Handbook of Economic Forecasting.
North-Holland: Amsterdam.
60
[149] Terasvirta, T., van Dijk, D., Medeiros, M.C., 2005, Smooth Transition Autoregressions,
Neural Networks, and Linear Models in Forecasting Macroeconomic Time Series: A Re-
examination. International Journal of Forecasting 21, 755-774.
[150] Tibshirani, R., 1996, Regression Shrinkage and Selection via the Lasso. Journal of the Royal
Statistical Society B 58, 267-288.
[151] Timmermann, A. 2006. Forecast Combinations. Pages 135-196 in G. Elliott, C.W.J. Granger,
A. Timmermann, eds. Handbook of Economic Forecasting. North-Holland: Amsterdam.
[152] Timmermann, A., 2007, An Evaluation of the World Economic Outlook Forecasts. Forthcom-
ing in IMF Staff Papers.
[153] Truman, B., 1994, Analyst Forecasts and Herding Behavior, Review of Financial Studies 7,
97-124.
[154] van Dijk, D., B. Strikholm and T. Teräsvirta, 2003, The effects of institutional and tech-
nological change and business cycle fluctuations on seasonal patterns in quarterly industrial
production series, Econometrics Journal 6, 79-98.
[155] Varian, H. R., 1974, A Bayesian Approach to Real Estate Assessment. In Studies in Bayesian
Econometrics and Statistics in Honor of Leonard J. Savage, eds. S.E. Fienberg and A. Zellner,
Amsterdam: North Holland, 195-208.
[156] Waggoner, D. and T. Zha 1999, Conditional Forecasts in Dynamic Multivariate Models,
Review of Economics and Statistics, 81, 639-651.
[157] Weiss, A.A., 1996, Estimating Time Series Models Using the Relevant Cost Function. Journal
of Applied Econometrics 11, 539-560.
[158] West, K.D., 1996, Asymptotic Inference about Predictive Ability. Econometrica 64, 1067-84.
[159] West, K.D., H.J. Edison and D. Cho, 1993, A Utility-based Comparison of Some Models of
Exchange Rate Volatility. Journal of International Economics 35, 23-46.
[160] West, K.D. and M.W. McCracken, 1998, Regression-Based Tests of Predictive Ability, Inter-
national Economic Review 39, 817-840.
[161] West, M. and J. Harrison, 1997, Bayesian Forecasting and Dynamic Models, second edition,
Springer Series in Statistics, Springer Verlag: New York.
[162] White, H., 2000, A Reality Check for Data Snooping. Econometrica 68, 1097-1127.
[163] White, H. 2001, Asymptotic Theory for Econometricians, 2nd Edition, Academic Press: New
York.
[164] Whiteman, C.H, 1996, Bayesian Prediction under Asymmetric Linear Loss: Forecasting State
Tax Revenues in Iowa. In W.O. Johnson, J.C. Lee and A. Zellner (eds.) Forecasting, Prediction
and Modeling in Statistics and Econometrics: Bayesian and non-Bayesian Approaches. New
York: Springer-Verlag.
61
[165] Zarnowitz, V., 1985, Rational Expectations and Macroeconomic Forecasts. Journal of Busi-
ness and Economic Statistics 3, 293-311.
[166] Zellner, A., 1986, Bayesian Estimation and Prediction Using Asymmetric Loss Functions.
Journal of the American Statistical Association, 81, 446-451.
[167] Zellner, A., and C. Hong, 1989, Forecasting International Growth Rates Using Bayesian
Shrinkage and other Procedures. Journal of Econometrics 40, 183-202.
62
Table 1: Out-of-sample Forecasting performance (annualized root mean squared error) for various
forecasting models, 1970 - 2003
Inflation SP500 Return
ModelName Expanding 10-year rolling Expanding 10-year rolling
window window window window
Autoregressive (AR) 0.78 0.77 15.9 15.9
Factor-augmented AR 0.78 0.80 15.9 15.9
BVAR - random walk prior 0.81 0.92 17.7 20.0
BVAR - white noise prior 0.81 0.94 17.4 19.1
Exponential smoothing 0.76 0.76 16.0 16.3
Double exp. smoothing 0.78 0.83 16.2 18.6
STAR 1 0.88 0.80 16.8 17.5
STAR 2 0.83 0.81 17.0 17.4
One Layer neural net 0.80 0.82 17.1 17.4
Two Layer neural net 0.78 0.77 16.0 17.5
BVAR - random walk prior 0.226 0.000 0.027 0.003 0.008 0.047 0.001 0.000 0.001
BVAR - white noise prior 0.000 0.013 0.001 0.003 0.021 0.000 0.000 0.000
Stock Returns
Autoregressive (AR) 0.937 0.000 0.000 0.073 0.000 0.010 0.001 0.009 0.060 0.150 0.177
Factor-augmented AR 0.000 0.000 0.050 0.000 0.004 0.000 0.004 0.047 0.054 0.155
BVAR - random walk prior 0.009 0.000 0.065 0.001 0.000 0.000 0.007 0.000 0.000
BVAR - white noise prior 0.000 0.528 0.048 0.022 0.034 0.100 0.000 0.002
Exponential smoothing 0.000 0.015 0.005 0.024 0.120 0.063 0.509
Double exp. smoothing 0.059 0.072 0.038 0.185 0.000 0.000
STAR 1 0.928 0.901 0.968 0.009 0.008
STAR 2 0.982 0.922 0.005 0.038
One Layer neural net 0.904 0.007 0.017
Two Layer neural net 0.098 0.176
-5
0
5
10
15
20
Ja 70
n-
Ja 71
n-
Ja 72
n-
Ja 73
n-
Ja 74
n-
Ja 75
n-
Ja 76
n-
Ja 77
n-
Ja 78
n-
Ja 79
n-
Ja 80
n-
Ja 81
n-
Ja 82
n-
Ja 83
n-
ExponentialSmoothing
Ja 84
n-
Ja 85
n-
Ja 86
n-
Ja 87
n-
Ja 88
n-
Ja 89
n-
Ja 90
FactorAugmentedAR
n-
Figure 1: Inflation Forecasts
Ja 91
n-
Ja 92
n-
Ja 93
n-
Ja 94
n-
Ja 95
n-
Ja 96
n-
TwoLayerNeuralNet
Ja 97
n-
Ja 98
n-
Ja 99
n-
Ja 00
n-
Ja 01
n-
Ja 02
n-
03
Ja
n-
-0.1
-0.08
-0.06
-0.04
-0.02
0
0.02
0.04
0.06
0.08
Ja 70
n-
Ja 71
n-
Ja 72
n-
Ja 73
n-
Ja 74
n-
Ja 75
n-
Ja 76
n-
Ja 77
n-
Ja 78
n-
Ja 79
n-
Ja 80
n-
Ja 81
n-
Ja 82
n-
Ja 83
n-
ExponentialSmoothing
Ja 84
n-
Ja 85
n-
Ja 86
n-
Ja 87
n-
Ja 88
n-
Ja 89
n-
Ja 90
TwoLayerNeuralNet
n-
Ja 91
n-
Figure 2: Forecasts of Stock Returns
Ja 92
n-
Ja 93
n-
Ja 94
n-
Ja 95
n-
Ja 96
n-
FactorAugmentedAR
Ja 97
n-
Ja 98
n-
Ja 99
n-
Ja 00
n-
Ja 01
n-
Ja 02
n-
03
Industry Life Cycle
Life cycle models are not just a phenomenon of the life sciences. Industries experience a similar
cycle of life. Just as a person is born, grows, matures, and eventually experiences decline and
ultimately death, so too do industries and product lines. The stages are the same for all
industries, yet every industry will experience these stages differently, they will last longer for some
and pass quickly for others. Even within the same industry, various firms may be at different life
cycle stages. A firm’s strategic plan is likely to be greatly influenced by the stage in the life cycle at
which the firm finds itself. Some companies or even industries find new uses for declining
products, thus extending their life cycle.
The growth of an industry's sales over time is used to chart the life cycle. The distinct stages of an
industry life cycle are: introduction, growth, maturity, and decline. Sales typically begin slowly at
the introduction phase, then take off rapidly during the growth phase. After levelling out at
maturity, sales then begin a gradual decline. In contrast, profits generally continue to increase
throughout the life cycle, as companies in an industry take advantage of expertise and economies
of scale and scope to reduce unit costs over time.
Introduction
In the introduction stage of the life cycle, an industry is in its infancy. Perhaps a new, unique
product offering has been developed and patented, thus beginning a new industry. Some analysts
even add an embryonic stage before introduction. At the introduction stage, the firm may be alone
in the industry. It may be a small entrepreneurial company or a proven company which used
research and development funds and expertise to develop something new. Marketing refers to
new product offerings in a new industry as "question marks" because the success of the product
and the life of the industry is unproven and unknown.
A firm will use a focused strategy at this stage to stress the uniqueness of the new product or
service to a small group of customers. These customers are typically referred to in the marketing
literature as the "innovators" and "early adopters." Marketing tactics during this stage are
intended to explain the product and its uses to consumers and thus create awareness for the
product and the industry. According to research by Hitt, Ireland, and Hoskisson, firms establish a
niche for dominance within an industry during this phase. For example, they often attempt to
establish early perceptions of product quality, technological superiority, or advantageous
relationships with vendors within the supply chain to develop a competitive advantage. Because it
costs money to create a new product offering, develop and test prototypes, and market the
product, the firm's and the industry’s profits are usually negative at this stage.
Any profits generated are typically reinvested into the company to solidify its position and help
fund continued growth. Introduction requires a significant cash outlay to continue to promote and
differentiate the offering and expand the production flow from a job shop to possibly a batch flow.
Market demand will grow from the introduction, and as the life cycle curve experiences growth at
an increasing rate, the industry is said to be entering the growth stage. Firms may also cluster
together in close proximity during the early stages of the industry life cycle to have access to key
materials or technological expertise, as in the case of the U.S. Silicon Valley computer chip
manufacturers.
Growth
Like the introduction stage, the growth stage also requires a significant amount of capital. The
goal of marketing efforts at this stage is to differentiate a firm's offerings from other competitors
within the industry. Thus the growth stage requires funds to launch a newly focused marketing
campaign as well as funds for continued investment in property, plant, and equipment to facilitate
the growth required by the market demands. However, the industry is experiencing more product
[Link]
standardization at this stage, which may encourage economies of scale and facilitate
development of a line-flow layout for production efficiency.
Research and development funds will be needed to make changes to the product or services to
better reflect customers' needs and suggestions. In this stage, if the firm is successful in the
market, growing demand will create sales growth. Earnings and accompanying assets will also
grow and profits will be positive for the firms. Marketing often refers to products at the growth
stage as "stars." These products have high growth and market share. The key issue in this stage is
market rivalry. Because there is industry-wide acceptance of the product, more new entrants join
the industry and more intense competition results.
The duration of the growth stage, as all the other stages, depends on the particular industry or
product line under study. Some items—like fad clothing, for example—may experience a very short
growth stage and move almost immediately into the next stages of maturity and decline. A hot toy
this holiday season may be non-existent or relegated to the back shelves of a deep-discounter the
following year. Because many new product introductions fail, the growth stage may be short or
non-existent for some products. However, for other products the growth stage may be longer due
to frequent product upgrades and enhancements that forestall movement into maturity. The
computer industry today is an example of an industry with a long growth stage due to upgrades in
hardware, services, and add-on products and features.
During the growth stage, the life cycle curve is very steep, indicating fast growth. Firms tend to
spread out geographically during this stage of the life cycle and continue to disperse during the
maturity and decline stages. As an example, the automobile industry in the United States was
initially concentrated in the Detroit area and surrounding cities. Today, as the industry has
matured, automobile manufacturers are spread throughout the country and internationally.
Maturity
As the industry approaches maturity, the industry life cycle curve becomes noticeably flatter,
indicating slowing growth. Some experts have labelled an additional stage, called expansion,
between growth and maturity. While sales are expanding and earnings are growing from these
"cash cow" products, the rate has slowed from the growth stage. In fact, the rate of sales
expansion is typically equal to the growth rate of the economy.
Some competition from late entrants will be apparent, and these new entrants will try to steal
market share from existing products. Thus, the marketing effort must remain strong and must
stress the unique features of the product or the firm to continue to differentiate a firm's offerings
from industry competitors. Firms may compete on quality to separate their product from other
lower-cost offerings, or conversely the firm may try a low-cost/low-price strategy to increase the
volume of sales and make profits from inventory turnover. A firm at this stage may have excess
cash to pay dividends to shareholders. But in mature industries, there are usually fewer firms, and
those that survive will be larger and more dominant.
While innovations continue they are not as radical as before and may be only a change in colour
or formulation to stress "new" or "improved" to consumers. Laundry detergents are examples of
mature products.
Decline
Declines are almost inevitable in an industry. If product innovation has not kept pace with other
competing products and/or service, or if new innovations or technological changes have caused
the industry to become obsolete, sales suffer and the life cycle experiences a decline. In this
phase, sales are decreasing at an accelerating rate. This is often accompanied by another, larger
shake-out in the industry as competitors who did not leave during the maturity stage now exit the
industry. Yet some firms will remain to compete in the smaller market. Mergers and
[Link]
consolidations will also be the norm as firms try other strategies to continue to be competitive or
grow through acquisition and/or diversification.
Management efficiency can help to prolong the maturity stage of the life cycle. Production
improvements, like just-in-time methods and lean manufacturing, can result in extra profits.
Technology, automation, and linking suppliers and customers in a tight supply chain are also
methods to improve efficiency.
New uses of a product can also revitalize an old brand. A prime example is Arm & Hammer baking
soda. In 1969, sales were dropping due to the introduction of packaged foods with baking soda
as an added ingredient and an overall decline in home baking. New uses for the product as a
deodorizer for refrigerators and later as a laundry additive, toothpaste additive, and carpet
freshener extended the life cycle of the baking soda industry.
Promoting new uses for old brands can increase sales by increasing usage frequency. In some
cases, this strategy is cheaper than trying to convert new users in a mature market.
To extend the growth phase as well as industry profits, firms approaching maturity can pursue
expansion into other countries and new markets. Expansion into another geographic region is an
effective response to declining demand. Because organizations have control over internal factors
and can often influence external factors, the life cycle does not have to end.
An example is feminine hygiene products. Sales in the United States have reached maturity due to
a number of external reasons, like the stable to declining population growth rate and the aging of
the baby boomers, who may no longer be consumers for these products. But when makers of
these products concentrated on foreign markets, sales grew and the maturity of the product was
prolonged. Often so-called "dog" products can find new life in other parts of the world. However,
once world saturation is reached, the eventual maturity and decline of the industry or product line
will result.
Just as industries experience life cycles, studies have documented life cycles in many other areas.
Countries have life cycles, for example, and we traditionally classify them as ranging from the First
World countries to Third World or developing countries, depending on their levels of capital,
technological change, infrastructure, or stability. Products also experience life cycles. Even within
an industry, various individual companies may be at different life cycle stages depending upon
when they entered the industry. The life cycle phenomenon is an important and universally
accepted concept to help managers better understand sales growth and change over time.
BIBLIOGRAPHY
Hitt, Michael A., R. Duane Ireland, and Robert E. Hoskisson. Strategic Management:
Competitiveness and Globalization Fourth Edition. South-Western College Publishing, 2001.
Wang, Zhu. "Learning, Diffusion, and Industry Life Cycle." Federal Reserve Bank of Kansas
City, Working Paper 04-01 Available from
[Link]/PUBLICAT/PSR/RWP/[Link] 15 January 2006.
Wansink, Brian, and Jennifer Marie Gilmore. "New Uses that Revitalize Old Brands." Journal of
Advertising Research. March 1999
[Link]
Across the Disciplines
Why This Chapter Matters To You
Accounting: You need to understand
interest rates and the various types of
bonds in order to be able to account prop-
erly for amortization of bond premiums
and discounts and for bond purchases and
LG2
Describe interest rate fundamentals, Operations: You need to understand how
the term structure of interest rates, the interest rate level may affect the firm’s
and risk premiums. ability to raise funds to maintain and
Review the legal aspects of bond increase the firm’s production capacity.
LG2
financing and bond cost.
228
229
1. These assumptions are made to describe the most basic interest rate, the real rate of interest. Subsequent discus-
sions relax these assumptions to develop the broader concept of the interest rate and required return.
230 PART 2 Important Financial Concepts
FIGURE 6.1
Supply–Demand D
k*1
S0
S1
D
S0 = D S1 = D
Funds Supplied/Demanded
money. The real rate of interest in the United States is assumed to be stable and
equal to around 1 percent.2 This supply–demand relationship is shown in Figure
6.1 by the supply function (labeled S0) and the demand function (labeled D). An
equilibrium between the supply of funds and the demand for funds (S0 ? D)
occurs at a rate of interest k0*, the real rate of interest.
Clearly, the real rate of interest changes with changing economic conditions,
tastes, and preferences. A trade surplus could result in an increased supply of
funds, causing the supply function in Figure 6.1 to shift to, say, S1. This could
result in a lower real rate of interest, k1*, at equilibrium (S1 ? D). Likewise, a
change in tax laws or other factors could affect the demand for funds, causing the
real rate of interest to rise or fall to a new equilibrium level.
2. Data in Stocks, Bonds, Bills and Inflation, 2001 Yearbook (Chicago: Ibbotson Associates, Inc., 2001), show that
over the period 1926–2000, U.S. Treasury bills provided an average annual real rate of return of about 0.7 percent.
Because of certain major economic events that occurred during the 1926–2000 period, many economists believe that
the real rate of interest during recent years has been about 1 percent.
231
FIGURE 6.2
15
Impact of Inflation
Relationship between annual
Inflationb
term structure
of interest rates
The relationship between the
interest rate or rate of return and Term Structure of Interest Rates
the time to maturity.
For any class of similar-risk securities, the term structure of interest rates relates
yield to maturity the interest rate or rate of return to the time to maturity. For convenience we will
Annual rate of return earned on a use Treasury securities as an example, but other classes could include securities
debt security purchased on a
that have similar overall quality or risk. The riskless nature of Treasury securities
given day and held to maturity.
also provides a laboratory in which to develop the term structure.
yield curve
A graph of the relationship
between the debt’s remaining
time to maturity (x axis) and its Yield Curves
yield to maturity (y axis); it
A debt security’s yield to maturity (discussed later in this chapter) represents the
shows the pattern of annual
returns on debts of equal quality annual rate of return earned on a security purchased on a given day and held to
and different maturities. maturity. At any point in time, the relationship between the debt’s remaining
Graphically depicts the term time to maturity and its yield to maturity is represented by the yield curve. The
structure of interest rates. yield curve shows the yield to maturity for debts of equal quality and different
inverted yield curve maturities; it is a graphical depiction of the term structure of interest rates. Fig-
A downward-sloping yield curve ure 6.3 shows three yield curves for all U.S. Treasury securities: one at May 22,
that indicates generally cheaper 1981, a second at September 29, 1989, and a third at March 15, 2002. Note
long-term borrowing costs than that both the position and the shape of the yield curves change over time. The
short-term borrowing costs.
yield curve of May 22, 1981, indicates that short-term interest rates at that time
normal yield curve were above longer-term rates. This curve is described as downward-sloping,
An upward-sloping yield curve reflecting long-term borrowing costs generally cheaper than short-term borrow-
that indicates generally cheaper
ing costs. Historically, the downward-sloping yield curve, which is often called
short-term borrowing costs than
long-term borrowing costs. an inverted yield curve, has been the exception. More frequently, yield curves
similar to that of March 15, 2002, have existed. These upward-sloping or
flat yield curve
normal yield curves indicate that short-term borrowing costs are below long-
A yield curve that reflects
relatively similar borrowing term borrowing costs. Sometimes, a flat yield curve, similar to that of September
costs for both short- and longer- 29, 1989, exists. It reflects relatively similar borrowing costs for both short- and
term loans. longer-term loans.
232 PART 2 Important Financial Concepts
FIGURE 6.3
The shape of the yield curve may affect the firm’s financing decisions. A
financial manager who faces a downward-sloping yield curve is likely to rely more
heavily on cheaper, long-term financing; when the yield curve is upward-sloping,
the manager is more likely to use cheaper, short-term financing. Although a vari-
ety of other factors also influence the choice of loan maturity, the shape of the
yield curve provides useful insights into future interest rate expectations.
Investors (lenders) tend to require a premium for tying up funds for longer
periods, whereas borrowers are generally willing to pay a premium to obtain
longer-term financing. These preferences of lenders and borrowers cause the yield
curve to tend to be upward-sloping. Simply stated, longer maturities tend to have
higher interest rates than shorter maturities.
market segmentation theory Market Segmentation Theory The market segmentation theory suggests
Theory suggesting that the that the market for loans is segmented on the basis of maturity and that the sup-
market for loans is segmented on
ply of and demand for loans within each segment determine its prevailing interest
the basis of maturity and that the
supply of and demand for loans rate. In other words, the equilibrium between suppliers and demanders of short-
within each segment determine term funds, such as seasonal business loans, would determine prevailing short-
its prevailing interest rate; the term interest rates, and the equilibrium between suppliers and demanders of
slope of the yield curve is long-term funds, such as real estate loans, would determine prevailing long-term
determined by the general
interest rates. The slope of the yield curve would be determined by the general
relationship between the prevail-
ing rates in each segment. relationship between the prevailing rates in each market segment. Simply stated,
low rates in the short-term segment and high rates in the long-term segment cause
the yield curve to be upward-sloping. The opposite occurs for high short-term
rates and low long-term rates.
All three theories of term structure have merit. From them we can conclude
that at any time, the slope of the yield curve is affected by (1) inflationary expec-
tations, (2) liquidity preferences, and (3) the comparative equilibrium of supply
and demand in the short- and long-term market segments. Upward-sloping yield
curves result from higher future inflation expectations, lender preferences for
shorter-maturity loans, and greater supply of short-term loans than of long-term
loans relative to demand. The opposite behaviors would result in a downward-
sloping yield curve. At any time, the interaction of these three forces determines
the prevailing slope of the yield curve.
3. Later in this chapter we demonstrate that debt instruments with longer maturities are more sensitive to changing
market interest rates. For a given change in market rates, the price or value of longer-term debts will be more signif-
icantly changed (up or down) than the price or value of debts with shorter maturities.
234 PART 2 Important Financial Concepts
Component Description
Default risk The possibility that the issuer of debt will not pay the contrac-
tual interest or principal as scheduled. The greater the uncer-
tainty as to the borrower’s ability to meet these payments, the
greater the risk premium. High bond ratings reflect low
default risk, and low bond ratings reflect high default risk.
Maturity risk The fact that the longer the maturity, the more the value of a
security will change in response to a given change in interest
rates. If interest rates on otherwise similar-risk securities sud-
denly rise as a result of a change in the money supply, the
prices of long-term bonds will decline by more than the prices
of short-term bonds, and vice versa.a
Contractual provision risk Conditions that are often included in a debt agreement or a
stock issue. Some of these reduce risk, whereas others may
increase risk. For example, a provision allowing a bond issuer
to retire its bonds prior to their maturity under favorable
terms increases the bond’s risk.
aA detailed discussion of the effects of interest rates on the price or value of bonds and other fixed-income
securities is presented later in this chapter.
Review Questions
6–1 What is the real rate of interest? Differentiate it from the risk-free rate of
interest for a 3-month U.S. Treasury bill.
6–2 What is the term structure of interest rates, and how is it related to the
yield curve?
6–3 For a given class of similar-risk securities, what does each of the following
yield curves reflect about interest rates: (a) downward-sloping; (b) upward-
sloping; and (c) flat? Which form has been historically dominant?
6–4 Briefly describe the following theories of the general shape of the yield
curve: (a) expectations theory; (b) liquidity preference theory; and (c) mar-
ket segmentation theory.
6–5 List and briefly describe the potential issuer- and issue-related risk compo-
nents that are embodied in the risk premium. Which are the purely debt-
specific risks?
235
Bond Indenture
bond indenture A bond indenture is a legal document that specifies both the rights of the bond-
A legal document that specifies holders and the duties of the issuing corporation. Included in the indenture are
both the rights of the bondhold-
descriptions of the amount and timing of all interest and principal payments, var-
ers and the duties of the issuing
corporation. ious standard and restrictive provisions, and, frequently, sinking-fund require-
ments and security interest provisions.
standard debt provisions Standard Provisions The standard debt provisions in the bond indenture
Provisions in a bond indenture specify certain record-keeping and general business practices that the bond issuer
specifying certain record- must follow. Standard debt provisions do not normally place a burden on a
keeping and general business
practices that the bond issuer
financially sound business.
must follow; normally, they do The borrower commonly must (1) maintain satisfactory accounting records
not place a burden on a in accordance with generally accepted accounting principles (GAAP); (2) periodi-
financially sound business. cally supply audited financial statements; (3) pay taxes and other liabilities when
due; and (4) maintain all facilities in good working order.
subordination Subordination means that subsequent creditors agree to wait until all claims
In a bond indenture, the stipula- of the senior debt are satisfied.
tion that subsequent creditors 5. Limit the firm’s annual cash dividend payments to a specified percentage or
agree to wait until all claims of
the senior debt are satisfied.
amount.
Trustee
trustee A trustee is a third party to a bond indenture. The trustee can be an individual, a
A paid individual, corporation, or corporation, or (most often) a commercial bank trust department. The trustee is
commercial bank trust depart- paid to act as a “watchdog” on behalf of the bondholders and can take specified
ment that acts as the third party
to a bond indenture and can take
actions on behalf of the bondholders if the terms of the indenture are violated.
specified actions on behalf of the
bondholders if the terms of the
indenture are violated. Cost of Bonds to the Issuer
The cost of bond financing is generally greater than the issuer would have to pay
for short-term borrowing. The major factors that affect the cost, which is the rate
of interest paid by the bond issuer, are the bond’s maturity, the size of the offer-
ing, the issuer’s risk, and the basic cost of money.
In Practice
FOCUS ON PRACTICE Ford Cruises the Debt Markets
Ford and Ford Motor Credit Co. maturities remained attractively one rating class. The lower ratings
(FMCC), its finance unit, were fre- low for corporations. Unlike some contributed to the higher yields on
quent visitors to the corporate other auto companies who limited Ford’s October debt. For example,
debt markets in 2001, selling over the size of their debt offerings, in April FMCC’s 10-year notes
$22 billion in long-term notes and FMCC decided to borrow as much yielded 7.1 percent, about 2 points
bonds. Despite the problems in as possible to lock in the very wide above U.S. Treasury bonds. In
the auto industry, investors ner- spread between its lower borrow- October, 10-year FMCC notes
vous about stock market volatility ing costs and what its auto loans yielded 7.3 percent, or 2.7 points
were willing to accept the credit yielded. above U.S. Treasury bonds.
risk to get higher yields. The com- All this debt came at a price, For corporations like Ford,
pany’s 2001 offerings had some- however. Both major bond-rating deciding when to issue debt and
thing for all types of investors, agencies—Moody’s Investors selecting the best maturities
ranging from 2- to 10-year notes Service and Standard & Poor’s requires knowledge of interest
to 30-year bonds. Demand for (S&P)—downgraded Ford’s debt rate fundamentals, risk premiums,
Ford’s debt was so high that in quality ratings in October 2001. issuance costs, ratings, and simi-
January the company increased Moody’s lowered Ford’s long-term lar features of corporate bonds.
the size of its issue from $5 billion debt rating by one rating class but
to $7.8 billion, and October’s plan did not change FMCC’s quality rat- Sources: Adapted from Jonathan Stempel,
to issue $3 billion turned into a ing. Ford spokesman Todd Nissen “‘Buy My Product, Buy My Bonds,’ [Link]-
$9.4 billion offering. was pleased that Moody’s con- panies Say,” Reuters, April 10, 2001, “Ford
Sells $9.4 Bln Bonds, Offers Big Yields,
The world’s second largest firmed the FMCC ratings. “It will Reuters, October 22, 2001, and “Moodys Cuts
auto manufacturer joined other help us keep our costs of borrow- Ford, but Not Ford Credit, Ratings,” Reuters
Business Report, October 18, 2001, all down-
corporate bond issuers to take ing down, which benefits Ford loaded from eLibrary, [Link]; Ed
advantage of strengthening bond Credit and ultimately Ford Motor,” Zwirn, “Ford to Issue $7.8 Billion and Count-
markets. Even though the Federal he said. S&P’s outlook for Ford ing,” [Link], January 24, 2001, and “Full
Speed Ahead for Auto Bonds,” [Link],
Reserve began cutting short-term was more negative; the agency cut January 19, 2001, both downloaded from
rates, interest rates for the longer ratings on all Ford and FMCC debt www. [Link].
the bondholders may increase, because larger offerings result in greater risk of
default.
FIGURE 6.4
Bond Quotations
Selected bond quotations for
April 22, 2002
IBM
the business sections of daily general newspapers. Here we focus on bond quota-
tions; stock quotations are reviewed in Chapter 7.
Figure 6.4 includes an excerpt from the New York Stock Exchange (NYSE)
bond quotations reported in the April 23, 2002, Wall Street Journal for trans-
actions through the close of trading on Monday, April 22, 2002. We’ll look at the
corporate bond quotation for IBM, which is highlighted in Figure 6.4. The
numbers following the company name—IBM—represent the bond’s coupon inter-
est rate and the year it matures: “7s25” means that the bond has a stated coupon
interest rate of 7 percent and matures sometime in the year 2025. This information
allows investors to differentiate between the various bonds issued by the corpora-
tion. Note that on the day of this quote, IBM had four bonds listed. The next col-
umn, labeled “Cur Yld.,” gives the bond’s current yield, which is found by dividing
its annual coupon (7%, or 7.000%) by its closing price (100.25), which in this case
turns out to be 7.0 percent (7.000 ? 100.25 ? 0.0698 ? 7.0%).
The “Vol” column indicates the actual number of bonds that traded on the
given day; 10 IBM bonds traded on Monday, April 22, 2002. The final two
columns include price information—the closing price and the net change in clos-
ing price from the prior trading day. Although most corporate bonds are issued
240 PART 2 Important Financial Concepts
with a par, or face, value of $1,000, all bonds are quoted as a percentage of par.
A $1,000-par-value bond quoted at 110.38 is priced at $1,103.80 (110.38% ?
$1,000). Corporate bonds are quoted in dollars and cents. Thus IBM’s closing
price of 100.25 for the day was $1,002.50—that is, 100.25% ? $1,000. Because
a “Net Chg.” of ?1.75 is given in the final column, the bond must have closed at
102 or $1,020 (102.00% ? $1,000) on the prior day. Its price decreased by 1.75,
or $17.50 (1.75% ? $1,000), on Tuesday, April 22, 2002. Additional informa-
tion may be included in a bond quotation, but these are the basic elements.
Bond Ratings
Independent agencies such as Moody’s and Standard & Poor’s assess the riski-
ness of publicly traded bond issues. These agencies derive the ratings by using
financial ratio and cash flow analyses to assess the likely payment of bond inter-
est and principal. Table 6.2 summarizes these ratings. Normally an inverse rela-
tionship exists between the quality of a bond and the rate of return that it must
provide bondholders: High-quality (high-rated) bonds provide lower returns
than lower-quality (low-rated) bonds. This reflects the lender’s risk-return trade-
off. When considering bond financing, the financial manager must be concerned
with the expected ratings of the bond issue, because these ratings affect salability
and cost.
Standard
Moody’s Interpretation & Poor’s Interpretation
Unsecured Bonds
Debentures Unsecured bonds that only creditworthy firms Claims are the same as those of any general
can issue. Convertible bonds are normally creditor. May have other unsecured bonds
debentures. subordinated to them.
Subordinated Claims are not satisfied until those of the Claim is that of a general creditor but not as good
debentures creditors holding certain (senior) debts have been as a senior debt claim.
fully satisfied.
Income bonds Payment of interest is required only when Claim is that of a general creditor. Are not in
earnings are available. Commonly default when interest payments are missed,
issued in reorganization of a failing firm. because they are contingent only on earnings
being available.
Secured Bonds
Mortgage bonds Secured by real estate or buildings. Claim is on proceeds from sale of mortgaged
assets; if not fully satisfied, the lender becomes a
general [Link] first-mortgage claim must be
fully satisfied before distribution of proceeds to
second-mortgage holders, and so on. A number
of mortgages can be issued against the same
collateral.
Collateral trust Secured by stock and (or) bonds that are owned Claim is on proceeds from stock and (or) bond
bonds by the issuer. Collateral value is generally 25% to collateral; if not fully satisfied, the lender becomes
35% greater than bond value. a general creditor.
Equipment trust Used to finance “rolling stock”—airplanes, trucks, Claim is on proceeds from the sale of the asset; if
certificates boats, railroad cars. A trustee buys such an asset proceeds do not satisfy outstanding debt, trust
with funds raised through the sale of trust cer- certificate lenders become general creditors.
tificates and then leases it to the firm, which,
after making the final scheduled lease payment,
receives title to the asset. A type of leasing.
Zero- (or low-) Issued with no (zero) or a very low coupon (stated interest) rate and sold at a large discount from par. A
coupon bonds significant portion (or all) of the investor’s return comes from gain in value (i.e., par value minus purchase
price). Generally callable at par value. Because the issuer can annually deduct the current year’s interest
accrual without having to pay the interest until the bond matures (or is called), its cash flow each year is
increased by the amount of the tax shield provided by the interest deduction.
Junk bonds Debt rated Ba or lower by Moody’s or BB or lower by Standard & Poor’s. Commonly used during the 1980s
by rapidly growing firms to obtain growth capital, most often as a way to finance mergers and takeovers.
High-risk bonds with high yields—often yielding 2% to 3% more than the best-quality corporate debt.
Floating-rate Stated interest rate is adjusted periodically within stated limits in response to changes in specified money
bonds market or capital market rates. Popular when future inflation and interest rates are uncertain. Tend to sell
at close to par because of the automatic adjustment to changing market conditions. Some issues provide
for annual redemption at par at the option of the bondholder.
Extendible notes Short maturities, typically 1 to 5 years, that can be renewed for a similar period at the option of holders.
Similar to a floating-rate bond. An issue might be a series of 3-year renewable notes over a period of
15 years; every 3 years, the notes could be extended for another 3 years, at a new rate competitive with
market interest rates at the time of renewal.
Putable bonds Bonds that can be redeemed at par (typically, $1,000) at the option of their holder either at specific dates
after the date of issue and every 1 to 5 years thereafter or when and if the firm takes specified actions, such
as being acquired, acquiring another company, or issuing a large amount of additional debt. In return for
its conferring the right to “put the bond” at specified times or when the firm takes certain actions, the
bond’s yield is lower than that of a nonputable bond.
aTheclaims of lenders (i.e., bondholders) against issuers of each of these types of bonds vary, depending on the bonds’ other features. Each of these
bonds can be unsecured or secured.
Review Questions
Key Inputs
There are three key inputs to the valuation process: (1) cash flows (returns), (2)
timing, and (3) a measure of risk, which determines the required return. Each is
described below.
EXAMPLE Celia Sargent, financial analyst for Groton Corporation, a diversified holding
company, wishes to estimate the value of three of its assets: common stock in
Michaels Enterprises, an interest in an oil well, and an original painting by a well-
known artist. Her cash flow estimates for each are as follows:
Stock in Michaels Enterprises Expect to receive cash dividends of $300 per
year indefinitely.
Oil well Expect to receive cash flow of $2,000 at the end of year 1, $4,000 at
the end of year 2, and $10,000 at the end of year 4, when the well is to be sold.
Original painting Expect to be able to sell the painting in 5 years for
$85,000.
244 PART 2 Important Financial Concepts
With these cash flow estimates, Celia has taken the first step toward placing a
value on each of the assets.
Timing
In addition to making cash flow estimates, we must know the timing of the cash
flows.4 For example, Celia expects the cash flows of $2,000, $4,000, and $10,000
for the oil well to occur at the ends of years 1, 2, and 4, respectively. The combina-
tion of the cash flow and its timing fully defines the return expected from the asset.
EXAMPLE Let’s return to Celia Sargent’s task of placing a value on Groton Corporation’s
original painting and consider two scenarios.
Scenario 1—Certainty A major art gallery has contracted to buy the paint-
ing for $85,000 at the end of 5 years. Because this is considered a certain sit-
uation, Celia views this asset as “money in the bank.” She thus would use the
prevailing risk-free rate of 9% as the required return when calculating the
value of the painting.
Scenario 2—High Risk The values of original paintings by this artist have
fluctuated widely over the past 10 years. Although Celia expects to be able to
get $85,000 for the painting, she realizes that its sale price in 5 years could
range between $30,000 and $140,000. Because of the high uncertainty sur-
rounding the painting’s value, Celia believes that a 15% required return is
appropriate.
These two estimates of the appropriate required return illustrate how this
rate captures risk. The often subjective nature of such estimates is also clear.
4. Although cash flows can occur at any time during a year, for computational convenience as well as custom, we
will assume they occur at the end of the year unless otherwise noted.
245
EXAMPLE Celia Sargent used Equation 6.2 to calculate the value of each asset (using present
value interest factors from Table A–2), as shown in Table 6.5. Michaels
Enterprises stock has a value of $2,500, the oil well’s value is $9,262, and the
original painting has a value of $42,245. Note that regardless of the pattern of
the expected cash flow from an asset, the basic valuation equation can be used to
determine its value.
Review Questions
Bond Fundamentals
As noted earlier in this chapter, bonds are long-term debt instruments used by
business and government to raise large sums of money, typically from a diverse
group of lenders. Most corporate bonds pay interest semiannually (every 6
months) at a stated coupon interest rate, have an initial maturity of 10 to
30 years, and have a par value, or face value, of $1,000 that must be repaid at
maturity.
EXAMPLE Mills Company, a large defense contractor, on January 1, 2004, issued a 10%
coupon interest rate, 10-year bond with a $1,000 par value that pays interest
semiannually. Investors who buy this bond receive the contractual right to two
cash flows: (1) $100 annual interest (10% coupon interest rate ? $1,000 par
value) distributed as $50 (1/2 ? $100) at the end of each 6 months, and (2) the
$1,000 par value at the end of the tenth year.
We will use data for Mills’s bond issue to look at basic bond valuation.
where
B0 ? value of the bond at time zero
I ? annual interest paid in dollars5
n ? number of years to maturity
M ? par value in dollars
kd ? required return on a bond
We can calculate bond value using Equation 6.3a and the appropriate financial
tables (A–2 and A–4) or by using a financial calculator.
EXAMPLE Assuming that interest on the Mills Company bond issue is paid annually and
that the required return is equal to the bond’s coupon interest rate, I ? $100, kd ?
10%, M ? $1,000, and n ? 10 years.
The computations involved in finding the bond value are depicted graphi-
cally on the following time line.
386.00
B0 = $1,000.50
Table Use Substituting the values noted above into Equation 6.3a yields
B0 ? $100 ? (PVIFA10%,10yrs) ? $1,000 ? (PVIF10%,10yrs)
? $100 ? (6.145) ? $1,000 ? (0.386)
? $614.50 ? $386.00 ? $ ? 1 ,0 0 0.??
50
??????
????? ?? ?
?
The bond therefore has a value of approximately $1,000.6
5. The payment of annual rather than semiannual bond interest is assumed throughout the following discussion.
This assumption simplifies the calculations involved, while maintaining the conceptual accuracy of the valuation
procedures presented.
6. Note that a slight rounding error ($0.50) results here from the use of the table factors, which are rounded to the
nearest thousandth.
248 PART 2 Important Financial Concepts
Calculator Use Using the Mills Company’s inputs shown at the left, you should
Input Function
find the bond value to be exactly $1,000. Note that the calculated bond value is
10 N
equal to its par value; this will always be the case when the required return is
10 I
equal to the coupon interest rate.
100 PMT
1000 FV
CPT
PV Bond Value Behavior
Solution In practice, the value of a bond in the marketplace is rarely equal to its par value.
1000
In bond quotations (see Figure 6.4), the closing prices of bonds often differ from
their par values of 100 (100 percent of par). Some bonds are valued below par
(quoted below 100), and others are valued above par (quoted above 100). A vari-
ety of forces in the economy, as well as the passage of time, tend to affect value.
Although these external forces are in no way controlled by bond issuers or
investors, it is useful to understand the impact that required return and time to
maturity have on bond value.
EXAMPLE The preceding example showed that when the required return equaled the
coupon interest rate, the bond’s value equaled its $1,000 par value. If for the
same bond the required return were to rise or fall, its value would be found as fol-
lows (using Equation 6.3a):
Table Use
Required Return ? 12% Required Return ? 8%
Calculator Use Using the inputs shown on the next page for the two different
required returns, you will find the value of the bond to be below or above par. At
249
FIGURE 6.5
Bond Values and
1,400
Required Returns
Bond values and required
Market Value of Bond, B0 ($)
1,300
returns (Mills Company’s
10% coupon interest rate,
1,200
10-year maturity, $1,000 par,
1,134
January 1, 2004, issue paying 1,100
annual interest) Premium
Par 1,000
Discount
900
887
800
700
0 2 4 6 8 10 12 14 16
Required Return, kd (%)
250 PART 2 Important Financial Concepts
Constant Required Returns When the required return is different from the
coupon interest rate and is assumed to be constant until maturity, the value of the
bond will approach its par value as the passage of time moves the bond’s value
closer to maturity. (Of course, when the required return equals the coupon inter-
est rate, the bond’s value will remain at par until it matures.)
EXAMPLE Figure 6.6 depicts the behavior of the bond values calculated earlier and pre-
sented in Table 6.6 for Mills Company’s 10% coupon interest rate bond paying
annual interest and having 10 years to maturity. Each of the three required
returns—12%, 10%, and 8%—is assumed to remain constant over the 10 years
to the bond’s maturity. The bond’s value at both 12% and 8% approaches and
ultimately equals the bond’s $1,000 par value at its maturity, as the discount (at
12%) or premium (at 8%) declines with the passage of time.
Changing Required Returns The chance that interest rates will change and
interest rate risk thereby change the required return and bond value is called interest rate risk.
The chance that interest rates (This was described as a shareholder-specific risk in Chapter 5, Table 5.1.) Bond-
will change and thereby change holders are typically more concerned with rising interest rates because a rise in
the required return and bond
value. Rising rates, which result
interest rates, and therefore in the required return, causes a decrease in bond
in decreasing bond values, are of value. The shorter the amount of time until a bond’s maturity, the less responsive
greatest concern. is its market value to a given change in the required return. In other words, short
maturities have less interest rate risk than long maturities when all other features
(coupon interest rate, par value, and interest payment frequency) are the same.
FIGURE 6.6
Time to Maturity
Market Value of Bond, B0 ($)
10 9 8 7 6 5 4 3 2 1 0
Time to Maturity (years)
251
This is because of the mathematics of time value; the present values of short-term
cash flows change far less than the present values of longer-term cash flows in
response to a given change in the discount rate (required return).
EXAMPLE The effect of changing required returns on bonds of differing maturity can be
illustrated by using Mills Company’s bond and Figure 6.6. If the required return
rises from 10% to 12% (see the dashed line at 8 years), the bond’s value
decreases from $1,000 to $901—a 9.9% decrease. If the same change in required
return had occurred with only 3 years to maturity (see the dashed line at 3 years),
the bond’s value would have dropped to just $952—only a 4.8% decrease. Simi-
lar types of responses can be seen for the change in bond value associated with
decreases in required returns. The shorter the time to maturity, the less the impact
on bond value caused by a given change in the required return.
EXAMPLE The Mills Company bond, which currently sells for $1,080, has a 10% coupon
interest rate and $1,000 par value, pays interest annually, and has 10 years to
maturity. Because B0 ? $1,080, I ? $100 (0.10 ? $1,000), M ? $1,000, and
n ? 10 years, substituting into Equation 6.3a yields
$1,080 ? $100 ? (PVIFAk ) ? $1,000 ? (PVIFk )
d,10yrs d,10yrs
Our objective is to solve the equation for kd, the YTM.
Trial and Error Because we know that a required return, kd, of 10% (which
equals the bond’s 10% coupon interest rate) would result in a value of $1,000,
the discount rate that would result in $1,080 must be less than 10%. (Remember
that the lower the discount rate, the higher the present value, and the higher the
discount rate, the lower the present value.) Trying 9%, we get
$100 ? (PVIFA9%,10yrs) ? $1,000 ? (PVIF9%,10yrs)
? $100 ? (6.418) ? $1,000 ? (0.422)
? $641.80 ? $422.00
? $1,063.80
252 PART 2 Important Financial Concepts
Because the 9% rate is not quite low enough to bring the value up to $1,080, we
next try 8% and get
$100 ? (PVIFA8%,10yrs) ? $1,000 ? (PVIF8%,10yrs)
? $100 ? (6.710) ? $1,000 ? (0.463)
? $671.00 ? $463.00
? $1,134.00
Input Function Because the value at the 8% rate is higher than $1,080 and the value at the 9%
10 N
rate is lower than $1,080, the bond’s yield to maturity must be between 8% and
?1080 PV
9%. Because the $1,063.80 is closer to $1,080, the YTM to the nearest whole
100 PMT
percent is 9%. (By using interpolation, we could eventually find the more precise
1000 FV YTM value to be 8.77%.)7
CPT
I Calculator Use [Note: Most calculators require either the present value (B0 in
Solution this case) or the future values (I and M in this case) to be input as negative num-
8.766 bers to calculate yield to maturity. That approach is employed here.] Using the
inputs shown at the left, you should find the YTM to be 8.766%.
I
? ? ? (PVIFAkd/2,2n) ? M ? (PVIFkd/2,2n) (6.4a)
2
EXAMPLE Assuming that the Mills Company bond pays interest semiannually and that the
required stated annual return, kd, is 12% for similar-risk bonds that also pay
semiannual interest, substituting these values into Equation 6.4a yields
$100
B0? ? ? (PVIFA12%/2,2?10yrs) ? $1,000 ? (PVIF12%/2,2?10yrs)
2
WW 7. For information on how to interpolate to get a more precise answer, see the book’s home page at www.
W
[Link]/gitman
253
Table Use
B0 ? $50 ? (PVIFA6%,20periods) ? $1,000 ? (PVIF6%,20periods)
Input Function ? $50 ? (11.470) ? $1,000 ? (0.312) ? $? 88 5 .5 0
20 N ????
??????
???
6 I
Calculator Use In using a calculator to find bond value when interest is paid
50 PMT semiannually, we must double the number of periods and divide both the
1000 FV required stated annual return and the annual interest by 2. For the Mills Com-
CPT pany bond, we would use 20 periods (2 ? 10 years), a required return of 6%
PV (12% ? 2), and an interest payment of $50 ($100 ? 2). Using these inputs, you
Solution should find the bond value with semiannual interest to be $885.30, as shown at
885.30 the left. Note that this value is more precise than the value calculated using the
rounded financial-table factors.
Comparing this result with the $887.00 value found earlier for annual com-
pounding (see Table 6.6), we can see that the bond’s value is lower when semian-
nual interest is paid. This will always occur when the bond sells at a discount. For
bonds selling at a premium, the opposite will occur: The value with semiannual
interest will be greater than with annual interest.
Review Questions
6–16 What basic procedure is used to value a bond that pays annual interest?
Semiannual interest?
6–17 What relationship between the required return and the coupon interest
rate will cause a bond to sell at a discount? At a premium? At its par
value?
6–18 If the required return on a bond differs from its coupon interest rate,
describe the behavior of the bond value over time as the bond moves
toward maturity.
6–19 As a risk-averse investor, would you prefer bonds with short or long peri-
ods until maturity? Why?
6–20 What is a bond’s yield to maturity (YTM)? Briefly describe both the trial-
and-error approach and the use of a financial calculator for finding YTM.
S U M M A RY
FOCUS ON VALUE
Interest rates and required returns embody the real cost of money, inflationary expecta-
tions, and issuer and issue risk. They reflect the level of return required by market partici-
pants as compensation for the risk perceived in a specific security or asset investment.
Because these returns are affected by economic expectations, they vary as a function of
254 PART 2 Important Financial Concepts
time, typically rising for longer-term maturities or transactions. The yield curve reflects such
market expectations at any point in time.
The value of an asset can be found by calculating the present value of its expected cash
flows, using the required return as the discount rate. Bonds are the easiest financial assets to
value, because both the amounts and the timing of their cash flows are known with certainty.
The financial manager needs to understand how to apply valuation techniques to bonds in
order to make decisions that are consistent with the firm’s share price maximization goal.
Definitions of variables
B0 ? bond value
CFt ? cash flow expected at the end of year t
I ? annual interest on a bond
k ? appropriate required return (discount rate)
kd ? required return on a bond
M ? par, or face, value of a bond
n ? relevant time period, or number of years to maturity
V0 ? value of the asset at time zero
Valuation formulas
Bond value:
n
1 1
B0 ? I ? ?? ?
t?1(1 ? k ) ?
d
?M ? ? ? ?
t (1 ? k ) d
n [Eq. 6.3]
Explain yield to maturity (YTM), its calcula- nually are valued by using the same procedure used
LG6
tion, and the procedure used to value bonds to value bonds paying annual interest, except that
that pay interest semiannually. Yield to maturity the interest payments are one-half of the annual in-
(YTM) is the rate of return investors earn if they terest payments, the number of periods is twice the
buy a bond at a specific price and hold it until ma- number of years to maturity, and the required re-
turity. YTM can be calculated by trial and error or turn is one-half of the stated annual required return
financial calculator. Bonds that pay interest semian- on similar-risk bonds.
LG6 ST 6–2 Yield to maturity Elliot Enterprises’ bonds currently sell for $1,150, have an
11% coupon interest rate and a $1,000 par value, pay interest annually, and
have 18 years to maturity.
a. Calculate the bonds’ yield to maturity (YTM).
b. Compare the YTM calculated in part a to the bonds’ coupon interest rate,
and use a comparison of the bonds’ current price and their par value to
explain this difference.
PROBLEMS
LG2 6–1 Yield curve A firm wishing to evaluate interest rate behavior has gathered yield
data on five U.S. Treasury securities, each having a different maturity and all
measured at the same point in time. The summarized data follow.
A 1 year 12.6%
B 10 years 11.2
C 6 months 13.0
D 20 years 11.0
E 5 years 11.4
LG2 6–2 Term structure of interest rates The following yield data for a number of high-
est quality corporate bonds existed at each of the three points in time noted.
Yield
Time to maturity (years) 5 years ago 2 years ago Today
a. On the same set of axes, draw the yield curve at each of the three given times.
b. Label each curve in part a with its general shape (downward-sloping,
upward-sloping, flat).
c. Describe the general inflationary and interest rate expectation existing at
each of the three times.
LG2 6–3 Risk-free rate and risk premiums The real rate of interest is currently 3%; the
inflation expectation and risk premiums for a number of securities follow.
CHAPTER 6 Interest Rates and Bond Valuation 257
Inflation expectation
Security premium Risk premium
A 6% 3%
B 9 2
C 8 2
D 5 4
E 11 1
LG2 6–4 Risk premiums Eleanor Burns is attempting to find the actual rate of interest
for each of two securities—A and B—issued by different firms at the same point
in time. She has gathered the following data:
a. If the real rate of interest is currently 2%, find the risk-free rate of interest
applicable to each security.
b. Find the total risk premium attributable to each security’s issuer and issue
characteristics.
c. Calculate the actual rate of interest for each security. Compare and discuss
your findings.
LG2 6–5 Bond interest payments before and after taxes Charter Corp. has issued 2,500
debentures with a total principal value of $2,500,000. The bonds have a coupon
interest rate of 7%.
a. What dollar amount of interest per bond can an investor expect to receive
each year from Charter Corp.?
b. What is Charter’s total interest expense per year associated with this bond
issue?
c. Assuming that Charter is in a 35% corporate tax bracket, what is the com-
pany’s net after-tax interest cost associated with this bond issue?
LG3 6–6 Bond quotation Assume that the following quote for the Financial Manage-
ment Corporation’s $1,000-par-value bond was found in the Wednesday,
November 8, issue of the Wall Street Journal.
Fin Mgmt 8.75 05 8.7 558 100.25 ?0.63
258 PART 2 Important Financial Concepts
LG4 6–7 Valuation fundamentals Imagine that you are trying to evaluate the economics
of purchasing an automobile. You expect the car to provide annual after-tax
cash benefits of $1,200 at the end of each year, and assume that you can sell the
car for after-tax proceeds of $5,000 at the end of the planned 5-year ownership
period. All funds for purchasing the car will be drawn from your savings, which
are currently earning 6% after taxes.
a. Identify the cash flows, their timing, and the required return applicable to
valuing the car.
b. What is the maximum price you would be willing to pay to acquire the car?
Explain.
LG4 6–8 Valuation of assets Using the information provided in the following table, find
the value of each asset.
Cash flow
Asset End of year Amount Appropriate required return
A 1 $ 5,000 18%
2 5,000
3 5,000
C 1 $ 0 16%
2 0
3 0
4 0
5 35,000
E 1 $ 2,000 14%
2 3,000
3 5,000
4 7,000
5 4,000
6 1,000
259
LG4 6–9 Asset valuation and risk Laura Drake wishes to estimate the value of an
asset expected to provide cash inflows of $3,000 per year at the end of years 1
through 4 and $15,000 at the end of year 5. Her research indicates that she
must earn 10% on low-risk assets, 15% on average-risk assets, and 22% on
high-risk assets.
a. Determine what is the most Laura should pay for the asset if it is classified as
(1) low-risk, (2) average-risk, and (3) high-risk.
b. Say Laura is unable to assess the risk of the asset and wants to be certain
she’s making a good deal. On the basis of your findings in part a, what is the
most she should pay? Why?
c. All else being the same, what effect does increasing risk have on the value of
an asset? Explain in light of your findings in part a.
LG5 6–10 Basic bond valuation Complex Systems has an outstanding issue of $1,000-
par-value bonds with a 12% coupon interest rate. The issue pays interest annu-
ally and has 16 years remaining to its maturity date.
a. If bonds of similar risk are currently earning a 10% rate of return, how much
should the Complex Systems bond sell for today?
b. Describe the two possible reasons why similar-risk bonds are currently earn-
ing a return below the coupon interest rate on the Complex Systems bond.
c. If the required return were at 12% instead of 10%, what would the current
value of Complex Systems’ bond be? Contrast this finding with your findings
in part a and discuss.
LG5 6–11 Bond valuation—Annual interest Calculate the value of each of the bonds
shown in the following table, all of which pay interest annually.
Bond Par value Coupon interest rate Years to maturity Required return
LG5 6–12 Bond value and changing required returns Midland Utilities has outstanding a
bond issue that will mature to its $1,000 par value in 12 years. The bond has a
coupon interest rate of 11% and pays interest annually.
a. Find the value of the bond if the required return is (1) 11%, (2) 15%, and
(3) 8%.
b. Plot your findings in part a on a set of “required return (x axis)–market value
of bond (y axis)” axes.
c. Use your findings in parts a and b to discuss the relationship between the
coupon interest rate on a bond and the required return and the market value
of the bond relative to its par value.
d. What two possible reasons could cause the required return to differ from the
coupon interest rate?
260 PART 2 Important Financial Concepts
LG5 6–13 Bond value and time—Constant required returns Pecos Manufacturing has just
issued a 15-year, 12% coupon interest rate, $1,000-par bond that pays interest
annually. The required return is currently 14%, and the company is certain it
will remain at 14% until the bond matures in 15 years.
a. Assuming that the required return does remain at 14% until maturity, find
the value of the bond with (1) 15 years, (2) 12 years, (3) 9 years, (4) 6 years,
(5) 3 years, and (6) 1 year to maturity.
b. Plot your findings on a set of “time to maturity (x axis)–market value of
bond (y axis)” axes constructed similarly to Figure 6.6.
c. All else remaining the same, when the required return differs from the coupon
interest rate and is assumed to be constant to maturity, what happens to the
bond value as time moves toward maturity? Explain in light of the graph in
part b.
LG5 6–14 Bond value and time—Changing required returns Lynn Parsons is considering
investing in either of two outstanding bonds. The bonds both have $1,000 par
values and 11% coupon interest rates and pay annual interest. Bond A has
exactly 5 years to maturity, and bond B has 15 years to maturity.
a. Calculate the value of bond A if the required return is (1) 8%, (2) 11%, and
(3) 14%.
b. Calculate the value of bond B if the required return is (1) 8%, (2) 11%, and
(3) 14%.
c. From your findings in parts a and b, complete the following table, and dis-
cuss the relationship between time to maturity and changing required returns.
8% ? ?
11 ? ?
14 ? ?
d. If Lynn wanted to minimize interest rate risk, which bond should she pur-
chase? Why?
LG6 6–15 Yield to maturity The relationship between a bond’s yield to maturity and
coupon interest rate can be used to predict its pricing level. For each of the
bonds listed, state whether the price of the bond will be at a premium to par, at
par, or at a discount to par.
A 6% 10% ?????????
B 8 8 ?????????
C 9 7 ?????????
D 7 9 ?????????
E 12 10 ?????????
261
LG6 6–16 Yield to maturity The Salem Company bond currently sells for $955, has a
12% coupon interest rate and a $1,000 par value, pays interest annually, and
has 15 years to maturity.
a. Calculate the yield to maturity (YTM) on this bond.
b. Explain the relationship that exists between the coupon interest rate
and yield to maturity and the par value and market value of a
bond.
LG6 6–17 Yield to maturity Each of the bonds shown in the following table pays interest
annually.
Bond Par value Coupon interest rate Years to maturity Current value
A $1,000 9% 8 $ 820
B 1,000 12 16 1,000
C 500 12 12 560
D 1,000 15 10 1,120
E 1,000 5 3 900
LG6 6–18 Bond valuation—Semiannual interest Find the value of a bond maturing in 6
years, with a $1,000 par value and a coupon interest rate of 10% (5% paid
semiannually) if the required return on similar-risk bonds is 14% annual interest
(7% paid semiannually).
LG6 6–19 Bond valuation—Semiannual interest Calculate the value of each of the bonds
shown in the following table, all of which pay interest semiannually.
A $1,000 10% 12 8%
B 1,000 12 20 12
C 500 12 5 14
D 1,000 14 10 10
E 100 6 4 14
Required
a. If the price of the common stock into which the bond is convertible rises to
$30 per share after 5 years and the issuer calls the bonds at $1,080, should
Annie let the bond be called away from her or should she convert it into com-
mon stock?
b. For each of the following required returns, calculate the bond’s value, assum-
ing annual interest. Indicate whether the bond will sell at a discount, at a pre-
mium, or at par value.
(1) Required return is 6%.
(2) Required return is 8%.
(3) Required return is 10%.
c. Repeat the calculations in part b, assuming that interest is paid semiannually
and that the semiannual required returns are one-half of those shown. Com-
pare and discuss differences between the bond values for each required return
calculated here and in part b under the annual versus semiannual payment
assumptions.
d. If Annie strongly believes that inflation will rise by 1% during the next 6
months, what is the most she should pay for the bond, assuming annual
interest?
e. If the Atilier bonds are downrated by Moody’s from Aa to A, and if such a
rating change will result in an increase in the required return from 8% to
8.75%, what impact will this have on the bond value, assuming annual
interest?
f. If Annie buys the bond today at its $1,000 par value and holds it for exactly
3 years, at which time the required return is 7%, how much of a gain or loss
will she experience in the value of the bond (ignoring interest already received
and assuming annual interest)?
g. Rework part f, assuming that Annie holds the bond for 10 years and sells it
when the required return is 7%. Compare your finding to that in part f, and
comment on the bond’s maturity risk.
263
h. Assume that Annie buys the bond at its current closing price of 98.38 and
holds it until maturity. What will her yield to maturity (YTM) be, assuming
annual interest?
i. After evaluating all of the issues raised above, what recommendation would
you give Annie with regard to her proposed investment in the Atilier Indus-
tries bonds?
WEB EXERCISE Go to the Web site [Link]. Click on Economy & Bonds. Then
WW click on Bond Calculator, which is located down the page under the column
W
Bond Tools. Read the instructions on how to use the bond calculator. Using the
bond calculator:
1. Calculate the yield to maturity (YTM) for a bond whose coupon rate is
7.5% with maturity date of July 31, 2030, which you bought for 95.
2. What is the YTM of the above bond if you bought it for 105? For 100?
3. Change the yield % box to 8.5. What would be the price of this bond?
4. Change the yield % box to 9.5. What is this bond’s price?
5. Change the maturity date to 2006 and reset yield % to 6.5. What is the price
of this bond?
6. Why is the price of the bond in Question 5 higher than the price of the bond
in Question 4?
7. Explore the other bond-related resources at the site. Using Bond Market
Update, comment on current interest rate levels and the yield curve.
• Find answers to the questions that confront the owners and managers of finance
companies and the financial directors of all kinds of companies in the
performance of their duties.
• Study in depth the changes that occur in the market and their effects on the
financial dimension of business activity.
All of these activities are programmed and carried out with the support of our
sponsoring companies. Apart from providing vital financial assistance, our sponsors
also help to define the Center’s research projects, ensuring their practical relevance.
[Link]
COMPANY VALUATION METHODS.
THE MOST COMMON ERRORS IN VALUATIONS
Pablo Fernández*
Abstract
In this paper, we describe the four main groups comprising the most widely used company
valuation methods: balance sheet-based methods, income statement-based methods, mixed
methods, and cash flow discounting-based methods. The methods that are conceptually “correct”
are those based on cash flow discounting. We will briefly comment on other methods since - even
though they are conceptually “incorrect” - they continue to be used frequently.
We also present a real-life example to illustrate the valuation of a company as the sum of the
value of different businesses, which is usually called the break-up value.
We conclude the paper with the most common errors in valuations: a list that contains the most
common errors that the author has detected in more than one thousand valuations he has had
access to in his capacity as business consultant or teacher.
Keywords: Value, Price, Free cash flow, Equity cash flow, Capital cash flow, Book value,
Market value, PER, Goodwill, Required return to equity, Working capital requirements.
COMPANY VALUATION METHODS.
THE MOST COMMON ERRORS IN VALUATIONS∗
For anyone involved in the field of corporate finance, understanding the mechanisms
of company valuation is an indispensable requisite. This is not only because of the importance
of valuation in acquisitions and mergers but also because the process of valuing the company
and its business units helps identify sources of economic value creation and destruction within
the company.
In this paper, we will briefly describe the four main groups comprising the most widely used
company valuation methods. Each of these groups is discussed in a separate section: balance
sheet-based methods (Section 2), income statement-based methods (Section 3), mixed methods
(Section 4), and cash flow discounting-based methods (Section 5).1
Section 7 uses a real-life example to illustrate the valuation of a company as the sum of the
value of different businesses, which is usually called the break-up value. Section 8 shows the
methods most widely used by analysts for different types of industry.
∗
Another version of this paper may be found in chapter 2 of the author's book “Valuation Methods and Shareholder
Value Creation,” Academic Press, San Diego, CA, 2002.
1
The reader interested in methods based on value creation measures can see Fernández (2002, chapters 1, 13 and
14). The reader interested in valuation using options theory can see Fernández (2001c).
The methods that are becoming increasingly popular (and are conceptually “correct”) are those
based on cash flow discounting. These methods view the company as a cash flow generator
and, therefore, assessable as a financial asset. We will briefly comment on other methods since
- even though they are conceptually “incorrect” - they continue to be used frequently.
Section 12 contains the most common errors in valuations: a list that contains the most
common errors that the author has detected in more than one thousand valuations he has had
access to in his capacity as business consultant or teacher.
Value should not be confused with price, which is the quantity agreed between the seller and
the buyer in the sale of a company. This difference in a specific company’s value may be due to
a multitude of reasons. For example, a large and technologically highly advanced foreign
company wishes to buy a well-known national company in order to gain entry into the local
market, using the reputation of the local brand. In this case, the foreign buyer will only value
the brand but not the plant, machinery, etc. as it has more advanced assets of its own.
However, the seller will give a very high value to its material resources, as they are able to
continue producing. From the buyer’s viewpoint, the basic aim is to determine the maximum
value it should be prepared to pay for what the company it wishes to buy is able to contribute.
From the seller’s viewpoint, the aim is to ascertain what should be the minimum value at which
it should accept the operation. These are the two figures that face each other across the table in
a negotiation until a price is finally agreed on, which is usually somewhere between the two
extremes.2 A company may also have different values for different buyers due to economies of
scale, economies of scope, or different perceptions about the industry and the company.
- For the buyer, the valuation will tell him the highest price he should pay.
- For the seller, the valuation will tell him the lowest price at which he should be
prepared to sell.
- The valuation is used to compare the value obtained with the share’s price on the
stock market and to decide whether to sell, buy or hold the shares.
- The valuation of several companies is used to decide the securities that the portfolio
should concentrate on: those that seem to it to be undervalued by the market.
- The valuation of several companies is also used to make comparisons between
companies. For example, if an investor thinks that the future course of GE’s share
price will be better than that of Amazon, he may buy GE shares and short-sell
Amazon shares. With this position, he will gain provided that GE’s share price does
better (rises more or falls less) than that of Amazon.
3. Public offerings:
- The valuation is used to justify the price at which the shares are offered to the
public.
- The valuation is used to compare the shares’ value with that of the other assets.
8. Strategic planning:
- The valuation of the company and the different business units is fundamental for
deciding what products/business lines/countries/customers… to maintain, grow or
abandon.
- The valuation provides a means for measuring the impact of the company’s possible
policies and strategies on value creation and destruction.
Some of these methods are the following: book value, adjusted book value, liquidation value,
and substantial value.
2.1. Book Value
A company’s book value, or net worth, is the value of the shareholders’ equity stated in the
balance sheet (capital and reserves). This quantity is also the difference between total assets and
liabilities, that is, the surplus of the company’s total goods and rights over its total debts with
third parties.
Let us take the case of a hypothetical company whose balance sheet is that shown in Table 1.
The shares’ book value (capital plus reserves) is 80 million dollars. It can also be calculated as
the difference between total assets (160) and liabilities (40 + 10 + 30), that is, 80 million dollars.
Table 1
Alfa Inc. Official Balance Sheet (Million Dollars)
ASSETS LIABILITIES
Cash 5 Accounts payable 40
Accounts receivable 10 Bank debt 10
Inventories 45 Long-term debt 30
Fixed assets 100 Shareholders’ equity 80
Total assets 160 Total liabilities 160
This value suffers from the shortcoming of its own definition criterion: accounting criteria are
subject to a certain degree of subjectivity and differ from “market” criteria, with the result that
the book value almost never matches the “market” value.
When the values of assets and liabilities match their market value, the adjusted net worth is
obtained. Continuing with the example of Table 1, we will analyze a number of balance sheet
items individually in order to adjust them to their approximate market value. For example, if
we consider that:
- Accounts receivable includes 2 million dollars of bad debt, this item should have a value
of 8 million dollars.
- Stock, after discounting obsolete, worthless items and revaluing the remaining items at
their market value, has a value of 52 million dollars.
- Fixed assets (land, buildings, and machinery) have a value of 150 million dollars,
according to an expert.
- The book value of accounts payable, bank debt and long-term debt is equal to their
market value.
ASSETS LIABILITIES
Cash 5 Accounts payable 40
Accounts receivable 8 Bank debt 10
Inventories 52 Long-term debt 30
Fixed assets 150 Capital and reserves 135
Total assets 215 Total liabilities 215
The adjusted book value is 135 million dollars: total assets (215) less liabilities (80). In this case,
the adjusted book value exceeds the book value by 55 million dollars.
Taking the example given in Table 2, if the redundancy payments and other expenses
associated with the liquidation of the company Alfa Inc. were to amount to 60 million dollars,
the shares’ liquidation value would be 75 million dollars (135-60).
Obviously, this method’s usefulness is limited to a highly specific situation, namely, when the
company is bought with the purpose of liquidating it at a later date. However, it always
represents the company’s minimum value as a company’s value, assuming it continues to
operate, is greater than its liquidation value.
It can also be defined as the assets’ replacement value, assuming the company continues to
operate, as opposed to their liquidation value. Normally, the substantial value does not include
those assets that are not used for the company’s operations (unused land, holdings in other
companies, etc.).
- Gross substantial value: this is the assets’ value at market price (in the example of Table
2: 215).
- Net substantial value or corrected net assets: this is the gross substantial value less
liabilities. It is also known as adjusted net worth, which we have already seen in the
previous section (in the example of Table 2: 135).
-5
- Reduced gross substantial value: this is the gross substantial value reduced only by the
value of the cost-free debt (in the example of Table 2: 175 = 215 - 40). The remaining
40 million dollars correspond to accounts payable.
Table 3
Market Value/Book Value (P/BV), PER and Dividend Yield (Div./P) of Different National Stock Markets
P/BV is the share’s price (P) divided by its book value (BV). PER is the share’s price divided by the earnings per share. Div/P is
the dividend per share divided by the price.
Figure 1 shows the evolution of the price/book value ratio of the British, German and United
States stock markets. It can be seen that the book value, in the 90’s, has lagged considerably
below the shares’ market price.
Figure 1
Evolution of the Price/Book Value Ratio on the British, German and United States Stock Mmarkets
0
1/75 1/77 1/79 1/81 1/83 1/85 1/87 1/89 1/91 1/93 1/95 1/97 1/99 1/01 1/03 1/05 1/07
The income statement for the company Alfa Inc. is shown in Table 4:
Table 4
Alfa Inc. Income Statement (Million Dollars)
Sales 300
Cost of sales 136
General expenses 120
Interest expense 4
Earnings before tax 40
Tax (35%) 14
Net income 26
Table 3 shows the mean PER of a number of different national stock markets in September
1992 and August 2000. Figure 2 shows the evolution of the PER for the German, English and
United States stock markets.
3
The PER (price earnings ratio) of a share indicates the multiple of the earnings per share that is paid on the stock
market. Thus, if the earnings per share in the last year has been $3 and the share’s price is $26, its PER will be 8.66
(26/3). On other occasions, the PER takes as its reference the forecast earnings per share for the next year, or the
mean earnings per share for the last few years. The PER is the benchmark used predominantly by the stock markets.
Note that the PER is a parameter that relates a market item (share price) with a purely accounting item (earnings).
-7
Figure 2
Evolution of the PER of the German, English and United States Stock Markets
35 PER Germany UK US
30
25
20
15
10
5
0
1/71 1/73 1/75 1/77 1/79 1/81 1/83 1/85 1/87 1/89 1/91 1/93 1/95 1/97 1/99 1/01 1/03 1/05 1/07
Sometimes, the relative PER is also used, which is simply the company’s PER divided by the
country’s PER.
Where: DPS = dividend per share distributed by the company in the last year; Ke = required
return to equity.
If, on the other hand, the dividend is expected to grow indefinitely at a constant annual rate g,
the above formula becomes the following:
Where DPS1 is the dividends per share for the next year.
Empirical evidence5 shows that the companies that pay more dividends (as a percentage of their
earnings) do not obtain a growth in their share price as a result. This is because when a
company distributes more dividends, normally it reduces its growth because it distributes the
money to its shareholders instead of plowing it back into new investments.
4
Other flows are share buy-backs and subscription rights. However, when capital increases take place that give rise
to subscription rights, the shares’ price falls by an amount approximately equal to the rights’ value.
5
There is an enormous and highly varied literature about the impact of dividend policies on equity value. Some
recommendable texts are to be found in Sorensen and Williamson (1985) and Miller (1986).
Figure 3
Evolution of the Dividend Yield of the German, Japanese and United States Stock Markets
Table 3 shows the dividend yield of several international stock markets in September 1992,
August 2000 and February 2007. In 2000, Japan was the country with the lowest dividend yield
(0.6%) and Spain had a dividend yield of 1.5%.
In order to analyze this method’s consistency, Smith Barney analyzed the relationship between
the price/sales ratio and the return on equity. The study was carried out in large corporations
(capitalization in excess of 150 million dollars) in 22 countries. He divided the companies into
five groups depending on their price/sales ratio: group 1 consisted of the companies with the
lowest ratio, and group 5 contained the companies with the highest price/sales ratio. The mean
return of each group of companies is shown in the following table:
Table 5
Relationship Between Return and the Price/Sales Ratio
-9
It can be seen from this table that, during the period December 84-December 89, the equity of
the companies with the lowest price/sales ratio in December 1984 on average provided a higher
return than that of the companies with a higher ratio. However, this ceased to apply during the
period December 89-September 97: there was no relationship between the price/sales ratio in
December 1989 and the return on equity during those years.
The price/sales ratio can be broken down into a further two ratios:
The first ratio (price/earnings) is the PER and the second (earnings/sales) is normally known as
return on sales.
- Value of the company / earnings before interest, taxes, depreciation and amortization
(EBITDA).
In March 2000, a French bank published its valuation of Terra based on the price/sales ratio of
comparable companies:
Applying the mean ratio (74) to Terra’s expected sales for 2001 (310 million dollars), they
estimated the value of Terra’s entire equity to be 19.105 billion dollars (68.2 dollars per share).
6
We could also list a number of other ratios which we could call “sui-generis”. One example of such ratios is the
value/owner. In the initial stages of a valuation that I was commissioned to perform by a family business that was
on sale, one of the brothers told me that he reckoned that the shares were worth about 30 million euros. When I
asked him how he had arrived at that figure, he answered, “We are three shareholder siblings and I want each one of
us to get 10 million”.
7
For a more detailed discussion of the multiples method, see Fernández (2001b).
8
See Fernández (2001a).
These methods apply a mixed approach: on the one hand, they perform a static valuation of the
company’s assets and, on the other hand, they try to quantify the value that the company will
generate in the future. Basically, these methods seek to determine the company’s value by
estimating the combined value of its assets plus a capital gain resulting from the value of its
future earnings: they start by valuing the company’s assets and then add a quantity related
with future earnings.
V = A + (n x B), or V = A + (z x F)
Where: A = net asset value; n = coefficient between 1.5 and 3; B = net income; z = percentage
of sales revenue; and F = turnover.
The first formula is mainly used for industrial companies, while the second is commonly used
for the retail trade.
When the first method is applied to the hypothetical company Alfa Inc., assuming that the
goodwill is estimated at three times the annual earnings, it would give a value for the
company’s equity amounting to 213 million dollars (135 + 3 x 26).
A variant of this method consists of using the cash flow instead of the net income.
9
The author feels duty bound to tell the reader that he does not like these methods at all but as they have been used
a lot in the past, and they are still used from time to time, a brief description of some of them is included. However,
we will not mention them again in the rest of the book. The reader can skip directly to Section 5. However, if he
continues to read this section, he should not look for much “science” in the methods that follow because they are
very arbitrary.
- 11
4.2. The Simplified "Abbreviated Goodwill Income" Method or the Simplified UEC10
Method
According to this method, a company’s value is expressed by the following formula:
V = A + an (B - iA)
Where:
B = net income for the previous year or that forecast for the coming year.
i = interest rate obtained by an alternative placement, which could be debentures, the return on
equities, or the return on real estate investments (after tax).
an (B - iA) = goodwill.
This formula could be explained in the following manner: the company’s value is the value of
its adjusted net worth plus the value of the goodwill. The value of the goodwill is obtained by
capitalizing, by application of a coefficient an, a "superprofit" that is equal to the difference
between the net income and the investment of the net assets “A” at an interest rate “i”
corresponding to the risk-free rate.
In the case of the company Alfa Inc., B = 26; A = 135. Let us assume that 5 years and 15% are
used in the calculation of an, which would give an = 3.352. Let us also assume that i = 10%.
With this hypothesis, the equity’s value would be: 135 + 3.352 (26 - 0.1 x 135) = 135 + 41.9 =
176.9 million dollars.
For the UEC, a company’s total value is equal to the substantial value (or revalued net assets)
plus the goodwill. This calculated by capitalizing at compound interest (using the factor an) a
superprofit which is the profit less the flow obtained by investing at a risk-free rate i a capital
equal to the company’s value V.
The difference between this method and the previous method lies in the value of the goodwill,
which, in this case, is calculated from the value V we are looking for, while in the simplified
method, it was calculated from the net assets A.
In the case of the company Alfa Inc., B = 26; A = 135, an = 3.352, i = 10%. With these
assumptions, the equity’s value would be: (135 + 3.352 x 26) / (1 + 0.1 x 3.352) = 222.1 /
1.3352 = 166.8 million dollars.
10
UEC: This is the acronym of “Union of European Accounting Experts”.
The rate i used is normally the interest rate paid on long-term Treasury bonds. As can be seen
in the first expression, this method gives equal weight to the value of the net assets (substantial
value) and the value of the return. This method has a large number of variants that are
obtained by giving different weights to the substantial value and the earnings’ capitalization
value.
In the case of the company Alfa Inc., B = 26; A = 135, i = 10%. With these assumptions, the
equity’s value would be 197.5 million dollars.
In this case, the value of the goodwill is obtained by restating for an indefinite duration the
value of the superprofit obtained by the company. This superprofit is the difference between the
net income and what would be obtained from placing at the interest rate i, a capital equal to
the value of the company’s assets. The rate tm is the interest rate earned on fixed-income
securities multiplied by a coefficient between 1.25 and 1.5 to adjust for the risk. In the case of
the company Alfa Inc., B = 26; A = 135, i = 10%. Let us assume that tm = 15%. With these
assumptions, the equity’s value would be 218.3 million dollars.
Here, the value of the goodwill is equal to a certain number of years of superprofits. The buyer
is prepared to pay the seller the value of the net assets plus m years of superprofits. The number
of years (m) normally used ranges between 3 and 5, and the interest rate (i) is the interest rate
for long-term loans.
In the case of the company Alfa Inc., B = 26; A = 135, i = 10%. With these assumptions, and if
m is 5 years, the equity’s value would be 197.5 million dollars.
The rate i is the rate of an alternative, risk-free placement; the rate t is the risk-bearing rate
used to restate the superprofit and is equal to the rate i increased by a risk ratio. According to
this method, a company’s value is equal to the net assets increased by the restated superprofit.
As can be seen, the formula is a variant of the UEC’s method when the number of years tends
towards infinity.
- 13
In the case of the company Alfa Inc., B = 26; A = 135, i = 10%. With these assumptions, if
t = 15%, the equity’s value would be 185 million dollars.
The mixed methods described previously have been used extensively in the past. However, they
are currently used increasingly less and it can be said that, nowadays, the cash flow
discounting method is generally used because it is the only conceptually correct valuation
method. In these methods, the company is viewed as a cash flow generator and the company’s
value is obtained by calculating these flows’ present value using a suitable discount rate.
Cash flow discounting methods are based on the detailed, careful forecast, for each period, of
each of the financial items related with the generation of the cash flows corresponding to the
company’s operations, such as, for example, collection of sales, personnel, raw materials,
administrative and sales expenses, loan repayments. Consequently, the conceptual approach is
similar to that of the cash budget.
In cash flow discounting-based valuations, a suitable discount rate is determined for each type
of cash flow. Determining the discount rate is one of the most important tasks and takes into
account the risk, historic volatilities; in practice, the minimum discount rate is often set by the
interested parties (the buyers or sellers are not prepared to invest or sell for less than a certain
return, etc.).
Although at first sight it may appear that the above formula is considering a temporary duration
of the flows, this is not necessarily so as the company’s residual value in the year n (Vn) can be
calculated by discounting the future flows after that period. A simplified procedure for
considering an indefinite duration of future flows after the year n is to assume a constant
growth rate (g) of flows after that period. Then the residual value in year n is VRn = CFn (1 + g) /
(k - g).
Although the flows may have an indefinite duration, it may be acceptable to ignore their value
after a certain period, as their present value decreases progressively with longer time horizons.
Furthermore, the competitive advantage of many businesses tends to disappear after a few
years.
5.2. Deciding the Appropriate Cash Flow for Discounting and the Company’s
Economic Balance Sheet
In order to understand what are the basic cash flows that can be considered in a valuation, the
following chart shows the different cash streams generated by a company and the appropriate
discount rates for each flow.
There are three basic cash flows: the free cash flow, the equity cash flow, and the debt cash
flow.
The easiest one to understand is the debt cash flow, which is the sum of the interest to be paid
on the debt plus principal repayments. In order to determine the present market value of the
existing debt, this flow must be discounted at the required rate of return to debt (cost of the
debt). In many cases, the debt’s market value shall be equivalent to its book value, which is
why its book value is often taken as a sufficient approximation to the market value.11
The free cash flow (FCF) enables the company’s total value12 (debt and equity: D + E) to be
obtained. The equity cash flow (ECF) enables the value of the equity to be obtained, which,
combined with the value of the debt, will also enable the company’s total value to be
determined. The discount rates that must be used for the FCF and the ECF are explained in the
following sections.
Figure 4 shows in simplified form the difference between the company’s full balance sheet and
its economic balance sheet. When we refer to the company’s (financial) assets, we are not
talking about its entire assets but about total assets less spontaneous financing (suppliers,
creditors...). To put it another way, the company’s (financial) assets consist of the net fixed
assets plus the working capital requirements.13 The company’s (financial) liabilities consist of
the shareholders’ equity (the shares) and its debt (short and long-term financial debt).14 In the
rest of the paper, when we talk about the company’s value, we will be referring to the value of
the debt plus the value of the shareholders’ equity (shares).
11
This is only valid if the required return to debt is equal to the debt’s cost.
12
The “company’s value” is usually considered to be the sum of the value of the equity plus the value of the
financial debt.
13
For an excellent discussion of working capital requirements (WCR), see Faus (1996).
14
The shareholders’ equity or capital can include, among others, common stock, preferred stock and convertible
preferred stock; and the different types of debt can include, among others, senior debt, subordinated debt, convertible
debt, fixed or variable interest debt, zero or regular coupon debt, short or long-term debt, etc.
- 15
Figure 4
Full and Economic Balance Sheet of a Company
The free cash flow (FCF) is the operating cash flow, that is, the cash flow generated by
operations, without taking into account borrowing (financial debt), after tax. It is the money
that would be available in the company after covering fixed asset investments and working
capital requirements, assuming that there is no debt and, therefore, there are no financial
expenses.
In order to calculate future free cash flows, we must forecast the cash we will receive and must
pay in each period. This is basically the approach used to draw up a cash budget. However, in
company valuation, this task requires forecasting cash flows further ahead in time than is
normally done in any cash budget.
Accounting cannot give us this information directly as, on one hand, it uses the accrual
approach and, on the other hand, it allocates its revenues, costs and expenses using basically
arbitrary mechanisms. These two features of accounting distort our perception of the
appropriate approach when calculating cash flows, which must be the “cash” approach, that is,
cash actually received or paid (collections and payments). However, when the accounting is
adjusted to this approach, we can calculate whatever cash flow we are interested in.
We will now try to identify the basic components of a free cash flow in the hypothetical
example of the company XYZ. The information given in the accounting statements shown in
Table 6 must be adjusted to give the cash flows for each period, that is, the sums of money
actually received and paid in each period.
Table 6 gives the income statement for the company XYZ, SA. Using this data, we shall
determine the company’s free cash flow, which we know by definition must not include any
Table 7 shows how the free cash flow is obtained from earnings before interest and tax (EBIT).
The tax payable on the EBIT must be calculated directly; this gives us the net income without
subtracting interest payments, to which we must add the depreciation for the period because it
is not a payment but merely an accounting entry. We must also consider the sums of money to
be allocated to new investments in fixed assets and new working capital requirements (WCR),
as these sums must be deducted in order to calculate the free cash flow.
Table 6
Income Statement for XYZ
Table 7
Free Cash Flow of XYZ, SA
In order to calculate the free cash flow, we must ignore financing for the company’s operations
and concentrate on the financial return on the company’s assets after tax, viewed from the
perspective of a going concern, taking into account in each period the investments required for
the business’s continued existence.
Finally, if the company had no debt, the free cash flow would be identical to the equity cash
flow, which is another cash flow variant used in valuations and will be analyzed below.
17
5.2.2. The Equity Cash Flow
The equity cash flow (ECF) is calculated by subtracting from the free cash flow the interest and
principal payments (after tax) made in each period to the debt holders and adding the new debt
provided. In short, it is the cash flow remaining available in the company after covering fixed
asset investments and working capital requirements and after paying the financial charges and
repaying the corresponding part of the debt’s principal (in the event that there exists debt). This
can be represented in the following expression:
ECF = FCF - [interest payments x (1- T)] - principal repayments + new debt
When making projections, the dividends and other expected payments to shareholders must
match the equity cash flows.
This cash flow assumes the existence of a certain financing structure in each period, by which
the interest corresponding to the existing debts is paid, the installments of the principal are
paid at the corresponding maturity dates and funds from new debt are received. After that there
remains a certain sum which is the cash available to the shareholders, which will be allocated
to paying dividends or buying back shares.
When we restate the equity cash flow, we are valuing the company’s equity (E), and, therefore,
the appropriate discount rate will be the required return to equity (Ke). To find the company’s
total value (D + E), we must add the value of the existing debt (D) to the value of the equity (E).
Capital cash flow (CCF) is the term given to the sum of the debt cash flow plus the equity cash
flow. The debt cash flow is composed by the sum of interest payments plus principal
repayments. Therefore:
It is important to not confuse the capital cash flow with the free cash flow.
5.3. Calculating the Value of the Company Using the Free Cash Flow
In order to calculate the value of the company using this method, the free cash flows are
discounted (restated) using the weighted average cost of debt and equity or weighted average
cost of capital (WACC):
E Ke + D Kd (1 - T )
E + D = present value [FCF; WACC] where WACC =
E+ D
D = market value of the debt. E = market value of the equity.
Kd = cost of the debt before tax = required return to debt. T = tax rate.
The WACC is calculated by weighting the cost of the debt (Kd) and the cost of the equity (Ke)
with respect to the company’s financial structure. This is the appropriate rate for this case as,
since we are valuing the company as a whole (debt plus equity), we must consider the required
return to debt and the required return to equity in the proportion to which they finance the
company.
The value of the company without debt is obtained by discounting the free cash flow, using the
rate of required return to equity that would be applicable to the company if it were to be
considered as having no debt. This rate (Ku) is known as the unlevered rate or required return
to assets. The required return to assets is smaller than the required return to equity if the
company has debt in its capital structure as, in this case, the shareholders would bear the
financial risk implied by the existence of debt and would demand a higher equity risk premium.
In those cases where there is no debt, the required return to equity (Ke = Ku) is equivalent to
the weighted average cost of capital (WACC), as the only source of financing being used is
capital.
The present value of the tax shield arises from the fact that the company is being financed with
debt, and it is the specific consequence of the lower tax paid by the company as a consequence
of the interest paid on the debt in each period. In order to find the present value of the tax
shield, we would first have to calculate the saving obtained by this means for each of the years,
multiplying the interest payable on the debt by the tax rate. Once we have obtained these flows,
we will have to discount them at the rate considered appropriate. Although the discount rate to
be used in this case is somewhat controversial, many authors suggest using the debt’s market
cost, which need not necessarily be the interest rate at which the company has contracted its
debt.
5.5. Calculating the Value of the Company’s Equity by Discounting the Equity
Cash Flow
The market value of the company’s equity is obtained by discounting the equity cash flow at
the rate of required return to equity for the company (Ke). When this value is added to the
market value of the debt, it is possible to determine the company’s total value.
The required return to equity can be estimated using any of the following methods:
Ke = [Div1 / P0] + g. Div1 = dividends to be received in the following period = Div0(1 + g).
15
This method is called APV (adjusted present value). For a more detailed discussion, the reader can see Fernández
(2002, chapters 19 and 21) and Fernández (2004).
- 19
For example, if a share’s price is 200 dollars, it is expected to pay a dividend of 10 dollars and
the dividend’s expected annual growth rate is 11%:
2. The capital asset pricing model (CAPM), which defines the required return to equity in the
following terms:
Ke = RF + ß (RM - RF)
Thus, given certain values for the equity’s beta, the risk-free rate and the market risk premium,
it is possible to calculate the required return to equity.17
5.6. Calculating the Company’s Value by Discounting the Capital Cash Flow
According to this model, the value of a company (market value of its equity plus market value
of its debt) is equal to the present value of the capital cash flows (CCF) discounted at the
weighted average cost of capital before tax (WACCBT):
There are more methods for valuing companies by discounting the expected cash flows.
Fernández (2004b) explains ten different methods for Valuing Companies by Cash Flow
Discounting and shows that all ten methods always give the same value. This result is logical,
as all the methods analyze the same reality under the same hypotheses; they differ only in the
cash flows taken as the starting point for the valuation.
16
The beta measures the systematic or market risk of a share. It indicates the sensitivity of the return on a share held
in the company to market movements. If the company has debt, the incremental risk arising from the leverage must
be added to the intrinsic systematic risk of the company’s business, thus obtaining the levered beta.
The classic finance textbooks provide a full discussion of the concepts analyzed here. For example, Brealey and
17
20 -
1. Historic and strategic analysis of the company and the industry
A. Financial analysis B. Strategic and competitive analysis
Evolution of income statements and balance sheets Evolution of the industry
Evolution of cash flows generated by the company Evolution of the company’s competitive position
Evolution of the company’s investments Identification of the value chain
Evolution of the company’s financing Competitive position of the main competitors
Analysis of the financial health Identification of the value drivers
Analysis of the business’s risk
Involvement of the company. The company’s managers must be involved in the analysis of the company, of the
industry and in the cash flow projections.
Multifunctional. The valuation is not a task to be performed solely by financial management. In order to obtain a good
valuation, it is vital that managers from other departments take part in estimating future cash flows and their risk.
Strategic. The cash flow restatement technique is similar in all valuations, but estimating the cash flows and calibrating the
risk must take into account each business unit’s strategy.
Compensation. The valuation’s quality is increased when it includes goals (sales, growth, market share, profits,
investments, ...) on which the managers’ future compensation will depend.
Real options. If the company has real options, these must be valued appropriately. Real options require a totally different
risk treatment from the cash flow restatements.
Historic analysis. Although the value depends on future expectations, a thorough historic analysis of the financial, strategic
and competitive evolution of the different business units helps assess the forecasts’ consistency.
Technically correct. Technical correction refers basically to: a) calculation of the cash flows; b) adequate treatment of the
risk, which translates into the discount rates; c) consistency of the cash flows used with the rates applied; d) treatment of the
residual value; e) treatment of inflation.
- 21
6. Which is the Best Method to Use?
Table 8 shows the value of the equity of the company Alfa Inc. obtained by different methods
based on shareholders’ equity, earnings and goodwill. The fundamental problem with these
methods is that some are based solely on the balance sheet, others are based on the income
statement, but none of them consider anything but historic data. We could imagine two
companies with identical balance sheets and income statements but different prospects: one
with high sales, earnings and margin potential, and the other in a stabilized situation with
fierce competition. We would all concur in giving a higher value to the former company than
to the latter, in spite of their historic balance sheets and income statements being equal.
The most suitable method for valuing a company is to discount the expected future cash flows,
as the value of a company’s equity - assuming it continues to operate - arises from the
company’s capacity to generate cash (flows) for the equity’s owners.
Table 8
Alfa Inc.
Value of the Equity According to Different Methods (Million Dollars)
Book value 80
Adjusted book value 135
Liquidation value 75
PER 173
Classic valuation method 213
Simplified UEC method 177
UEC method 167
Indirect method 197
Direct or Anglo-Saxon method 218
Annual profit purchase method 197
Risk-bearing and risk-free rate method 185
The best way to explain this method is with an example. Table 9 shows the valuation of a
North American company performed in early 1980. The company in question had three separate
divisions: household products, shipbuilding, and car accessories.
18
For a more detailed discussion of this type of valuation, we recommend Chapter 14 of the book “Valuation:
measuring and managing the value of companies”, by Copeland et al., edited by Wiley, 2000.
Table 9 shows that the investment bank valued the company’s equity between 430 and 479
million dollars (or, to put it another way, between 35 and 39 dollars per share). But let us see
how it arrived at that value. First of all, it projected each division’s net income and then
allocated a (maximum and minimum) PER to each one. Using a simple multiplication (earnings
x PER), it calculated the value of each division. The company’s value is simply the sum of the
three divisions’ values.
We can call this value (between 387 and 436 million dollars) the value of the earnings
generated by the company. We must now add to this figure the company’s cash surplus, which
the investment bank estimated at 77.5 million dollars. However, the company’s pension plan
was not fully funded (it was short by 34.5 million dollars), and consequently, this quantity had
to be subtracted from the company’s value.
After performing these operations, the conclusion reached is that each share is worth between
35 and 39 dollars, which is very close to the offer made of 38 dollars per share.
Table 9
Valuation of a Company as the Sum of the Value of its Divisions
Individual Valuation of Each Business Using the PER Criterion
*Cash surplus: 103.1 million dollars in cash, less 10 million dolars for operations and less 15.6 million dollars of financial debt.
The growth of utility companies is usually fairly stable. In developed countries, the rates
charged for their services are usually indexed to the CPI, or they are calculated in accordance
with a legal framework. Therefore, it is simpler to extrapolate their operating statement and
then discount the cash flows. In these cases, particular attention must be paid to regulatory
changes, which may introduce uncertainties.
- 23
In the case of banks, the focus of attention is the operating profit (financial margin less
commissions less operating expenses), adjusting basically for bad debts. Their industry portfolio
is also analyzed. Valuations such as the PER are used, or the net worth method (shareholders’
equity adjusted for provision surpluses/deficits, and capital gains or losses on assets such as the
industry portfolio).
Industrial and commercial companies. In these cases, the most commonly used valuations -
apart from restated cash flows - are those based on financial ratios (PER, price/sales, price/cash
flow).
These issues are discussed in greater detail in Fernández (2002, chapters 3 and 4).
Table 10
Factors Influencing the Equity’s Value (Value Drivers)
VALUE OF EQUITY
Expectations of future cash flows Required return to equity
Adquisitions / disposals
Control of operations
Managers. People.
Risk management
Corporate culture
Assets in place
Buyer / target
Profit margin
Real options
Technology
Financing
Liquidity
Taxes
Size
24 -
Table 10 shows that the equity’s value depends on three primary factors (value drivers):
These factors can be subdivided in turn into return on the investment, company growth, risk-
free interest rate, market risk premium, operating risk and financial risk. However, these factors
are still very general. It is very important that a company identify the fundamental parameters
that have most influence on the value of its shares and on value creation. Obviously, each
factor’s importance will vary for the different business units.
The speculative bubble theory can be derived from fundamental analysis and occupies a middle
ground between the above two theories, which seek to account for the behavior and evolution
of share prices. The MIT professor Olivier Blanchard developed the algebraic expression of the
speculative bubble, and it can be obtained from the same equation that gives the formula
normally used by the fundamentalists. It simply makes use of the fact that the equation has
several solutions, one of which is the fundamental solution and another is the fundamental
solution with a speculative bubble tacked onto it. By virtue of the latter solution, a share’s price
can be greater than its fundamental value (Net Present Value of all future dividends) if a bubble
develops simultaneously, which at any given time may: a) continue to grow, or b) burst and
vanish. To avoid tiring ourselves with equations, we can imagine the bubble as an equity
overvaluation: an investor will pay today for a share a quantity that is greater than its
fundamental value if he hopes to sell it tomorrow for a higher price, that is, if he hopes that the
bubble will continue growing. This process can continue so long as there are investors who
trust that the speculative bubble will continue to grow, that is, investors who expect to find in
19
The communication with the market factor not only refers to communication and transparency with the markets
in the strict sense but also to communication with: analysts, rating companies, regulatory agencies, board of
directors, employees, customers, distribution channels, partner companies, suppliers, financial institutions, and
shareholders.
- 25
the future other trusting investors to whom they can sell the bubble (share) for a price that is
greater than the price they have paid. Bubbles tend to grow during periods of euphoria, when it
seems that the market’s only possible trend is upwards. However, there comes a day when there
are no more trusting investors left and the bubble bursts and vanishes: shares return to their
fundamental value.
This theory is attractive because it enables fundamental theory to be synthesized with the
existence of anomalous behaviors (for the fundamentalists) in the evolution of share prices.
Many analysts have used this theory to account for the tremendous drop in share prices on the
New York stock market and on the other world markets on 19 October 1987. According to this
explanation, the bursting of a bubble that had been growing over the previous months caused
the stock market crash. A recent study performed by the Yale professor Shiller provides further
evidence in support of this theory. Shiller interviewed 1000 institutional and private investors.
The investors who sold before the Black Monday said that they sold because they thought that
the stocks were already overvalued. However, the most surprising finding is that more than
90% of the institutional investors who did not sell said that they too believed that the market
was overvalued, but hoped that they would be able to sell before the inevitable downturn. In
other words, it seems that more than 90% of the institutional investors were aware that a
speculative bubble was being formed - the stock was being sold for more than its fundamental
value -, but trusted that they would be able to sell before the bubble burst. Among the private
investors who did not sell before 19 October, more than 60% stated that they also believed that
the stocks were overvalued.
Figure 5
The 1929 American Stock Market Crisis
300
250
200
150
100
50
0
1-26 1-27 1-28 1-29 1-30 1-31 1-32 1-33 1-34 1-35 1-36 1-37
Figure 6
The Spanish Stock Market Crisis of October 1987
3.600
3.400
3.200
3.000
IBEX 35
2.800
2.600
2.400
2.200
2.000
01-87 03-87 05-87 07-87 09-87 11-87 01-88 03-88 05-88 07-88 09-88 11-88
Speculative bubbles can also develop outside of the stock market. One often-quoted example is
that of the Dutch tulips in the 17th Century. An unusual strain of tulips began to become
increasingly sought after and its price rose continuously... In the end, the tulips’ price returned
to normal levels and many people were ruined. There have also been many speculative bubbles
in the real estate business. The story is always the same: prices temporarily rocket upwards and
then return to “normal” levels. In the process, many investors who trust that the price will
continue to rise lose a lot of money. The problem with this theory, as with many of the
economic interpretations, is that it provides an ingenious explanation to account for events a
posteriori but it is not very useful for providing forecasts about the course that share prices will
follow in the future. For this, we would need to know how to detect the bubble and predict its
future course. This means being able to separate the share price into two components (the
fundamental value and the bubble) and knowing the number of investors who trust that the
bubble will continue to grow (here many chartists can be included). What the theory does
remind us is that the bubble can burst at any time. History shows that, so far, all the bubbles
have eventually burst.
The only sure recipe to avoid being trapped in a speculative bubble is to not enter it: to never
buy what seems to be expensive, even if advised to do so by certain "experts", who appeal to
esoteric tendencies and the foolishness or rashness of other investors.
- 27
The most common errors are in italic
1. Errors in the discount rate calculation and concerning the company’s riskiness
A. Wrong risk-free rate used for the valuation
1. Using the historical average of the risk-free rate.
2. Using the short-term Government rate.
3. Wrong calculation of the real risk-free rate.
B. Wrong beta used for the valuation
1. Using the historical industry beta, or the average of the betas of similar companies, when the result
goes against common sense.
2. Using the historical beta of the company when the result goes against common sense.
3. Assuming that the beta calculated from historical data captures the country risk.
4. Using the wrong formulae for levering and unlevering the beta.
5. Arguing that the best estimation of the beta of a company from an emerging market is the beta of the
company with respect to the S&P 500.
6. When valuing an acquisition, using the beta of the acquiring company.
C. Wrong market risk premium used for the valuation
1. The required market risk premium is equal to the historical equity premium.
2. The required market risk premium is equal to zero.
3. Assume that the required market risk premium is the expected risk premium.
D. Wrong calculation of WACC
1. Wrong definition of WACC.
2. The debt to equity ratio used to calculate the WACC is different from the debt to equity ratio resulting
from the valuation.
3. Using discount rates lower than the risk-free rate.
4. Using the statutory tax rate, instead of the effective tax rate of the levered company.
5. Valuing all the different businesses of a diversified company using the same WACC (same leverage
and same Ke).
6. Considering that WACC / (1-T) is a reasonable return for the company’s stakeholders.
7. Using the wrong formula for the WACC when the value of debt is not equal to its book value.
8. Calculating the WACC assuming a certain capital structure and deducting the outstanding debt from
the enterprise value.
9. Calculating the WACC using book values of debt and equity.
10. Calculating the WACC using strange formulae.
E. Wrong calculation of the value of tax shields
1. Discounting the tax shield using the cost of debt or the required return to unlevered equity.
2. Odd or ad-hoc formulae.
F. Wrong treatment of country risk
1. Not considering the country risk, arguing that it is diversifiable.
2. Assuming that a disaster in an emerging market will increase the beta of the country’s companies
calculated with respect to the S&P 500.
3. Assuming that an agreement with a government agency eliminates country risk.
4. Assuming that the beta provided by Market Guide with the Bloomberg adjustment incorporates the
illiquidity risk and the small cap premium.
5. Odd calculations of the country risk premium.
G. Including an illiquidity, small-cap, or specific premium when it is not appropriate
1. Including an odd small-cap premium.
2. Including an odd illiquidity premium.
3. Including a small-cap premium equal for all companies.
28 -
3. Errors in the calculation of the residual value
A. Inconsistent cash flow used to calculate perpetuity.
B. The debt to equity ratio used to calculate the WACC to discount the perpetuity is different from the debt to equity ratio resulting from the
valuation.
C. Using ad hoc formulas that have no economic meaning.
D. Using arithmetic averages instead of geometric averages to assess growth.
E. Calculating the residual value using the wrong formula.
F. Assume that a perpetuity starts a year before it really starts.
6. Organizational errors
A. Making a valuation without checking the forecasts made by the client.
B. Commissioning a valuation from an investment bank without having any involvement in it.
C. Involving only the finance department in valuing a target company.
- 29
References
Brealey, R.A. and S.C. Myers (2000), “Principles of Corporate Finance,” 6th edition, McGraw-
Hill, New York.
Copeland, T. E., T. Koller and J. Murrin (2000), “Valuation: Measuring and Managing the Value
of Companies”, 3rd edition, Wiley, New York.
Copeland and Weston (1988), “Financial Theory and Corporate Policy,” 3rd edition, Addison-
Wesley, Reading, Massachusetts.
Faus, Josep (1996), “Finanzas operativas,” Biblioteca IESE de Gestión de Empresas, Ediciones
Folio.
Fernández, Pablo (2001a), “Internet Valuations: The Case of Terra-Lycos”, SSRN Working Paper
n. 265608.
Fernandez, Pablo (2001b), “Valuation using multiples. How do analysts reach their
conclusions?,” SSRN Working Paper n. 274972.
Fernández, Pablo (2001c), “Valuing real options: frequently made errors,” SSRN Working Paper
n. 274855.
Fernández, Pablo (2002), “Valuation Methods and Shareholder Value Creation,” Academic
Press, San Diego, CA.
Fernández, Pablo (2004), “The Value of Tax Shields is NOT Equal to the Present Value of Tax
Shields,” Journal of Financial Economics, Vol. 73/1, (July), pp. 145-165.
Fernández, Pablo (2004b), “Valuing Companies by Cash Flow Discounting: Ten Methods and
Nine Theories," SSRN Working Paper n. 256987.
Fernández, Pablo and Jose Maria Carabias (2006), “96 Common and Uncommon Errors in
Company Valuation,” SSRN Working Paper n. 895151.
Miller, M.H. (1986), “Behavioral Rationality in Finance: The Case of Dividends,” Journal of
Business, No. 59, october, pp. 451-468.
Sorensen, E. H. and D.A. Williamson (1985), “Some evidence on the value of the dividend
discount model,” Financial Analysts Journal, 41, pp. 60-69.
30 -