Election
Election
Electoral Studies
journal homepage: [Link]/locate/electstud
a r t i c l e i n f o a b s t r a c t
Article history: From the 1970s onwards, a wide range of forecasting techniques have been developed in the literature on
Received 26 November 2014 electoral forecasting. However, these models have primarily been applied in two-party, presidential
Received in revised form democracies, with the US being by far the most popular country to investigate. The question thus arises
27 April 2015
whether the same techniques that have proved successful in this context can also be applied to the more
Accepted 15 June 2015
complex, multiparty democracies in northern Europe. This paper seeks to answer this question and in
Available online 3 July 2015
the process makes two main contributions. Firstly, the popular dynamic linear model (Jackman, 2005) is
tried and tested in Germany and Sweden where it is shown that reasonable forecasts can be made
Keywords:
Election forecasting
despite the complexity of the systems and the emergence of new parties. A novelty is then introduced
Multiparty systems when cyclical changes in party support are modelled through a seasonal component. This extension of the
Dynamic linear model dynamic linear model helps to significantly lower the error in early forecasts and is thus something that
Political polling could be useful in future applications of the model.
© 2015 The Author. Published by Elsevier Ltd. This is an open access article under the CC BY license
([Link]
1. Introduction faced by many scientists working with other types of data survey
data. The main difference with political polling is that here we get a
Predicting election results is a relatively recent and increas- perfectly unbiased measurement, the national election, which
ingly popular part of political science research. Competitive makes it possible to test our estimations. This allows us to pro-
elections are the hallmark of modern democracy and being able gressively develop techniques that can be used also in many other
to foreshadow who wins them is a tantalizing skill that has disciplines as well.
garnered significant scientific attention (Fisher et al., 2011; Lewis- The goal of the present paper is to test whether it is possible to
Beck and Bellucci, 1982; Lock and Gelman, 2010; Gibson and predict elections also in difficult parliamentary systems where a
Lewis-Beck, 2011; Jackman, 2005). Election forecasting stands out wide range of parties are competing for power, and if this can be
from many other types of political science research in a number done with reasonable lead time. For this purpose, two countries
of ways. It is highly data-driven, focused on a very concrete and with a long tradition of multiparty competition, namely Germany
delimited task, and in most studies the goal is not to explain and Sweden, have been selected. These cases provide a compelling
election outcomes but to describe and predict them. In that sense, tests since most of the previous studies have focused on more
the question ‘how’ rather than the standard scientific question stable two-party systems and the methods have also been devel-
‘why’ is in focus. oped to fit such political environments.
The question ‘how’ is still highly relevant from a scientific Sweden, in contrast, has 8 parties represented in parliament and
perspective. In order to answer it with reasonable accuracy you has seen large shifts in the electoral fortunes of the parties during
need to make the most of limited and flawed polling data, while the first decade of the 21st century. In the latest German elections
controlling for seasonal fluctuations in public opinion, variability in in 2013, 6 parties received more than 4% of the vote. Such a set-up
measurements and bias associated with particular polling houses. requires our models to take many more active players into
In overcoming these problems we can shed more and better light consideration.
on public opinion by overcoming the flaws inherent in individual Previous studies have focused on a wide range of countries,
polls. In addition, the circumstances facing election polling are also but the U.S. has received the lion's share of scholarly attention.
Other countries, though, such as the U.K.(Fisher et al., 2011),
France (Foucault and Nadeau, 2012), Australia (Carlsen, 2000),
and Italy (Lewis-Beck and Bellucci, 1982), have also been studied.
E-mail address: [Link]@[Link]. The most popular technique used in these studies (see Bartels
[Link]
0261-3794/© 2015 The Author. Published by Elsevier Ltd. This is an open access article under the CC BY license ([Link]
2 D. Walther / Electoral Studies 40 (2015) 1e13
and Zaller (2001); Hibbs Jr (2000) for early overviews) is some In the 70s and 80s, economists and political scientists also
type of structural model. This approach treats incumbent vote started taking an interest in forecasting and a number of
share as the dependent variable and a range of economic and competing models emerged. The polling companies had, perhaps
political measures are used as explanatory variables that are unsurprisingly, relied mainly on their own estimations of vote
believed to have a bearing on the outcome (e.g. Lewis-Beck intentions that they got from their pre-election polls. The scien-
(2005). tists, in contrast, introduced regression-based approaches. These
In this paper, in contrast, it is argued that pure structural relied on economic and political variables assumed to influence
models are difficult to apply in the multiparty context. To accu- government popularity and inserted them into structural models
rately forecast the results of all parliamentary parties it appears that were used to forecast election outcomes. OLS regression was
necessary to include some kind of polling data1 in the model. An the dominant approach.
increasingly popular way of doing this is through a dynamic linear
model that belongs to the general family of space state models
(Jackman, 2005; Harrison and West, 1997; Pickup and Johnston, 2.1. Predicting elections through structural models
2007). A latent trend of popular support is estimated by aggre-
gating polling results into a time series through Kalman filtering. Most applications of structural models in electoral forecasting
The election forecast is then made by extrapolating this trend into have been of the type:
the future.
Government election result ¼ political measures
The main contributions of this paper are firstly the demonstra-
tion that reasonable election forecasts in multiparty systems can be þ economic performance þ error
made through a dynamic linear model. Structural models appear to
be of limited utility though, at least if not complemented by polling The models have mainly differed in how they have oper-
data. The average error of the DLM predictions is found be around ationalized political and economic performance. The political
0.69 percentage points per party in Germany and 0.78 i Sweden for situation has been captured for example through measures of
the last three elections. These results are on par, or better, than popularity of the president/PM, general left- or right-wing
what many of the previous studies have achieved in the more stable sentiment among the populace as well as current length of stay
two party systems. in office (Bartels and Zaller, 2001; Foucault and Nadeau, 2012;
A second contribution is the introduction of what has been Abramowitz, 2008). The belief is that such variables can capture
termed a seasonal component (Kitagawa and Gersch, 1984; Bell the general political mood in the country rather than just the
and Hillmer, 1984). In the economics literature a seasonal current support for a specific political party. By tapping into the
component is used to capture predictable seasonal trends (such general mood we can learn what the political backdrop to the
as a boom in sales just before Christmas), and it has been argued election will look like and this situation is then modified (exac-
that the development in support of political parties shows similar erbated or ameliorated) through the government's economic
predictable patterns (Sanders, 1991). The stronger these reoc- performance.
curring trends are, the more the standard DLM benefits from the Measures of the government's economic track record have also
inclusion of the seasonal component. The seasonal component is been plentiful. The most popular ones include changes in GDP per
empirically tested in both Sweden and Germany. It is found to capita, inflation and unemployment, but growth in real income and
help us make reasonable forecasts at an earlier stage and it re- the subjective beliefs among voters about the government's eco-
duces the error compared to the standard DLM one month before nomic performance have also often been used (Anderson, 2000;
the election with around 17% in the Swedish case and 8% in Bartels and Zaller, 2001). The causal theory here is clear: a gov-
Germany. Adding the seasonal component makes a model that is ernment that has handled state finances successfully and improved
already good at nowcasting better at the more difficult art of the prosperity of its citizens will be rewarded come Election Day,
forecasting. and a government that has failed on these measures will be pun-
This paper starts with an overview of the main forecasting ished (Healy and Lenz, 2014; Saalfeld, 2008).
techniques that have been used in the past. I then discuss appli- Estimates for the coefficients used in the structural models are
cations of structural models and find that we have theoretical generated by applying the model to as many past elections as are
reasons and empirical findings from the field of economic voting available to learn what effect the variables have had in the past.
that suggest that such techniques hold little utility in the multi- Predicting the next election therefore becomes a simple matter of
party case. The paper then moves on to the dynamic linear model inserting the relevant values for the political and economic in-
and applies it to both the German and Swedish cases. In the final dicators as they stand in the election year and then multiplying
section I apply the model with the addition of a seasonal compo- them with the coefficients that previous elections tell us provide
nent. The conclusion is that the dynamic linear works well but that the best fit.
more work is needed to incorporate explanatory factors into the One illustrative (and very Spartan) example of such an
model to further improve early forecasting. approach is the famous bread and peace model developed by
Douglas Hibbs (see Hibbs Jr (2000) for an early overview). Hibbs
2. The main approaches to election forecasting argues that only two measures are needed to predict the outcome
of the US presidential election: namely the weighted cumulative
Historically speaking, election forecasts are a relatively recent income growth during the full term of the government and the
phenomenon. The polling company Gallup attempted forecasts of number of US soldiers killed in foreign wars (particularly in Korea
the US presidential elections in the 30s and 40s (with very modest and Vietnam). Using only those two indicators he manages to
success) and by the 50s more pollsters, not just in the US, took a account for 90% of the variation in election results. Controlling for
stab at predicting the election (see Lewis-Beck (2005) for an other variables that have been suggested in the literature does
overview). nothing to improve predictive capacity, Hibbs argues (Hibbs Jr,
2000).
Hibbs' diminutive model nicely illustrates the two compelling
1
The word data is throughout this paper treated as a singular noun. advantages structural models have over polls. Firstly, the
D. Walther / Electoral Studies 40 (2015) 1e13 3
information you need to make your prediction is usually readily Predictions using polls range from the more simplistic
available well before the next election. Once the model has been weighted poll averages (aggregating the polls while accounting for
applied to previous elections and the coefficients are at hand, poll size) to some highly complex dynamic linear models (Jackman
making the prediction becomes a simple matter of inserting the (2005) provides a good introduction). The dynamic linear model
latest data. In this view, what political parties do in the run up to (DLM) has become the staple horse technique for polling based
the election is primarily to inform and convince people of the forecasts (Pickup and Johnston, 2007; Fisher et al., 2011; Linzer,
underlying economic and political reality. But political scientists 2013) and it is therefore worth going over the technique in some
who know where to look can find the relevant data well before detail.
the parties go into campaign mode and consequently know how
the campaigns will eventually come to change public opinion.
A second advantage of structural models is that they deal with 2.3. The dynamic linear model
actual causes of the outcome of the election. It makes intuitive
sense that an unpopular president from a party that has been in DLMs rely on Markov chains with random walk and a Kalman
office for a long time and has a poor economic track record fares filter (Harrison and West, 1997) to estimate the underlying public
worse at the polls. Structural models can test exactly how robustly support for each party. Each poll that comes in is taken to be a
such factors are associated with the election outcomes and can thus slightly flawed measure of the real support for the party at time t,
test what matters the most to voters. and the polls are then pooled into a time series that tracks party
support over time. By doing so the hope is to overcome the limi-
tations and biases of a single poll to establish a more realistic
2.2. Election forecasts through polling-based methods aggregate measure. So two of the main advantages of applying the
DLM are firstly that the model constitutes a highly effective way to
The second popular, and increasingly dominant, method for combine many polls over time into an estimate that conforms with
predicting elections generally disregards objective economic and the laws of probability. Secondly, the model allows us to perform a
political indicators and instead looks at people's subjectively real-time tracking of party support, which gives us more contin-
stated voting intentions. Opinion polls are becoming ever more uous information than the one off estimate provided by structural
plentiful in most Western countries and now provide frequent models.
insights into public opinion. Polls are now common not just in the The core of a DLM is defined by the following set of equations
run-up to an election but throughout the electoral cycle and, (Petris et al., 2009):
thanks to technological advances, now also have more
respondents. Pi ¼ mi þ s2i ; s2 N 0; s2i (1)
In Sweden, for example, the average number of polls was fairly
stable between the 1970s and the late 90s, but then increased
manifold during the first decade of the 21st century. An overview of mi ¼ mi1 þ di ; d N 0; d2i (2)
this trend is available in Fig. 1. Apart from the notable spikes in
election years, we can see a clear gradual progression in the past
m0 Nðm0 ; C0 Þ (3)
decade towards the current figure of close to 8 polls per month
(which rose to over 10 in 2014). This gives us a lot more data to Where equation (1) defines the observed time trend (i.e. the
work with than has been available in the past and makes polling actual polling data) and (2) the assumed ‘real’ underlying trend. (3)
data more feasible as the basis for forecasting. is the starting value that sets the Markov chains in motion. We can
In Germany the number of national polls was already quite high see that the mean of the polling data (ui ) is the estimation from
in the first years of the 21st century. Still, since then the number of the underlying trend, which in turn is a result of the previous es-
polls has continued to rise from an average of 16 per month in 2002 timations plus a variance term (di ). The actual polls (Pi ) are then
to just over 21 in the 2013. simply the estimated latent trend with the addition of some
The process works by continuously trying to predict the next Party R2 DLM with seasonal component R2 standard DLM Country
point in the time series. So a prediction is made, and once the next CDU/CSU 0.872 0.872 Germany
result is actually measured (i.e. new polling data comes in) we SPD 0.440 0.429 Germany
calculate how far off we were, update our estimation of the true GRUENE 0.823 0.820 Germany
latent state, and then make a new prediction for the next obser- FDP 0.515 0.503 Germany
LINKE 0.325 0.316 Germany
vation. The actual election is then simply another upcoming
S 0.226 0.217 Sweden
observation we are trying to predict. V 0.542 0.540 Sweden
The scholars that have applied DLMs have adapted the standard MP 0.223 0.209 Sweden
model outlined above in ways particularly suited to election fore- M 0.783 0.783 Sweden
FP 0.042 0.020 Sweden
casts. There are a couple of things one might wish to control for,
C 0.225 0.202 Sweden
such as the size of the poll, bias associated with a particular houses KD 0.133 0.110 Sweden
(Pickup and Johnston, 2007; Fisher et al., 2011). These factors
The best fit DLM model and the best fit DLM model with a seasonal component were
modify how much confidence we should have in a new poll that
computed for each party. The time series ran for a total of 24 months.
comes our way (Silver, 2012).
Various techniques to achieve this have been suggested. We can
control for the size of the poll by adding a prior to the variance of close to the parliamentary threshold in the polls also tend to slump
the sigma term in equation (1): in the middle of the election period only to regain support in the
run-up to the election. This is because voters hesitate to let them
pð1 pÞ
s2 ¼ (4) disappear completely, especially if they are part of a larger coalition
N (Freden, 2014). Also, following a scandal or another unpopular
Pickup and Johnston (2007) suggested a way to control for bias event, parties can temporarily lose support. If we were to base our
among polling houses by using the median polling house as an forecast on the party's standing just after the event we likely un-
anchor (reference category) and then calculating whether certain derestimate their eventual result. In all of these cases, the devel-
houses differ significantly from this. You can run a series of re- opment of party support can be seen as seasonal or cyclical since it
gressions with dummy variables for the polling houses, according tends to revert back from temporary outliers.
to a variation of equation (1): So if we do believe that there are seasonal trends in how party
support develops, how can the dynamic linear model be extended
Pi ¼ mi þ bi Xi (5) to incorporate this information before we see it in the polls? One
possible solution is to extend the time series of polling data to
Where the X's are dummy variables representing each polling incorporate a greater number of polls and then model the reoc-
house. The question in the regression thus becomes: which polls curring patterns through what in the economics literature is
differ significantly from the median estimate once we have known as a ‘seasonal component’.2 This is often used to capture
controlled for the DLM estimate? If a particular house has a co- predictable economic events such as a slump during summer or a
efficient that shows that it consistently differs from the median sales boom in the days leading up to Christmas. Mathematically,
estimate, its impact on our forecast can be weighted down in this can be modelled through (Kitagawa and Gersch, 1984; Scott,
proportion to how far off it is on average. The problem is that we 2014):
need to use one house as the reference point (the median in
Pickup and Johnston's model) and there are clearly no guarantees
that the polling industry as a whole gets it right on average. It X
L1
could very well be that the extreme outlier is in fact the best poll St ¼ St1 þ εt Nð0; sÞ (6)
out there. i¼1
Still, if the polling industry as a whole gets it significantly
Where S is the season and L the number of periods in each season.3
wrong on average our poll based model is unlikely to be of much
Seasonal adjustment of time series dates back to the 1920s and
use anyway. We are already assuming that aggregating the polls
a wide range of different techniques have been suggested (see Bell
gives us a better estimate and that we improve our estimate by
and Hillmer (1984) for an early overview). The techniques, natu-
treating the polls as a continuous time-series. If those assumptions
rally, have different strengths and weaknesses, and the method
hold, this technique should be a useful way to measure the running
developed by Kitigawa and Gersch was selected for two main
bias associated with particular houses and will thus be applied
reasons. First, it is highly flexible and can be applied to long time
here.
series of different durations. This can be contrasted with the
popular X12-Arima model (used e.g. by the US census bureau)
2.4. Adding a seasonal component
which works best for monthly or quarterly data (Hyndman and
Athanasopoulos, 2014; Jain, 2001). Second, this way to model
One problem with the DLM, even with the modifications of the
seasonality can easily be added as a new layer in the DLM model
basic model listed above, is that it is built exclusively on the in-
and can thus be calculated separately alongside the other com-
formation that has been present in the polls up until that point in
ponents. Unlike some other techniques, this makes it possible to
time. This means that even if we know that certain changes in party
set priors for some of the terms and to calculate their individual
support are likely to happen in the coming months, the basic model
variance (Jain, 2001). This ensures that the seasonal component
cannot take this into account.
For example, it is well known that the support of government
parties tends to drop in the middle of its term in office only to be 2
Technically, a seasonal component is of fixed duration (e.g. every six months),
partly or completely regained when the next election is whereas a cyclical component has a more flexible life cycle. Here the terms are used
approaching (Sanders, 1991). This is the so-called political business more or less interchangeably.
3
cycle. Similarly, at least in some systems, smaller parties that are See the appendix for a discussion of how this is applied here.
D. Walther / Electoral Studies 40 (2015) 1e13 5
can be successfully integrated into the overarching Bayesian set- with translating particular economic and political results to precise
up. vote shares for all the different parties. These problems seem
But even though we have theoretical reasons to expect sea- inherent in the set-up of multiparty systems and each will be dis-
sonal swings, how can we be sure that adjusting for it will cussed briefly in turn.
improve the model? Testing for the extent to which seasonality is
present in the party time series here is less straightforward than 3.1. Lack of a clear dependent variable
in normal economic time series, since the trends are not neces-
sarily as predictable as those that occur on a daily or monthly As shown above, almost all applications of structural models in
basis. Still, one possible solution, proposed by Hyndman and election forecasting use incumbent vote share as the dependent
Athanasopoulos (2014), is to fit two models: one that controls variable and then use various political and economic indicators as
for seasonality and one that doesn't and then compare their one- predictors. This works well when you have two parties, since if
step-ahead prediction errors. Since the DLM assumes linearity, party A is in government and you can forecast its result, party B
the total ability of the models to pick up the variance of the time gets 100 minus the score of party A. This is useful, because it
series can be measured through R2 . The result of this exercise can means that you only have to run a single regression on your
be seen in Table 1. sample since this allows you to estimate the result of both main
The results indicate that the predictive accuracy improves, parties.
slightly but reliably, when controlling for seasonality. In percent- Even in other cases, where there are two main parties but also
ages the improvement was slightly larger in the Swedish case. other minor parties in the competition (e.g. the UK, France), the
Moreover, the test employed here (accuracy in predicting the next regression estimates can still answer the question of how the
poll in the time series) should be the standard DLM's strongest suits competition for the reins of governments will end. For example, in
since the polls are temporally close. It is likely that controlling for Foucault and Nadeau's attempt at election forecasting in France
seasonality will be even more useful when predicting elections that (Foucault and Nadeau, 2012) they used the results of the conser-
are further away and this expectation will be put to the test in vative party in the 2nd round as the dependent variable.
Section 6. In systems with proportional representation though, where
parties usually have to enter into a coalition after the election to
2.5. The reasons for adopting a Bayesian approach form a government (Mitchell and Nyblade, 2008), the precise re-
sults of all the parties matter. For example, even if you could
The main reason for adopting a model based on Kalman filtering accurately forecast how the German conservatives and social
and a general Bayesian recursive model is accuracy. The Kalman democrats will do in the election, you still would not be able to
filter can be proved to be optimal when trying to model a linear predict who would eventually be in power. That depends too much
time series subject to Gaussian noise, and it can easily be extended on how the liberals, greens, left party and others do.
within a Bayesian framework to make it more adapted to electoral This means that we cannot simply use incumbent party vote
forecasting (Harrison and West, 1997). share as our dependent variable since there is no way of knowing
The main benefit in our case is that our confidence in the polling how the share of votes will be distributed among the various op-
estimates can be quantified. We can control for both the size of position parties. Instead we need to have separate models for each
polls and the reliability of the polling houses through priors to party and thus generate unique coefficients for each player in the
determine exactly how much influence each poll should have on party system. But adopting this approach immediately leads to
our estimations. This way to explicitly model confidence is difficult new, even more damaging problems.
to emulate in classical frequentist time series models and it offers a
highly flexible way of including other information in our model not 3.2. Difficulty assigning responsibility to individual parties
directly available in the polls.
That being said, Bayesian statistics will be utilized in a fairly Using incumbent vote share as the dependent variable is ad-
pragmatic fashion here. For example, the estimates of the variance vantageous not only for practical reasons, as discussed above, but
that we get from the DLM will be used to calculate confidence in- also because of our understanding of how the causal processes
tervals even though this practice is sometimes shunned in Bayesian work here. Incumbent parties should be affected by the economic
circles (but is done e.g. by Jackman (2005)). and political situation of the country, since these parties are
responsible for it.
3. Why structural models are problematic in the multiparty However, using every party in the country as the dependent
case variable in a series of independent regressions, as we need to do if
we want to estimate the support of many players in a multiparty
Both structural models and DLM models been tried and tested in system, neglects this logical link between results and responsibility.
a wide range of studies and we know that both of them produce If e.g. the Social Democrats are in power and the unemployment
reasonable results. At least in stable settings with few competing rate increases by 1.5%, how does this impact the Christian Demo-
parties. When trying to translate the techniques to a multiparty crats, the Greens and the Left parties in opposition? Will all op-
framework a number of interesting challenges emerge for the position parties be affected in the same way? Even if we assume
structural models. There are three main reasons for this, namely: that the voters want to punish the incumbent social democrats,
there is little reason to expect that all of the opposition parties will
1. Lack of a clear dependent variable be affected in a proportional and predictable manner as would be
2. Difficulty in assigning economic and political responsibility to the case in a two party system.
individual parties Moreover, if there is a coalition government in power, are all
3. Difficulty in dealing with new parties parties held equally responsible by voters? This seems unlikely
given that the coalition members have had different re-
The common theme in these three factors is that multiparty sponsibilities in the governance of the country and generally are
systems display greater flux, more frequent emergence of new ac- associated with different overarching policy areas in the eyes of
tors, and, given the greater number of active players here, problems voters.
6 D. Walther / Electoral Studies 40 (2015) 1e13
The argument that this causal link matters receives both theo- current government. The incumbent coalition is treated as one
retical and empirical support from the field of economic voting (e.g, unitary actor which means that you have one agent responsible for
Paldam (1991); Lewis-Beck and Paldam (2000); Anderson (2000); economic and political developments.
Bengtsson (2004); Nadeau et al. (2013); Goodhart (2009); Narud If the horse race for power is the key question, this could be a
and Valen (2008)). Anderson summarized the main findings of the suitable strategy. The ‘chancellor model’ offered by Norpoth and
field in 2000 by stating that three key features of the political Gschwend, for example, has been remarkably successful since
system mediate the effect of the economy on the government's 2002. Still, in many cases we are interested in the electoral fate not
election result, namely: only of the government parties, but also of the parties in the op-
position. The current Swedish government consists of two parties,
Institutional clarity with six in opposition. The models employed in these studies, even
Governing party size if successful, would only be able to predict the results of a small
Availability of alternatives subset of the parties. And even then we would only get the
aggregate result of the entire government, not the results of the
These three factors have one feature in common that can be individual coalition members.
termed ‘clarity of responsibility’. If the institutional set-up makes it Additionally, sometimes the estimates are simply off the mark.
clear who is in charge, if there is one dominant governing party The prediction offered by Sundell and Lewis-Beck was for the
and it is apparent what the alternative is to the incumbent ruler, Swedish right-wing coalition government to get 49.7% of the votes
then responsibility for policy outcomes is clear (Goodhart, 2009). in the 2014 election. This was 10.3 percentage points away from the
When there is clarity of responsibility the voters know who to actual result of 39.4, which can be compared with the DLM pre-
blame or credit for the economic conditions and then the gov- diction from the same time of 37.8.
ernment's track record plays an important role in predicting the Another strategy, pursued by Je ro
^ me et al. (2013), is to combine
election result. estimates from a structural model with some polling data. Graefe
This shows why economic performance plays an important role (2015) goes even further and combines a weighted combination
in presidential systems where one party government is the norm. In of regression estimates, polls, expert judgements and betting
such systems it is abundantly clear who is responsible for national markets. The study by Jerome et al. is interesting in its parsimony,
economic management and usually there is one main opposition because it uses economic and political variables for the two larger
party that provides a clear alternative to the current government. In parties in the German system, and then combines this with polling
multiparty systems where coalition government is the norm and data (primarily focussing on voting intentions) for the smaller
there are numerous opposition parties, the relationship is likely to parties. They produce a prediction that is both reasonable and has
be far less clear-cut. good lead time, and, as I will argue below, such an integrated model
has the potential to combine the strengths of both polling and
3.3. Dealing with new parties structural models.
In general, then, the existing scholarship suggests that the
One additional problem that makes structural models prob- problems associated with structural models in multiparty system
lematic in multiparty systems is that the number of parties in such are difficult to overcome. You either have to focus on a subset of
systems tends to change over time. In Sweden, three parties have parties or reinforce your model through polling data. So how much,
gained and kept parliamentary representation since the late 80s, exactly, can we do with polling data in multiparty systems?
and many other countries with proportional systems, especially in
Eastern Europe, have seen even swifter changes of the political
landscape (Jungerstam-Mulders, 2006). In Germany, one new party 4. The dynamic linear model in the multiparty case
has gained a parliamentary foothold and a few others (e.g. the
Pirate Party and Alternative für Deutschland) have been close in Many of the problems associated with structural models can, at
national elections and have successfully secured representation in least seemingly, be overcome by utilizing polling data instead. The
regional parliaments. problems with the lack of a sensible dependent variable, too few
But when new parties emerge we have no previous elections to data points, and how to handle new parties all disappear when we
estimate model coefficients on and thus no reasonable way to instead look at polls. Indeed, systems where structural models
gauge how certain political and economic conditions should in- work well (e.g. the US, the UK), do not necessarily have any
fluence a given party. Thus, even if we could get around the prob- advantage when it comes to poll based techniques, other than that
lems with what to use as a dependent variable and how to model there are fewer parties to estimate.
the causal chain of responsibility, the added flux in multiparty Some of the key challenges of the DLM have been touched upon
systems compared to their majoritarian equivalents means that above and have also already been discussed in more detail else-
forecasting techniques relying exclusively on some kind of struc- where (e.g. Jackman (2005); Linzer (2013); Fisher et al. (2011)). But
tural model are bound to run into problems. before we proceed to applying the DLM, we should first deal with a
few challenges to the standard model that arise in particular in the
3.4. How previous studies have dealt with these issues multiparty case.
One statistical problem at the outset is that the DLM assumes a
Five recent articles have applied some kind of structural model normal distribution. But technically speaking, the underlying data
to the German or Swedish cases (Je ro
^me et al., 2013; Norpoth and generating process in the multiparty case should follow the
Gschwend, 2010; Graefe, 2015; Kayser and Leininger, 2013; multinomial distribution, since the starting point of the data is the
Sundell and Lewis-Beck, 2014). The question then inescapably ari- individual survey respondent that chooses between a discrete
ses how these have dealt with the theoretical problems outlined number of different parties. As has been noted in the statistical
above. One popular solution, opted for by Kayser and Leininger literature (Severini, 2005), though, the multinomial has a normal
(2013); Norpoth and Gschwend (2010); Sundell and Lewis-Beck approximation. Jackman (2005) comes to the same conclusion
(2014), is to focus only on the subset of parties that make up the when he argues that the binomial distribution in the Australian
D. Walther / Electoral Studies 40 (2015) 1e13 7
party system can be approximated through the univariate normal and thus why it is advantageous to estimate support as a contin-
in that case. uous trend.
Moreover, in a recent attempt to apply the DLM in the With this in mind, let us turn to an application of the DLM to the
multiparty system of Norway (Stoltenberg, 2013), Stoltenberg elections that took place in 2005, 2009 and 2013. In Fig. 3 we can
develops a new DLM assuming a multinomial distribution. He see the point predictions for Election Day in those three elections.
applies both the standard, Gaussian model and his own multi- These predictions are a snapshot from the time trends in Fig. 2
nomial model to the Norwegian election of 2013 and gets very where have zoomed in on the relevant dates but only used data
similar results. This suggests that we have both theoretical and up until the day before the election. The added confidence intervals
practical reasons to assume that the DLM can be applied also in show the uncertainty of the estimate.
multiparty systems. The overall impression is that the model works well. All results
From a practical standpoint we need to deal with unreliability (represented by the grey dots) fit within the 95% confidence in-
in the polls stemming from bias associated with particular polling tervals (although there are a few close cases, such as FDP in 2005
houses according to the regression method outlined in equation and CDU/CSU in 2013), and the average error is around 0.69 per-
(5). This was done in both the Swedish and German cases. In total centage points. The model seems to work well also for parties that
two polling companies in Sweden (Skop and United Minds) and have displayed large variation over time (such as CDU and the
three in Germany (FGW, Forsa and INSA) were found to deviate Greens) and performs equally well in 2005 and 2013 despite the
consistently from our DLM estimate and were weighted down in slightly lower availability of polls in 2005. 2009 was the best
accordance with their average deviance (generally between 5 and election for the model with an average error of only 0.5.
10%). Another promising result is that the two new parties in 2013 did
not seem to be more difficult to predict. For AfD there was only
4.1. Applying the DLM to Germany polling available from April 2013 onwards, which meant that the
model had only around 5 months of solid data to create the esti-
Let us now turn to an empirical test of whether our theoretical mates. The actual result of AfD still fit comfortably within the 95%
belief in the suitability of the DLM in the multiparty case is also confidence interval which suggests that the DLM works reasonably
practically justified. We start by using the standard DLM that well even when a new party emerges late in the process and data is
doesn't control for seasonality. Here the predictions were made on scarce.
the day before the elections, whereas forecasts with better lead A final thing to note is that the uncertainty in the estimates
time will be made in the next section. stems not only from the number of polls available, but also from
First, let us turn to the three German elections that took place how much estimates from different polling houses vary and from
in 2005, 2009 and 2013. There were five main parties competing how much support levels of the party have changed over time. For
in 2005 and 2009 and two more (die Piraten and AfD) in 2013. example, the confidence intervals of the Social Democrats are
Fig. 2 gives an overview of how popular support for the five main noticeably wider in 2013 than in 2005 since in 2013 the party was
parties has developed since 2005 according to the DLM fluctuating both up and down and the polling companies were
estimates. less in agreement about the direction in which the party was
The developments in party support over time show that there going.
has been fairly extensive fluctuations. CDU/CSU went from a high of
close to 50% a few months before the 2005 election to just over 30% 4.2. Applying the DLM to Sweden
in 2007. Some of the smaller parties, especially the Green party,
show even larger fluctuations relative to their size. This nicely il- With these results at hand, let us now undertake a second
lustrates the large variations in party support in multiparty systems empirical test by applying the model to Sweden. The Swedish party
system is even more complicated with 9 viable parties competing The party predictions in the three latest elections are available
for power in the latest election in 2014. in Fig. 4. Again all the actual results fit within the 95% confidence
To make matters worse, the four right-wing parties have since intervals, but with a few close calls such as the Sweden Democrats
2004 stood jointly in the elections as one right-wing bloc known (SD) and Greens (MP) in 2014 and the Centre Party (C) in 2006.
as “the Alliance”. This means that the parties have entered elec- The average prediction error was 0.78 and thus slightly higher
tions both as individual actors and as parts of larger pre-electoral than in the German case. However, the error for the two first
coalitions. Such a set-up made it necessary for voters to also think elections was only 0.63 percentage points on average so the
strategically about who they favoured in the left vs. right horse slightly poorer performance here is in large parts driven by the
race for power since two clear government alternatives were unexpected results in 2014 from the Sweden Democrats and the
present. In order to capture this two forecasts have been made in Greens.
the Swedish case e one at the party level, the other at the level of Interestingly, many of the prediction errors were in the same
the blocs. direction for each election. The large parties were always
D. Walther / Electoral Studies 40 (2015) 1e13 9
underestimated, whereas most of the smaller ones tended to be 5. Adding a seasonal component to capture re-occurring
overestimated. Moreover, the trend was consistent for many of the trends
individual parties. For example, the results of the Greens and the
Liberals (FP) were noticeably lower than the forecasts in all the The applications above show that the DLM is a feasible option
elections while the Social Democrats performed significantly bet- for making forecasts in multiparty systems. The average error is
ter. All in all though, the standard DLM with the set-up and weights relatively small and all the information needed to make the forecast
used here seems to work reasonably well even with 9 active players is available before the election. But the forecasts were made using
to forecast. polling data up until the day before the election, and in real world
The model generally works even better at the bloc level, as we applications we would ideally be able to say something meaningful
can see in Fig. 5. Here the average error was only 0.6, even though at an earlier stage.
the confidence intervals for the 2014 election left-wing side missed As argued above, one way of improving the early forecasts is
the actual result by a whisker. This means that the total error in the to include a seasonal component in the model. If we know that
left vs. right horse race for power is far lower than the sum of the there are predictable swings and that the parties are likely to
individual errors. This is an encouraging result since the issue of revert to a particular level, we should be able to foreshadow this
who will get control of the reins of the government after the trend through seasonal adjustment before it is visible in the polls
election is usually a key question to answer in pre-electoral (Jain, 2001). But the extent to which this boosts accuracy is tied
forecasts. to the magnitude of the seasonal trends. As we saw in Table 1,
The greater precision at the bloc level here suggests that at least controlling for seasonality appears to improve model accuracy
some of the uncertainty in the model comes from voters who for most parties generally, but slightly more so in the Swedish
switch between parties within the same bloc. It is likely that some case.
voters first make a general decision about which overarching bloc In the models shown in Figs. 6 and 7, a second prediction of the
to support, and then at a later stage make decisions about which German election in 2013 and Swedish election in 2014 has been
specific party to give their vote to.4 made, but this time one month before the elections took place. For
Pre-electoral coalitions are quite common in parliamentary comparison, the model with the seasonal component is compared
democracies. Golder (2006) finds that out of the 292 elections in and contrasted with the prediction from the standard DLM model
her study, 44% had at least one pre-electoral coalition. Around 1/4 also used in the preceding section. The standard DLM basically
of the governments that formed were based on some kind of pre- employs its current estimation of the true level of the support for
electoral agreement. If the DLM generally works better when in- each party as its election forecast, and is thus the ‘nowcast’ as it
dividual party forecast errors cancel out on the aggregate, bloc stands one month before the election.
level, this could be a useful strength that could help improve In Fig. 6 we can see that controlling for seasonality leads to a
predictions. measurable, but inconsistent, improvement in the accuracy of the
early forecast in Germany. The average error of 1.52 percentage
points is about 8% lower than the error of the standard DLM. The
main improvement came from the seasonal model's ability to
correctly foreshadow that the Greens would revert to a lower level
4
At least in the 2014 Swedish election, there was quite a lot of movement among than polling up until that point had suggested. For most other
voters between the parties within the two blocs, but hardly any movement be-
parties the difference between the models was small and for the
tween the blocs: [Link]
ValuResultat_riksdagsval_2014_PK_0914.pdf. Left party the seasonal model nudged the prediction slightly in the
10 D. Walther / Electoral Studies 40 (2015) 1e13
Fig. 6. Comparison of predictions with and without a seasonal component for the German election in 2013.
wrong direction compared to the standard DLM. So overall we can election, the DLM with the seasonal component offered notice-
see a clear, but not resounding, improvement when incorporating able improvements and was able to predict the result of the
the seasonal adjustment into the model. seven parties with an error of just over 1 percentage point per
So will forecasts for the Swedish parties, that on average party.
displayed more noticeable seasonal trends, benefit more from Taken together, the results in Germany and Sweden suggest that
adjusting for seasonality? The results in Fig. 7 seem to suggest the DLM does benefit from seasonal adjustment of the time series.
that they do. The model with the seasonal component performs And the stronger the historical seasonal trends in the party system,
noticeably better. The mean error for the standard DLM is 1.28 the larger the benefit. Including a seasonal component thus seems
percentage points and this is cut by around 17% by the model to be one way of improving the early forecasting capacity of the
with the seasonal component to 1.06. Like in the German case, DLM.
seasonal adjustment seems to be particularly useful in cases However, there are also many other ways to improve the stan-
where the standard DLM was far off, e.g. for the Greens (MP) and dard DLM. In cases where economic and political fundamentals
the Moderates. For the Greens, for example, the standard DLM play a predictable role, the DLM can be combined with regression
missed by 3.5 percentage points, but this is cut to 2.1 when we ro
terms for forecasts with better lead time (Je ^ me et al., 2013). This
control for seasonality. In total then, one month before the should be true especially for the larger parties.
Fig. 7. Comparison of predictions with and without a seasonal component for the Swedish election in 2014.
D. Walther / Electoral Studies 40 (2015) 1e13 11
Another way to improve the forecast could be to tie the pre- Appendix
dictions for each party more closely together and carry out joint
estimations. Since the results for all parties, and the ‘other’ cate- 1. Technical model details
gory, have to sum to 100 this can be placed as an overarching
constraint so that the parties are modelled through a joint multi- The models here were estimated in R using the Bayesian Struc-
variate time series. Significant improvements are thus still waiting tural Time Series Package (Scott, 2014; Scott and Varian, 2014) and
to be made to squeeze even more knowledge out of the data Markov Chain Monte Carlo simulations. 10 000 simulations were
available to us. made for each party, with the first 1000 being discarded as ‘burn-in’
(Jackman, 2009; Kruschke, 2010). The estimated level of party sup-
6. Conclusion port at any particular time point is taken to be the mean of the
posterior distribution for the 10 000 simulations. One prior was
The goal of this paper has been to apply methods and insights consciously specified at the outset, namely the sigma term in
from the extensive literature on electoral forecasting to the ‘diffi- equation (1). This was given an inverse gamma prior corresponding
cult’ multiparty cases of Germany and Sweden. Having a wide range to the square root of the sample size in the poll (after having adjusted
of competing actors and significant changes in electoral fortunes for polling house reliability). The other terms outlined in equations
between elections, these countries provide fruitful testing grounds (2), (3) and (6) (the seasonal adjustment), all assumed to be normal,
for popular forecasting techniques. were given vague inverse gamma priors that were uninformative at
Theoretical arguments were presented against the use of the outset. Trace plots and Geweke diagnostics were calculated to
structural models in the multiparty case. Unlike systems with fewer check that the chains converged. These are available below.
competing parties, in multiparty systems it is often difficult to Another technical decision is what length to assign to the sea-
assign economic and political responsibility to individual cabinet sonal component. In equation (6) we have both the season (S) and
parties (Anderson, 2000; Goodhart, 2009) and it is unclear how, if the sub-period (L) that need to be specified. S is simply the term of
at all, the various opposition parties will be affected. The frequent office which is the same for all parties. The smaller sub-periods in
emergence of completely new parties for which we have no prior contrast are what we use to capture the seasonal fluctuations and
data also make it difficult to accurately forecast elections using this these are calculated based on how quickly cyclical changes in party
approach. support are estimated to occur. Like with any time series (unless we
Instead the polling based dynamic linear model was applied. have strong theoretical reasons to expect a certain dynamic), the
With average errors of 0.68 percentage points in Germany and precise length of the seasonal component should be empirically
0.78 in Sweden in the three latest elections, the application of the estimated on past data to see which level fits best with the actual
DLM shows that forecasts with reasonable accuracy can be made development of the series. Here seasonal components were chosen
also in volatile multiparty settings. In the Swedish case, where that maximized the historical performance of the model up until
the parties in most cases entered the elections as parts of pre- the time when the prediction was made (i.e. the components that
electoral coalitions, separate forecasts were made on the bloc maximized R2 in Table 1). The periods were found to generally last
level. Here a significant share of the prediction errors for indi- between 2 and 6 months.
vidual parties cancelled out since some of the uncertainty in the
predictions is between parties within the same ideological bloc. 2. Abbreviations of party names
This approach can also be adopted in other multiparty systems
where pre-electoral coalitions take place. Germany
Finally, in order to improve the early predictions from the DLM,
inspiration was drawn from the economics literature and a seasonal CDU/CSU ¼ Christian Democratic Union/Christian Social Union
component was added to the model. This term can capture pre- SPD ¼ Social Democrats (Sozialdemokratische Partei
dictable cyclical fluctuations in party support, and incorporating Deutschlands)
seasonal adjustment into the DLM helped improve the forecast in Gruene ¼ Green Party (die Grünen)
both Sweden and Germany. A forecast one month before the 2014 FDP ¼ Liberals (Freie Demokratische Partei)
election from a standard DLM and from a model with the seasonal Linke ¼ Left Party (die Linke)
component added shows that the component taps into seasonal Piraten ¼ Pirate Party
patterns and cuts the error rate by 8% in Germany and 17% in AfD ¼ Alternative for Germany (Alternative für Deutschland)
Sweden.
All in all, the results above are encouraging since they Sweden
demonstrate that reasonable forecasts are possible also in a
multiparty system when a DLM is used. This is an interesting S ¼ Social Democrats
finding in its own right that can be applied not only to party V ¼ Left party (V€ansterpartiet)
support, but also to other types of regularly occurring polls MP ¼ Greens (Miljo € partiet)
where we want to make as much as possible of the scarce data M ¼ Moderates
available to us. To make the model even more practically inter- FP ¼ Liberals (Folkpartiet)
esting, though, in future iterations it needs to do more in the way C ¼ Centrist party
of explaining why these particular results are likely to occur by KD ¼ Christian Democrats
also looking into causes, potentially by incorporating regression SD ¼ Sweden Democrats
terms. FI ¼ Feminist Initiative
Acknowledgements
3. Traceplots and Geweke diagnostics
I gratefully acknowledge the support of the Marianne and
Marcus Wallenberg Foundation (MMW 2011.0030) that made this The traceplots are from the final draws of the MCMC chains in
research possible. the 2013 and 2014 elections.
12 D. Walther / Electoral Studies 40 (2015) 1e13
References Foucault, M., Nadeau, R., 2012. Forecasting the 2012 french presidential election. PS:
Polit. Sci. Polit. 45 (02), 218e222.
Freden, A., 2014. Threshold insurance voting in pr systems: a study of voters stra-
Abramowitz, A.I., 2008. Forecasting the 2008 presidential election with the time-
tegic behavior in the 2010 swedish general election. J. Elect. Public Opin. Parties
eforechange model. PS: Polit. Sci. Polit. 41 (04), 691e695.
24 (4), 473e492.
Anderson, C.J., 2000. Economic voting and political context: a comparative
Gibson, R., Lewis-Beck, M.S., 2011. Methodologies of election forecasting: calling the
perspective. Elect. Stud. 19 (2), 151e170.
2010 uk hung parliament. Elect. Stud. 30 (2), 247e249.
Bartels, L.M., Zaller, J., 2001. Presidential vote models: a recount. Polit. Sci. Polit. 34
Golder, S.N., 2006. Pre-electoral coalition formation in parliamentary democracies.
(01), 9e20.
Br. J. Polit. Sci. 36 (02), 193e212.
Bell, W.R., Hillmer, S.C., 1984. Issues involved with the seasonal adjustment of
Goodhart, L., 2009. Context, Clarity, and Signals: Economic Voting for Political
economic time series. J. Bus. Econ. Stat. 2 (4), 98e127.
Parties (Unpublished manuscript).
Bengtsson, A., 2004. Economic voting: the effect of political context, volatility and
Graefe, A., 2015. German election forecasting: comparing and combining methods
turnout on voters assignment of responsibility. Eur. J. Polit. Res. 43 (5), 749e767.
for 2013. German Polit. 24 (2), 195e204.
Carlsen, F., 2000. Unemployment, inflation and government popularityare there
Harrison, J., West, M., 1997. Bayesian Forecasting and Dynamic Models. Springer,
partisan effects? Elect. Stud. 19 (2), 141e150.
New York.
Fisher, S.D., Ford, R., Jennings, W., Pickup, M., Wlezien, C., 2011. From polls to votes
Healy, A., Lenz, G.S., 2014. Substituting the end for the whole: why voters respond
to seats: forecasting the 2010 british general election. Elect. Stud. 30 (2),
primarily to the electionyear economy. Am. J. Polit. Sci. 58 (1), 31e47.
250e257.
D. Walther / Electoral Studies 40 (2015) 1e13 13
Hibbs Jr., D.A., 2000. Bread and peace voting in us presidential elections. Public Nadeau, R., Lewis-Beck, M.S., Blanger, E., 2013. Economics and elections revisited.
Choice 104 (1e2), 149e180. Comp. Polit. Stud. 46 (5), 551e573.
Hyndman, R.J., Athanasopoulos, G., 2014. Forecasting: Principles and Practice. Narud, H.M., Valen, H., 2008. Coalition membership and electoral performance. In:
OTexts. Strøm, K., Müller, W., Bergman, T. (Eds.), Cabinets and Coalition Bargaining - the
Jackman, S., 2005. Pooling the polls over an election campaign. Aust. J. Polit. Sci. 40 Democratic Life Cycle in Western Europe. Oxford University Press, Oxford.
(4), 499e517. [Link] Norpoth, H., Gschwend, T., 2010. The chancellor model: forecasting german elec-
Jackman, S., 2009. Bayesian Analysis for the Social Sciences, vol. 846. John Wiley & tions. Int. J. Forecast. 26 (1), 42e53.
Sons. Paldam, M., 1991. How robust is the vote function? a study of seventeen nations
Jain, R.K., 2001. State space model-based method of seasonal adjustment. a. Mon. over four decades. Econ. Polit. Calc. Support 9e31.
Lab. Rev. 124, 37. Petris, G., Petrone, S., Campagnoli, P., 2009. Dynamic Linear Models with R. Springer,
ro
Je ^ me, B., Je
ro^me-Speziari, V., Lewis-Beck, M.S., 2013. A political-economy forecast New York.
for the 2013 german elections: who to rule with angela merkel? PS: Polit. Sci. Pickup, M., Johnston, R., 2007. Campaign trial heats as electoral information: evi-
Polit. 46 (03), 479e480. dence from the 2004 and 2006 canadian federal elections. Elect. Stud. 26 (2),
Jungerstam-Mulders, S., 2006. Post-communist EU Member States: Parties and 460e476.
Party Systems. Ashgate Publishing Company. Saalfeld, T., 2008. Institutions, chance and choices: the dynamics of cabinet survival.
Kayser, M.A., Leininger, A., 2013. A Benchmarking Forecast of the 2013 Bundestag In: Strøm, K., Müller, W., Bergman, T. (Eds.), Cabinets and Coalition Bargaining -
Election. Unpublished manuscript. Available at: [Link] the Democratic Life Cycle in Western Europe. Oxford University Press, Oxford.
Kitagawa, G., Gersch, W., 1984. A smoothness priorsstate space modeling of time Sanders, D., 1991. Government popularity and the next general election. Polit. Q. 62
series with trend and seasonality. J. Am. Stat. Assoc. 79 (386), 378e389. (2), 235e261.
Kruschke, J., 2010. Doing Bayesian Data Analysis: a Tutorial Introduction with R. Scott, S.L., 2014. Bayesian Structural Time Series in R. [Link]
Academic Press. packages/bsts/[Link].
Lewis-Beck, M.S., 2005. Election forecasting: principles and practice. Br. J. Polit. Int. Scott, S.L., Varian, H.R., 2014. Predicting the present with bayesian structural time
Relat. 7 (2), 145e164. series. Int. J. Math. Model. Numer. Optim. 5 (1), 4e23.
Lewis-Beck, M.S., Bellucci, P., 1982. Economic influences on legislative elections in Severini, T.A., 2005. Elements of Distribution Theory, vol. 17. Cambridge University
multiparty systems: France and italy. Polit. Behav. 4 (1), 93e107. Press, Cambridge.
Lewis-Beck, M.S., Paldam, M., 2000. Economic voting: an introduction. Elect. Stud. Silver, N., 2012. The Signal and the Noise: Why So Many Predictions Fail-but Some
19 (2), 113e121. Don't. Penguin.
Linzer, D.A., 2013. Dynamic bayesian forecasting of presidential elections in the Stoltenberg, E., 2013. Bayesian Forecasting of Election Results in Multiparty Sys-
states. J. Am. Stat. Assoc. 108 (501), 124e134. tems. Unpublished manuscript. Available at: [Link]
Lock, K., Gelman, A., 2010. Bayesian combination of state polls and election fore- 10852/36975?show¼full.
casts. Polit. Anal. 18 (3), 337e348. Sundell, A., Lewis-Beck, M.S., 2014. Forecasting the 2014 Parliamentary Election in
Mitchell, P., Nyblade, B., 2008. Government formation and cabinet type. In: Sweden. Available at: SSRN 2450229.
Strøm, K., Müller, W., Bergman, T. (Eds.), Cabinets and Coalition Bargaining - the
Democratic Life Cycle in Western Europe. Oxford University Press, Oxford.