Volatility Modeling with ARCH and GARCH
Volatility Modeling with ARCH and GARCH
Introduction
In basic univariate time series analysis we check whether the series is stationary or not. i.e. whether
the unconditional mean variance and covariance are time invariant or not. If it is stationary we
formulate different models like AR, MA, ARMA, otherwise we try to make the series meaningfully
stationary for modeling the conditional mean. So in univariate time series modeling we consider
the conditional mean changes over time. For predicting conditional mean we consider the lagged
values of the series as predictors and or other series. We estimate and predict the conditional mean
of the series assuming the disturbance is pure white noise and conditionally homoscedastic. These
assumptions are also true for a disturbance in standard regression model. In practice however
disturbances are heteroscedastic then we detect the nature of heteroscedasticity and estimate the
parameter addressing heteroscedasticity in cross section study. However, in time series we are
interested to estimate and forecast the conditional variance of a series which changes over time.
i.e. in time series analysis we study conditional heteroscedasticity which is auto correlated.
What is volatility?
The changes of fluctuation in the series are referred to as volatility. It is a situation where variance
in the process is not constant and keep changing over time. So, volatility means the changing
variance conditional on time or lagged values of the variable. It is evident that short run (local)
variance of economic and financial time series changes over time. Moreover, in many cases we
observe a period of relative tranquility followed by a period of relatively high volatility The
evidence of the phase of high volatility followed by the phase of low volatility is called volatility
clustering. We see the graph of weekly oil price changes during January 1986 to March, 2020 in
figure 1 as an example of volatility clustering.
Figure 1 Change of volatility for the rate change of oil spot price (weekly data)
6
rate of change of oil price
4
2
($ per barrel)
0
-2
-4
-6
Jan Jan Jan Jan Jan Jan Jan Jan Jan Jan Jan Jan Jan Jan Jan Jan Jan Jan
03, 03, 03, 03, 03, 03, 03, 03, 03, 03, 03, 03, 03, 03, 03, 03, 03, 03,
1986 1988 1990 1992 1994 1996 1998 2000 2002 2004 2006 2008 2010 2012 2014 2016 2018 2020
Dates
1
Rate of change of oil spot price =100*(log (Pt)-log (Pt-1)) =gt
𝑙𝑛(𝑥+1)
[We know lim = 1 ⇒ ln(𝑥 + 1) = 𝑥 𝑤ℎ𝑒𝑛 𝑥 𝑖𝑠 𝑠𝑚𝑎𝑙𝑙 𝑒𝑛𝑜𝑢𝑔ℎ
𝑥→0 𝑥
𝑝𝑡 −𝑝𝑡−1 𝑝𝑡
By definition 𝑔𝑡 = =𝑝 − 1 = 𝑥(𝑠𝑎𝑦)
𝑝𝑡−1 𝑡−1
𝑝𝑡 𝑝𝑡
Now ln(𝑥 + 1) = 𝑥 ; ⇒ ln (𝑝 )=𝑝 − 1= gt (approximate growth rate)
𝑡−1 𝑡−1
2
In example of the graph of weekly return of google, we observe the volatility clustering and we see
that the distribution is approximately symmetric but it has high peak and fat tails relative to the
normal distribution. Therefore, thick tails, and leptokurtosis is stylized fact of time series.
Fig-3 Weekly values of the spot prices of Oil (May 15, 1987 –November 1, 2013
Spot
160
140
120
100
80
60
40
20
0
May 15, 1988
In some cases shocks to a series can display a high degree of persistence. In figure 4 shows the
movements of short term and long term interest rates in US during 1960Q1- 2012 Q4. Neither of the
rates has a clear upward nor downward trend. Nevertheless both show a high degree of persistence. Notice
that T-bill rate experienced two upward surges in the 1970s and remained at these high level for several
years. Similarly after a sharp decrease in in the late 1980s, the rate never again displayed the level attained
in the early 1980s. Moreover, in the figure we have observed a clear co-movement of the series. Therefore,
we can expect there is co-movement of the volatility of the series.
Figure 4 Short term and long term interest rates in US during 1960Q1- 2012 Q4
16
14
12
10
0
60 65 70 75 80 85 90 95 00 05 10
TBILL R5
3
iv) Leverage Effects, it is a common feature of financial series like return of an asset, refers to the
tendency for changes in stock prices to be negatively correlated with changes in volatility. i.e if
stick prices increase today volatility will reduce tomorrow and vice versa.
Therefore, in the process to estimation and forecasting volatility we should be careful about these
feature of the series.
In order explain these nature of the economic time series econometricians have been developing
different model for estimating and forecasting volatility of the series. Besides, estimation and
forecasting of the local variance of many financial time series like return of assets and of many
macroeconomic series like inflation rate, interest rates, and exchange rates are of interest for several
reasons.
First, the variance of an asset price is a measure of the risk of holding the asset. The larger variance
of daily stock price changes indicates the higher risk of holding the asset, i.e. an asset holder may
gain or lose heavily in a particular day. So a risk averse investor would be less likely to participate
in the in the stock market during a period of high, rather than low, volatility.
Second: forecasting the variance makes it possible to have accurate forecast intervals for mean
prediction. For example, firms and labour union needs to forecast the inflation rate over the
duration of labour contract. Economic theory suggests that terms of wage contract will depend on
the inflation forecast and the uncertainty concerning the accuracy of these forecasts. If the parties
to the contract have rational expectation, the terms of contract depends on conditional mean and
conditional variance of inflation rate as opposed to the unconditional mean and unconditional
variance.
Thirdly, the ability to forecast financial market volatility is important for portfolio selection and
asset management as well as for the pricing of primary and derivative assets (Engle and Ng, 1993).
4
However, in time series analysis there may be reason to believe that the variance of disturbance
changes over time, particularly depends on the magnitude of past errors. In univariate time series
analysis we have believed that series evolve depending on its own past values and past shocks.
Even in some cases time series like daily changed of oil prices, rate of daily return of an assets, are
completely unpredictable. So there is no question of explanatory variable/predictors that are used
for estimation and predict the conditional mean. In these applications there is often evidence of a
clumping of large and small errors over time. In figure we have observed period of relatively high
volatility followed by periods of relative tranquility. Thus in time series heteroscedasticity refers
to the dependency of the variance of disturbance on the volatility of the errors in the recent
past. It is called conditional heteroscedasticity. Here, heteroscedasticity is auto correlated ie. We
deal with the auto correlated heteroscedasticity. We are interested to model and forecast
conditional volatility because in time series analysis long run or unconditional forecast is inferior
to conditional forecasts. Besides in some cases short run analysis is appropriate. Example includes
that a risk averse asset holder would be interested to forecast of the rate of return and its variance
over the holding period. The unconditional variance (i.e., the long-run forecast of the variance)
would be unimportant if he plans to buy the asset at time t and sell at t + 1. If the labour union
and firm have rational expectation, the terms of labour contract depends on conditional mean and
conditional variance of inflation rate not the unconditional mean and unconditional variance of
inflation rate.
5
If 𝑓𝑜𝑟 𝑎𝑙𝑙 𝑡, 𝑥𝑡−1= constant, the {𝜀𝑡 } sequence is the familiar white-noise process with a constant
variance. However, when the realizations of the {xt} sequence are not all equal, the variance of
𝜀𝑡 conditional on the observable value of 𝑥𝑡−1 is var(𝜀𝑡 |𝑥𝑡−1) =𝜎 2 𝑥𝑡−1
2
Here the conditional variance of 𝜀𝑡 is dependent on the realized value of 𝑥𝑡−1 . Since 𝑥𝑡−1at time
period t is observed, we can form the variance of 𝜀𝑡 conditionally on the realized value of 𝑥𝑡−1 . If
2
the magnitude 𝑥𝑡−1 is large (small), the variance of 𝜀𝑡 will be large (small) as well. Furthermore, if
the successive values of {xt} exhibit positive serial correlation (so that a large value of xt tends to
be followed by a large value of xt+1, the conditional variance of the {𝜀𝑡 } sequence will exhibit
positive serial correlation as well. In this way, the introduction of the {xt} sequence can explain
periods of volatility in the {𝜀𝑡 } sequence.
𝜀
Moreover, 𝐸( 𝑡⁄𝜀𝑡−1 , 𝜀𝑡−2 , … ) =0 ∵ 𝐸(𝑣𝑡 ) = 0
𝜀
and 𝑉𝑎𝑟( 𝑡⁄𝜀𝑡−1 , 𝜀𝑡−2 , … ) = 𝜎 2 𝑥𝑡−12
𝜀
Therefore, 𝑉𝑎𝑟( 𝑡⁄𝜀𝑡−1 , 𝜀𝑡−2 , … ) changes over time, but does not change when 𝜀𝑡−1changes,
rather changes when 𝑥𝑡−1 changes.
The major difficulty of the strategy is that the variance of the series 𝜀𝑡 given the past values does
not change when 𝜀𝑡−1changes. So the variance does not evolve over time with the process itself.
Oftentimes, we might not have a firm theoretical reason for selecting one candidate for the {xt}
sequence over other reasonable choices. Was it the oil price shocks, complete demonetisation,
and/or the introduction of GST that was responsible for the volatility of real investment in India
during the late 2010s?
Bilinear model
Let 𝜀𝑡 = 𝑣𝑡 𝜀𝑡−1 where 𝑣𝑡 is a white-noise disturbance term with constant variance 𝜎 2
𝜀
In this model 𝐸( 𝑡⁄𝜀𝑡−1 , 𝜀𝑡−2 , … ) =0 ∵ 𝐸(𝑣𝑡 ) = 0
𝜀
and 𝑉𝑎𝑟( 𝑡⁄𝜀𝑡−1 , 𝜀𝑡−2 , … ) = 𝜎 2 𝜀𝑡−1
2
𝜀
Therefore, 𝑉𝑎𝑟( 𝑡⁄𝜀𝑡−1 , 𝜀𝑡−2 , … ) changes over time, when 𝜀𝑡−1changes. So the variance evolves
over time with the process itself.
The problem with this model is that the unconditional variance of 𝜀𝑡 is either zero or infinite.
6
Check: Var (𝜀𝑡 )= 𝜎 2 𝑉𝑎𝑟(𝜀𝑡−1 )
Var(𝜀𝑡 )(1 − 𝜎 2 ) = 0
= infinite when 𝜎 2 = 1
This result does not associated with practical nature of time series process. That is why this strategy
is not a popular for modelling volatility.
7
Var(𝜀𝑡 ) = 𝐸(𝜀𝑡2 )
𝑛𝑜𝑤, 𝐸(𝜀𝑡2 ) = 𝐸(𝑣𝑡2 )𝐸(𝛼0 + 𝛼1 𝜀𝑡−1
2
)
Or, 𝐸(𝜀𝑡2 ) = 𝛼0 + 𝛼1 𝐸(𝜀𝑡−1
2
)
Since unconditional variance of 𝜀𝑡 is identical to that of 𝜀𝑡−1 , 𝐸(𝜀𝑡2 ) = 𝐸(𝜀𝑡−1
2 )
𝛼
∴ 𝐸(𝜀𝑡2 ) = 1−𝛼0 >0
1
Therefore, under the specification (2) 𝜀𝑡 is pure white noise. The specification of 𝜀𝑡 does not affect
the properties of pure white noise disturbance of the regression (mean) model. At this point one
may think that the specification (2) does not affect the properties of 𝜀𝑡 . However, under this
specification unconditional distribution of 𝜀𝑡 does not follow normal distribution. It has
leptokurtosis. We know that kurtosis of normal distribution is 3. Here,
2
𝜀𝑡 = 𝑣𝑡 √𝛼0 + 𝛼1 𝜀𝑡−1
3𝛼02 (1 + 𝛼1 )
⇒ 𝑚4 =
(1 − 𝛼1 )(1 − 3𝛼12 )
It implies the fourth moment of 𝜀𝑡 is finite if 3𝛼12 < 1.
𝑚4 3(1−𝛼12 )
Therefore, kurtosis = = >3 ∵ (1 − 3𝛼12 ) <(1 − 𝛼12 ) as 0 < 𝛼1 ≤ 1
𝑚22 (1−3𝛼12 )
It means that the unconditional distribution of 𝜀𝑡 is fatter tailed and leptokurtic than the normal
distribution. Therefore, ARCH specification is allowed in the process which follow t distribution
because t distribution has fatter tail. One of the stylized fact of financial time series is that
distribution of the series have a fatter tail compared to that of normal distribution.
Let us now check the conditional distribution of 𝜀𝑡 . The conditional mean of 𝜀𝑡 given the past
𝜀
value of the series is zero. That is, 𝐸( 𝑡⁄𝜀𝑡−1 , 𝜀𝑡−2 , … ) = 0 since 𝐸𝑡−1 (𝑣𝑡 ) = 0
8
However, the influence of (2) entirely fall on the conditional variance of 𝜀𝑡
Engle (1982) shows that conditional variance of 𝜀𝑡
𝜀2
𝜎𝑡2 = ℎ𝑡 = 𝐸 ( 𝑡 ⁄𝜀𝑡−1 , 𝜀𝑡−2 , … ) = 𝛼0 + 𝛼1 𝜀𝑡−1
2
(3)
where 𝑤𝑡 = 𝜀𝑡2 − ℎ𝑡
Equation (3a) is nothing but a AR(1) model of 𝜀𝑡2 , therefore model (3) or (3a) is called ARCH(1).
We should check whether 𝑤𝑡 in PWN or not. Here, 𝑤𝑡 is i.i.d with E (𝑤𝑡 )=0 and Var (𝑤𝑡 )= 𝜆2
and Cov (𝑤𝑡 , 𝑤𝑡−1 ) = 0 and 𝑤𝑡 ≥ −𝛼0
We have 𝑤𝑡 = ℎ𝑡 (𝑣𝑡2 − 1)
𝑤𝑡2⁄ 2 2 2 2
Var (𝑤𝑡 )= 𝐸 ( 𝜀𝑡−1 , 𝜀𝑡−2 , … ) = 𝐸(ℎ𝑡 )𝐸(𝑣𝑡 − 1) = 𝜎𝑤 (𝑠𝑎𝑦)
∴ 𝐸(ℎ𝑡2 )= 𝛼0 2 + 𝛼1 2 𝐸(𝜀𝑡−1
4 2
) + 2𝛼0 𝛼1𝐸(𝜀𝑡−1 )
3 𝛼 2 (1+𝛼 ) 2𝛼02 𝛼1
𝛼0 2 + 𝛼1 2 (1−𝛼0 )(1−3𝛼1 2) +
1 1 1−𝛼1
9
= 3+1-2 =2
(𝟏+𝜶𝟏 )
∴ Var (𝒘𝒕 )= 𝝈𝟐𝒘 =2 𝜶𝟎 𝟐 [(𝟏−𝜶 𝟐 ] Constant
𝟏 )(𝟏−𝟑𝜶𝟏 )
Cov (𝒘𝒕 , 𝒘𝒕−𝟏 )= 𝑬[(𝒉𝒕 , 𝒉𝒕−𝟏 ) [𝟏 − 𝟏 − 𝟏 + 𝟏]=0 ∵ 𝑣𝑡 𝑎𝑛𝑑 𝑣𝑡−1 𝑎𝑟𝑒 𝑖𝑛𝑑𝑒𝑝𝑒𝑛𝑑𝑒𝑛𝑡 𝑆𝑁𝐷,
2 2
𝑣𝑡2 𝑎𝑛𝑑 𝑣𝑡−1
2
𝑎𝑟𝑒 𝑖𝑛𝑑𝑒𝑝𝑒𝑛𝑑𝑒𝑛𝑡 𝜒(1) 𝑣𝑎𝑟𝑖𝑎𝑡𝑒, 𝑎𝑛𝑑 𝐸(𝜒(𝑛) )=𝑛
Therefore, 𝑤𝑡 is a white noise process. It ensures that 𝜀𝑡2 is a AR(1) process. That is why; equation
(3) and (3a) is called ARCH (1) (Autoregressive Conditionally Heteroscedastic model of order1).
In an ARCH model, the conditional and unconditional expectations of the error terms are equal
to zero. Moreover, the {𝜀t} sequence is serially uncorrelated because, for all s ≠ 0, E𝜀t𝜀t−s = 0.
The key point is that the errors are not correlated but dependent since they are related through
their second moment (recall that correlation is a linear relationship). Thus the dependency non-
linear by nature.
The conditional variance itself is an autoregressive process resulting in conditionally
2
heteroskedastic errors. When the realized value of 𝜀t−1 is far from zero—so that 𝛼1𝜀𝑡−1 is relatively
large—the variance of 𝜀t will tend to be large.
We can also prove that 𝜀𝑡2 is stationary under the specification in (3).
10
We have already proved that
𝛼
𝐸(𝜀𝑡2 ) = 1−𝛼0 it says unconditional mean of 𝜀𝑡2 is constant.
1
𝛼 2 𝜎2 2 𝛼0 2
Var (𝜀𝑡2 ) = 𝐸 (𝜀𝑡2 − 1−𝛼0 ) = 𝐸(𝑤𝑡2 (1 + 𝛼1 𝐿 + 𝛼12 𝐿2 … … )2 = 1−𝛼
𝑤
2=[(1−𝛼 2 2 ]
1 1 1) ((1−3𝛼1 )
(1+𝛼1)
Since, 𝜎𝑤2 =2 𝛼0 2 [(1−𝛼 2 ]
1)(1−3𝛼1 )
So the unconditional variance of 𝜀𝑡2 is time invariant and finite if 3𝛼12 < 1
Now, covariance of 𝜀𝑡2 𝑎𝑛𝑑 𝜀𝑡−1
2
𝛼0 𝛼0
𝐸 (𝜀𝑡2 − 2
) (𝜀𝑡−1 − )
1 − 𝛼1 1 − 𝛼1
= 𝐸(𝑤𝑡 + 𝛼1 𝑤𝑡−1 + 𝛼12 𝑤𝑡−2 + 𝛼13 𝑤𝑡−3 …)(𝑤𝑡−1 + 𝛼1 𝑤𝑡−2 + 𝛼12 𝑤𝑡−3 + 𝛼13 𝑤𝑡−4 …)
= 𝐸(𝑤𝑡 (𝑤𝑡−1 + 𝛼1 𝑤𝑡−2 + 𝛼12 𝑤𝑡−3 + 𝛼13 𝑤𝑡−4 …)+𝐸𝛼1 (𝑤𝑡−1 + 𝛼1 𝑤𝑡−2 + 𝛼12 𝑤𝑡−3 +
𝛼13 𝑤𝑡−4 … )2
2 2 𝛼1 𝛼0 2
= 0+𝛼1var (𝜀𝑡−1 )= (1−𝛼1 )2 ((1−3𝛼12 )
2 𝛼1 𝛼0 2
∴ Cov(𝜀𝑡2 , 𝜀𝑡−1
2
)= (1−𝛼 2 2
1 ) ((1−3𝛼1 )
2
2 𝛼1𝑠 𝛼0
Similarly, we can say Cov(𝜀𝑡2 , 𝜀𝑡−𝑠
2
)= (1−𝛼 2 2
1 ) (1−3𝛼1 )
Since the variance of 𝜀𝑡 in equation (3) depends only on last period’s volatility, we called the model
as ARCH (1). More generally, the variance could depend on any number of lagged volatilities.
Then ARCH (q) model can be written as
2 2 2 2
ℎ𝑡 = 𝛼0 + 𝛼1 𝜀𝑡−1 + 𝛼2 𝜀𝑡−2 +𝛼3 𝜀𝑡−3 + ⋯ + 𝛼𝑞 𝜀𝑡−𝑞 (4)
Equation (4) says that conditional variance of 𝜀𝑡 depends on shocks from 𝜀𝑡−1 𝑡𝑜 𝜀𝑡−𝑞 Note that
the conditional variance of 𝜀𝑡 depends on magnitude of the shocks from 𝜀𝑡−1 𝑡𝑜 𝜀𝑡−𝑞 not on their
signs as the equation 4 considers the squares of the past errors.
11
Moreover, the conditional heteroscedasticity in {𝜀t} will result in {yt} being heteroskedastic itself.
𝑦 𝑦 2
Conditional variance of 𝑦𝑡 𝑉𝑎𝑟( 𝑡⁄𝑦𝑡−1 , 𝑦𝑡−2 , … ) = 𝐸(𝑦𝑡 − 𝐸( 𝑡⁄𝑦𝑡−1 , 𝑦𝑡−2 , … ) =
𝜀2 2 2 2 2
𝐸 ( 𝑡 ⁄𝜀𝑡−1 , 𝜀𝑡−2 , … ) = 𝛼0 + 𝛼1 𝜀𝑡−1 + 𝛼2 𝜀𝑡−2 +𝛼3 𝜀𝑡−3 + ⋯ + 𝛼𝑞 𝜀𝑡−𝑞
Thus, the ARCH model is able to capture periods of tranquility and volatility in the {yt} series.
Therefore, ARCH types of model suitable for modelling volatility of financial series like returns
which can capture three important stylized facts of economic time series. (i) There is almost
no correlation between returns for different days. (ii) There is positive dependence between
absolute returns on nearby days and thus for squared returns (Taylor, 2005). (iii) returns do not
follow a normal distribution, in most cases the distribution is symmetric (Taylor, 2005), however,
in some cases, distribution is characterized by a asymmetric shape, normally skewned to the left
and high Kurtosis which implies that returns have a high peak (Mandelbrot, 1963) and fat tails
(Fama, 1965).
12
𝜌2
𝑄 = 𝑇(𝑇 + 2) ∑𝑛𝑖=1 𝑇−𝑖
𝑖 2
~𝜒(𝑛) Where, n= lag length (usually T/4)
we confirm that the 𝑦𝑡2 (𝜀𝑡2 ) is significantly auto correlated up to lag length n. then we suggest to
formulate ARCH(m).
Formal test
A formal Lagrange multiplier test for ARCH errors is the test by McLeod and Li (1983). The
methodology involves the following two steps:
STEP 1: Use OLS to estimate the most appropriate regression equation or ARMA model and let
{ 𝜀̂𝑡2} denote the squares of the fitted errors.
STEP 2: Regress these squared residuals on a constant and on the q lagged values i.e estimate the
following regression 𝜀̂𝑡2 = 𝛼0 + 𝛼1 𝜀̂𝑡−1
2 2
+ 𝛼2 𝜀̂𝑡−2 2
+𝛼3 𝜀̂𝑡−3 2
+ ⋯ + 𝛼𝑞 𝜀̂𝑡−𝑞 (5)
If there are no ARCH or GARCH effects, the estimated values of 𝛼1 through 𝛼q should be zero.
Hence, this regression will have little explanatory power so that the coefficient of determination
(i.e., the usual𝑅2 ) will be quite low. Using a sample of T residuals, under the null hypothesis of no
ARCH errors, the test statistic 𝑇𝑅2 converges to a 𝜒2 distribution with q degrees of freedom. If
𝑇𝑅 2 is sufficiently large, rejection of the null hypothesis that 𝛼1 through 𝛼q are jointly equal to zero
is equivalent to rejection of the null hypothesis of no ARCH errors. On the other hand, if 𝑇𝑅2 is
sufficiently low, it is possible to conclude that there are no ARCH effects. In the small sample
sizes typically used in applied work, an F-test for the null hypothesis 𝛼1 = · · · = 𝛼q = 0 has been
shown to be superior to a 𝜒2 test. Compare the sample value of F to the values in an F-table with
q degrees of freedom in the numerator and T − q degrees of freedom in the denominator.
13
a parsimonious model. Engle (1982), however, circumvented this problem by specifying an
arbitrary linearly declining lag length on an ARCH (4) model for quarterly data series as follows.
𝜎𝑡2 = 𝛾0 + 𝛾1 (0.4𝜀̂𝑡−1
2 2
+ 0.3𝜀̂𝑡−2 2
+ 0.2𝜀̂𝑡−3 2
+ 0.1𝜀̂𝑡−4 ) here only two parameters are required
in the conditional variance equation 𝛾0 and 𝛾1 , rather than the five which could be required for an
unrestricted ARCH(4) model.
Third, ARCH models over predict the volatility because they response slowly to large isolated
shocks to the series of interest.
Fourth, non-negative constraint might be violated, everything else equal, the more parameters
there are in the conditional variance equation, the more likely it is that one or more of them will
have negative estimated value.
Fifth, in ARCH model the conditional variance of 𝜀𝑡 depends on magnitude of the shocks from
𝜀𝑡−1 𝑡𝑜 𝜀𝑡−𝑞 not on their signs as the equation 4 considers the squares of the past errors. Thus
leverage effect, if it exists, cannot be explained with the basic ARCH model.
GARCH model reduces these problems, this is why GARCH models are widely used in practice.
The ARCH model has been generalised to allow linear dependence of the conditional variance, ℎ𝑡
on past values of ℎ𝑡 as well as on past squared values of the series. The GARCH model was
developed by Bollerslev (1986) and Taylor (1986).
𝑦𝑡 = 𝜇 + 𝛽𝑥𝑡 + 𝜀𝑡 (6)
2
Such that 𝜖𝑡 = 𝑣𝑡 (𝛼0 + 𝛼1 𝜖𝑡−1 + 𝛽1 ℎ𝑡−1 )0.5 it is called GARCH error process.
Using the GARCH (1,1) model it is possible to interpret the current fitted variance ℎ𝑡 as a weighted
function of a long run average value (dependent on 𝛼0), information about volatility during the
2
previous period 𝛼1 𝜖𝑡−1 (the ARCH term) and the fitted variance from the model during the
previous period 𝛽1 ℎ𝑡−1(the GARCH term). Note that the value of 𝛼1 must be strictly positive,
14
because if𝛼1 = 0, there is no way for the 𝜀𝑡 series to affect ℎ𝑡 . And value of 𝛽1 is expected to be
positive. Finally, 0 < 𝛼1 + 𝛽1 < 1 for stationarity condition of GRACH process. Large values of
both 𝛼1 𝑎𝑛𝑑 𝛽1 act to increase the conditional volatility, but they do so in different ways. The
larger is 𝛼1, the larger is the response of ℎ𝑡 to new information; clearly if 𝛼1is large, a 𝑣𝑡 shock has
a sizeable effect on 𝜖𝑡2 and ℎ𝑡+1. On the other hand, large 𝛽1 displays the response of ℎ𝑡 current
conditional variance, to the volume of past fitted conditional variance which captured the effects
of past shocks.
Note that GARCH (1, 1) model is effectively an ARMA (1,1)model for the squared residuals. To
see this consider squared residuals relative to the conditional variance is
𝑢𝑡 = 𝜖𝑡2 − ℎ𝑡
⇒ ℎ𝑡 = 𝜖𝑡2 − 𝑢𝑡
𝜖𝑡2 − 𝑢𝑡 = 𝛼0 + 𝛼1 𝜖𝑡−1
2
+ 𝛽1 ℎ𝑡−1
⇒ 𝜖𝑡2 = 𝛼0 + 𝛼1 𝜖𝑡−1
2
+ 𝛽1 ℎ𝑡−1 + 𝑢𝑡
⇒ 𝜖𝑡2 = 𝛼0 + 𝛼1 𝜖𝑡−1
2 2
+ 𝛽1 (𝜖𝑡−1 − 𝑒𝑡−1 ) + 𝑢𝑡
∴ 𝜖𝑡2 is ARMA(1,1) for GARCH (1, 1) error process with AR coefficient (𝛼1 + 𝛽1 ) (It is a good
exercise to show 𝑢 pure white noise. (Home task)
Let us now see that GARCH is more parsimonious compared to ARCH. It avoids overfitting.
Consequently, the model is less likely to breach non-negativity constraints. Consider the
conditional variance equations for GARCH (1,1) model
2
ℎ𝑡 = 𝛼0 + 𝛼1 𝜖𝑡−1 + 𝛽1 ℎ𝑡−1 .....................................(i)
2
⇒ ℎ𝑡−1 = 𝛼0 + 𝛼1 𝜖𝑡−2 + 𝛽1 ℎ𝑡−2 .......................(ii) for t=t-1
2
ℎ𝑡−2 = 𝛼0 + 𝛼1 𝜖𝑡−3 + 𝛽1 ℎ𝑡−3 ...............................(iii) for t=t-2
2 2
ℎ𝑡 = 𝛼0 + 𝛼1 𝜖𝑡−1 + 𝛽1 (𝛼0 + 𝛼1 𝜖𝑡−2 + 𝛽1 ℎ𝑡−2 )
2 2
= 𝛼0 + 𝛼1 𝜖𝑡−1 + 𝛽1 𝛼0 + 𝛽1 𝛼1 𝜖𝑡−2 + 𝛽12 ℎ𝑡−2
15
2 2
ℎ𝑡 = 𝛼0 + 𝛼1 𝜖𝑡−1 + 𝛽1 𝛼0 + 𝛽1 𝛼1 𝜖𝑡−2 + 𝛽12 (𝛼0 + 𝛼1 𝜖𝑡−3
2
+ 𝛽1 ℎ𝑡−3 )
2 2
= 𝛼0 + 𝛼1 𝜖𝑡−1 + 𝛽1 𝛼0 + 𝛽1 𝛼1 𝜖𝑡−2 + 𝛽12 𝛼0 + 𝛽12 𝛼1 𝜖𝑡−3
2
+ 𝛽13 ℎ𝑡−3
2 2
ℎ𝑡 = 𝛾0 + 𝛾1 𝜖𝑡−1 + 𝛾2 𝜖𝑡−2 + ⋯ … … … … .. ⇒ ARCH (∞) (8)
Thus GARCH (1,1) model containing only three parameters in the conditional variance equation,
is a very parsimonious model, that allows an infinite number of past squared errors to influence
the current conditional variance.
On the other side, we can show that ARCH (∞) can be written as GARCH (1, 1)
If we recognise that equation (8) is simply a distributed lag model for ℎ𝑡
2 2 2
ARCH (∞) ⇒ ℎ𝑡 = 𝛾0 + 𝛾1 𝜀𝑡−1 + 𝛾2 𝜀𝑡−2 + 𝛾3 𝜀𝑡−3 + ⋯ … … … ….
2
Let 𝛾𝑖 = 𝜆𝛾𝑖−1 ⇒ ℎ𝑡 = 𝛾0 + 𝛾0 𝜆𝜀𝑡−1 + 𝛾0 𝜆2 𝜀𝑡−2
2
+ 𝛾0 𝜆3 𝜀𝑡−3
2
+ ⋯………
2
Then ℎ𝑡−1 = 𝛾0 + 𝛾0 𝜆𝜀𝑡−2 + 𝛾0 𝜆2 𝜀𝑡−3
2
+ 𝛾0 𝜆3 𝜀𝑡−4
2
+ ⋯………
2
Finally, ℎ𝑡 − λℎ𝑡−1 = 𝛾0 (1 − 𝜆) + 𝛾0 𝜆𝜀𝑡−1
2
⇒ ℎ𝑡 = 𝛾0 (1 − 𝜆) + 𝛾0 𝜆𝜀𝑡−1 + λℎ𝑡−1
2
= 𝛾0 − 𝛾1 + 𝛾1 𝜀𝑡−1 + λℎ𝑡−1
2
= 𝑎0 + 𝛾1 𝜀𝑡−1 + λℎ𝑡−1
It is called a GARCH (1,1) process. It is a model of conditional variance where ℎ𝑡 depends on past
error’s squared and past values of conditional variance. Here we have only three coefficient
parameters to be estimated. Hence, GARCH is parsimonious than ARCH.
We can extend the GARCH (1, 1) model for any number (p) of GARCH terms and any number
(q) of ARCH terms. When we estimate a GARCH (p, q) process, we estimate the two interrelated
equations.
16
𝑦𝑡 = 𝑎0 + 𝛽𝑥𝑡 + 𝜀𝑡 ...............................................(a)
2 2
and 𝜖𝑡 = 𝑣𝑡 (𝛼0 + 𝛼1 𝜖𝑡−1 + ⋯ … + 𝛼𝑞 𝜖𝑡−𝑞 + 𝛽1 ℎ𝑡−1 + 𝛽2 ℎ𝑡−2 + ⋯ . . +𝛽𝑝 ℎ𝑡−𝑝 )0.5..........(b)
where 𝑥𝑡 can be an ARMA process of order (𝑝𝑚 , 𝑞𝑚 ). Moreover 𝑥𝑡 can contain exogenous
variables.
The first equation is a model of the conditional mean and the second yields the model for the
conditional variance. The systems 𝑝𝑚 and 𝑞𝑚 are used to denote the order of the ARMA process
for the mean need not equal the order of the GARCH (p, q) equation. ℎ𝑡 is the conditional variance
of the mean equation. From (b) we get-
2
Or ℎ𝑡 = 𝛼0 + 𝛼1 𝜖𝑡−1 + 𝛽1 ℎ𝑡−1
Unconditional variance of 𝜀𝑡
2
= 𝛼0 + 𝛼1 𝐸(𝜖𝑡−1 ) + 𝛽1 𝐸(ℎ𝑡−1 ) [Since 𝐸(𝑣𝑡2 ) = 1]
(1 − 𝛼1 )𝐸(𝜖𝑡2 ) = 𝛼0 + 𝛽1 𝐸(ℎ𝑡−1 )
17
[Alternative way: From the law of iterated expectations we know
⇒ 𝐸(𝜖𝑡2 ) = 𝐸(ℎ𝑡 )
∴ (1 − 𝛼1 − 𝛽1 )𝐸(𝜖𝑡2 ) = 𝛼0
𝛼0
⇒ 𝐸(𝜖𝑡2 ) = > 0 𝑤ℎ𝑒𝑛 𝛼0 > 0, 𝑎𝑛𝑑 𝛼1 + 𝛽1 < 1
(1 − 𝛼1 − 𝛽1 )
For more general model GARCH (𝑝, 𝑞) it follows that the variance will be finite if
𝑞 𝑝
1 − ∑ 𝛼𝑖 − ∑ 𝛽𝑖 > 0
𝑖=1 𝑖=1
1 1
𝐸(𝜖𝑡 𝜖𝑡−𝑖 ) = 𝐸 [𝑣𝑡 ℎ𝑡2 𝑣𝑡−𝑖 ℎ𝑡−𝑖
2
]=0
Therefore, under the GARCH specification error is pure white noise. So the errors are non auto
correlated. But we can’t say they are independent.
Moreover, under this GARCH specification unconditional distribution of 𝜀𝑡 does not follow
normal distribution. It has leptokurtosis. We know that kurtosis of normal distribution is 3. Here,
2
𝜀𝑡 = 𝑣𝑡 √𝛼0 + 𝛼1 𝜀𝑡−1 + 𝛽1 ℎ𝑡−1
4 4 2
Further, 𝜀𝑡−1 = 𝑣𝑡−1 ℎ𝑡−1
4 4 2 2 2 𝑚4
⇒ 𝐸𝜀𝑡−1 = 𝐸𝑣𝑡−1 𝐸ℎ𝑡−1 ⇒ 𝑚4 = 3 𝐸ℎ𝑡−1 ⇒ 𝐸ℎ𝑡−1 = 3
18
3𝛼02 [(1 − 𝛼1 − 𝛽1 )2 + 2(𝛼1 + 𝛽1 )(1 − 𝛼1 − 𝛽1 ) + 2𝛼1 𝛽1]
⇒ 𝑚4 =
(1 − 3𝛼12 − 𝛽12 )(1 − 𝛼1 − 𝛽1 )2
It implies the fourth moment of 𝜀𝑡 is finite if 3𝛼12 + 𝛽12 < 1.
𝑚 3[(1−𝛼1 −𝛽1 )2 +2(𝛼1 +𝛽1 )(1−𝛼1 −𝛽1 )+2𝛼1𝛽1 ]
Therefore, kurtosis = 𝑚42=
2 (1−3𝛼12−𝛽12 )
It means that the unconditional distribution of 𝜀𝑡 is fatter tailed and high peak than that of the
normal distribution.
This simple result is the essential feature of GARCH modelling. The conditional variance of the
error process is not constant. With the appropriate specification of the parameters of ℎ𝑡 , it is
possible to model and forecast the conditional variance.
Volatility persistence: In a GARCH the errors are uncorrelated as 𝐸(𝜖𝑡 𝜖𝑡−𝑖 ) = 0, However, the
squared errors of a GARCH (1,1) process are correlated. We know that squared residuals of
GARCH (1,1) is nothing but a ARMA (1,1) with AR coefficient (𝛼1 + 𝛽1 ). Thus from the
properties of ARMA (1,1) we say that ACF of the squared residuals is decaying (𝛼1 + 𝛽1 ) rate if
(𝛼1 + 𝛽1 ) <1. However, if it is close to unity ACF will not decay quickly. [See ARMA Processes].
So if (𝛼1 + 𝛽1 ) is almost equal to 1 we say the current volatility is persistent. In other words, we
say current level of conditional volatility is expected to prevail into the future.
It ensures that the GARCH model capture the important stylized fact of the financial time series.
19
time series analysis. First consider the sum of squared residuals (SSR) as a measure of the goodness
of fit. In a GARCH model, a reasonable measure of the goodness of fit is the sum of squares of
the {vt} sequence, i.e. sum of squares of the standardized residuals of the fitted GARCH model.
More concretely we can construct the AIC and SBC using
AIC=−2 ln L + 2n
SBC=−2 ln L + n ln(T)
where L is likelihood function and n is the number of estimated parameters. Smaller the value of
value of AIC or SBC higher the goodness of fit of the model.
In order to estimate the ARCH or GARCH models we apply maximum likelihood estimation
technique. We apply ML estimation technique since the ARCH and GARCH models are not in
linear form where OLS is not usable. Further to estimate the mean and variance equation
simultaneously. If we apply OLS first in the mean equation then estimate variance equation
considering the residuals then we cannot get efficient estimates of the parameters because in the
application of OLS we assume that variance of the disturbance term is constant. Let us now see
the ML method works for estimating ARCH or GARCH models.
Suppose the mean equation 𝑦𝑡 = 𝜇 + 𝛽𝑥𝑡 + 𝜀𝑡
2
Where 𝜀𝑡 = 𝑣𝑡 √ℎ𝑡 and ℎ𝑡 = 𝛼0 + 𝛼1 𝜖𝑡−1 + 𝛽1 ℎ𝑡−1
Initially suppose that values of {𝜀𝑡 } are drawn from a normal distribution having mean zero and
variance ℎ𝑡 which is not constant. Then from standard distribution theory the likelihood of any
realisation of 𝜀𝑡 is
1 −𝜖𝑡2
𝐿𝑡 = 𝑒𝑥𝑝 Where 𝐿𝑡 is the likelihood of 𝜀𝑡
√2𝜋ℎ𝑡 2ℎ𝑡
Since the realisation of 𝜀𝑡 are independent, the likelihood of the joint realisations of
𝜀1 𝜀2 … … … . 𝜀𝑇 is the product of the individual likelihoods. Hence likelihood of the joint
realisation is
𝑇
1 −𝜀𝑡2
𝐿=∏ exp( )
√2𝜋ℎ𝑡 2ℎ𝑡
𝑖=1
20
It is far easier to work with a sum than with a product. As such, it is convenient to take the natural
log of each side so as to obtain
𝑇 1 1 𝜖2
ln 𝐿 = − 2 ln(2𝜋) − 2 ∑𝑇𝑡=1 ln ℎ𝑡 − 2 ∑𝑇𝑡=1 𝑡⁄ℎ
𝑡
2 2
Where 𝜀𝑡 = 𝑦𝑡 − 𝜇 − 𝛽𝑥𝑡 and ℎ𝑡 = 𝛼0 + 𝛼1 𝜖𝑡−1 + 𝛽1 ℎ𝑡−1 for ARCH(1) ℎ𝑡 = 𝛼0 + 𝛼1 𝜖𝑡−1
Noe that as the models of variance need some lag value so we need some adjustment regarding
the sample size. Once we adjust the sample size it is possible to maximise ln 𝐿 with respect
to 𝜇, 𝛼0 , 𝛼1 𝛽1𝑎𝑛𝑑 𝛽. It is easy to check that there are no simple solutions to the first order
conditions for a maximum. Actually normal equations are found to be non-linear. Thus numerical
technique i.e iteration is required to maximise ln 𝐿, where first we assume 𝛼0 given find out
𝑡ℎ𝑒 𝑣𝑎𝑙𝑢𝑒𝑠 𝑜𝑓 𝑡ℎ𝑒𝑟 𝑝𝑎𝑟𝑎𝑚𝑒𝑡𝑒𝑟𝑠. Again the others parameters assuming given find out 𝛼0. The
process will continue until two successive values of the parameters converges. For ARCH &
GARCH model iterative technique is necessary to estimate the parameters. It is trickier process
actually for a higher order GARCH model where variance is time varying. In EViews we have two
algorithms available for optimisation; Berndt Hall, Hall and Hausman (1974) method and
Marquardt algorithm.
However, in GARCH model standardised residuals do not follow normal distribution. If the
normality assumption does not hold, the parameter estimation will still be consistent if the mean
and variance equation are correctly specified. However, in the context of non-normality, the usual
standard error estimates will be inappropriate. In that case we can use the alternative distribution
of the errors like t distribution or generalised exponential distribution option from EViews. Finally
to get the robust variance-covariance matrix estimator in the non-normality cases we apply
Bollerslev Wooldridge (1992) standard errors. It is also available in Eviews option. This procedure
i.e maximum likelihood with Bollerslev Wooldridge (1992) standard errors is known as quasi-
maximum likelihood or QLM.
21
ℎ̂𝑡0.5in order to obtain an estimate of what we have been calling the {vt} sequence. Since 𝜀t has a
𝜀̂𝑡
zero mean and a variance of ht, we can think of vt = ̂𝑡0.5 as the standardized value of 𝜀t. The
ℎ
resulting series, which we will call st, should have a mean of zero and a variance of unity. If there
is any serial correlation in the {st} sequence, the model of the mean is not properly specified. To
test the adequacy of the model of the mean, form the Ljung–Box Q-statistics for the {st} sequence.
If we cannot reject the null hypothesis that the autocorrelation coefficients of vt are equal to zero.
To test for remaining GARCH effects, form the Ljung–Box Q-statistics of the squared
𝜀𝑡2
standardized residuals (i.e., 𝑠𝑡2 ). The basic idea is that 𝑠𝑡2 is an estimate of =𝑣𝑡2 . Hence, the
ℎ𝑡
If there are no remaining GARCH effects, we should not be able to reject the null hypothesis that
the auto correlation coefficients of 𝑣𝑡2 are equal to zero. Otherwise, there is remaining conditional
volatility.
If you assumed normality, you should check to determine whether the estimated {vt} series actually
follows a normal distribution. Here we can apply Jarque-Bera test for normality. If we can reject
the null hypothesis of normal distribution we have to change the distribution of error in the
process of estimation.
Once we have obtained a satisfactory model, we can forecast future values of yt and its conditional
variance. As conditional variance of 𝑦𝑡 is same as the conditional variance of 𝜀𝑡 we proceed to
understand the process of forecasting conditional variance of 𝜀𝑡 . The one-step-ahead forecast of
the conditional variance is easy to obtain once we know the values of 𝛼0 , 𝛼1 𝑎𝑛𝑑 𝛽1. If we update
ℎ𝑡 by one period, we find ℎ𝑡+1 = 𝛼0 + 𝛼1 𝜖𝑡2 + 𝛽1 ℎ𝑡 . Since 𝜖𝑡2 and ℎ𝑡 are known in period t,
one period ahead forecast is simply 𝛼0 + 𝛼1 𝜖𝑡2 + 𝛽1 ℎ𝑡 . It is, however, somewhat more difficult
to obtain the jth step ahead forecasts.
2 2
To begin, use the fact that 𝜖𝑡2 = 𝑣𝑡2 ℎ𝑡 ⇒ 𝜖𝑡+𝑗 = 𝑣𝑡+𝑗 ℎ𝑡+𝑗
22
2 2 2 2
𝐸𝑡 𝜖𝑡+𝑗 = 𝐸𝑡 (𝑣𝑡+𝑗 ℎ𝑡+𝑗 ) [take expectation at t] ⇒ 𝐸𝑡 𝜖𝑡+𝑗 = 𝐸𝑡 (ℎ𝑡+𝑗 ) [Since 𝐸(𝑣𝑡+𝑗 )=
1 𝑎𝑛𝑑 𝑣𝑡+𝑗 𝑖𝑠 𝑖𝑛𝑑𝑒𝑝𝑒𝑛𝑑𝑒𝑛𝑡 𝑜𝑓 ℎ𝑡+𝑗 ]
Note that lower the value of (𝛼1 + 𝛽1 ) higher will be the speed of convergence and vice versa.
Now if we wish we can generalise the result for GARCH(p,q) model.
23
At period t, we have all of the information necessary to calculate the value of ℎ𝑡+1for any ARCH
process. Now if we update (1) by two periods and take the conditional expectation we get
2
𝐸𝑡 ℎ𝑡+2 = 𝛼0 + 𝛼1 𝐸𝑡 𝜖𝑡+1
2
Since 𝐸𝑡 𝜖𝑡+1 = ℎ𝑡+1, it follows that
𝐸𝑡 ℎ𝑡+2 = 𝛼0 + 𝛼1 ℎ𝑡+1 = 𝛼0 + 𝛼1 (𝛼0 + 𝛼1 𝜖𝑡2 )= 𝛼0 + 𝛼1 𝛼0 + 𝛼12 𝜖𝑡2
𝐸𝑡 ℎ𝑡+3 = 𝛼0 + 𝛼1 ℎ𝑡+2 = 𝛼0 + 𝛼1 (𝛼0 + 𝛼1 𝛼0 + 𝛼12 𝜖𝑡2 ) = 𝛼0 + 𝛼1 𝛼0 + 𝛼12 𝛼0 + 𝛼13 𝜖𝑡2
𝑗−1 𝑗
𝐸𝑡 ℎ𝑡+𝑗 = 𝛼0 + 𝛼1 𝛼0 + 𝛼12 𝛼0 + 𝛼13 𝛼0 + ⋯ + 𝛼1 𝛼0 +𝛼1 𝜖𝑡2
Thus it is possible to obtain the j step ahead forecasts of the conditional variance recursively. As
the value of 𝑗 → ∞, the forecasts of ℎ𝑡+𝑗 should converge to the unconditional variance
𝛼0
𝐸(𝜖𝑡2 ) =
1 − 𝛼1
The necessary condition for convergence is 𝛼1 < 1. To ensure that the Variance is always positive,
we also require that 𝛼0 > 0 and 𝛼1 ≤ 1. We can easily generalise the jth step ahead forecasting
for an ARCH (q) model.
Let us now see how the forecast volatility using ARCH or GARCH model can be used to
determine the accurate confidence bands around the forecast mean using the estimates of
2
conditional standard deviation. Since 𝐸𝑡 𝜀𝑡+1 = ℎ𝑡+1, a two-standard deviation confidence interval
for the forecasted mean can be constructed using 𝐸𝑡 𝑦𝑡+1 ± 2 √ℎ𝑡+1 The result is quite general;
since the mean of each value of {𝜀t} is zero, the optimal j-step-ahead forecast of 𝑦𝑡+𝑗 does not affected by
the presence of GARCH errors. However, the size of any confidence interval surrounding the forecasts
mean does depend on the conditional volatility. Clearly, in times when there is substantial
conditional volatility (i.e., when ℎ𝑡+1 is large), the variance of the forecast error will be large. And
thereby confidence interval will be large. Then we cannot be as confident of our mean forecasts
in periods when conditional volatility is high.
24
conditional variance. This class of model, called the ARCH in mean (ARCH-M) model, is
particularly suited to the study of asset markets. The basic insight is that risk-averse agents will
require compensation for holding a risky asset. Given that an asset’s riskiness can be measured by
the variance of returns, the risk premium will be an increasing function of the conditional variance
of returns. Engle, Lilien, and Robins express this idea by writing the excess return from holding a
risky asset as
𝑦𝑡 = 𝜇 + 𝛿ℎ𝑡 + 𝜀𝑡 𝛿>0 (1)
2 2 2
𝜀𝑡 = 𝑣𝑡 √ℎ𝑡 where ℎ𝑡 = 𝛼0 + 𝛼1 𝜖𝑡−1 + 𝛼2 𝜖𝑡−2 … . +𝛼1 𝜖𝑡−𝑞
where: yt = excess return from holding a long-term asset relative to a one-period treasury bill
𝜇 + 𝛿ℎ𝑡 = excess risk premium necessary to induce the risk-averse agent to hold the long-term
asset rather than the one-period bond.
ht is the conditional variance of 𝜀t. The conditional variance is an ARCH(q) process
𝜀𝑡 = unforecastable shock to the excess return on the long-term asset
To explain (1), note that the expected excess return from holding the long-term asset must be
just equal to the risk premium: 𝐸𝑡−1 𝑦𝑡 = 𝜇 + 𝛿ℎ𝑡
Engle, Lilien, and Robins assume that the risk premium is an increasing function of the conditional
variance of 𝜀t; in other words, the greater the conditional variance of returns, the greater the
compensation necessary to induce the agent to hold the long-term asset. It should be pointed out
that, if the conditional variance is constant (i.e., if 𝛼1 = 𝛼2 = · · · = 𝛼q = 0), the ARCH-M model
degenerates into the more traditional case of a constant risk premium.
We can use the GARCH specification for the variance equation the model is called GARCH-M.
In current literature GARCH-M is more popular compared to ARCH-M.
𝑦𝑡 = 𝜇 + 𝛽𝑥𝑡 + 𝜖𝑡
2
Such that ℎ𝑡 = 𝛼0 + 𝛼1 𝜖𝑡−1 + 𝛽1 ℎ𝑡−1
25
Where 𝛼1 + 𝛽1 = 1
IGARCH (1,1) has some very interesting properties. From the analysis of forecasting using
GARCH (1,1) we have
When 𝛼1 + 𝛽1 = 1,
𝐸𝑡 ℎ𝑡+𝑗 = 𝑗𝛼0 + ℎ𝑡
Thus, except for the intercept term 𝛼0, the forecast of the conditional variance for the next period
is the current value of the conditional variance. However, the unconditional variance of 𝜀𝑡 is clearly
infinite when𝛼1 + 𝛽1 = 1. Nevertheless, Nelson (1990) showed that the similarity between the
IGARCH process and an ARIMA process with a unit root is not perfect. Given that 𝛼1 + 𝛽1 = 1
and that ℎ𝑡−1 = 𝐿ℎ𝑡 we can write the conditional variance as
2
ℎ𝑡 = 𝛼0 + (1 − 𝛽1 )𝜖𝑡−1 + 𝛽1 𝐿ℎ𝑡
𝛼
0
⇒ℎ𝑡 = 1−𝛽 + (1 − 𝛽1 )[1 + 𝐿𝛽1 + 𝐿2 𝛽12 + ⋯ … … ]𝜖𝑡−1
2
1
𝛼
= 1−𝛽0 + (1 − 𝛽1 )[∑∞ 𝑖 2
𝑖=0 𝛽1 𝜖𝑡−1−𝑖 ] It is nothing but an ARCH (∞)
1
Thus, unlike a true non-stationary process, the conditional variance is a geometrically decaying
function of the current and past realisation of the {𝜖𝑡2 } sequence. As such an IGARCH model can
be estimated like any other GARCH (1, 1) model, because ARCH (∞) can be expressed as
GARCH (1, 1).
26
𝑦𝑡 = 𝛽𝑥𝑡 + 𝜀𝑡
2
And ℎ𝑡 = 𝛼0 + 𝛼1 𝜖𝑡−1 + 𝛽1 ℎ𝑡−1 + 𝛾𝐷𝑡
Now if it is found that 𝛾 > 0 , it is possible to conclude that sudden demonetisation increased
the mean of the conditional volatility returns. Note that we can consider a time series as exogenous
predictor say (oil price changes) of the variance of 𝜀𝑡 (return of an asset).
We measure ‘new information’ by the size of 𝜀𝑡 . If 𝜀𝑡 = 0 , the expected volatility in one period
ahead, 𝐸𝑡 ℎ𝑡+1 , is the distance Oa. Any news increases volatility; however, if the news is good i.e.
if 𝜀𝑡 > 0 volatility increases along ab. If the news is ‘bad’ volatility increases along ac. Since
segment ac is steeper than ab, a positive 𝜀𝑡 shock will have a smaller effect on future volatility than
a negative shock of the same magnitude.
27
Glosten, Jaganathan and Runkle (1994) showed how to allow the effects of good and bad news to
have different effects on volatility in GARCH framework. In a sense, 𝜀𝑡−1 = 0 is a threshold such
that shocks greater than the threshold have different effects than shocks below the threshold. With
this idea Glosten, Jaganathan and Runkle (1994) have proposed a GARCH model which is known
as GJR model or threshold GARCH (TGARCH) model. The variance equation of the TGARCH
(1,1) process can be written as.
2 2
ℎ𝑡 = 𝛼0 + 𝛼1 𝜀𝑡−1 + 𝜆1 𝑑𝑡−1 𝜀𝑡−1 + 𝛽1 ℎ𝑡−1
The intuition behind the TGARCH model is that positive values of 𝜖𝑡−1 are associated with a zero
2
value of 𝑑𝑡−1. Hence, if𝜀𝑡−1 ≥ 0, the effect of an 𝜀𝑡−1 shock on ℎ𝑡 is 𝛼1 𝜀𝑡−1 . When 𝜀𝑡−1 < 0,
2
𝑑𝑡−1 = 1 and the effect of an 𝜀𝑡−1 shock on ℎ𝑡 is (𝛼1 + 𝜆1 )𝜀𝑡−1 . If 𝜆1 > 0, the negative shock
will have larger effects on volatility than positive shocks. We can easily create a dummy variable
2
𝑑𝑡−1 and the product𝑑𝑡−1 𝜀𝑡−1 from the available information. If we find 𝜆1is statistically different
from zero, we can conclude that data contain a threshold effect.
The variance equation for EGARCH (1,1) model can be written as follows.
There are three interesting features to be discussed about the EGARCH model.
1. The equation for the conditional variance is in log-linear form. Regardless of the magnitude
of ln( ℎ𝑡 ), the implied value of ℎ𝑡 can never be negative. Hence it is permissible for the
28
coefficients to be negative. Actually if the relationship between volatility and return is
negative, the sign of 𝛼1 would be negative.
2
2. Instead of using the value of 𝜖𝑡−1 , the EGARCH model uses the level of standardised
value of 𝜖𝑡−1 [i.e. 𝜖𝑡−1 divided by √ℎ𝑡−1 ]. Nelson argues that this standardization allows
for a more natural interpretation of the size and persistence of shocks. After all, the
standardised value of 𝜖𝑡−1 is a unit free measure.
3. The EGARCH model allows for leverage effects. If 𝜖𝑡−1 ⁄√ℎ𝑡−1 is positive, the effect of
the shock on the log of the conditional variance is 𝛼1 + 𝜆1. If 𝜖𝑡−1 ⁄√ℎ𝑡−1 is negative, the
effect of the shock on the log of the conditional variance is−𝛼1 + 𝜆1. Note that if the
relationship between volatility and return is negative, i.e. in case of leverage effect the sign
of 𝛼1 would be negative. Therefore, magnitude of the effect of bad news would be greater
than the magnitude of the effect of good news.
Although the EGARCH model has some advantages over TGARCH model, it is difficult to
forecast the conditional variance of an EGARCH model. In his original formulation
Nelson(1991) assumed a generalised exponential distribution(GED) for the errors. This option
is available in EViews, but for easy computation and interpretation in most of the practical
application of EGARCH normality assumption for the errors have been considered.
One way to test for leverage is to estimate the TGARCH or EGARCH model and perform a
t test for the null hypothesis 𝜆1 = 0 . However, there is a specific diagnostic test that allows
to determine whether there are any leverage effects in residuals.
̂
𝜖𝑡
Step 1: Estimate an ARCH or GARCH model, form a standardised residuals series 𝑠𝑡 = ̂𝑡
√ℎ
Thus the { 𝑠𝑡 } sequence consists of each residual divided by its standard deviation. To test for
leverage effect, estimate a regression of the form
If there are no leverage effects, the squared errors should be uncorrelated with the level of the
error terms. Hence, we can conclude there are leverage effects if the sample value of F for the null
hypothesis 𝑎1 = 𝑎2 = ⋯ … … … . . = ⋯. exceeds the critical value obtained from an F table.
29
Engle and Nag (1993) developed a second way to determine whether positive and negative shocks
have different effects on the conditional variance. Again let 𝑑𝑡−1 be the dummy variable that is
equal to 1 if 𝜖̂
𝑡−1 < 0 and is equal to zero if 𝜖̂
𝑡−1 ≥ 0 . The test is to determine whether the
If t test confirms 𝑎1 is statistically significant, the sign of the current period shock is helpful in
predicting the conditional volatility. More concretely, if 𝑎1 is positive and statistically significant
we conclude that bad news compared to good news have greater impact on future volatility. To
generalise the test we can estimate the regression
The presence of 𝑑𝑡−1 𝑠𝑡−1 and (1 − 𝑑𝑡−1 )𝑠𝑡−1 is designed to determine whether the effects of
positive and negative shocks also depends on their size. We can use an F statistic to test the null
hypothesis 𝑎1 = 𝑎2 = 𝑎3 = 0 . If we reject 𝐻0 we conclude leverage effect. If we conclude that
there is a leverage effect, we can estimate a specific form of the TGARCH or EGARCH model.
30
For sake of simplicity suppose that there are two variables 𝑦1𝑡 𝑎𝑛𝑑 𝑦2𝑡 . For now we are not
interested in the means of the series so we can consider only the two error processes
𝜖1𝑡 = 𝑣1𝑡 √ℎ11𝑡
31
2
𝜖1𝑡 𝜖1𝑡 𝜖2𝑡
=( 2 )
𝜖1𝑡 𝜖2𝑡 𝜖2𝑡
2
𝜖1𝑡
Hence Vech (𝜀𝑡 𝜀𝑡 ′ ) = (𝜖1𝑡 𝜖2𝑡 )
2
𝜖2𝑡
𝛼11 𝛼12 𝛼13
If we now let 𝐴 = ( 21 𝛼22 𝛼23 )
𝛼
𝛼31 𝛼32 𝛼33
𝛽11 𝛽12 𝛽13
𝐵 = (𝛽21 𝛽22 𝛽23 )
𝛽31 𝛽32 𝛽33
𝑐10
𝐶 = (𝑐20 )
𝑐30
Then Vech model denoting (1) to (3) can be written as
𝑉𝑒𝑐ℎ (𝐻𝑡 ) = 𝐶 + 𝐴 𝑉𝑒𝑐ℎ(𝜀𝑡−1 𝜀𝑡−1 ′ ) + 𝐵 𝑉𝑒𝑐ℎ (𝐻𝑡−1 ) ....................(4)
Although simple to conceptualise, MGARCH models in the form of (1) to (3) or (4) can be very
difficult to estimate a MGARCH (p,q) for n variables. The problems with this general model are
as follows.
1. The number of parameters necessary to estimate can get quite large. In the two variable
cases we have 21 parameters. The number of parameter grows very quickly as more
variables are added to the system and as the order of the GARCH process increases. For
example, GARCH (2, 1) model needs the estimation of 9 additional parameters. For three
variable case MGARCH (1, 1) needs to estimate six equations for
ℎ11𝑡 ℎ12𝑡 ℎ13𝑡 ℎ22𝑡 ℎ33𝑡 ℎ32𝑡 respectively and the each equation 12 coefficients and
1 constant. Therefore, 78 parameters need to estimate. Moreover, we have not begun to
specify the models of the mean. If we have two variables 𝑦1𝑡 𝑎𝑛𝑑 𝑦2𝑡 it is possible to
estimate the means by specifying 𝑦1𝑡 − 𝑢1 = 𝜖1𝑡 and 𝑦2𝑡 − 𝑢2 = 𝜖2𝑡 . Once lagged values
of 𝑦1𝑡 𝑎𝑛𝑑 𝑦2𝑡 and/or explanatory variables are added to the mean equation estimation
problem can be complicated.
2. As in the univariate case, there is not an analytic solution to the maximisation of log
likelihood. As such, it is necessary to use numerical method to find the parameter values
that maximise the function L. Unfortunately, such method may not be able to find a
maximum value if the model is over parameterised. To explain, if the coefficient is small
relative to the SE, it necessarily has a large confidence interval. As such there is large range
32
in which the coefficient may lie and slight changes in the coefficients will have little
influence on the value of L.
The numerical hill climbing techniques that computers use in their maximisation routines
will have difficulty pinning down the value of such a coefficient. Hence, in the case of
estimating a over parameterised model, it is typical for a software package to indicate that
its search algorithm did not converge.
3. Since conditional variances are necessarily positive, the restrictions for the multivariate case
are far more complicated than for the univariate case. The results of such maximisation
problem must be such that every one of the conditional variances is always positive and
ℎ𝑖𝑗
the implied correlation coefficients 𝜌𝑖𝑗 = ⁄ are between -1 and +1.
√ℎ𝑖𝑖 ℎ𝑗𝑗
2
ℎ11𝑡 = 𝑐10 + 𝛼11 𝜖1𝑡−1 + 𝛽11 ℎ11𝑡−1 ...............................(4)
ℎ12𝑡 = 𝑐20 + 𝛼22 𝜖1𝑡−1 𝜖2𝑡−1 + 𝛽22 ℎ12𝑡−1 .......................(5)
2
ℎ22𝑡 = 𝑐30 + 𝛼33 𝜖2𝑡−1 + 𝛽33 ℎ22𝑡−1 ................................(6)
Given the large number of restrictions, the model is relatively easy to estimate. Each conditional
variance is equivalent to that of a univariate GARCH process and the conditional covariance is
quite parsimonious as well. The problem is that setting all 𝛼𝑖𝑗 = 𝛽𝑖𝑗 = 0 ∀ 𝑖 ≠ 𝑗 means that there
are no direct interactions among the variances. An 𝜖1𝑡−1 shock, for example, affect ℎ11𝑡 and ℎ12𝑡
, but does not affect the conditional variance ℎ22𝑡 . This restriction implies that there are no direct
volatility spill overs from one series to another Notice that the system wide estimation does have
the advantage of controlling for the contemporaneous correlation of the residuals across
equations.
33
BEK or BEKK Model
Baba, Engle and Kroner (1991) popularised a MGARCH model which is known as BEK or BEKK
model. In order to ensures the conditional variances are positive this model set the parameters to
enter the model via quadratic forms. Although there are several different variants of the model,
consider the specification
𝐻𝑡 = 𝐶 ′ 𝐶 + 𝐴′ 𝜖𝑡−1 𝜖𝑡−1 ′ 𝐴 + 𝐵 ′ 𝐻𝑡−1 𝐵 ......................................(7)
In our two variable case
ℎ ℎ12𝑡 𝑐11 𝑐12
𝐻𝑡 = ( 11𝑡 ); 𝐶 = (𝑐 𝑐22 );
ℎ21𝑡 ℎ22𝑡 21
34
Maximum Likelihood Estimation for Multivariate GARCH Model
To explain the method of estimation first we consider a two variable model (1) to (3). Suppose
that 𝜖1𝑡 and 𝜖2𝑡 are zero mean random variables that are jointly normally distributed. Initially we
assume that the variance and the covariance terms are constant. As such, we can drop the time
subscripts on the ℎ𝑖𝑗𝑡 . In such a circumstance the log likelihood function for the joint realisation
of 𝜖1𝑡 and 𝜖2𝑡 is
1 1 𝜖2 𝜖2 2𝜌 𝜖 𝜖
𝐿𝑡 = exp [− 2(1−𝜌2 (ℎ1𝑡 + ℎ2𝑡 − (ℎ 12ℎ 1𝑡)0.5
2𝑡
)] .........................(8)
2 )
2𝜋√ℎ11 ℎ22 (1−𝜌12 12 ) 11 22 11 22
ℎ12
Where 𝜌12 is the correlation coefficient between 𝜖1𝑡 and 𝜖2𝑡 ; 𝜌12 = ⁄(ℎ ℎ )0.5
11 22
ℎ ℎ12
Now if we define 𝐻 = [ 11 ] the likelihood function can be written as
ℎ21 ℎ22
1 1
𝐿𝑡 = 1 exp [− 2 𝜖𝑡 ′ 𝐻−1 𝜖𝑡 ] .....................................(9)
2𝜋|𝐻|2
Where 𝜖𝑡 = (𝜖1𝑡 𝜖2𝑡 )′ and |𝐻| is the determinant of H. To see that the two representations
2
given by (8) and (9) are equivalent. Note that |𝐻| = ℎ11 ℎ22 − ℎ12 . Since ℎ12 = 𝜌12 (ℎ11 ℎ22 )0.5
2
, it follows that |𝐻| = [(1 − 𝜌12 ) ℎ11ℎ22] . Moreover,
2 2
𝜖1𝑡 ℎ22 − 2𝜖1𝑡 𝜖2𝑡 ℎ12 + 𝜖2𝑡 ℎ11
𝜖𝑡 ′ 𝐻−1 𝜖𝑡 = 2
ℎ11 ℎ22 − ℎ12
Since ℎ12 = 𝜌12 (ℎ11 ℎ22 )0.5 , it follows that
2 2
1 𝜖1𝑡 𝜖2𝑡 2𝜌12 𝜖1𝑡 𝜖2𝑡
𝜖𝑡 ′ 𝐻 −1 𝜖𝑡 = [ 2 ( + − )]
(1 − 𝜌12 ) ℎ11 ℎ22 (ℎ11 ℎ22 )0.5
Since the realisation of {𝜖𝑡 } are conditionally dependent, the conditional likelihood of the joint
realisations of 𝜖1 𝜖2 … … … … . 𝜖 𝑇 is the product in the individual likelihood. Hence, if all have the
same variance, the conditional likelihood of the joint realisation is
𝑇
1 1 ′ −1
𝐿=∏ exp [− 𝜖 𝐻 𝜖𝑡 ]
1 2 𝑡
𝑡=1 2𝜋|𝐻|2
It is far easier to work with a sum than with a product. As such it is convenient to take the natural
log of each side so as to obtain
𝑇
𝑇 𝑇 1
ln 𝐿 = − ln(2𝜋) − ln|𝐻| − ∑ 𝜖𝑡 ′ 𝐻−1 𝜖𝑡
2 2 2
𝑡=1
The procedure used in MLE is to select the distributional parameters so as to maximise the
likelihood of drawing the observed sample. Given the realisation in 𝜖𝑡 , it is possible to select
ℎ11 , ℎ12 𝑎𝑛𝑑 ℎ22 so as to maximise the likelihood function.
35
Let us now allow the values of ℎ𝑖𝑗 to be time-varying. Then
𝑇
𝑇 1
ln 𝐿 = − ln(2𝜋) − ∑(ln|𝐻𝑡 | + 𝜖𝑡 ′ 𝐻−1 𝜖𝑡 )
2 2
𝑡=1
References
Enders, Walter, Applied Econometric Time Series, 4th ed., Wiley 2015.
Brooks, Chris, Introductory Econometrics for Finance, Cambridge University Press,
2002.
Hamilton, James D., Time Series Analysis, Princeton University Press, 1994.
36
In univariate GARCH models, conditional covariance is not explicitly modeled, as the focus is on the conditional variance of a single time series. However, in multivariate GARCH models, the conditional covariance between series is explicitly modeled to capture the co-movements and potential correlation in volatilities between different time series. This is achieved by modeling conditional covariance as a function of past covariances and variances of the involved series, which allows for a more comprehensive understanding of the relationships and potential volatility spillovers among the series, which are critical in multivariate contexts like portfolio management .
For a GARCH(1,1) model to remain stationary, the sum of \( \alpha_1 \) and \( \beta_1 \) must be less than one. This condition is critical because it ensures that the influence of shocks to conditional variance diminishes over time, allowing the model to eventually revert to a long-term mean and preventing it from exploding to infinity. Without this condition, the model could produce non-stationary processes that do not accurately reflect the real-world data .
Estimating parameters for a Multivariate GARCH (p,q) model is challenging due to the large number of parameters, as each additional variable significantly increases the complexity. This can lead to overparameterization and difficulties in ensuring positive definiteness of variance-covariance matrices. Typically, numerical optimization techniques are used for estimation, though these can struggle with convergence if the model is too complex. To address these issues, constraints are often imposed to reduce parameter space, and simpler models like the diagonal BEKK are employed to ensure tractability while still capturing essential cross-series interactions .
The GARCH(1,1) model addresses volatility clustering by modeling the current conditional variance as a function of its past values and past squared observations. This is captured through the equation \( h_t = \alpha_0 + \alpha_1 \epsilon_{t-1}^2 + \beta_1 h_{t-1} \). Here, \( \alpha_1 \) represents the impact of the last period's squared innovation (the ARCH term), and \( \beta_1 \) represents the influence of the last period's forecasted variance (the GARCH term). This allows the model to increase conditional variance when large scale innovations are followed by large scale conditional variances, reflecting periods of clustering volatility .
The leverage effect, which posits that negative asset returns often lead to higher subsequent volatility than positive returns of the same magnitude, has significant implications in volatility modeling. Traditional GARCH models may not always capture this asymmetry effectively, as they typically assume symmetric response to innovations. This has led to extensions like the EGARCH or TGARCH models that are specifically designed to incorporate the leverage effect by allowing different responses to positive and negative shocks. This can improve model fit and forecasting accuracy for assets that exhibit significant asymmetric volatility behavior .
Ensuring positivity of conditional variances in MGARCH models is crucial because negative variances would be non-sensical in a financial context where variance represents squared deviations, which are inherently non-negative. To maintain this, models such as the BEKK impose parameter structures that guarantee positive definiteness through quadratic forms, ensuring all variances remain positive. This structural adjustment helps avoid negative or undefined standard deviations and correlations, thus maintaining the model's viability and interpretability .
The GARCH(1,1) model is more parsimonious than the ARCH model because it requires fewer parameters to adequately capture the volatility dynamics in time series data. Specifically, while the ARCH model generally requires many lagged squared terms to capture long memory in volatility, the GARCH(1,1) model captures this by incorporating lagged conditional variances, which reduces the number of parameters needed and minimizes the risk of overfitting. This parsimonious nature of GARCH models makes them more efficient in terms of parameter estimation and less likely to violate non-negativity constraints .
The BEKK model improves upon the basic GARCH structure in multivariate settings by ensuring that all conditional variances are positive. It achieves this by parameterizing the model in a way that the variables enter through quadratic terms, which inherently are non-negative. This approach not only maintains the positivity constraint but also allows individual variances and covariances to influence each other, effectively capturing volatility spillovers and co-movements among multiple financial time series .
Multivariate GARCH models, such as the Vech or BEKK models, handle volatility spillover by specifying equations that account for interactions between the variances of each series, and between variances and covariances themselves. In these models, the conditional variance of each series depends not only on its past squared innovations and previous variances but also on past innovations and covariances of other series in the system. This setup allows shocks to one series to impact the volatility of another series, reflecting the interconnected nature of financial markets and accurately capturing the dynamics of co-volatility among multiple time series .
In a GARCH(1,1) model, the conditional variance is mean-reverting, which means it has a tendency to revert to a long-term average over time. The rate of this mean reversion is determined by the sum \( \alpha_1 + \beta_1 \); a smaller sum indicates faster reversion and quicker stabilization of variance forecasts towards the unconditional variance \( \alpha_0/(1 - \alpha_1 - \beta_1) \). This property ensures that any shock to volatility is transient, and long-term forecasts of volatility will eventually stabilize at the long-run variance, adjusting only gradually to new information .