0% found this document useful (0 votes)
20 views36 pages

Volatility Modeling with ARCH and GARCH

Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
20 views36 pages

Volatility Modeling with ARCH and GARCH

Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Modelling Volatility: Use of ARCH and GARCH models

Introduction

In basic univariate time series analysis we check whether the series is stationary or not. i.e. whether
the unconditional mean variance and covariance are time invariant or not. If it is stationary we
formulate different models like AR, MA, ARMA, otherwise we try to make the series meaningfully
stationary for modeling the conditional mean. So in univariate time series modeling we consider
the conditional mean changes over time. For predicting conditional mean we consider the lagged
values of the series as predictors and or other series. We estimate and predict the conditional mean
of the series assuming the disturbance is pure white noise and conditionally homoscedastic. These
assumptions are also true for a disturbance in standard regression model. In practice however
disturbances are heteroscedastic then we detect the nature of heteroscedasticity and estimate the
parameter addressing heteroscedasticity in cross section study. However, in time series we are
interested to estimate and forecast the conditional variance of a series which changes over time.
i.e. in time series analysis we study conditional heteroscedasticity which is auto correlated.

What is volatility?
The changes of fluctuation in the series are referred to as volatility. It is a situation where variance
in the process is not constant and keep changing over time. So, volatility means the changing
variance conditional on time or lagged values of the variable. It is evident that short run (local)
variance of economic and financial time series changes over time. Moreover, in many cases we
observe a period of relative tranquility followed by a period of relatively high volatility The
evidence of the phase of high volatility followed by the phase of low volatility is called volatility
clustering. We see the graph of weekly oil price changes during January 1986 to March, 2020 in
figure 1 as an example of volatility clustering.

Figure 1 Change of volatility for the rate change of oil spot price (weekly data)
6
rate of change of oil price

4
2
($ per barrel)

0
-2
-4
-6
Jan Jan Jan Jan Jan Jan Jan Jan Jan Jan Jan Jan Jan Jan Jan Jan Jan Jan
03, 03, 03, 03, 03, 03, 03, 03, 03, 03, 03, 03, 03, 03, 03, 03, 03, 03,
1986 1988 1990 1992 1994 1996 1998 2000 2002 2004 2006 2008 2010 2012 2014 2016 2018 2020
Dates

Data source: [Link]

1
Rate of change of oil spot price =100*(log (Pt)-log (Pt-1)) =gt
𝑙𝑛(𝑥+1)
[We know lim = 1 ⇒ ln(𝑥 + 1) = 𝑥 𝑤ℎ𝑒𝑛 𝑥 𝑖𝑠 𝑠𝑚𝑎𝑙𝑙 𝑒𝑛𝑜𝑢𝑔ℎ
𝑥→0 𝑥
𝑝𝑡 −𝑝𝑡−1 𝑝𝑡
By definition 𝑔𝑡 = =𝑝 − 1 = 𝑥(𝑠𝑎𝑦)
𝑝𝑡−1 𝑡−1

𝑝𝑡 𝑝𝑡
Now ln(𝑥 + 1) = 𝑥 ; ⇒ ln (𝑝 )=𝑝 − 1= gt (approximate growth rate)
𝑡−1 𝑡−1

⇒ ln(𝑝𝑡 ) − ln(𝑝𝑡−1 ) = 𝑔𝑡]

Stylized Facts of Economic Time Series relevant for Modelling Volatility


Before going to modelling volatility we review some stylized facts of economic time series which
encourage for modelling volatility. First we note the facts then illustrate them with the help of data.
i) The first stylized fact of economic time series, like rate of returns of an asset, is that there is
almost no correlation between changes of returns for different days. But there is positive
dependence between absolute changes returns on nearby days and thus for squared of changes of
returns. Thus volatility of many series is not constant over time, there are many evidence of
volatility clustering.
ii) The distribution of many series have thick tails, they tend to be leptokurtic.
iii) Some of the series exhibits breaks
iii) Shocks to a series can display a high degree of persistence
iv) Co-movements in volatility. There is considerable evidence that volatility is positively correlated
across assets in a market and even across markets
iv)Leverage Effects, it is a common features financial series like return of an asset, refers to the
tendency for changes in stock prices to be negatively correlated with changes in volatility.

2
In example of the graph of weekly return of google, we observe the volatility clustering and we see
that the distribution is approximately symmetric but it has high peak and fat tails relative to the
normal distribution. Therefore, thick tails, and leptokurtosis is stylized fact of time series.

Fig-3 Weekly values of the spot prices of Oil (May 15, 1987 –November 1, 2013

Spot
160
140
120
100
80
60
40
20
0
May 15, 1988

May 15, 1995

May 15, 1997

May 15, 2012


May 15, 1987

May 15, 1989


May 15, 1990
May 15, 1991
May 15, 1992
May 15, 1993
May 15, 1994

May 15, 1996

May 15, 1998


May 15, 1999
May 15, 2000
May 15, 2001
May 15, 2002
May 15, 2003
May 15, 2004
May 15, 2005
May 15, 2006
May 15, 2007
May 15, 2008
May 15, 2009
May 15, 2010
May 15, 2011

May 15, 2013


Some of the series exhibits break. The financial crises of 2007-08 caused a number of time series
breaks. Figure 3 shows that how the oil prices in US dips at the time of financial crises.

In some cases shocks to a series can display a high degree of persistence. In figure 4 shows the
movements of short term and long term interest rates in US during 1960Q1- 2012 Q4. Neither of the
rates has a clear upward nor downward trend. Nevertheless both show a high degree of persistence. Notice
that T-bill rate experienced two upward surges in the 1970s and remained at these high level for several
years. Similarly after a sharp decrease in in the late 1980s, the rate never again displayed the level attained
in the early 1980s. Moreover, in the figure we have observed a clear co-movement of the series. Therefore,
we can expect there is co-movement of the volatility of the series.

Figure 4 Short term and long term interest rates in US during 1960Q1- 2012 Q4
16

14

12

10

0
60 65 70 75 80 85 90 95 00 05 10

TBILL R5

3
iv) Leverage Effects, it is a common feature of financial series like return of an asset, refers to the
tendency for changes in stock prices to be negatively correlated with changes in volatility. i.e if
stick prices increase today volatility will reduce tomorrow and vice versa.

Therefore, in the process to estimation and forecasting volatility we should be careful about these
feature of the series.

In order explain these nature of the economic time series econometricians have been developing
different model for estimating and forecasting volatility of the series. Besides, estimation and
forecasting of the local variance of many financial time series like return of assets and of many
macroeconomic series like inflation rate, interest rates, and exchange rates are of interest for several
reasons.

First, the variance of an asset price is a measure of the risk of holding the asset. The larger variance
of daily stock price changes indicates the higher risk of holding the asset, i.e. an asset holder may
gain or lose heavily in a particular day. So a risk averse investor would be less likely to participate
in the in the stock market during a period of high, rather than low, volatility.

Second: forecasting the variance makes it possible to have accurate forecast intervals for mean
prediction. For example, firms and labour union needs to forecast the inflation rate over the
duration of labour contract. Economic theory suggests that terms of wage contract will depend on
the inflation forecast and the uncertainty concerning the accuracy of these forecasts. If the parties
to the contract have rational expectation, the terms of contract depends on conditional mean and
conditional variance of inflation rate as opposed to the unconditional mean and unconditional
variance.
Thirdly, the ability to forecast financial market volatility is important for portfolio selection and
asset management as well as for the pricing of primary and derivative assets (Engle and Ng, 1993).

Conditional Heteroscedasticity in time series Analysis


In our classical regression analysis we consider the problem of heteroscedasticity in such a way
that variance of disturbance depends on the explanatory variable. If there is heteroscedasticity
problem, the application of OLS is not robust. So to get robust estimate of the parameter in the
presence of heteroscedasticity we correct the heteroscedasticity problem.

4
However, in time series analysis there may be reason to believe that the variance of disturbance
changes over time, particularly depends on the magnitude of past errors. In univariate time series
analysis we have believed that series evolve depending on its own past values and past shocks.
Even in some cases time series like daily changed of oil prices, rate of daily return of an assets, are
completely unpredictable. So there is no question of explanatory variable/predictors that are used
for estimation and predict the conditional mean. In these applications there is often evidence of a
clumping of large and small errors over time. In figure we have observed period of relatively high
volatility followed by periods of relative tranquility. Thus in time series heteroscedasticity refers
to the dependency of the variance of disturbance on the volatility of the errors in the recent
past. It is called conditional heteroscedasticity. Here, heteroscedasticity is auto correlated ie. We
deal with the auto correlated heteroscedasticity. We are interested to model and forecast
conditional volatility because in time series analysis long run or unconditional forecast is inferior
to conditional forecasts. Besides in some cases short run analysis is appropriate. Example includes
that a risk averse asset holder would be interested to forecast of the rate of return and its variance
over the holding period. The unconditional variance (i.e., the long-run forecast of the variance)
would be unimportant if he plans to buy the asset at time t and sell at t + 1. If the labour union
and firm have rational expectation, the terms of labour contract depends on conditional mean and
conditional variance of inflation rate not the unconditional mean and unconditional variance of
inflation rate.

Strategies for Modelling Volatility


No doubt that most of the cases volatility is not constant over time. There seems to be no plausible
explanation for the changes in volatility, however there is a small part that can be explained through
modelling volatility.
Use of exogenous information
Modeling volatility is not new in literature. To get a rough idea of volatility of a series we can
consider the historical variance of the series. Still it is useful as a benchmark for comparing the
forecasting ability of more complex models. One approach to forecasting the variance of a series
is the use of exogenous information. Suppose there is an independent series 𝑥𝑡 which helps to
predict the volatility in𝜀𝑡 .
Consider the simplest strategy and let 𝜀𝑡 = 𝑣𝑡 𝑥𝑡−1
where: 𝜀𝑡 is the variable of interest
𝑣𝑡 is a white-noise disturbance term with variance 𝜎 2
𝑥𝑡−1 is an independent variable that can be observed at period t-1

5
If 𝑓𝑜𝑟 𝑎𝑙𝑙 𝑡, 𝑥𝑡−1= constant, the {𝜀𝑡 } sequence is the familiar white-noise process with a constant
variance. However, when the realizations of the {xt} sequence are not all equal, the variance of
𝜀𝑡 conditional on the observable value of 𝑥𝑡−1 is var(𝜀𝑡 |𝑥𝑡−1) =𝜎 2 𝑥𝑡−1
2

Here the conditional variance of 𝜀𝑡 is dependent on the realized value of 𝑥𝑡−1 . Since 𝑥𝑡−1at time
period t is observed, we can form the variance of 𝜀𝑡 conditionally on the realized value of 𝑥𝑡−1 . If
2
the magnitude 𝑥𝑡−1 is large (small), the variance of 𝜀𝑡 will be large (small) as well. Furthermore, if
the successive values of {xt} exhibit positive serial correlation (so that a large value of xt tends to
be followed by a large value of xt+1, the conditional variance of the {𝜀𝑡 } sequence will exhibit
positive serial correlation as well. In this way, the introduction of the {xt} sequence can explain
periods of volatility in the {𝜀𝑡 } sequence.
𝜀
Moreover, 𝐸( 𝑡⁄𝜀𝑡−1 , 𝜀𝑡−2 , … ) =0 ∵ 𝐸(𝑣𝑡 ) = 0
𝜀
and 𝑉𝑎𝑟( 𝑡⁄𝜀𝑡−1 , 𝜀𝑡−2 , … ) = 𝜎 2 𝑥𝑡−12

𝜀
Therefore, 𝑉𝑎𝑟( 𝑡⁄𝜀𝑡−1 , 𝜀𝑡−2 , … ) changes over time, but does not change when 𝜀𝑡−1changes,
rather changes when 𝑥𝑡−1 changes.

The major difficulty of the strategy is that the variance of the series 𝜀𝑡 given the past values does
not change when 𝜀𝑡−1changes. So the variance does not evolve over time with the process itself.
Oftentimes, we might not have a firm theoretical reason for selecting one candidate for the {xt}
sequence over other reasonable choices. Was it the oil price shocks, complete demonetisation,
and/or the introduction of GST that was responsible for the volatility of real investment in India
during the late 2010s?

Bilinear model
Let 𝜀𝑡 = 𝑣𝑡 𝜀𝑡−1 where 𝑣𝑡 is a white-noise disturbance term with constant variance 𝜎 2
𝜀
In this model 𝐸( 𝑡⁄𝜀𝑡−1 , 𝜀𝑡−2 , … ) =0 ∵ 𝐸(𝑣𝑡 ) = 0
𝜀
and 𝑉𝑎𝑟( 𝑡⁄𝜀𝑡−1 , 𝜀𝑡−2 , … ) = 𝜎 2 𝜀𝑡−1
2

𝜀
Therefore, 𝑉𝑎𝑟( 𝑡⁄𝜀𝑡−1 , 𝜀𝑡−2 , … ) changes over time, when 𝜀𝑡−1changes. So the variance evolves
over time with the process itself.
The problem with this model is that the unconditional variance of 𝜀𝑡 is either zero or infinite.

6
Check: Var (𝜀𝑡 )= 𝜎 2 𝑉𝑎𝑟(𝜀𝑡−1 )

Var(𝜀𝑡 ) - 𝜎 2 𝑉𝑎𝑟(𝜀𝑡 )=0 since unconditional variance of 𝜀𝑡 is identical to that of 𝜀𝑡−1

Var(𝜀𝑡 )(1 − 𝜎 2 ) = 0

Therefore Var (𝜀𝑡 ) = 0 𝑤ℎ𝑒𝑛 𝜎 2 < 1 and

= infinite when 𝜎 2 = 1

This result does not associated with practical nature of time series process. That is why this strategy
is not a popular for modelling volatility.

Autoregressive Conditionally Heteroscedastic (ARCH) models


Autoregressive volatility models are a relatively simple example from the class of stochastic
volatility specification. For modeling volatility ARCH and GARCH models have been proposed
by proposed by Robert Engle (1982), an American Statistician who won Nobel Memorial Prize
in Economic Sciences,2003 , sharing the award with Clive Granger, "for methods of analyzing
economic time series with time-varying Volatility (ARCH). Engle (1982) proposed the class of
multiplicative conditionally heteroscedastic model for the disturbance in connection with a
conditional mean model. Mean model may be an ARMA or ARIMA model even it may be a
multivariate regression model. We consider the conditional heteroscedasticity for the disturbance
term because from time series any trend, seasonal effects, autoregressive effect and influence of
predictors/ other variables should be removed. Usually modeling volatility is related to a high
frequency time series which are almost unpredictable. So it is prudent to consider a mean model
without any predictors/autoregressive components as follows.

The simplest form of the model is


𝑦𝑡 = 𝜇 + 𝜀𝑡 mean equation (1)
Such that
2
𝜀𝑡 = 𝑣𝑡 𝜎𝑡 ; (2) 𝑤𝑖𝑡ℎ 𝜎𝑡 = √𝛼0 + 𝛼1 𝜀𝑡−1 where 𝑣𝑡 is a white-noise
disturbance term with constant variance 1 and follows normal distribution. 𝜎𝑡 denotes the
conditional standard deviation of 𝜀𝑡 . Note that 𝑣𝑡 and 𝜀𝑡−1 are independent of each other and 𝛼0
and 𝛼1 are constants such that 𝛼0 > 0 𝑎𝑛𝑑 0 ≤ 𝛼1 ≤ 1
This model is popular because it has some nice properties. First consider the properties of 𝜀𝑡
process under this specification.
E(𝜀𝑡 ) = 0 , since E(𝑣𝑡 ) = 0

7
Var(𝜀𝑡 ) = 𝐸(𝜀𝑡2 )
𝑛𝑜𝑤, 𝐸(𝜀𝑡2 ) = 𝐸(𝑣𝑡2 )𝐸(𝛼0 + 𝛼1 𝜀𝑡−1
2
)
Or, 𝐸(𝜀𝑡2 ) = 𝛼0 + 𝛼1 𝐸(𝜀𝑡−1
2
)
Since unconditional variance of 𝜀𝑡 is identical to that of 𝜀𝑡−1 , 𝐸(𝜀𝑡2 ) = 𝐸(𝜀𝑡−1
2 )

𝛼
∴ 𝐸(𝜀𝑡2 ) = 1−𝛼0 >0
1

Therefore, unconditional variance of 𝜀𝑡 is time invariant.


The covariance of 𝜀𝑡 𝑎𝑛𝑑 𝜀𝑡−𝑠
E(𝜀𝑡 , 𝜀𝑡−𝑠 ) = 𝐸(𝑣𝑡 , 𝑣𝑡−1 )𝐸(𝜎𝑡 , 𝜎𝑡−1 ) = 0 ∵ 𝑣𝑡 is a white-noise disturbance term with constant
variance 1 𝑎𝑛𝑑 𝑣𝑡 and 𝜀𝑡−1 are independent of each other.

Therefore, under the specification (2) 𝜀𝑡 is pure white noise. The specification of 𝜀𝑡 does not affect
the properties of pure white noise disturbance of the regression (mean) model. At this point one
may think that the specification (2) does not affect the properties of 𝜀𝑡 . However, under this
specification unconditional distribution of 𝜀𝑡 does not follow normal distribution. It has
leptokurtosis. We know that kurtosis of normal distribution is 3. Here,

2
𝜀𝑡 = 𝑣𝑡 √𝛼0 + 𝛼1 𝜀𝑡−1

𝜀𝑡4 = 𝑣𝑡4 (𝛼0 + 𝛼1 𝜀𝑡−1


2
)2
⇒ 𝐸𝜀𝑡4 = 𝐸(𝑣𝑡4 )𝐸(𝛼0 + 𝛼1 𝜀𝑡−1
2
)2
⇒ 𝑚4 = 3𝐸[𝛼02+𝛼12 𝜀𝑡−1
4 2
+ 2𝛼0 𝛼1 𝜀𝑡−1 ]
𝛼2𝛼 𝛼
⇒ 𝑚4 = 3[𝛼02 +𝛼12 𝑚4 + 2 1−𝛼
0 1
] We have 𝑚2 = 𝐸(𝜀𝑡2 ) = 𝐸(𝜀𝑡−1
2 )
= 1−𝛼0
1 1

3𝛼02 (1 + 𝛼1 )
⇒ 𝑚4 =
(1 − 𝛼1 )(1 − 3𝛼12 )
It implies the fourth moment of 𝜀𝑡 is finite if 3𝛼12 < 1.
𝑚4 3(1−𝛼12 )
Therefore, kurtosis = = >3 ∵ (1 − 3𝛼12 ) <(1 − 𝛼12 ) as 0 < 𝛼1 ≤ 1
𝑚22 (1−3𝛼12 )

It means that the unconditional distribution of 𝜀𝑡 is fatter tailed and leptokurtic than the normal
distribution. Therefore, ARCH specification is allowed in the process which follow t distribution
because t distribution has fatter tail. One of the stylized fact of financial time series is that
distribution of the series have a fatter tail compared to that of normal distribution.

Let us now check the conditional distribution of 𝜀𝑡 . The conditional mean of 𝜀𝑡 given the past
𝜀
value of the series is zero. That is, 𝐸( 𝑡⁄𝜀𝑡−1 , 𝜀𝑡−2 , … ) = 0 since 𝐸𝑡−1 (𝑣𝑡 ) = 0

8
However, the influence of (2) entirely fall on the conditional variance of 𝜀𝑡
Engle (1982) shows that conditional variance of 𝜀𝑡

𝜀2
𝜎𝑡2 = ℎ𝑡 = 𝐸 ( 𝑡 ⁄𝜀𝑡−1 , 𝜀𝑡−2 , … ) = 𝛼0 + 𝛼1 𝜀𝑡−1
2
(3)

because 𝐸𝑡−1 (𝑣𝑡 )2 = 1


Thus equation (3) indicates the short run variance of the error term (because we estimate the model
disturbances are not observable, only errors are observable) responses to the amount of volatility
observed in recent periods. Equation (3) shows that conditional variance 𝜀𝑡 , ℎ𝑡 has two
components: a constant and last period’s news about volatility, which is modelled as last period’s
squared residuals (ARCH term). It implies that 𝜀𝑡 is heteroscedastic, conditional on 𝜀𝑡−1. In (3),
2
the conditional variance of 𝜀𝑡 is dependent on the realized value of 𝜀𝑡−1 . If the realized value of
2
𝜀𝑡−1 is large, the conditional variance in t will be large as well. Therefore, ARCH(1) is able to
capture the volatility cluster of the series.

Alternatively equation (3) can be written as 𝜀𝑡2 = 𝛼0 + 𝛼1 𝜀𝑡−1


2
+ 𝑤𝑡 (3a)

where 𝑤𝑡 = 𝜀𝑡2 − ℎ𝑡

Equation (3a) is nothing but a AR(1) model of 𝜀𝑡2 , therefore model (3) or (3a) is called ARCH(1).

We should check whether 𝑤𝑡 in PWN or not. Here, 𝑤𝑡 is i.i.d with E (𝑤𝑡 )=0 and Var (𝑤𝑡 )= 𝜆2
and Cov (𝑤𝑡 , 𝑤𝑡−1 ) = 0 and 𝑤𝑡 ≥ −𝛼0

We have 𝑤𝑡 = ℎ𝑡 (𝑣𝑡2 − 1)

𝑬(𝒘𝒕 ) = 𝑬(𝒉𝒕 )[𝑬(𝒗𝟐𝒕 ) − 𝟏) = 𝟎 ; since 𝐸(𝑣𝑡2 ) = 1

𝑤𝑡2⁄ 2 2 2 2
Var (𝑤𝑡 )= 𝐸 ( 𝜀𝑡−1 , 𝜀𝑡−2 , … ) = 𝐸(ℎ𝑡 )𝐸(𝑣𝑡 − 1) = 𝜎𝑤 (𝑠𝑎𝑦)

We have ℎ𝑡2 = (𝛼0 + 𝛼1 𝜀𝑡−1


2
)2= 𝛼0 2 + 𝛼1 2 𝜀𝑡−1
4 2
+ 2𝛼0 𝛼1 𝜀𝑡−1

∴ 𝐸(ℎ𝑡2 )= 𝛼0 2 + 𝛼1 2 𝐸(𝜀𝑡−1
4 2
) + 2𝛼0 𝛼1𝐸(𝜀𝑡−1 )

3 𝛼 2 (1+𝛼 ) 2𝛼02 𝛼1
𝛼0 2 + 𝛼1 2 (1−𝛼0 )(1−3𝛼1 2) +
1 1 1−𝛼1

Now, 𝐸(𝑣𝑡2 − 1)2 = 𝐸(𝑣𝑡4 ) + 1 − 2𝐸(𝑣𝑡 )2 as 𝑣𝑡 ~𝑁(0,1)

9
= 3+1-2 =2

(𝟏+𝜶𝟏 )
∴ Var (𝒘𝒕 )= 𝝈𝟐𝒘 =2 𝜶𝟎 𝟐 [(𝟏−𝜶 𝟐 ] Constant
𝟏 )(𝟏−𝟑𝜶𝟏 )

Cov (𝑤𝑡 , 𝑤𝑡−1 ) = 𝐸[(ℎ𝑡 , ℎ𝑡−1 ) (𝑣𝑡2 − 1)(𝑣𝑡−1


2
− 1)]

∵ ℎ𝑡 𝑎𝑛𝑑 𝑣𝑡 𝑎𝑟𝑒 𝑖𝑛𝑑𝑒𝑝𝑒𝑛𝑑𝑒𝑛𝑡 𝑎𝑛𝑑 𝑣𝑡 𝑎𝑛𝑑 𝜀𝑡 𝑎𝑟𝑒 𝑖𝑛𝑑𝑒𝑝𝑒𝑛𝑑𝑒𝑛𝑡

𝐸[(ℎ𝑡 , ℎ𝑡−1 ) [𝐸(𝑣𝑡2 , 𝑣𝑡−1


2 )
− 𝐸(𝑣𝑡2 ) − 𝐸(𝑣𝑡−1
2
) + 1)]

= 𝐸[(ℎ𝑡 , ℎ𝑡−1 ) [𝐸(𝑣𝑡2 )𝐸(𝑣𝑡−1


2 )
− 1 − 1 + 1]

Cov (𝒘𝒕 , 𝒘𝒕−𝟏 )= 𝑬[(𝒉𝒕 , 𝒉𝒕−𝟏 ) [𝟏 − 𝟏 − 𝟏 + 𝟏]=0 ∵ 𝑣𝑡 𝑎𝑛𝑑 𝑣𝑡−1 𝑎𝑟𝑒 𝑖𝑛𝑑𝑒𝑝𝑒𝑛𝑑𝑒𝑛𝑡 𝑆𝑁𝐷,

2 2
𝑣𝑡2 𝑎𝑛𝑑 𝑣𝑡−1
2
𝑎𝑟𝑒 𝑖𝑛𝑑𝑒𝑝𝑒𝑛𝑑𝑒𝑛𝑡 𝜒(1) 𝑣𝑎𝑟𝑖𝑎𝑡𝑒, 𝑎𝑛𝑑 𝐸(𝜒(𝑛) )=𝑛

Therefore, 𝑤𝑡 is a white noise process. It ensures that 𝜀𝑡2 is a AR(1) process. That is why; equation
(3) and (3a) is called ARCH (1) (Autoregressive Conditionally Heteroscedastic model of order1).

In (3), the conditional variance is a first-order Auto Regressive Conditionally Heteroskedastic


process denoted by ARCH(1). As opposed to a usual autoregression, the coefficients 𝛼0 and 𝛼1
have to be restricted. In order to ensure that the conditional variance is never negative, it is
necessary to assume that both 𝛼0 and 𝛼1 are positive. After all, if 𝛼0 is negative, a sufficiently
small realization of 𝜀t−1 will mean that conditional variance is negative. Similarly, if 𝛼1 is negative,
a sufficiently large realization of 𝜀t−1 can render a negative value for the conditional variance.
Moreover, to ensure the stationarity/stability of the process, it is necessary to restrict 𝛼1 such that
0 ≤ 𝛼1 ≤ 1. We check these stationarity condition just after a moment

In an ARCH model, the conditional and unconditional expectations of the error terms are equal
to zero. Moreover, the {𝜀t} sequence is serially uncorrelated because, for all s ≠ 0, E𝜀t𝜀t−s = 0.
The key point is that the errors are not correlated but dependent since they are related through
their second moment (recall that correlation is a linear relationship). Thus the dependency non-
linear by nature.
The conditional variance itself is an autoregressive process resulting in conditionally
2
heteroskedastic errors. When the realized value of 𝜀t−1 is far from zero—so that 𝛼1𝜀𝑡−1 is relatively
large—the variance of 𝜀t will tend to be large.
We can also prove that 𝜀𝑡2 is stationary under the specification in (3).

10
We have already proved that
𝛼
𝐸(𝜀𝑡2 ) = 1−𝛼0 it says unconditional mean of 𝜀𝑡2 is constant.
1

𝛼 2 𝜎2 2 𝛼0 2
Var (𝜀𝑡2 ) = 𝐸 (𝜀𝑡2 − 1−𝛼0 ) = 𝐸(𝑤𝑡2 (1 + 𝛼1 𝐿 + 𝛼12 𝐿2 … … )2 = 1−𝛼
𝑤
2=[(1−𝛼 2 2 ]
1 1 1) ((1−3𝛼1 )

(1+𝛼1)
Since, 𝜎𝑤2 =2 𝛼0 2 [(1−𝛼 2 ]
1)(1−3𝛼1 )

So the unconditional variance of 𝜀𝑡2 is time invariant and finite if 3𝛼12 < 1
Now, covariance of 𝜀𝑡2 𝑎𝑛𝑑 𝜀𝑡−1
2

𝛼0 𝛼0
𝐸 (𝜀𝑡2 − 2
) (𝜀𝑡−1 − )
1 − 𝛼1 1 − 𝛼1
= 𝐸(𝑤𝑡 + 𝛼1 𝑤𝑡−1 + 𝛼12 𝑤𝑡−2 + 𝛼13 𝑤𝑡−3 …)(𝑤𝑡−1 + 𝛼1 𝑤𝑡−2 + 𝛼12 𝑤𝑡−3 + 𝛼13 𝑤𝑡−4 …)
= 𝐸(𝑤𝑡 (𝑤𝑡−1 + 𝛼1 𝑤𝑡−2 + 𝛼12 𝑤𝑡−3 + 𝛼13 𝑤𝑡−4 …)+𝐸𝛼1 (𝑤𝑡−1 + 𝛼1 𝑤𝑡−2 + 𝛼12 𝑤𝑡−3 +
𝛼13 𝑤𝑡−4 … )2
2 2 𝛼1 𝛼0 2
= 0+𝛼1var (𝜀𝑡−1 )= (1−𝛼1 )2 ((1−3𝛼12 )

2 𝛼1 𝛼0 2
∴ Cov(𝜀𝑡2 , 𝜀𝑡−1
2
)= (1−𝛼 2 2
1 ) ((1−3𝛼1 )
2
2 𝛼1𝑠 𝛼0
Similarly, we can say Cov(𝜀𝑡2 , 𝜀𝑡−𝑠
2
)= (1−𝛼 2 2
1 ) (1−3𝛼1 )

Therefore, covariance of 𝜀𝑡2 , 𝜀𝑡−𝑠


2
is time invariant but geometrically decay as lag length increases.
We can say that autocorrelation function of AR series 𝜀𝑡2 will decay exponentially as lag length
increases iff |𝛼1 | < 1. (Home task) Note that we can also prove that conditional covariance of
𝜀𝑡2 , 𝜀𝑡−𝑠
2
is zero (Home task).
This analysis establish that under the specification of model (3) errors are not linearly correlated
but they are correlated at their second moment. That is why ARCH model is called non-linear
model.

Since the variance of 𝜀𝑡 in equation (3) depends only on last period’s volatility, we called the model
as ARCH (1). More generally, the variance could depend on any number of lagged volatilities.
Then ARCH (q) model can be written as

2 2 2 2
ℎ𝑡 = 𝛼0 + 𝛼1 𝜀𝑡−1 + 𝛼2 𝜀𝑡−2 +𝛼3 𝜀𝑡−3 + ⋯ + 𝛼𝑞 𝜀𝑡−𝑞 (4)

Equation (4) says that conditional variance of 𝜀𝑡 depends on shocks from 𝜀𝑡−1 𝑡𝑜 𝜀𝑡−𝑞 Note that
the conditional variance of 𝜀𝑡 depends on magnitude of the shocks from 𝜀𝑡−1 𝑡𝑜 𝜀𝑡−𝑞 not on their
signs as the equation 4 considers the squares of the past errors.

11
Moreover, the conditional heteroscedasticity in {𝜀t} will result in {yt} being heteroskedastic itself.

𝑦 𝑦 2
Conditional variance of 𝑦𝑡 𝑉𝑎𝑟( 𝑡⁄𝑦𝑡−1 , 𝑦𝑡−2 , … ) = 𝐸(𝑦𝑡 − 𝐸( 𝑡⁄𝑦𝑡−1 , 𝑦𝑡−2 , … ) =

𝜀2 2 2 2 2
𝐸 ( 𝑡 ⁄𝜀𝑡−1 , 𝜀𝑡−2 , … ) = 𝛼0 + 𝛼1 𝜀𝑡−1 + 𝛼2 𝜀𝑡−2 +𝛼3 𝜀𝑡−3 + ⋯ + 𝛼𝑞 𝜀𝑡−𝑞

Thus, the ARCH model is able to capture periods of tranquility and volatility in the {yt} series.

Therefore, ARCH types of model suitable for modelling volatility of financial series like returns
which can capture three important stylized facts of economic time series. (i) There is almost
no correlation between returns for different days. (ii) There is positive dependence between
absolute returns on nearby days and thus for squared returns (Taylor, 2005). (iii) returns do not
follow a normal distribution, in most cases the distribution is symmetric (Taylor, 2005), however,
in some cases, distribution is characterized by a asymmetric shape, normally skewned to the left
and high Kurtosis which implies that returns have a high peak (Mandelbrot, 1963) and fat tails
(Fama, 1965).

Identification of ARCH model


ARCH models were created in the context of econometric and finance problems where we are
interested in investments or stocks increase (or decrease) per time period. For that reason, we
suggest that the variable of interest in these problems might either be growth rate, the proportion
gained or lost since the last time, or log(𝑦𝑡 /𝑦𝑡−1 )=log(𝑦𝑡 )−log(𝑦𝑡−1 ), the logarithm of the ratio of
this time’s value to last time’s value. However, it is not necessary that one of these be the only
primary variable of interest. An ARCH model could be used for any series that has periods of
increased or decreased variance. This might, for example, be a property of residuals after an
ARIMA model has been fit to the data.
The best identification tool may be a time series plot of the series. It’s usually easy to spot periods
of increased variation scattered through the series. We can look at the ACF and PACF of the
residuals and squared residuals after estimating the mean equation for yt (say) and for instance,
if 𝜀̂𝑡 appears to be white noise and 𝜀̂𝑡2 appears to be AR(1), then an ARCH(1) model for the
variance is suggested. If the PACF of the 𝜀̂𝑡2 suggests AR(m), then ARCH(m) may work. In this
case we check the Lung-Box Q statistic

12
𝜌2
𝑄 = 𝑇(𝑇 + 2) ∑𝑛𝑖=1 𝑇−𝑖
𝑖 2
~𝜒(𝑛) Where, n= lag length (usually T/4)

for up to lag length (T/4). If we reject the null hypothesis of 𝐻0 : 𝜌1 = 𝜌2 = 𝜌3 = ⋯ = 𝜌𝑚

we confirm that the 𝑦𝑡2 (𝜀𝑡2 ) is significantly auto correlated up to lag length n. then we suggest to
formulate ARCH(m).

Formal test
A formal Lagrange multiplier test for ARCH errors is the test by McLeod and Li (1983). The
methodology involves the following two steps:
STEP 1: Use OLS to estimate the most appropriate regression equation or ARMA model and let
{ 𝜀̂𝑡2} denote the squares of the fitted errors.
STEP 2: Regress these squared residuals on a constant and on the q lagged values i.e estimate the
following regression 𝜀̂𝑡2 = 𝛼0 + 𝛼1 𝜀̂𝑡−1
2 2
+ 𝛼2 𝜀̂𝑡−2 2
+𝛼3 𝜀̂𝑡−3 2
+ ⋯ + 𝛼𝑞 𝜀̂𝑡−𝑞 (5)

If there are no ARCH or GARCH effects, the estimated values of 𝛼1 through 𝛼q should be zero.
Hence, this regression will have little explanatory power so that the coefficient of determination
(i.e., the usual𝑅2 ) will be quite low. Using a sample of T residuals, under the null hypothesis of no
ARCH errors, the test statistic 𝑇𝑅2 converges to a 𝜒2 distribution with q degrees of freedom. If
𝑇𝑅 2 is sufficiently large, rejection of the null hypothesis that 𝛼1 through 𝛼q are jointly equal to zero
is equivalent to rejection of the null hypothesis of no ARCH errors. On the other hand, if 𝑇𝑅2 is
sufficiently low, it is possible to conclude that there are no ARCH effects. In the small sample
sizes typically used in applied work, an F-test for the null hypothesis 𝛼1 = · · · = 𝛼q = 0 has been
shown to be superior to a 𝜒2 test. Compare the sample value of F to the values in an F-table with
q degrees of freedom in the numerator and T − q degrees of freedom in the denominator.

Limitations of ARCH model


However recently ARCH models themselves have rarely been used in literature, since it has some
limitations. First, it is not easy to determine the value of q, the number of lags of the squared
residuals in the model. Second, q may be very large to capture all of the dependence in the
conditional variance. The problem in this case is that a large number of parameters must be
estimated and this may be difficult to do with any precision. Then the ARCH model could not be

13
a parsimonious model. Engle (1982), however, circumvented this problem by specifying an
arbitrary linearly declining lag length on an ARCH (4) model for quarterly data series as follows.

𝜎𝑡2 = 𝛾0 + 𝛾1 (0.4𝜀̂𝑡−1
2 2
+ 0.3𝜀̂𝑡−2 2
+ 0.2𝜀̂𝑡−3 2
+ 0.1𝜀̂𝑡−4 ) here only two parameters are required
in the conditional variance equation 𝛾0 and 𝛾1 , rather than the five which could be required for an
unrestricted ARCH(4) model.

Third, ARCH models over predict the volatility because they response slowly to large isolated
shocks to the series of interest.

Fourth, non-negative constraint might be violated, everything else equal, the more parameters
there are in the conditional variance equation, the more likely it is that one or more of them will
have negative estimated value.

Fifth, in ARCH model the conditional variance of 𝜀𝑡 depends on magnitude of the shocks from
𝜀𝑡−1 𝑡𝑜 𝜀𝑡−𝑞 not on their signs as the equation 4 considers the squares of the past errors. Thus
leverage effect, if it exists, cannot be explained with the basic ARCH model.

GARCH model reduces these problems, this is why GARCH models are widely used in practice.

GARCH (Generalised ARCH) Model

The ARCH model has been generalised to allow linear dependence of the conditional variance, ℎ𝑡
on past values of ℎ𝑡 as well as on past squared values of the series. The GARCH model was
developed by Bollerslev (1986) and Taylor (1986).

The simplest GARCH (1,1) can be written as

𝑦𝑡 = 𝜇 + 𝛽𝑥𝑡 + 𝜀𝑡 (6)

2
Such that 𝜖𝑡 = 𝑣𝑡 (𝛼0 + 𝛼1 𝜖𝑡−1 + 𝛽1 ℎ𝑡−1 )0.5 it is called GARCH error process.

or, 𝐸𝑡−1 (𝜖𝑡2 ) = ℎ𝑡 = 𝛼0 + 𝛼1 𝜖𝑡−1


2
+ 𝛽1 ℎ𝑡−1 (7)

Using the GARCH (1,1) model it is possible to interpret the current fitted variance ℎ𝑡 as a weighted
function of a long run average value (dependent on 𝛼0), information about volatility during the
2
previous period 𝛼1 𝜖𝑡−1 (the ARCH term) and the fitted variance from the model during the
previous period 𝛽1 ℎ𝑡−1(the GARCH term). Note that the value of 𝛼1 must be strictly positive,

14
because if𝛼1 = 0, there is no way for the 𝜀𝑡 series to affect ℎ𝑡 . And value of 𝛽1 is expected to be
positive. Finally, 0 < 𝛼1 + 𝛽1 < 1 for stationarity condition of GRACH process. Large values of
both 𝛼1 𝑎𝑛𝑑 𝛽1 act to increase the conditional volatility, but they do so in different ways. The
larger is 𝛼1, the larger is the response of ℎ𝑡 to new information; clearly if 𝛼1is large, a 𝑣𝑡 shock has
a sizeable effect on 𝜖𝑡2 and ℎ𝑡+1. On the other hand, large 𝛽1 displays the response of ℎ𝑡 current
conditional variance, to the volume of past fitted conditional variance which captured the effects
of past shocks.

Note that GARCH (1, 1) model is effectively an ARMA (1,1)model for the squared residuals. To
see this consider squared residuals relative to the conditional variance is
𝑢𝑡 = 𝜖𝑡2 − ℎ𝑡

⇒ ℎ𝑡 = 𝜖𝑡2 − 𝑢𝑡

𝜖𝑡2 − 𝑢𝑡 = 𝛼0 + 𝛼1 𝜖𝑡−1
2
+ 𝛽1 ℎ𝑡−1

⇒ 𝜖𝑡2 = 𝛼0 + 𝛼1 𝜖𝑡−1
2
+ 𝛽1 ℎ𝑡−1 + 𝑢𝑡

⇒ 𝜖𝑡2 = 𝛼0 + 𝛼1 𝜖𝑡−1
2 2
+ 𝛽1 (𝜖𝑡−1 − 𝑒𝑡−1 ) + 𝑢𝑡

⇒ 𝜖𝑡2 = 𝛼0 + (𝛼1 + 𝛽1 )𝜖𝑡−1


2
− 𝛽1 𝑒𝑡−1 + 𝑢𝑡 (7a)

∴ 𝜖𝑡2 is ARMA(1,1) for GARCH (1, 1) error process with AR coefficient (𝛼1 + 𝛽1 ) (It is a good
exercise to show 𝑢 pure white noise. (Home task)

Let us now see that GARCH is more parsimonious compared to ARCH. It avoids overfitting.
Consequently, the model is less likely to breach non-negativity constraints. Consider the
conditional variance equations for GARCH (1,1) model

2
ℎ𝑡 = 𝛼0 + 𝛼1 𝜖𝑡−1 + 𝛽1 ℎ𝑡−1 .....................................(i)

2
⇒ ℎ𝑡−1 = 𝛼0 + 𝛼1 𝜖𝑡−2 + 𝛽1 ℎ𝑡−2 .......................(ii) for t=t-1

2
ℎ𝑡−2 = 𝛼0 + 𝛼1 𝜖𝑡−3 + 𝛽1 ℎ𝑡−3 ...............................(iii) for t=t-2

Putting (3) in (1) we get

2 2
ℎ𝑡 = 𝛼0 + 𝛼1 𝜖𝑡−1 + 𝛽1 (𝛼0 + 𝛼1 𝜖𝑡−2 + 𝛽1 ℎ𝑡−2 )

2 2
= 𝛼0 + 𝛼1 𝜖𝑡−1 + 𝛽1 𝛼0 + 𝛽1 𝛼1 𝜖𝑡−2 + 𝛽12 ℎ𝑡−2

15
2 2
ℎ𝑡 = 𝛼0 + 𝛼1 𝜖𝑡−1 + 𝛽1 𝛼0 + 𝛽1 𝛼1 𝜖𝑡−2 + 𝛽12 (𝛼0 + 𝛼1 𝜖𝑡−3
2
+ 𝛽1 ℎ𝑡−3 )

2 2
= 𝛼0 + 𝛼1 𝜖𝑡−1 + 𝛽1 𝛼0 + 𝛽1 𝛼1 𝜖𝑡−2 + 𝛽12 𝛼0 + 𝛽12 𝛼1 𝜖𝑡−3
2
+ 𝛽13 ℎ𝑡−3

=𝛼0 (1 + 𝛽1 + 𝛽12 + ⋯ … ) + 𝛼1 𝜖𝑡−1


2 (1
+ 𝛽1 𝐿 + 𝛽12 𝐿 + ⋯ … . . ) + ⋯ … + 𝛽1∞ ℎ0

As number of observations tends to infinite 𝛽1∞ → 0

Hence GARCH (1,1) can be written as

2 2
ℎ𝑡 = 𝛾0 + 𝛾1 𝜖𝑡−1 + 𝛾2 𝜖𝑡−2 + ⋯ … … … … .. ⇒ ARCH (∞) (8)

Thus GARCH (1,1) model containing only three parameters in the conditional variance equation,
is a very parsimonious model, that allows an infinite number of past squared errors to influence
the current conditional variance.

On the other side, we can show that ARCH (∞) can be written as GARCH (1, 1)
If we recognise that equation (8) is simply a distributed lag model for ℎ𝑡
2 2 2
ARCH (∞) ⇒ ℎ𝑡 = 𝛾0 + 𝛾1 𝜀𝑡−1 + 𝛾2 𝜀𝑡−2 + 𝛾3 𝜀𝑡−3 + ⋯ … … … ….

2
Let 𝛾𝑖 = 𝜆𝛾𝑖−1 ⇒ ℎ𝑡 = 𝛾0 + 𝛾0 𝜆𝜀𝑡−1 + 𝛾0 𝜆2 𝜀𝑡−2
2
+ 𝛾0 𝜆3 𝜀𝑡−3
2
+ ⋯………

2
Then ℎ𝑡−1 = 𝛾0 + 𝛾0 𝜆𝜀𝑡−2 + 𝛾0 𝜆2 𝜀𝑡−3
2
+ 𝛾0 𝜆3 𝜀𝑡−4
2
+ ⋯………

Now 𝜆ℎ𝑡−1 = 𝜆𝛾0 + 𝛾0 𝜆2 𝜀𝑡−2


2
+ 𝛾0 𝜆3 𝜀𝑡−3
2
+ 𝛾0 𝜆4 𝜀𝑡−4
2
+ ⋯………

2
Finally, ℎ𝑡 − λℎ𝑡−1 = 𝛾0 (1 − 𝜆) + 𝛾0 𝜆𝜀𝑡−1

2
⇒ ℎ𝑡 = 𝛾0 (1 − 𝜆) + 𝛾0 𝜆𝜀𝑡−1 + λℎ𝑡−1

2
= 𝛾0 − 𝛾1 + 𝛾1 𝜀𝑡−1 + λℎ𝑡−1

2
= 𝑎0 + 𝛾1 𝜀𝑡−1 + λℎ𝑡−1

It is called a GARCH (1,1) process. It is a model of conditional variance where ℎ𝑡 depends on past
error’s squared and past values of conditional variance. Here we have only three coefficient
parameters to be estimated. Hence, GARCH is parsimonious than ARCH.

We can extend the GARCH (1, 1) model for any number (p) of GARCH terms and any number
(q) of ARCH terms. When we estimate a GARCH (p, q) process, we estimate the two interrelated
equations.

16
𝑦𝑡 = 𝑎0 + 𝛽𝑥𝑡 + 𝜀𝑡 ...............................................(a)

2 2
and 𝜖𝑡 = 𝑣𝑡 (𝛼0 + 𝛼1 𝜖𝑡−1 + ⋯ … + 𝛼𝑞 𝜖𝑡−𝑞 + 𝛽1 ℎ𝑡−1 + 𝛽2 ℎ𝑡−2 + ⋯ . . +𝛽𝑝 ℎ𝑡−𝑝 )0.5..........(b)

where 𝑥𝑡 can be an ARMA process of order (𝑝𝑚 , 𝑞𝑚 ). Moreover 𝑥𝑡 can contain exogenous
variables.

The first equation is a model of the conditional mean and the second yields the model for the
conditional variance. The systems 𝑝𝑚 and 𝑞𝑚 are used to denote the order of the ARMA process
for the mean need not equal the order of the GARCH (p, q) equation. ℎ𝑡 is the conditional variance
of the mean equation. From (b) we get-

𝜖𝑡2 = 𝑣𝑡2 (𝛼0 + 𝛼1 𝜖𝑡−1


2 2
+ ⋯ … + 𝛼𝑞 𝜖𝑡−𝑞 + 𝛽1 ℎ𝑡−1 + 𝛽2 ℎ𝑡−2 + ⋯ . . +𝛽𝑝 ℎ𝑡−𝑝 )

ℎ𝑡 = 𝐸𝑡−1 (𝜖𝑡2 ) = 𝛼0 + 𝛼1 𝜖𝑡−1


2 2
+ ⋯ … + 𝛼𝑞 𝜖𝑡−𝑞 + 𝛽1 ℎ𝑡−1 + 𝛽2 ℎ𝑡−2 + ⋯ . . +𝛽𝑝 ℎ𝑡−𝑝

Properties of GARCH (1, 1) Error Process

Under GARCH (1,1)


2
𝜖𝑡 = 𝑣𝑡 (𝛼0 + 𝛼1 𝜖𝑡−1 + 𝛽1 ℎ𝑡−1 )0.5

⇒ 𝐸𝑡−1 (𝜖𝑡2 ) = 𝛼0 + 𝛼1 𝜖𝑡−1


2
+ 𝛽1 ℎ𝑡−1

2
Or ℎ𝑡 = 𝛼0 + 𝛼1 𝜖𝑡−1 + 𝛽1 ℎ𝑡−1

The unconditional mean of 𝜀𝑡 = 0

𝑆𝑖𝑛𝑐𝑒 𝐸 (𝑣𝑡 ) = 0 & 𝐸(ℎ𝑡−1 𝑣𝑡 ) = 0 𝐸(𝑣𝑡 𝜖𝑡−1 ) = 0

Unconditional variance of 𝜀𝑡

𝜖𝑡2 = 𝑣𝑡2 (𝛼0 + 𝛼1 𝜖𝑡−1


2
+ 𝛽1 ℎ𝑡−1 )

𝐸(𝜖𝑡2 ) = 𝐸(𝑣𝑡2 )𝐸(𝛼0 + 𝛼1 𝜖𝑡−1


2
+ 𝛽1 ℎ𝑡−1 )

2
= 𝛼0 + 𝛼1 𝐸(𝜖𝑡−1 ) + 𝛽1 𝐸(ℎ𝑡−1 ) [Since 𝐸(𝑣𝑡2 ) = 1]

(1 − 𝛼1 )𝐸(𝜖𝑡2 ) = 𝛼0 + 𝛽1 𝐸(ℎ𝑡−1 )

[ 𝜖𝑡2 = 𝑣𝑡2 ℎ𝑡 ⇒ 𝜖𝑡−1


2 2
= 𝑣𝑡−1 2 )
ℎ𝑡−1 ⇒ 𝐸(𝜖𝑡−1 = 𝐸(ℎ𝑡−1 ) 2 )
𝑠𝑖𝑛𝑐𝑒 𝐸(𝑣𝑡−1 =1]

17
[Alternative way: From the law of iterated expectations we know

𝐸(𝜖𝑡2 ) = 𝐸[𝐸𝑡−1 (𝜖𝑡2 )]

⇒ 𝐸(𝜖𝑡2 ) = 𝐸(ℎ𝑡 )

⇒ Unconditional expectation of conditional variance is unconditional variance.]

∴ (1 − 𝛼1 − 𝛽1 )𝐸(𝜖𝑡2 ) = 𝛼0

𝛼0
⇒ 𝐸(𝜖𝑡2 ) = > 0 𝑤ℎ𝑒𝑛 𝛼0 > 0, 𝑎𝑛𝑑 𝛼1 + 𝛽1 < 1
(1 − 𝛼1 − 𝛽1 )

For more general model GARCH (𝑝, 𝑞) it follows that the variance will be finite if

𝑞 𝑝

1 − ∑ 𝛼𝑖 − ∑ 𝛽𝑖 > 0
𝑖=1 𝑖=1

The auto covariance function

1 1
𝐸(𝜖𝑡 𝜖𝑡−𝑖 ) = 𝐸 [𝑣𝑡 ℎ𝑡2 𝑣𝑡−𝑖 ℎ𝑡−𝑖
2
]=0

Since 𝐸(𝑣𝑡 ) = 0 ∀ 𝑡 𝑎𝑛𝑑 𝐸(𝑣𝑡 ℎ𝑡−𝑖 ) = 0 ∀ 𝑖 ≥ 1

Therefore, under the GARCH specification error is pure white noise. So the errors are non auto
correlated. But we can’t say they are independent.

Moreover, under this GARCH specification unconditional distribution of 𝜀𝑡 does not follow
normal distribution. It has leptokurtosis. We know that kurtosis of normal distribution is 3. Here,

2
𝜀𝑡 = 𝑣𝑡 √𝛼0 + 𝛼1 𝜀𝑡−1 + 𝛽1 ℎ𝑡−1

𝜀𝑡4 = 𝑣𝑡4 (𝛼0 + 𝛼1 𝜀𝑡−1


2
+ 𝛽1 ℎ𝑡−1 )2
⇒ 𝐸𝜀𝑡4 = 𝐸(𝑣𝑡4 )𝐸(𝛼0 + 𝛼1 𝜀𝑡−1
2
+ 𝛽1 ℎ𝑡−1 )2
⇒ 𝑚4 = 3𝐸[𝛼02+𝛼12 𝜀𝑡−1
4
+𝛽12 ℎ𝑡−1
2 2
+ 2𝛼0 𝛼1 𝜀𝑡−1 2
+ 2𝛼0 𝛽1ℎ𝑡−1 + 2𝛼1 𝛽1 ℎ𝑡−1 𝜀𝑡−1 ]
𝑚4 𝛼2𝛼 𝛼 𝛼
⇒ 𝑚4 = 3[𝛼02 +𝛼12 𝑚4 +𝛽12 1
+ 2 1−𝛼0 −𝛽 + 2𝛼0 𝛽1 1−𝛼 0−𝛽 + 2𝛼1 𝛽1 (1−𝛼 0−𝛽 )2]
3 1 1 1 1 1 1
𝛼0
We have 𝑚2 = 𝐸(𝜀𝑡2 ) = 𝐸(𝜀𝑡−1
2 )
= 𝐸(ℎ𝑡−1 ) = 1−𝛼
1 −𝛽1

4 4 2
Further, 𝜀𝑡−1 = 𝑣𝑡−1 ℎ𝑡−1
4 4 2 2 2 𝑚4
⇒ 𝐸𝜀𝑡−1 = 𝐸𝑣𝑡−1 𝐸ℎ𝑡−1 ⇒ 𝑚4 = 3 𝐸ℎ𝑡−1 ⇒ 𝐸ℎ𝑡−1 = 3

18
3𝛼02 [(1 − 𝛼1 − 𝛽1 )2 + 2(𝛼1 + 𝛽1 )(1 − 𝛼1 − 𝛽1 ) + 2𝛼1 𝛽1]
⇒ 𝑚4 =
(1 − 3𝛼12 − 𝛽12 )(1 − 𝛼1 − 𝛽1 )2
It implies the fourth moment of 𝜀𝑡 is finite if 3𝛼12 + 𝛽12 < 1.
𝑚 3[(1−𝛼1 −𝛽1 )2 +2(𝛼1 +𝛽1 )(1−𝛼1 −𝛽1 )+2𝛼1𝛽1 ]
Therefore, kurtosis = 𝑚42=
2 (1−3𝛼12−𝛽12 )

3[1−(𝛼1+𝛽1 )2 +2𝛼1 𝛽1 ] 3[(1−𝛼 2 −𝛽2 ]


= = (1−3𝛼21−𝛽21) >3 ∵ (1 − 3𝛼12 −𝛽12 ) <(1 − 𝛼12 −𝛽12 ) as 0 < 𝛼1 ≤ 1
(1−3𝛼12 −𝛽12 ) 1 1

It means that the unconditional distribution of 𝜀𝑡 is fatter tailed and high peak than that of the
normal distribution.

The conditional variance

𝐸𝑡−1 (𝜖𝑡2 ) = ℎ𝑡 = 𝛼0 + 𝛼1 𝜖𝑡−1


2
+ 𝛽1 ℎ𝑡−1

This simple result is the essential feature of GARCH modelling. The conditional variance of the
error process is not constant. With the appropriate specification of the parameters of ℎ𝑡 , it is
possible to model and forecast the conditional variance.

Volatility persistence: In a GARCH the errors are uncorrelated as 𝐸(𝜖𝑡 𝜖𝑡−𝑖 ) = 0, However, the
squared errors of a GARCH (1,1) process are correlated. We know that squared residuals of
GARCH (1,1) is nothing but a ARMA (1,1) with AR coefficient (𝛼1 + 𝛽1 ). Thus from the
properties of ARMA (1,1) we say that ACF of the squared residuals is decaying (𝛼1 + 𝛽1 ) rate if
(𝛼1 + 𝛽1 ) <1. However, if it is close to unity ACF will not decay quickly. [See ARMA Processes].
So if (𝛼1 + 𝛽1 ) is almost equal to 1 we say the current volatility is persistent. In other words, we
say current level of conditional volatility is expected to prevail into the future.

It ensures that the GARCH model capture the important stylized fact of the financial time series.

Identification and Assessing the Fit


First, we test whether there is ARCH effect or not considering the residual of appropriate mean
equation. Next plotting the ACF and PACF of the squared residuals of mean equation we
determine the order of GARCH model. However, Bollerslev suggested that GARCH (1,1) model
is a good candidate for starting modelling the conditional variance. One way to assess the adequacy
of a GARCH model is to see how well it fits the data. It is now standard to assess the fit of a
GARCH model using model selection criteria such as the AIC and SBC discussed in univariate

19
time series analysis. First consider the sum of squared residuals (SSR) as a measure of the goodness
of fit. In a GARCH model, a reasonable measure of the goodness of fit is the sum of squares of
the {vt} sequence, i.e. sum of squares of the standardized residuals of the fitted GARCH model.
More concretely we can construct the AIC and SBC using
AIC=−2 ln L + 2n
SBC=−2 ln L + n ln(T)
where L is likelihood function and n is the number of estimated parameters. Smaller the value of
value of AIC or SBC higher the goodness of fit of the model.

Estimation of ARCH (1)/GARCH (1, 1) Model

In order to estimate the ARCH or GARCH models we apply maximum likelihood estimation
technique. We apply ML estimation technique since the ARCH and GARCH models are not in
linear form where OLS is not usable. Further to estimate the mean and variance equation
simultaneously. If we apply OLS first in the mean equation then estimate variance equation
considering the residuals then we cannot get efficient estimates of the parameters because in the
application of OLS we assume that variance of the disturbance term is constant. Let us now see
the ML method works for estimating ARCH or GARCH models.
Suppose the mean equation 𝑦𝑡 = 𝜇 + 𝛽𝑥𝑡 + 𝜀𝑡

2
Where 𝜀𝑡 = 𝑣𝑡 √ℎ𝑡 and ℎ𝑡 = 𝛼0 + 𝛼1 𝜖𝑡−1 + 𝛽1 ℎ𝑡−1

Initially suppose that values of {𝜀𝑡 } are drawn from a normal distribution having mean zero and
variance ℎ𝑡 which is not constant. Then from standard distribution theory the likelihood of any
realisation of 𝜀𝑡 is

1 −𝜖𝑡2
𝐿𝑡 = 𝑒𝑥𝑝 Where 𝐿𝑡 is the likelihood of 𝜀𝑡
√2𝜋ℎ𝑡 2ℎ𝑡

Since the realisation of 𝜀𝑡 are independent, the likelihood of the joint realisations of
𝜀1 𝜀2 … … … . 𝜀𝑇 is the product of the individual likelihoods. Hence likelihood of the joint
realisation is

𝑇
1 −𝜀𝑡2
𝐿=∏ exp( )
√2𝜋ℎ𝑡 2ℎ𝑡
𝑖=1

20
It is far easier to work with a sum than with a product. As such, it is convenient to take the natural
log of each side so as to obtain

𝑇 1 1 𝜖2
ln 𝐿 = − 2 ln(2𝜋) − 2 ∑𝑇𝑡=1 ln ℎ𝑡 − 2 ∑𝑇𝑡=1 𝑡⁄ℎ
𝑡

2 2
Where 𝜀𝑡 = 𝑦𝑡 − 𝜇 − 𝛽𝑥𝑡 and ℎ𝑡 = 𝛼0 + 𝛼1 𝜖𝑡−1 + 𝛽1 ℎ𝑡−1 for ARCH(1) ℎ𝑡 = 𝛼0 + 𝛼1 𝜖𝑡−1

Noe that as the models of variance need some lag value so we need some adjustment regarding
the sample size. Once we adjust the sample size it is possible to maximise ln 𝐿 with respect
to 𝜇, 𝛼0 , 𝛼1 𝛽1𝑎𝑛𝑑 𝛽. It is easy to check that there are no simple solutions to the first order
conditions for a maximum. Actually normal equations are found to be non-linear. Thus numerical
technique i.e iteration is required to maximise ln 𝐿, where first we assume 𝛼0 given find out
𝑡ℎ𝑒 𝑣𝑎𝑙𝑢𝑒𝑠 𝑜𝑓 𝑡ℎ𝑒𝑟 𝑝𝑎𝑟𝑎𝑚𝑒𝑡𝑒𝑟𝑠. Again the others parameters assuming given find out 𝛼0. The
process will continue until two successive values of the parameters converges. For ARCH &
GARCH model iterative technique is necessary to estimate the parameters. It is trickier process
actually for a higher order GARCH model where variance is time varying. In EViews we have two
algorithms available for optimisation; Berndt Hall, Hall and Hausman (1974) method and
Marquardt algorithm.

However, in GARCH model standardised residuals do not follow normal distribution. If the
normality assumption does not hold, the parameter estimation will still be consistent if the mean
and variance equation are correctly specified. However, in the context of non-normality, the usual
standard error estimates will be inappropriate. In that case we can use the alternative distribution
of the errors like t distribution or generalised exponential distribution option from EViews. Finally
to get the robust variance-covariance matrix estimator in the non-normality cases we apply
Bollerslev Wooldridge (1992) standard errors. It is also available in Eviews option. This procedure
i.e maximum likelihood with Bollerslev Wooldridge (1992) standard errors is known as quasi-
maximum likelihood or QLM.

Diagnostic Checks for Model Adequacy


In addition to providing a good fit, an estimated GARCH model should capture all dynamic
aspects of the model of the mean and the model of the variance. The estimated residuals should
be serially uncorrelated and should not display any remaining conditional volatility. We test to
ensure that the model has captured these properties, by standardizing the residuals dividinĝ 𝜀̂𝑡 by

21
ℎ̂𝑡0.5in order to obtain an estimate of what we have been calling the {vt} sequence. Since 𝜀t has a
𝜀̂𝑡
zero mean and a variance of ht, we can think of vt = ̂𝑡0.5 as the standardized value of 𝜀t. The

resulting series, which we will call st, should have a mean of zero and a variance of unity. If there
is any serial correlation in the {st} sequence, the model of the mean is not properly specified. To
test the adequacy of the model of the mean, form the Ljung–Box Q-statistics for the {st} sequence.
If we cannot reject the null hypothesis that the autocorrelation coefficients of vt are equal to zero.
To test for remaining GARCH effects, form the Ljung–Box Q-statistics of the squared
𝜀𝑡2
standardized residuals (i.e., 𝑠𝑡2 ). The basic idea is that 𝑠𝑡2 is an estimate of =𝑣𝑡2 . Hence, the
ℎ𝑡

properties of the 𝑠𝑡2 sequence should mimic those of 𝑣𝑡2 .

If there are no remaining GARCH effects, we should not be able to reject the null hypothesis that
the auto correlation coefficients of 𝑣𝑡2 are equal to zero. Otherwise, there is remaining conditional
volatility.

If you assumed normality, you should check to determine whether the estimated {vt} series actually
follows a normal distribution. Here we can apply Jarque-Bera test for normality. If we can reject
the null hypothesis of normal distribution we have to change the distribution of error in the
process of estimation.

Forecasting the Conditional Variance Estimating GARCH Model

Once we have obtained a satisfactory model, we can forecast future values of yt and its conditional
variance. As conditional variance of 𝑦𝑡 is same as the conditional variance of 𝜀𝑡 we proceed to
understand the process of forecasting conditional variance of 𝜀𝑡 . The one-step-ahead forecast of
the conditional variance is easy to obtain once we know the values of 𝛼0 , 𝛼1 𝑎𝑛𝑑 𝛽1. If we update
ℎ𝑡 by one period, we find ℎ𝑡+1 = 𝛼0 + 𝛼1 𝜖𝑡2 + 𝛽1 ℎ𝑡 . Since 𝜖𝑡2 and ℎ𝑡 are known in period t,
one period ahead forecast is simply 𝛼0 + 𝛼1 𝜖𝑡2 + 𝛽1 ℎ𝑡 . It is, however, somewhat more difficult
to obtain the jth step ahead forecasts.
2 2
To begin, use the fact that 𝜖𝑡2 = 𝑣𝑡2 ℎ𝑡 ⇒ 𝜖𝑡+𝑗 = 𝑣𝑡+𝑗 ℎ𝑡+𝑗

Taking conditional expectation of each side we get-

22
2 2 2 2
𝐸𝑡 𝜖𝑡+𝑗 = 𝐸𝑡 (𝑣𝑡+𝑗 ℎ𝑡+𝑗 ) [take expectation at t] ⇒ 𝐸𝑡 𝜖𝑡+𝑗 = 𝐸𝑡 (ℎ𝑡+𝑗 ) [Since 𝐸(𝑣𝑡+𝑗 )=
1 𝑎𝑛𝑑 𝑣𝑡+𝑗 𝑖𝑠 𝑖𝑛𝑑𝑒𝑝𝑒𝑛𝑑𝑒𝑛𝑡 𝑜𝑓 ℎ𝑡+𝑗 ]

First we consider the GARCH(1,1)


2
ℎ𝑡 = 𝛼0 + 𝛼1 𝜖𝑡−1 + 𝛽1 ℎ𝑡−1 (1.a)
Updating by 𝑗𝑡ℎ period we get
2
ℎ𝑡+𝑗 = 𝛼0 + 𝛼1 𝜖𝑡+𝑗−1 + 𝛽1 ℎ𝑡+𝑗−1 (1.b)
Taking conditional expectation
2
𝐸𝑡 (ℎ𝑡+𝑗 ) = 𝛼0 + 𝛼1 𝐸𝑡 (𝜖𝑡+𝑗−1 ) + 𝛽1 𝐸𝑡 (ℎ𝑡+𝑗−1 )
= 𝛼0 + (𝛼1 + 𝛽1 )𝐸𝑡 (ℎ𝑡+𝑗−1 ) ..............................(1.c)
Thus 1.c can be viewed as a first order difference equation in the 𝐸𝑡 (ℎ𝑡+𝑗 ) sequence with the
initial condition for ℎ𝑡 , Given ℎ𝑡 we can use (1.c) to forecast all subsequent values of the
conditional variance as
𝐸𝑡 (ℎ𝑡+𝑗 ) = 𝛼0 + (𝛼1 + 𝛽1 )[𝛼0 + (𝛼1 + 𝛽1 )𝐸𝑡 (ℎ𝑡+𝑗−2 )]
= 𝛼0 + (𝛼1 + 𝛽1 )𝛼0 + (𝛼1 + 𝛽1 )2 𝐸𝑡 (ℎ𝑡+𝑗−2 )
= 𝛼0 + (𝛼1 + 𝛽1 )𝛼0 + (𝛼1 + 𝛽1 )2 [𝛼0 + (𝛼1 + 𝛽1 )𝐸𝑡 (ℎ𝑡+𝑗−3 )]
= 𝛼0 + (𝛼1 + 𝛽1 )𝛼0 + (𝛼1 + 𝛽1 )2 𝛼0 + (𝛼1 + 𝛽1 )3 𝐸𝑡 (ℎ𝑡+𝑗−3 )
= 𝛼0 [1 + (𝛼1 + 𝛽1 ) + (𝛼1 + 𝛽1 )2 + ⋯ … . +(𝛼1 + 𝛽1 )𝑗−1 ] + (𝛼1 + 𝛽1 )𝑗 ℎ𝑡
If (𝛼1 + 𝛽1 ) < 1, the conditional forecast of ℎ𝑡+𝑗 will converse to the long run variance of the
error(unconditional variance.
𝛼0
𝐸(ℎ𝑡 ) =
1 − 𝛼1 − 𝛽1

Note that lower the value of (𝛼1 + 𝛽1 ) higher will be the speed of convergence and vice versa.
Now if we wish we can generalise the result for GARCH(p,q) model.

Forecasting of Conditional Variance of the ARCH Process

Similarly we can forecast the conditional variance of an ARCH(1) model.


2
Here ℎ𝑡 = 𝛼0 + 𝛼1 𝜖𝑡−1 .................................(1)
If we update (1) by one period we obtain
ℎ𝑡+1 = 𝛼0 + 𝛼1 𝜖𝑡2

23
At period t, we have all of the information necessary to calculate the value of ℎ𝑡+1for any ARCH
process. Now if we update (1) by two periods and take the conditional expectation we get
2
𝐸𝑡 ℎ𝑡+2 = 𝛼0 + 𝛼1 𝐸𝑡 𝜖𝑡+1
2
Since 𝐸𝑡 𝜖𝑡+1 = ℎ𝑡+1, it follows that
𝐸𝑡 ℎ𝑡+2 = 𝛼0 + 𝛼1 ℎ𝑡+1 = 𝛼0 + 𝛼1 (𝛼0 + 𝛼1 𝜖𝑡2 )= 𝛼0 + 𝛼1 𝛼0 + 𝛼12 𝜖𝑡2
𝐸𝑡 ℎ𝑡+3 = 𝛼0 + 𝛼1 ℎ𝑡+2 = 𝛼0 + 𝛼1 (𝛼0 + 𝛼1 𝛼0 + 𝛼12 𝜖𝑡2 ) = 𝛼0 + 𝛼1 𝛼0 + 𝛼12 𝛼0 + 𝛼13 𝜖𝑡2
𝑗−1 𝑗
𝐸𝑡 ℎ𝑡+𝑗 = 𝛼0 + 𝛼1 𝛼0 + 𝛼12 𝛼0 + 𝛼13 𝛼0 + ⋯ + 𝛼1 𝛼0 +𝛼1 𝜖𝑡2

Thus it is possible to obtain the j step ahead forecasts of the conditional variance recursively. As
the value of 𝑗 → ∞, the forecasts of ℎ𝑡+𝑗 should converge to the unconditional variance
𝛼0
𝐸(𝜖𝑡2 ) =
1 − 𝛼1
The necessary condition for convergence is 𝛼1 < 1. To ensure that the Variance is always positive,
we also require that 𝛼0 > 0 and 𝛼1 ≤ 1. We can easily generalise the jth step ahead forecasting
for an ARCH (q) model.

Let us now see how the forecast volatility using ARCH or GARCH model can be used to
determine the accurate confidence bands around the forecast mean using the estimates of
2
conditional standard deviation. Since 𝐸𝑡 𝜀𝑡+1 = ℎ𝑡+1, a two-standard deviation confidence interval
for the forecasted mean can be constructed using 𝐸𝑡 𝑦𝑡+1 ± 2 √ℎ𝑡+1 The result is quite general;

since the mean of each value of {𝜀t} is zero, the optimal j-step-ahead forecast of 𝑦𝑡+𝑗 does not affected by
the presence of GARCH errors. However, the size of any confidence interval surrounding the forecasts
mean does depend on the conditional volatility. Clearly, in times when there is substantial
conditional volatility (i.e., when ℎ𝑡+1 is large), the variance of the forecast error will be large. And
thereby confidence interval will be large. Then we cannot be as confident of our mean forecasts
in periods when conditional volatility is high.

Extensions of basic GARCH Model

The ARCH –M Model


The ARCH in mean (ARCH-M) model is a model where the conditional mean of a series say yt
depends on the conditional variance of the disturbance, ht. Engle, Lilien, and Robins (1987)
extended the basic ARCH framework to allow the mean of a sequence to depend on its own

24
conditional variance. This class of model, called the ARCH in mean (ARCH-M) model, is
particularly suited to the study of asset markets. The basic insight is that risk-averse agents will
require compensation for holding a risky asset. Given that an asset’s riskiness can be measured by
the variance of returns, the risk premium will be an increasing function of the conditional variance
of returns. Engle, Lilien, and Robins express this idea by writing the excess return from holding a
risky asset as
𝑦𝑡 = 𝜇 + 𝛿ℎ𝑡 + 𝜀𝑡 𝛿>0 (1)
2 2 2
𝜀𝑡 = 𝑣𝑡 √ℎ𝑡 where ℎ𝑡 = 𝛼0 + 𝛼1 𝜖𝑡−1 + 𝛼2 𝜖𝑡−2 … . +𝛼1 𝜖𝑡−𝑞
where: yt = excess return from holding a long-term asset relative to a one-period treasury bill
𝜇 + 𝛿ℎ𝑡 = excess risk premium necessary to induce the risk-averse agent to hold the long-term
asset rather than the one-period bond.
ht is the conditional variance of 𝜀t. The conditional variance is an ARCH(q) process
𝜀𝑡 = unforecastable shock to the excess return on the long-term asset
To explain (1), note that the expected excess return from holding the long-term asset must be
just equal to the risk premium: 𝐸𝑡−1 𝑦𝑡 = 𝜇 + 𝛿ℎ𝑡
Engle, Lilien, and Robins assume that the risk premium is an increasing function of the conditional
variance of 𝜀t; in other words, the greater the conditional variance of returns, the greater the
compensation necessary to induce the agent to hold the long-term asset. It should be pointed out
that, if the conditional variance is constant (i.e., if 𝛼1 = 𝛼2 = · · · = 𝛼q = 0), the ARCH-M model
degenerates into the more traditional case of a constant risk premium.
We can use the GARCH specification for the variance equation the model is called GARCH-M.
In current literature GARCH-M is more popular compared to ARCH-M.

The IGARCH model:


In financial time series, the conditional volatility is typically persistent. In fact, if we estimate a
GARCH (1, 1) model using a long time series of stock returns, we will find that 𝛼1 + 𝛽1 is very
close to unity. Nelson (1990) argued that constraining 𝛼1 + 𝛽1 = 1 can yield a very parsimonious
representation of the distribution of an asset’s return. In some respects, this constraint forces the
conditional variance to act like a process with unit root. Then the model is called integrated
GARCH model. The IGARCH (1,1) model can be written as

𝑦𝑡 = 𝜇 + 𝛽𝑥𝑡 + 𝜖𝑡

2
Such that ℎ𝑡 = 𝛼0 + 𝛼1 𝜖𝑡−1 + 𝛽1 ℎ𝑡−1

25
Where 𝛼1 + 𝛽1 = 1

IGARCH (1,1) has some very interesting properties. From the analysis of forecasting using
GARCH (1,1) we have

𝐸𝑡 ℎ𝑡+𝑗 = 𝛼0 + (𝛼1 + 𝛽1 )𝐸𝑡 ℎ𝑡+𝑗−1

Or, 𝐸𝑡 ℎ𝑡+𝑗 = 𝛼0 [1 + (𝛼1 + 𝛽1 ) + (𝛼1 + 𝛽1 )2 + ⋯ … . +(𝛼1 + 𝛽1 )𝑗−1 ] + (𝛼1 + 𝛽1 )𝑗 ℎ𝑡

When 𝛼1 + 𝛽1 = 1,

𝐸𝑡 ℎ𝑡+𝑗 = 𝑗𝛼0 + ℎ𝑡

Thus, except for the intercept term 𝛼0, the forecast of the conditional variance for the next period
is the current value of the conditional variance. However, the unconditional variance of 𝜀𝑡 is clearly
infinite when𝛼1 + 𝛽1 = 1. Nevertheless, Nelson (1990) showed that the similarity between the
IGARCH process and an ARIMA process with a unit root is not perfect. Given that 𝛼1 + 𝛽1 = 1
and that ℎ𝑡−1 = 𝐿ℎ𝑡 we can write the conditional variance as

2
ℎ𝑡 = 𝛼0 + (1 − 𝛽1 )𝜖𝑡−1 + 𝛽1 𝐿ℎ𝑡

𝛼
0
⇒ℎ𝑡 = 1−𝛽 + (1 − 𝛽1 )[1 + 𝐿𝛽1 + 𝐿2 𝛽12 + ⋯ … … ]𝜖𝑡−1
2
1

𝛼
= 1−𝛽0 + (1 − 𝛽1 )[∑∞ 𝑖 2
𝑖=0 𝛽1 𝜖𝑡−1−𝑖 ] It is nothing but an ARCH (∞)
1

Thus, unlike a true non-stationary process, the conditional variance is a geometrically decaying
function of the current and past realisation of the {𝜖𝑡2 } sequence. As such an IGARCH model can
be estimated like any other GARCH (1, 1) model, because ARCH (∞) can be expressed as
GARCH (1, 1).

GARCH Models with Explanatory Variables


Just as the model of the mean can contain explanatory variables, the specification of ℎ𝑡 also allows
for exogenous variables. Suppose that we want to determine whether the sudden demonetisation
of Nov, 8, 2016, increased the volatility of asset returns. One way to accomplish the task would be
to create a dummy variable 𝐷𝑡 = 0 before demonetisation and 𝐷𝑡 = 1 thereafter. Then the
modification of GARCH (1,1) model will be

26
𝑦𝑡 = 𝛽𝑥𝑡 + 𝜀𝑡

2
And ℎ𝑡 = 𝛼0 + 𝛼1 𝜖𝑡−1 + 𝛽1 ℎ𝑡−1 + 𝛾𝐷𝑡

Now if it is found that 𝛾 > 0 , it is possible to conclude that sudden demonetisation increased
the mean of the conditional volatility returns. Note that we can consider a time series as exogenous
predictor say (oil price changes) of the variance of 𝜀𝑡 (return of an asset).

Asymmetric GARCH Models: TGARCH & EGARCH


One major limitation of basic GARCH models is that they enforce a symmetric response of
volatility to positive and negative shocks. This arises since the conditional variance in basic
GARCH models depends on the magnitudes of the lagged residuals and not on their signs. By
squaring lagged residuals in the variance equation their sign is lost. However, an interesting feature
of asset prices is that “bad” news seems to have a more pronounced effect on volatility than does
“good” news. For many stocks, there is a strong negative correlation between the current return
and the future volatility. The falling tendency of volatility when returns rise and mounting tendency
of volatility when returns fall is often called the leverage effect. The idea of the leverage effect is
captured in the following figure.

We measure ‘new information’ by the size of 𝜀𝑡 . If 𝜀𝑡 = 0 , the expected volatility in one period
ahead, 𝐸𝑡 ℎ𝑡+1 , is the distance Oa. Any news increases volatility; however, if the news is good i.e.
if 𝜀𝑡 > 0 volatility increases along ab. If the news is ‘bad’ volatility increases along ac. Since
segment ac is steeper than ab, a positive 𝜀𝑡 shock will have a smaller effect on future volatility than
a negative shock of the same magnitude.

27
Glosten, Jaganathan and Runkle (1994) showed how to allow the effects of good and bad news to
have different effects on volatility in GARCH framework. In a sense, 𝜀𝑡−1 = 0 is a threshold such
that shocks greater than the threshold have different effects than shocks below the threshold. With
this idea Glosten, Jaganathan and Runkle (1994) have proposed a GARCH model which is known
as GJR model or threshold GARCH (TGARCH) model. The variance equation of the TGARCH
(1,1) process can be written as.

2 2
ℎ𝑡 = 𝛼0 + 𝛼1 𝜀𝑡−1 + 𝜆1 𝑑𝑡−1 𝜀𝑡−1 + 𝛽1 ℎ𝑡−1

Where 𝑑𝑡−1 is a dummy variable such that

𝑑𝑡−1 = 1 𝑖𝑓 𝜀𝑡−1 < 0

and 𝑑𝑡−1 = 0 𝑖𝑓 𝜀𝑡−1 ≥ 0

The intuition behind the TGARCH model is that positive values of 𝜖𝑡−1 are associated with a zero
2
value of 𝑑𝑡−1. Hence, if𝜀𝑡−1 ≥ 0, the effect of an 𝜀𝑡−1 shock on ℎ𝑡 is 𝛼1 𝜀𝑡−1 . When 𝜀𝑡−1 < 0,

2
𝑑𝑡−1 = 1 and the effect of an 𝜀𝑡−1 shock on ℎ𝑡 is (𝛼1 + 𝜆1 )𝜀𝑡−1 . If 𝜆1 > 0, the negative shock
will have larger effects on volatility than positive shocks. We can easily create a dummy variable
2
𝑑𝑡−1 and the product𝑑𝑡−1 𝜀𝑡−1 from the available information. If we find 𝜆1is statistically different
from zero, we can conclude that data contain a threshold effect.

Exponential GARCH Model (EGARCH)


One problem with a standard GARCH model is that it is necessary to ensure that all of the
estimated coefficients are positive. And the standard GARCH model fails to capture the leverage
effect. With this end in view Nelson 1991 proposed a specification, called exponential GARCH
(EGARCH) model, under the GARCH framework that does not require non-negativity restriction
for the coefficients and can capture the asymmetric effects of good news and bad news.

The variance equation for EGARCH (1,1) model can be written as follows.

ln( ℎ𝑡 ) = 𝛼0 + 𝛼1 (𝜖𝑡−1 ⁄√ℎ𝑡−1 ) + 𝜆1 |𝜖𝑡−1 ⁄√ℎ𝑡−1 | + 𝛽1 𝑙𝑛ℎ𝑡−1 ..........................(#)

There are three interesting features to be discussed about the EGARCH model.

1. The equation for the conditional variance is in log-linear form. Regardless of the magnitude
of ln( ℎ𝑡 ), the implied value of ℎ𝑡 can never be negative. Hence it is permissible for the
28
coefficients to be negative. Actually if the relationship between volatility and return is
negative, the sign of 𝛼1 would be negative.
2
2. Instead of using the value of 𝜖𝑡−1 , the EGARCH model uses the level of standardised

value of 𝜖𝑡−1 [i.e. 𝜖𝑡−1 divided by √ℎ𝑡−1 ]. Nelson argues that this standardization allows
for a more natural interpretation of the size and persistence of shocks. After all, the
standardised value of 𝜖𝑡−1 is a unit free measure.
3. The EGARCH model allows for leverage effects. If 𝜖𝑡−1 ⁄√ℎ𝑡−1 is positive, the effect of

the shock on the log of the conditional variance is 𝛼1 + 𝜆1. If 𝜖𝑡−1 ⁄√ℎ𝑡−1 is negative, the
effect of the shock on the log of the conditional variance is−𝛼1 + 𝜆1. Note that if the
relationship between volatility and return is negative, i.e. in case of leverage effect the sign
of 𝛼1 would be negative. Therefore, magnitude of the effect of bad news would be greater
than the magnitude of the effect of good news.

Although the EGARCH model has some advantages over TGARCH model, it is difficult to
forecast the conditional variance of an EGARCH model. In his original formulation
Nelson(1991) assumed a generalised exponential distribution(GED) for the errors. This option
is available in EViews, but for easy computation and interpretation in most of the practical
application of EGARCH normality assumption for the errors have been considered.

Testing for Leverage Effects

One way to test for leverage is to estimate the TGARCH or EGARCH model and perform a
t test for the null hypothesis 𝜆1 = 0 . However, there is a specific diagnostic test that allows
to determine whether there are any leverage effects in residuals.

̂
𝜖𝑡
Step 1: Estimate an ARCH or GARCH model, form a standardised residuals series 𝑠𝑡 = ̂𝑡
√ℎ

Thus the { 𝑠𝑡 } sequence consists of each residual divided by its standard deviation. To test for
leverage effect, estimate a regression of the form

𝑠𝑡2 = 𝑎0 + 𝑎1 𝑠𝑡−1 + 𝑎2 𝑠𝑡−2 + ⋯ … … … ….

If there are no leverage effects, the squared errors should be uncorrelated with the level of the
error terms. Hence, we can conclude there are leverage effects if the sample value of F for the null
hypothesis 𝑎1 = 𝑎2 = ⋯ … … … . . = ⋯. exceeds the critical value obtained from an F table.

29
Engle and Nag (1993) developed a second way to determine whether positive and negative shocks
have different effects on the conditional variance. Again let 𝑑𝑡−1 be the dummy variable that is
equal to 1 if 𝜖̂
𝑡−1 < 0 and is equal to zero if 𝜖̂
𝑡−1 ≥ 0 . The test is to determine whether the

estimated squared residuals can be predicted using the {𝑑𝑡−1 } sequence.

The sign-bias test uses the regression equation of the form

𝑠𝑡2 = 𝑎0 + 𝑎1 𝑑𝑡−1 + 𝜖1𝑡

Where 𝜖1𝑡 is a regression disturbance

If t test confirms 𝑎1 is statistically significant, the sign of the current period shock is helpful in
predicting the conditional volatility. More concretely, if 𝑎1 is positive and statistically significant
we conclude that bad news compared to good news have greater impact on future volatility. To
generalise the test we can estimate the regression

𝑠𝑡2 = 𝑎0 + 𝑎1 𝑑𝑡−1 + 𝑎2 𝑑𝑡−1 𝑠𝑡−1 + 𝑎3 (1 − 𝑑𝑡−1 )𝑠𝑡−1 + 𝜖1𝑡

The presence of 𝑑𝑡−1 𝑠𝑡−1 and (1 − 𝑑𝑡−1 )𝑠𝑡−1 is designed to determine whether the effects of
positive and negative shocks also depends on their size. We can use an F statistic to test the null
hypothesis 𝑎1 = 𝑎2 = 𝑎3 = 0 . If we reject 𝐻0 we conclude leverage effect. If we conclude that
there is a leverage effect, we can estimate a specific form of the TGARCH or EGARCH model.

Multivariate GARCH Model (MGARCH)


We have seen that one important stylized fact of economic time series is the co-movement of the
series over time. If two or more series move together, it is reasonable to guess the presence of co-
volatility among the series. If we have a data set with several variables, it often makes sense to
estimate the conditional volatility of the variables simultaneously. MGARCH models take
advantage of the fact that the contemporaneous shocks to variables can be correlated with each
other. In addition to the variance equations, MGARCH models specify equations for how the
covariance move over time. Moreover, MGARCH models allows for volatility spill over such that
volatility shocks to one variable might affect the volatility of other related variables. For example,
suppose we want to model the NIFTY instead of simply modelling the SENSEX. Although we
could separately model the variance of each index, we might expect the volatility of the two series
to be interrelated. After all shocks that increases the uncertainty of one index are likely to increase
the uncertainty of the other.

30
For sake of simplicity suppose that there are two variables 𝑦1𝑡 𝑎𝑛𝑑 𝑦2𝑡 . For now we are not
interested in the means of the series so we can consider only the two error processes
𝜖1𝑡 = 𝑣1𝑡 √ℎ11𝑡

and 𝜖2𝑡 = 𝑣2𝑡 √ℎ22𝑡


As in the univariate case, if we assume (𝑣1𝑡 ) = 𝑉𝑎𝑟(𝑣2𝑡 ) = 1 , we can think of ℎ11𝑡 and ℎ22𝑡 as
the conditional variance of 𝜖1𝑡 and 𝜖2𝑡 respectively. Since we want to allow for the possibility
that the shocks are correlated denote ℎ12𝑡 as the conditional covariance of 𝜖1𝑡 𝜖2𝑡 , i.e.
𝐸𝑡−1 𝜖1𝑡 𝜖2𝑡 = ℎ12𝑡 . A natural way to construct a MGARCH (1,1) process is to allow all of the
volatility terms to interact with each other. Consider the so called Vech model
2 2
ℎ11𝑡 = 𝑐10 + 𝛼11 𝜖1𝑡−1 + 𝛼12 𝜖1𝑡−1 𝜖2𝑡−1 + 𝛼13 𝜖2𝑡−1 + 𝛽11 ℎ11𝑡−1 + 𝛽12 ℎ12𝑡−1 + 𝛽13 ℎ22𝑡−1
,.............................(1)
2 2
ℎ12𝑡 = 𝑐20 + 𝛼21 𝜖1𝑡−1 + 𝛼22 𝜖1𝑡−1 𝜖2𝑡−1 + 𝛼23 𝜖2𝑡−1 + 𝛽21 ℎ11𝑡−1 + 𝛽22 ℎ12𝑡−1 + 𝛽23 ℎ22𝑡−1
..............................(2)
2 2
ℎ22𝑡 = 𝑐30 + 𝛼31 𝜖1𝑡−1 + 𝛼32 𝜖1𝑡−1 𝜖2𝑡−1 + 𝛼33 𝜖2𝑡−1 + 𝛽31 ℎ11𝑡−1 + 𝛽32 ℎ12𝑡−1 + 𝛽33 ℎ22𝑡−1
............................... (3)
Here conditional variance of each variable ℎ11𝑡 and ℎ22𝑡 depends on its own past, the conditional
2 2
covariance between the two variables ℎ12𝑡 , the lagged squared errors 𝜖1𝑡−1 and 𝜖2𝑡−1 and the
product of lagged errors 𝜖1𝑡−1 𝜖2𝑡−1 . Notice that conditional covariance depends on the same set
of variables. Clearly there is rich interaction between the variables. For example a 𝜖1𝑡−1 shock
affect ℎ11𝑡 , ℎ22𝑡 and ℎ12𝑡 .

Matrix Specification of Vech model as follows


The Vech operator transforms the upper (lower) triangle of a symmetric matrix into column
vector.
Consider the symmetric covariance matrix of 𝜖1𝑡 and 𝜖2𝑡 as
ℎ ℎ12𝑡
𝐻𝑡 = [ 11𝑡 ]
ℎ21𝑡 ℎ22𝑡
ℎ11𝑡
So that Vech (𝐻𝑡 ) = (ℎ12𝑡 )
ℎ22𝑡
𝜖1𝑡
Now consider the vector 𝜀𝑡 = (𝜖 )
2𝑡
𝜖1𝑡
The product 𝜀𝑡 𝜀𝑡 ′ = (𝜖 ) (𝜖1𝑡 𝜖2𝑡 )
2𝑡

31
2
𝜖1𝑡 𝜖1𝑡 𝜖2𝑡
=( 2 )
𝜖1𝑡 𝜖2𝑡 𝜖2𝑡
2
𝜖1𝑡
Hence Vech (𝜀𝑡 𝜀𝑡 ′ ) = (𝜖1𝑡 𝜖2𝑡 )
2
𝜖2𝑡
𝛼11 𝛼12 𝛼13
If we now let 𝐴 = ( 21 𝛼22 𝛼23 )
𝛼
𝛼31 𝛼32 𝛼33
𝛽11 𝛽12 𝛽13
𝐵 = (𝛽21 𝛽22 𝛽23 )
𝛽31 𝛽32 𝛽33
𝑐10
𝐶 = (𝑐20 )
𝑐30
Then Vech model denoting (1) to (3) can be written as
𝑉𝑒𝑐ℎ (𝐻𝑡 ) = 𝐶 + 𝐴 𝑉𝑒𝑐ℎ(𝜀𝑡−1 𝜀𝑡−1 ′ ) + 𝐵 𝑉𝑒𝑐ℎ (𝐻𝑡−1 ) ....................(4)
Although simple to conceptualise, MGARCH models in the form of (1) to (3) or (4) can be very
difficult to estimate a MGARCH (p,q) for n variables. The problems with this general model are
as follows.
1. The number of parameters necessary to estimate can get quite large. In the two variable
cases we have 21 parameters. The number of parameter grows very quickly as more
variables are added to the system and as the order of the GARCH process increases. For
example, GARCH (2, 1) model needs the estimation of 9 additional parameters. For three
variable case MGARCH (1, 1) needs to estimate six equations for
ℎ11𝑡 ℎ12𝑡 ℎ13𝑡 ℎ22𝑡 ℎ33𝑡 ℎ32𝑡 respectively and the each equation 12 coefficients and
1 constant. Therefore, 78 parameters need to estimate. Moreover, we have not begun to
specify the models of the mean. If we have two variables 𝑦1𝑡 𝑎𝑛𝑑 𝑦2𝑡 it is possible to
estimate the means by specifying 𝑦1𝑡 − 𝑢1 = 𝜖1𝑡 and 𝑦2𝑡 − 𝑢2 = 𝜖2𝑡 . Once lagged values
of 𝑦1𝑡 𝑎𝑛𝑑 𝑦2𝑡 and/or explanatory variables are added to the mean equation estimation
problem can be complicated.

2. As in the univariate case, there is not an analytic solution to the maximisation of log
likelihood. As such, it is necessary to use numerical method to find the parameter values
that maximise the function L. Unfortunately, such method may not be able to find a
maximum value if the model is over parameterised. To explain, if the coefficient is small
relative to the SE, it necessarily has a large confidence interval. As such there is large range

32
in which the coefficient may lie and slight changes in the coefficients will have little
influence on the value of L.
The numerical hill climbing techniques that computers use in their maximisation routines
will have difficulty pinning down the value of such a coefficient. Hence, in the case of
estimating a over parameterised model, it is typical for a software package to indicate that
its search algorithm did not converge.

3. Since conditional variances are necessarily positive, the restrictions for the multivariate case
are far more complicated than for the univariate case. The results of such maximisation
problem must be such that every one of the conditional variances is always positive and
ℎ𝑖𝑗
the implied correlation coefficients 𝜌𝑖𝑗 = ⁄ are between -1 and +1.
√ℎ𝑖𝑖 ℎ𝑗𝑗

Diagonal Vech Model


In order to circumvent the problem of Vech model recent works on MGARCH involve finding
suitable restrictions on the general model (1) to (3). On set of restrictions that become popular in
the early literature is the so-called diagonal Vech model. The idea is to diagonalise the system such
that ℎ𝑖𝑗𝑡 contains only lags of itself and the cross products of 𝜖𝑖𝑡 𝜖𝑗𝑡 . For example, the diagonalised
version of (1) to (3) is

2
ℎ11𝑡 = 𝑐10 + 𝛼11 𝜖1𝑡−1 + 𝛽11 ℎ11𝑡−1 ...............................(4)
ℎ12𝑡 = 𝑐20 + 𝛼22 𝜖1𝑡−1 𝜖2𝑡−1 + 𝛽22 ℎ12𝑡−1 .......................(5)
2
ℎ22𝑡 = 𝑐30 + 𝛼33 𝜖2𝑡−1 + 𝛽33 ℎ22𝑡−1 ................................(6)
Given the large number of restrictions, the model is relatively easy to estimate. Each conditional
variance is equivalent to that of a univariate GARCH process and the conditional covariance is
quite parsimonious as well. The problem is that setting all 𝛼𝑖𝑗 = 𝛽𝑖𝑗 = 0 ∀ 𝑖 ≠ 𝑗 means that there
are no direct interactions among the variances. An 𝜖1𝑡−1 shock, for example, affect ℎ11𝑡 and ℎ12𝑡
, but does not affect the conditional variance ℎ22𝑡 . This restriction implies that there are no direct
volatility spill overs from one series to another Notice that the system wide estimation does have
the advantage of controlling for the contemporaneous correlation of the residuals across
equations.

33
BEK or BEKK Model
Baba, Engle and Kroner (1991) popularised a MGARCH model which is known as BEK or BEKK
model. In order to ensures the conditional variances are positive this model set the parameters to
enter the model via quadratic forms. Although there are several different variants of the model,
consider the specification
𝐻𝑡 = 𝐶 ′ 𝐶 + 𝐴′ 𝜖𝑡−1 𝜖𝑡−1 ′ 𝐴 + 𝐵 ′ 𝐻𝑡−1 𝐵 ......................................(7)
In our two variable case
ℎ ℎ12𝑡 𝑐11 𝑐12
𝐻𝑡 = ( 11𝑡 ); 𝐶 = (𝑐 𝑐22 );
ℎ21𝑡 ℎ22𝑡 21

𝛼11 𝛼12 𝛽 𝛽12


𝐴 = (𝛼 𝛼22 ); 𝐵 = ( 11 )
21 𝛽21 𝛽22
2 2 ) 2 2 2 2
Here ℎ11𝑡 = (𝑐11 + 𝑐12 + (𝛼11 𝜖1𝑡−1 + 2𝛼11𝛼21𝜖1𝑡−1 𝜖2𝑡−1 + 𝛼21 𝜖2𝑡−1 )
2 2
+(𝛽11 ℎ11𝑡−1 + 2𝛽11 𝛽21 ℎ12𝑡−1 + 𝛽21 ℎ22𝑡−1 )
In general, ℎ𝑖𝑗𝑡 , will depend on the squared residuals, cross products of the residuals, and the
conditional variances and conditional co-variances of all variables in the system. As such, the
model allows for shocks to the variance of one variable to spill over to the others.
The problem is that the BEK formulation can be quite difficult to estimate. The model has a large
number of parameters that are not globally identified. Changing the sign of all elements of A, B,
C will have no effect on the value of the likelihood function. As such, convergence can be quite
difficult to achieve.

Constant Conditional Correlation Model


CCC model is another MGARCH specification. In this model restrictions are that the correlation
coefficients of conditional variance to be constant. As such for each ≠ 𝑗 ℎ𝑖𝑗𝑡 = 𝜌𝑖𝑗 (ℎ𝑖𝑖𝑡 ℎ𝑗𝑗𝑡 )0.5.
In a sense, the CCC model does not compromise with the variance term but covariance terms are
always proportional to (ℎ𝑖𝑖𝑡 ℎ𝑗𝑗𝑡 )0.5. In our initial model of two variable ℎ11𝑡 and ℎ22𝑡 equations
are same in CCC model but ℎ12𝑡 = 𝜌12 (ℎ11𝑡 ℎ22𝑡 )0.5. Hence ℎ12𝑡 equation entails only one
parameter instead of the seven parameters appearing in (2). CCC model has an important
advantage over the separate estimation of each equation. CCC model captures the
contemporaneous correlation between the various error terms. As such, the coefficients estimates
of the CCC model are more efficient than those from a single equation estimations.

34
Maximum Likelihood Estimation for Multivariate GARCH Model
To explain the method of estimation first we consider a two variable model (1) to (3). Suppose
that 𝜖1𝑡 and 𝜖2𝑡 are zero mean random variables that are jointly normally distributed. Initially we
assume that the variance and the covariance terms are constant. As such, we can drop the time
subscripts on the ℎ𝑖𝑗𝑡 . In such a circumstance the log likelihood function for the joint realisation
of 𝜖1𝑡 and 𝜖2𝑡 is
1 1 𝜖2 𝜖2 2𝜌 𝜖 𝜖
𝐿𝑡 = exp [− 2(1−𝜌2 (ℎ1𝑡 + ℎ2𝑡 − (ℎ 12ℎ 1𝑡)0.5
2𝑡
)] .........................(8)
2 )
2𝜋√ℎ11 ℎ22 (1−𝜌12 12 ) 11 22 11 22

ℎ12
Where 𝜌12 is the correlation coefficient between 𝜖1𝑡 and 𝜖2𝑡 ; 𝜌12 = ⁄(ℎ ℎ )0.5
11 22

ℎ ℎ12
Now if we define 𝐻 = [ 11 ] the likelihood function can be written as
ℎ21 ℎ22
1 1
𝐿𝑡 = 1 exp [− 2 𝜖𝑡 ′ 𝐻−1 𝜖𝑡 ] .....................................(9)
2𝜋|𝐻|2

Where 𝜖𝑡 = (𝜖1𝑡 𝜖2𝑡 )′ and |𝐻| is the determinant of H. To see that the two representations
2
given by (8) and (9) are equivalent. Note that |𝐻| = ℎ11 ℎ22 − ℎ12 . Since ℎ12 = 𝜌12 (ℎ11 ℎ22 )0.5
2
, it follows that |𝐻| = [(1 − 𝜌12 ) ℎ11ℎ22] . Moreover,
2 2
𝜖1𝑡 ℎ22 − 2𝜖1𝑡 𝜖2𝑡 ℎ12 + 𝜖2𝑡 ℎ11
𝜖𝑡 ′ 𝐻−1 𝜖𝑡 = 2
ℎ11 ℎ22 − ℎ12
Since ℎ12 = 𝜌12 (ℎ11 ℎ22 )0.5 , it follows that
2 2
1 𝜖1𝑡 𝜖2𝑡 2𝜌12 𝜖1𝑡 𝜖2𝑡
𝜖𝑡 ′ 𝐻 −1 𝜖𝑡 = [ 2 ( + − )]
(1 − 𝜌12 ) ℎ11 ℎ22 (ℎ11 ℎ22 )0.5
Since the realisation of {𝜖𝑡 } are conditionally dependent, the conditional likelihood of the joint
realisations of 𝜖1 𝜖2 … … … … . 𝜖 𝑇 is the product in the individual likelihood. Hence, if all have the
same variance, the conditional likelihood of the joint realisation is
𝑇
1 1 ′ −1
𝐿=∏ exp [− 𝜖 𝐻 𝜖𝑡 ]
1 2 𝑡
𝑡=1 2𝜋|𝐻|2

It is far easier to work with a sum than with a product. As such it is convenient to take the natural
log of each side so as to obtain
𝑇
𝑇 𝑇 1
ln 𝐿 = − ln(2𝜋) − ln|𝐻| − ∑ 𝜖𝑡 ′ 𝐻−1 𝜖𝑡
2 2 2
𝑡=1

The procedure used in MLE is to select the distributional parameters so as to maximise the
likelihood of drawing the observed sample. Given the realisation in 𝜖𝑡 , it is possible to select
ℎ11 , ℎ12 𝑎𝑛𝑑 ℎ22 so as to maximise the likelihood function.

35
Let us now allow the values of ℎ𝑖𝑗 to be time-varying. Then
𝑇
𝑇 1
ln 𝐿 = − ln(2𝜋) − ∑(ln|𝐻𝑡 | + 𝜖𝑡 ′ 𝐻−1 𝜖𝑡 )
2 2
𝑡=1

For k variable case


𝑇
𝑇𝑘 1
ln 𝐿 = − ln(2𝜋) − ∑(ln|𝐻𝑡 | + 𝜖𝑡 ′ 𝐻−1 𝜖𝑡 )
2 2
𝑡=1

Here 𝐻𝑡 is symmetric 𝑘 x 𝑘 matrix, 𝜖𝑡 is a 𝑘 x 1 column vector.

References
Enders, Walter, Applied Econometric Time Series, 4th ed., Wiley 2015.
Brooks, Chris, Introductory Econometrics for Finance, Cambridge University Press,
2002.
Hamilton, James D., Time Series Analysis, Princeton University Press, 1994.

36

Common questions

Powered by AI

In univariate GARCH models, conditional covariance is not explicitly modeled, as the focus is on the conditional variance of a single time series. However, in multivariate GARCH models, the conditional covariance between series is explicitly modeled to capture the co-movements and potential correlation in volatilities between different time series. This is achieved by modeling conditional covariance as a function of past covariances and variances of the involved series, which allows for a more comprehensive understanding of the relationships and potential volatility spillovers among the series, which are critical in multivariate contexts like portfolio management .

For a GARCH(1,1) model to remain stationary, the sum of \( \alpha_1 \) and \( \beta_1 \) must be less than one. This condition is critical because it ensures that the influence of shocks to conditional variance diminishes over time, allowing the model to eventually revert to a long-term mean and preventing it from exploding to infinity. Without this condition, the model could produce non-stationary processes that do not accurately reflect the real-world data .

Estimating parameters for a Multivariate GARCH (p,q) model is challenging due to the large number of parameters, as each additional variable significantly increases the complexity. This can lead to overparameterization and difficulties in ensuring positive definiteness of variance-covariance matrices. Typically, numerical optimization techniques are used for estimation, though these can struggle with convergence if the model is too complex. To address these issues, constraints are often imposed to reduce parameter space, and simpler models like the diagonal BEKK are employed to ensure tractability while still capturing essential cross-series interactions .

The GARCH(1,1) model addresses volatility clustering by modeling the current conditional variance as a function of its past values and past squared observations. This is captured through the equation \( h_t = \alpha_0 + \alpha_1 \epsilon_{t-1}^2 + \beta_1 h_{t-1} \). Here, \( \alpha_1 \) represents the impact of the last period's squared innovation (the ARCH term), and \( \beta_1 \) represents the influence of the last period's forecasted variance (the GARCH term). This allows the model to increase conditional variance when large scale innovations are followed by large scale conditional variances, reflecting periods of clustering volatility .

The leverage effect, which posits that negative asset returns often lead to higher subsequent volatility than positive returns of the same magnitude, has significant implications in volatility modeling. Traditional GARCH models may not always capture this asymmetry effectively, as they typically assume symmetric response to innovations. This has led to extensions like the EGARCH or TGARCH models that are specifically designed to incorporate the leverage effect by allowing different responses to positive and negative shocks. This can improve model fit and forecasting accuracy for assets that exhibit significant asymmetric volatility behavior .

Ensuring positivity of conditional variances in MGARCH models is crucial because negative variances would be non-sensical in a financial context where variance represents squared deviations, which are inherently non-negative. To maintain this, models such as the BEKK impose parameter structures that guarantee positive definiteness through quadratic forms, ensuring all variances remain positive. This structural adjustment helps avoid negative or undefined standard deviations and correlations, thus maintaining the model's viability and interpretability .

The GARCH(1,1) model is more parsimonious than the ARCH model because it requires fewer parameters to adequately capture the volatility dynamics in time series data. Specifically, while the ARCH model generally requires many lagged squared terms to capture long memory in volatility, the GARCH(1,1) model captures this by incorporating lagged conditional variances, which reduces the number of parameters needed and minimizes the risk of overfitting. This parsimonious nature of GARCH models makes them more efficient in terms of parameter estimation and less likely to violate non-negativity constraints .

The BEKK model improves upon the basic GARCH structure in multivariate settings by ensuring that all conditional variances are positive. It achieves this by parameterizing the model in a way that the variables enter through quadratic terms, which inherently are non-negative. This approach not only maintains the positivity constraint but also allows individual variances and covariances to influence each other, effectively capturing volatility spillovers and co-movements among multiple financial time series .

Multivariate GARCH models, such as the Vech or BEKK models, handle volatility spillover by specifying equations that account for interactions between the variances of each series, and between variances and covariances themselves. In these models, the conditional variance of each series depends not only on its past squared innovations and previous variances but also on past innovations and covariances of other series in the system. This setup allows shocks to one series to impact the volatility of another series, reflecting the interconnected nature of financial markets and accurately capturing the dynamics of co-volatility among multiple time series .

In a GARCH(1,1) model, the conditional variance is mean-reverting, which means it has a tendency to revert to a long-term average over time. The rate of this mean reversion is determined by the sum \( \alpha_1 + \beta_1 \); a smaller sum indicates faster reversion and quicker stabilization of variance forecasts towards the unconditional variance \( \alpha_0/(1 - \alpha_1 - \beta_1) \). This property ensures that any shock to volatility is transient, and long-term forecasts of volatility will eventually stabilize at the long-run variance, adjusting only gradually to new information .

You might also like