0% found this document useful (0 votes)
16 views45 pages

Understanding Volatility in Finance

The document discusses volatility as a risk measure, defining it through standard deviation and highlighting its implications in option pricing and market behavior. It emphasizes that implied volatility tends to be overpriced and explores the non-normal distribution of daily changes in exchange rates, suggesting power-law distributions as better alternatives. Additionally, it covers volatility estimation methods, including the Exponentially Weighted Moving Average Model (EWMA), which is useful for forecasting volatility in financial markets.

Uploaded by

jwlin2002
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
16 views45 pages

Understanding Volatility in Finance

The document discusses volatility as a risk measure, defining it through standard deviation and highlighting its implications in option pricing and market behavior. It emphasizes that implied volatility tends to be overpriced and explores the non-normal distribution of daily changes in exchange rates, suggesting power-law distributions as better alternatives. Additionally, it covers volatility estimation methods, including the Exponentially Weighted Moving Average Model (EWMA), which is useful for forecasting volatility in financial markets.

Uploaded by

jwlin2002
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

– Volatility as a risk measure

• Definition of Volatility
– Suppose that Si is the value of a variable on the day i. The volatility per
day is the standard deviation of ln(Si /Si-1)
» Normally the days when markets are closed are ignored in volatility
calculations [volatility is most generated by trading itself]
– The volatility per year is 252 times the daily volatility
» The variance rate is the square of the volatility
– Volatility measures the average risk (not specifically for tail risks)
• Implied Volatility
– Of the variables needed to price an option, the one that cannot be
observed directly is the volatility
– We can imply volatilities from market prices -- One can get a better feel
for whether the option is over/undervalued from IV
– IV depends on its strike price and time to expiration, leading to an
implied volatility surface (it captures both the average and tail risks)
– Volatility smile/skew (due to fat tails of the return distribution)
• VIX Index: A Measure of the Implied Volatility of the S&P 500
90.00

80.00

70.00

60.00

50.00

40.00

30.00

20.00

10.00

0.00

– Implied volatility on average is overpriced


• Selling implied volatility is one of the most popular trading strategies for
equity derivatives
• In developed markets, volatility (implied volatility) has a negative
correlation with the stock market
– The current futures price of volatility > future realized volatility
– Implied volatility on average is overpriced
• ETN like VXX (based on short-term VIX futures) is a popular instrument
to trade, but should not be used as a long-term hedge.

• Volatility overpricing is unlikely to disappear


– Demand for put protection
» Variable annuity products --- total market value of more than 1T
USD
» Index implied volatility demand from structured products
» Market makers charge a premium for net selling volatility
– Are Daily Changes in Exchange Rates Normally Distributed?
Real World (%) Normal Model (%)
>1SD 23.32 31.73
>2SD 4.67 4.55
>3SD 1.3 0.27
>4SD 0.49 0.01
>5SD 0.24 0
>6SD 0.13 0

• Daily exchange rate changes are


not normally distributed
• Heavy Tails: The distribution has
heavier tails than the normal distribution
• It is also more peaked
• Many market variables have the
property, known as excess kurtosis
• In the mid-1980s few traders were aware of heavy tails and volatility
smiles in foreign currency options-- OTM options were under-priced.
These profit opportunities disappeared later

– Alternatives to Normal Distributions: The Power Law (a special case of Pareto
distribution from EVT)
Prob(v > x) = Kx-α

This seems to fit the behavior of the returns on many market variables better
than the normal distribution
Cubic law for stock market returns, 𝛼𝛼 = 3
– Log-Log Test for Exchange Rate Data (v is the number of standard deviations
which the exchange rate moves)
0
ln[Prob(v > x)] = ln(K) –α ln(x) 0 0.5 1 1.5 2
-1
ln(x)
-2

ln(Prob(v>x)
-3

-4

-5

-6

-7
On a Side Note
For a normal distribution
Distribution of heights of men in China: 172 ± 7 cm (2.31m: 7σ deviation)
Prob( 𝑥𝑥 > 𝑥𝑥̅ + 3𝜎𝜎) : 0.27%
Prob(𝑥𝑥 > 𝑥𝑥̅ + 7𝜎𝜎) : 2.5x10-10
Prob(𝑥𝑥 > 𝑥𝑥̅ + 20𝜎𝜎) : 2.7x10-87 (black swan events)
not a good distribution for risk management
– We do have 20σ black swan events:
1987 Black Monday, 2015 Swiss Franc’s 30% move
Power-law tail distribution is a better choice.
Gutenberg-Richter law: P(>E) : E-b, valid for an enormous range of energies.
The magnitude of 4: 1000 tons of TNT (happens 6000 times a year)
The magnitude of 8: 32 billion tons (every ten years)
no typical-size earthquake can be defined!
Thinking in terms of a normal distribution ~ thinking in terms of a fixed
typical/maximum value (could it contribute to the Fukushima nuclear
disaster?)
• Can the stock market return be described by the normal
distribution? Date SP500 Return
1/2/2003 909.03
𝑟𝑟̅ = 0.00036, 𝜎𝜎 = 0.012
1/3/2003 908.59 -0.00048
Return of SP500
1/6/2003 929.01 0.022474
60
1/7/2003 922.93 -0.00654
40
1/8/2003 909.93 -0.01409
20 1/9/2003 927.58 0.019397
0 1/10/2003 927.57 -1.1E-05
-0.12 -0.07 -0.02 0.03 0.08 1/13/2003 926.26 -0.00141
-20
1/14/2003 931.66 0.00583
Probability Density Normal Distribution
1/15/2003 918.22 -0.01443
Return of SP500 1/16/2003 914.6 -0.00394
100
SP500
6000
0.1
-0.12 -0.07 -0.02 0.03 0.08
0.0001 5000

0.0000001
4000
1E-10
3000
1E-13

1E-16 2000

1E-19
1000
1E-22
0
Probability Density Normal Distribution 1/9/2002 28/5/2005 22/2/2008 18/11/2010 14/8/2013 10/5/2016 4/2/2019 31/10/2021
• The empirical cubic law for the tail distribution
𝑪𝑪
Right tail: 𝑷𝑷𝑷𝑷𝑷𝑷𝑷𝑷 𝒓𝒓 > 𝒙𝒙 = 𝒙𝒙𝟑𝟑𝟏𝟏
𝑪𝑪
Left tail: 𝑷𝑷𝑷𝑷𝑷𝑷𝑷𝑷 𝒓𝒓 < 𝒙𝒙 = (−𝒙𝒙)
𝟐𝟐
𝟑𝟑

Cumulative Probability
0.07

0.06

0.05

0.04

0.03

0.02

0.01

0
-0.12 -0.11 -0.1 -0.09 -0.08 -0.07 -0.06 -0.05 -0.04 -0.03 -0.02
-0.01

CPD CPD (Normal) CPD(Cubic Law)


• Can the stock market return be described by the normal
distribution?
– A closer look at the left tail (in log scale)
Cumulative Probability Cumulative Probability
1 1
-0.12 -0.1 -0.08 -0.06 -0.04 -0.02 -0.12 -0.1 -0.08 -0.06 -0.04 -0.02
0.001

1E-06 0.1

1E-09

1E-12 0.01

1E-15

1E-18 0.001

1E-21

1E-24 0.0001

CPD CPD (Normal) CPD(Cubic Law) CPD CPD (Normal) CPD(Cubic Law)

– Other financial tail risks can also be described by power-laws


– Volatility Estimating and Forecasting
• Estimating volatility from historical data
The standard way
𝑆𝑆𝑖𝑖
– 𝑢𝑢𝑖𝑖 = 𝑙𝑙𝑙𝑙
𝑆𝑆𝑖𝑖−1
– An unbiased estimate of the variance rate per day
𝑚𝑚 𝑚𝑚
1 1
𝜎𝜎𝑛𝑛 2 = � 2,
�(𝑢𝑢𝑛𝑛−𝑖𝑖 − 𝑢𝑢) 𝑢𝑢� = � 𝑢𝑢𝑛𝑛−𝑖𝑖
𝑚𝑚 − 1 𝑚𝑚
𝑖𝑖=1 𝑖𝑖=1
The following is often used in practice
𝑆𝑆𝑖𝑖 −𝑆𝑆𝑖𝑖−1
1. 𝑢𝑢𝑖𝑖 = 𝑆𝑆𝑖𝑖−1
2. 𝑢𝑢� is assumed to be zero
3. m-1 is replaced by m (maximum likelihood estimate)

𝑚𝑚
1
𝜎𝜎𝑛𝑛 2 = � 𝑢𝑢𝑛𝑛−𝑖𝑖 2
𝑚𝑚
𝑖𝑖=1
• Weighting Schemes
Compared to stock returns, volatility is easier to forecast
– Volatility clustering, Mean-reverting
It makes sense to give more weight to recent data (volatility clustering)
𝜎𝜎𝑛𝑛 2 = ∑𝑚𝑚 2 𝑚𝑚
𝑖𝑖=1 𝛼𝛼𝑖𝑖 𝑢𝑢𝑛𝑛−𝑖𝑖 , ∑𝑖𝑖=1 𝛼𝛼𝑖𝑖 = 1, 𝛼𝛼𝑖𝑖 < 𝛼𝛼𝑗𝑗 when i > j
An extension of this is to include a long-term average variance rate
𝜎𝜎𝑛𝑛 2 = 𝛾𝛾𝑉𝑉𝐿𝐿 + ∑𝑚𝑚 2
𝑖𝑖=1 𝛼𝛼𝑖𝑖 𝑢𝑢𝑛𝑛−𝑖𝑖 , 𝛾𝛾 + ∑𝑚𝑚𝑖𝑖=1 𝛼𝛼𝑖𝑖 = 1
• The Exponentially Weighted Moving Average Model (EWMA)
𝜎𝜎𝑛𝑛 2 = 𝜆𝜆𝜎𝜎𝑛𝑛−1 2 + (1 − 𝜆𝜆) 𝑢𝑢𝑛𝑛−1 2 (0<𝜆𝜆 < 1) --- online algorithm
This can be applied iteratively
𝜎𝜎𝑛𝑛 2 = 𝜆𝜆(𝜆𝜆𝜎𝜎𝑛𝑛−2 2 + 1 − 𝜆𝜆 𝑢𝑢𝑛𝑛−2 2 ) + (1 − 𝜆𝜆) 𝑢𝑢𝑛𝑛−1 2
= 𝜆𝜆2 (𝜆𝜆𝜎𝜎𝑛𝑛−3 2 + 1 − 𝜆𝜆 𝑢𝑢𝑛𝑛−3 2 ) + 1 − 𝜆𝜆 𝜆𝜆𝑢𝑢𝑛𝑛−2 2 + (1 − 𝜆𝜆) 𝑢𝑢𝑛𝑛−1 2
= (1 − 𝜆𝜆) ∑𝑚𝑚 𝑖𝑖=1 𝜆𝜆
𝑖𝑖−1 𝑢𝑢 2 𝑚𝑚
𝑛𝑛−𝑖𝑖 + 𝜆𝜆 𝜎𝜎𝑛𝑛−𝑚𝑚
2

For large m, the last term is very small. It is the weighted average with
𝛼𝛼𝑖𝑖 = 1 − 𝜆𝜆 𝜆𝜆𝑖𝑖−1 = 1 − 𝜆𝜆 𝑒𝑒 − 𝑖𝑖−1 𝑙𝑙𝑙𝑙𝑙𝑙
• EWMA (cont..)
𝜎𝜎𝑛𝑛 2 = 𝜆𝜆𝜎𝜎𝑛𝑛−1 2 + (1 − 𝜆𝜆) 𝑢𝑢𝑛𝑛−1 2 (0<𝜆𝜆 < 1)
» 𝜆𝜆 = 0.94 has been found to be a good choice across a wide
range of market variables
𝜎𝜎𝑛𝑛 2 = 0.94𝜎𝜎𝑛𝑛−1 2 + 0.06 𝑢𝑢𝑛𝑛−1 2
» Little data need to be stored
2
Day Price Return Return σ2(EWMA)
0 20
1 20.1 0.005 2.5E-05 2.5E-05
2 19.9 -0.01 9.901E-05 2.5E-05
3 20 0.00503 2.525E-05 2.94404E-05
4 20.5 0.025 0.000625 2.91891E-05
5 20.25 -0.0122 0.0001487 6.49378E-05
6 20.9 0.0321 0.0010303 6.99648E-05
7 20.9 0 0 0.000127587
8 20.9 0 0 0.000119932
9 20.6 -0.0144 0.000206 0.000112736
• The GARCH(1,1) Model
Combining EWMA and a long-run average
𝜎𝜎𝑛𝑛 2 = 𝛾𝛾𝑉𝑉𝐿𝐿 + 𝛼𝛼𝑢𝑢𝑛𝑛−1 2 + 𝛽𝛽𝜎𝜎𝑛𝑛−1 2 = 𝜔𝜔 + 𝛼𝛼𝑢𝑢𝑛𝑛−1 2 + 𝛽𝛽𝜎𝜎𝑛𝑛−1 2
𝛾𝛾 + 𝛼𝛼 + 𝛽𝛽 = 1, 𝛼𝛼 + 𝛽𝛽 < 1
– EWMA is a particular case with 𝛾𝛾=0, 𝛼𝛼=1-𝜆𝜆, 𝛽𝛽=𝜆𝜆
– 𝛽𝛽 describes the decay rate
– GARCH(1,1) model has the property that over time the variance
tends to get pulled back to a long-run average level
– The parameters can be estimated using maximum likelihood
methods
using data for the S&P index between July 18, 2005 and Aug. 13,
2010, one can estimate
𝛼𝛼=0.083394, 𝛽𝛽=0.910116, 𝜔𝜔=0.0000013465
VL =𝜔𝜔/(1-𝛼𝛼-𝛽𝛽)=0.0002075, this gives the long-term volatility of
1.4404% per day (~23% per annum)
• Using Garch(1,1) to Forecast Future Volatility
From
𝜎𝜎𝑛𝑛 2 = 1 − 𝛼𝛼 − 𝛽𝛽 𝑉𝑉𝐿𝐿 + 𝛼𝛼𝑢𝑢𝑛𝑛−1 2 + 𝛽𝛽𝜎𝜎𝑛𝑛−1 2
we have
𝜎𝜎𝑛𝑛 2 − 𝑉𝑉𝐿𝐿 = 𝛼𝛼(𝑢𝑢𝑛𝑛−1 2 − 𝑉𝑉𝐿𝐿 ) + 𝛽𝛽(𝜎𝜎𝑛𝑛−1 2 − 𝑉𝑉𝐿𝐿 )
On day n+t in the future
𝜎𝜎𝑛𝑛+𝑡𝑡 2 − 𝑉𝑉𝐿𝐿 = 𝛼𝛼(𝑢𝑢𝑛𝑛+𝑡𝑡−1 2 − 𝑉𝑉𝐿𝐿 ) + 𝛽𝛽(𝜎𝜎𝑛𝑛+𝑡𝑡−1 2 − 𝑉𝑉𝐿𝐿 )

Since E[𝑢𝑢𝑛𝑛+𝑡𝑡−1 2 ] is E[𝜎𝜎𝑛𝑛+𝑡𝑡−1 2 ]


𝐸𝐸 𝜎𝜎𝑛𝑛+𝑡𝑡 2 − 𝑉𝑉𝐿𝐿 = 𝛼𝛼 + 𝛽𝛽 𝐸𝐸(𝜎𝜎𝑛𝑛+𝑡𝑡−1 2 − 𝑉𝑉𝐿𝐿 )= 𝛼𝛼 + 𝛽𝛽 𝑡𝑡
(𝜎𝜎𝑛𝑛 2 − 𝑉𝑉𝐿𝐿 )
Thus,
𝐸𝐸 𝜎𝜎𝑛𝑛+𝑡𝑡 2 = 𝑉𝑉𝐿𝐿 + 𝛼𝛼 + 𝛽𝛽 𝑡𝑡 (𝜎𝜎𝑛𝑛 2 − 𝑉𝑉𝐿𝐿 )
1
Define V t = 𝐸𝐸 𝜎𝜎𝑛𝑛+𝑡𝑡 2 and 𝑎𝑎 = 𝑙𝑙𝑙𝑙 𝛼𝛼+𝛽𝛽 > 0, we have
V(t) = VL+e-at[V(0)-VL]
The forecast of the future variance rate tends towards VL as we look
further and further ahead.
• Volatility Term Structures
The average variance rate per day between today and time T is given by
1 𝑇𝑇 1 − 𝑒𝑒 −𝑎𝑎𝑎𝑎
� 𝑉𝑉 𝑡𝑡 𝑑𝑑𝑑𝑑 = 𝑉𝑉𝐿𝐿 + [𝑉𝑉 0 − 𝑉𝑉𝐿𝐿 ]
𝑇𝑇 0 𝑎𝑎𝑎𝑎
Define σ(T) as the volatility per annum that should be used to price a T-
day option under GARCH(1,1)
2
1 − 𝑒𝑒 −𝑎𝑎𝑎𝑎
𝜎𝜎(𝑇𝑇) = 252(𝑉𝑉𝐿𝐿 + [𝑉𝑉 0 − 𝑉𝑉𝐿𝐿 ])
𝑎𝑎𝑎𝑎
This is the volatility term structure predicted from GARCH(1,1), It can be
used to understand the general feature of the term structure, and how the
term structure responds to the volatility change.
For the example of the S&P 500 data we considered earlier, we have
a=0006511, VL=0.0002075. Suppose current variance V(0)=0.0003 per day
S&P 500 volatility term structure from GARCH(1,1) [ σ𝐿𝐿 = 252𝑉𝑉𝐿𝐿 =22.9%]
Option life(days) 10 30 50 100 500
Volatility (% per annum) 27.36 27.10 26.87 26.35 24.32
– Maximum Likelihood Methods
• In maximum likelihood methods, we choose parameters that
maximize the likelihood of the observations occurring
– Example 1:
We observe that a certain event happens one time in ten trials.
What is our estimate of the proportion of the time, p, that it
happens?
The probability of the outcome is p(1-p)9
We maximize this to obtain a maximum likelihood estimate:
p = 0.1 [(1-p)9 – 9p(1-p)8 = 0 1-p-9p = 0 p=1/10]
– Example 2: Estimate the variance of observations from a normal
distribution with a mean of zero
1 −u2i 1 u2i
Maximize: ∏𝑛𝑛𝑖𝑖=1 exp 2ν ; same as maximizing ∑𝑛𝑛𝑖𝑖=1 − ln 𝜈𝜈 −
2πν 2 2ν
1 𝑛𝑛
The maximum value reached when 𝜈𝜈 = ∑𝑖𝑖=1 𝑢𝑢𝑖𝑖2
𝑛𝑛
– EWMA and GARCH parameter estimation can be done using the
same objective functions (numerically)
• S&P 500 (July 2005 – Aug 2010)
1800
1600
1400
1200
1000
800
600
400
200
0
Jul-05 Jul-06 Jul-07 Jul-08 Jul-09 Jul-10

• The GARCH Estimate of Volatility of the S&P 500


ω=0.0000013465, α=0.083394, β=0.910116
• Variance Targeting
– One way of implementing GARCH(1,1) that increases stability is by
using variance targeting
– The long-run average variance sets to equal the sample variance
– Only two other parameters then have to be estimated
• How Good is the Model?
– The Ljung-Box statistic tests for autocorrelation
– We compare the autocorrelation of the
ui2 with the autocorrelation of the ui2/σi2
Time Autocorrelation Autocorrelation
Lag for u2 u2/σ2
1 0.183 -0.063
2 0.385 -0.004
3 0.16 -0.007
4 0.301 0.022
5 0.339 0.014
6 0.308 0.014
– Note: These are from the in-sample analysis.
• Measuring Risk
– Correlation and Copulas (important for evaluating systematic
risk)
• The coefficient of correlation between two variables V1 and V2
is defined as
𝐸𝐸 𝑉𝑉1 𝑉𝑉2 − 𝐸𝐸 𝑉𝑉1 𝐸𝐸 𝑉𝑉2
𝑆𝑆𝑆𝑆 𝑉𝑉1 𝑆𝑆𝑆𝑆(𝑉𝑉2 )
• The covariance is
𝐸𝐸 𝑉𝑉1 𝑉𝑉2 − 𝐸𝐸 𝑉𝑉1 𝐸𝐸 𝑉𝑉2
• Independence
– V1 and V2 are independent if the knowledge of one does not
affect the probability distribution for the other
𝑓𝑓 𝑉𝑉2 𝑉𝑉1 = 𝑥𝑥 = 𝑓𝑓(𝑉𝑉2 )
where f(.) denotes the probability density function
• Independence is not the same as zero correlation
– Suppose V1 = –1, 0, or +1 (equally likely)
– If V1 = -1 or V1 = +1 then V2 = 1
– If V1 = 0 then V2 = 0
V2 is clearly dependent on V1 (and vice versa) but the
coefficient of correlation is zero
• Examples: Types of Dependence
E(V2) E(V2)
V1 V1

(a) (b)
E(V2)

V1

(c)
• Monitoring Correlation Between Two Variables X and Y
Define xi=(Xi−Xi-1)/Xi-1 and yi=(Yi−Yi-1)/Yi-1
Also
varx,n: daily variance of X
vary,n: daily variance of Y
covn: covariance between x and y
These are calculated based on the data on day n-1 and before

Covn = E(xy)−E(x)E(y)
---- It is usually approximated as E(xy) when the
averages of the variables are small
For a simple estimate
𝑚𝑚
1
𝑐𝑐𝑐𝑐𝑐𝑐𝑛𝑛 = � 𝑥𝑥𝑛𝑛−𝑖𝑖 𝑦𝑦𝑛𝑛−𝑖𝑖
𝑚𝑚
𝑖𝑖=1
The correlation is
𝒄𝒄𝒄𝒄𝒄𝒄𝒏𝒏
𝒗𝒗𝒗𝒗𝒗𝒗𝒙𝒙,𝒏𝒏 𝒗𝒗𝒗𝒗𝒗𝒗𝒚𝒚,𝒏𝒏
• Estimating Covariance
EWMA: covn = 𝜆𝜆covn-1 + (1-𝜆𝜆)xn-1yn-1
GARCH(1,1): covn = 𝜔𝜔 + 𝛼𝛼xn-1yn-1 + 𝛽𝛽covn-1
• Positive Finite Definite Condition
A variance-covariance matrix Ω, is internally consistent if the positive
semi-definite condition
wTΩw ≥ 0 holds for all vectors w
wTΩw is the variance of the portfolio with weights given by w
This ensures that variances are nonnegative for all principal
components
– Example:
1 0 0.9
0 1 0.9 is not internally consistent (try w=[1,1,-1]T); this cannot
0.9 0.9 1
be a covariance matrix. (𝑤𝑤 𝑇𝑇 Ω𝑤𝑤 = −0.6)
When estimating variance/covariance, the same numerical procedure
should be used to ensure consistency (for example, use the same 𝜆𝜆 for
EWMA for variance and covariance)
– Estimation of covariance -- Factor Models
1. When there are N variables, Vi (i = 1, 2,..N), there are N(N−1)/2
covariance/correlations
2. In risk management, one never estimates all covariance directly
3. We can reduce the number of correlation parameters that have to be
estimated with a factor model and express covariance in terms of
factor covariance.
• A simple example: One-Factor Model (Gaussian variables)
– Correlations among variables following a normal distribution can be
defined easily. If {Ui } follows the standard normal distribution we can
set Ui = ai F + 1 − 𝑎𝑎𝑖𝑖2 Zi , where the common factor F and the
idiosyncratic component Zi have independent standard normal
distributions (𝐸𝐸 𝑍𝑍𝑖𝑖 𝑍𝑍𝑗𝑗 = 𝛿𝛿𝑖𝑖𝑖𝑖 , etc.)
Note var(Ui) = 1 because E(F2) =1, E(Zi2) =1, E(FZi) = 0

𝑣𝑣𝑣𝑣𝑣𝑣(𝑈𝑈𝑖𝑖 ) = 𝐸𝐸 𝑈𝑈𝑖𝑖2 = 𝑎𝑎𝑖𝑖2 𝐸𝐸 𝐹𝐹 2 + 2𝑎𝑎𝑖𝑖 1 − 𝑎𝑎𝑖𝑖2 𝐸𝐸 𝐹𝐹𝑍𝑍𝑖𝑖 + 1 − 𝑎𝑎𝑖𝑖2 𝐸𝐸(𝑍𝑍𝑖𝑖2 )

Correlation between Ui and Uj (i≠j), 𝐸𝐸 𝑈𝑈𝑖𝑖 𝑈𝑈𝑗𝑗 = 𝑎𝑎𝑖𝑖 𝑎𝑎𝑗𝑗 𝐸𝐸 𝐹𝐹 2 = 𝑎𝑎𝑖𝑖 𝑎𝑎𝑗𝑗
• What should we do if the variable does not follow the normal distribution?
– Gaussian Copula Model:

-0.2 0 0.2 0.4 0.6 0.8 1 1.2 -0.2 0 0.2 0.4 0.6 0.8 1 1.2

V1 V2

- -

-6 -4 -2 0 2 4 6 -6 -4 -2 0 2 4 6

U1 U2

N(x)

F(V1) = N(U1) : V1  U1
• Examples:

V1 V2
• V1 Mapping to U1 and V2 Mapping to U2

V1 Percentile U1 V2 Percentile U2

0.2 20 -0.84 0.2 8 −1.41


0.4 55 0.13 0.4 32 −0.47
0.6 80 0.84 0.6 68 0.47
0.8 95 1.64 0.8 92 1.41
• Example of Calculation of Joint Cumulative Distribution
– The probability that V1 and V2 are both less than 0.2 is the
probability that U1 < −0.84 and U2 < −1.41
– When the copula correlation is 0.5, this probability is
M( −0.84, −1.41, 0.5) = 0.043.
where M is the cumulative distribution function for the bivariate
normal distribution
1 𝑈𝑈12 + 𝑈𝑈22 − 2𝜌𝜌𝑈𝑈1 𝑈𝑈2
exp −
2𝜋𝜋 1 − 𝜌𝜌 2 2(1 − 𝜌𝜌2 )
– 𝜌𝜌 is often estimated using the maximum-likelihood estimation
• 5000 Random Samples from the Bivariate Normal Dist. ρ=0.5.
5
4
3
2
1
0
-5 -4 -3 -2 -1 0 1 2 3 4 5
-1
-2
-3
-4
-5
• Multivariate Gaussian Copula
– We can similarly define a correlation structure between
V1, V2,…Vn
– We transform each variable Vi to a new variable Ui that
has a standard normal distribution on a “percentile-to-
percentile” basis.
– The U’s are assumed to have a multivariate normal
distribution
• Factor Copula Model (an alternative way to define the
correlation structure)
– In a factor copula model, the correlation structure between the
U’s is generated by assuming one or more factors.
– Value at Risk and Expected Shortfall
• The VaR is measuring tail risk
– With an X% confidence, the N-day loss will not exceed theVaR. ”

Gain Loss

-VaR VaR

• VaR and Regulatory Capital


– Regulators have traditionally used VaR to calculate the capital
they require banks to keep
– The market-risk capital is based on a 10-day VaR estimated
where the confidence level is 99%
– Credit risk and operational risk capital --- one-year 99.9% VaR
• Expected Shortfall vs. VaR
– VaR is the loss level that will not be exceeded with a specified
probability
– Expected shortfall (ES) is the expected loss given that the loss is
greater than the VaR level (also called C-VaR and Tail Loss)
𝐸𝐸𝐸𝐸 = [∑𝐿𝐿>𝑉𝑉𝑉𝑉𝑉𝑉 𝐿𝐿 ∗ 𝑝𝑝(𝐿𝐿)]/𝑃𝑃(𝐿𝐿 > 𝑉𝑉𝑉𝑉𝑉𝑉)
– Regulators have started to move from using VaR to using ES for
determining risk capital for the reason that ES is a coherent risk
measure while VaR is not.
• Coherent Risk Measures
– What properties does a coherent risk measure have?
1. If one portfolio always produces a worse outcome than another
its risk measure should be greater (monotonicity)
2. If we add an amount of cash K to a portfolio its risk measure
should go down by K (translation invariance)
3. Changing the size of a portfolio by a factor of λ should result in
the risk measure being multiplied by λ (homogeneity)
4. The risk measures for two portfolios after they have been
merged should be no greater than the sum of their risk
measures before they were merged (subadditivity)
• Example
– A bank has two $10 million one-year loans. Possible outcomes
are
Outcome Probability
Neither Loan Defaults 97.50%
Loan 1 defaults, loan 2 does not default 1.25%
Loan 2 defaults, loan 1 does not default 1.25%
Both loans default 0.00%
» Note the defaults are not independent in this example
– If a default occurs, losses between 0% and 100% are equally
likely. If a loan does not default, a profit of 0.2 million is made.
– What is the 99% VaR and the ES of each project?
» VaR=$2M [10M x (1.25%-1.0%)/1.25%] and
ES=($10+$2)/2= $6M
0.01
-2
0.0125
-10
– What are the 99% VaR and the ES for the portfolio?
» 0.025 probability the loss between -0.2 and $(10-0.2)M 
0.01 probability the loss between $(6-0.2)M to $(10-0.2)M
» VaR = 6-0.2 = $5.8M; ES = 8-0.2 = $7.8M
0.01 0.025

-5.8

-9.8

– VaR does not satisfy the subadditivity condition


5.8 > 2 + 2
– ES satisfies the subadditivity condition
7.8 < 6 + 6
• Formulas for Normal Distribution
– When losses (gains) are normally distributed with mean µ and standard
deviation σ (Given X the confidence level)
(loss – μ )/σ follows the standard normal distribution
• VaR = μ + σz, z = N-1(X) [ Prob(loss<VaR)=(N[(VaR- μ )/σ] = X]
Considering μ=0, VaR(X=99.9%)/VaR(X=99%)
=N-1(0.999)/N-1(0.99)=3.09/2.33=1.32!
if a power-law tail distribution is used, the answer makes better sense.
2
𝑒𝑒 −𝑧𝑧 /2
• ES = 𝜇𝜇 + 𝜎𝜎 2𝜋𝜋(1−𝑋𝑋)
This can be derived (considering the positive tail)
2 2
∞ 𝑒𝑒 −𝑦𝑦 /2 ∞ 𝑒𝑒 −𝑦𝑦 /2
ES = ∫𝑧𝑧 (𝜇𝜇 + 𝜎𝜎𝑦𝑦) 2𝜋𝜋 𝑑𝑑𝑑𝑑/ ∫𝒛𝒛 𝑑𝑑𝑑𝑑
2𝜋𝜋

z loss
• Changing the Time Horizon
– If losses in successive days are independent, normally distributed, and
have a mean of zero
T-day VaR = 1-day VaR x √T
T-day ES = 1-day ES x √T

• Aggregating VaRs
An approximate approach that seems to work well is

𝑉𝑉𝑉𝑉𝑉𝑉𝑡𝑡𝑡𝑡𝑡𝑡𝑡𝑡𝑡𝑡 = � � 𝑉𝑉𝑉𝑉𝑉𝑉𝑖𝑖 𝑉𝑉𝑉𝑉𝑉𝑉𝑗𝑗 𝜌𝜌𝑖𝑖𝑖𝑖


𝑖𝑖 𝑗𝑗

where VaRi is the VaR for the ith segment, VaRtotal is the total VaR, and ρij is
the coefficient of correlation between losses from the ith and jth segments
• Back-testing to check for the risk measure
– Back Test for VaR is relatively easy to do than ES.
– Back-testing a VaR involves looking at how often exceptions (loss >
VaR) occur --- Intuitively, for 99% VaR, 7% occurrence means the VaR
is likely to be under-estimated; on the other hand, 0.2% occurrence
means that VaR is likely to be over-estimated.
– A more quantitative statistical test
• Suppose that the theoretical probability of an exception is p (=1−X),
and the number of occurrences (out of n) is m with m/n > p.
– When can we reject the model for producing VaR being too
low? Consider the probability of m or more exceptions
𝑛𝑛 𝑚𝑚−1
𝑛𝑛! 𝑛𝑛!
� 𝑝𝑝𝑘𝑘 (1 − 𝑝𝑝)𝑛𝑛−𝑘𝑘 = 1 − � 𝑝𝑝𝑘𝑘 (1 − 𝑝𝑝)𝑛𝑛−𝑘𝑘
𝑘𝑘! 𝑛𝑛 − 𝑘𝑘 ! 𝑘𝑘! 𝑛𝑛 − 𝑘𝑘 !
𝑘𝑘=𝑚𝑚 𝑘𝑘=0
– A significant level often used is 5%. If the resulting probability is
less than 5%, the model can be rejected.
– Example: 600 days of data, consider 1 day 99%VaR, we can
calculate that if m>10, the probability is less than 5%, the model
should be rejected --- VaR is too low.
• Suppose that the number of occurrences (out of n) is m with m/n < p.
– When can we reject the model for producing VaR being high?
Consider the probability of m or fewer exceptions
𝑚𝑚
𝑛𝑛!
� 𝑝𝑝𝑘𝑘 (1 − 𝑝𝑝)𝑛𝑛−𝑘𝑘
𝑘𝑘! 𝑛𝑛 − 𝑘𝑘 !
𝑘𝑘=0
– Again, if the resulting probability is less than 5%, the model can be
rejected.
– The same example: 600 days of data, consider 1 day 99%VaR, we
can calculate that if m=1, the probability is 1.7%, less than 5%, the
model should be rejected --- VaR is too high. But m > 1 (m/n<p),
the probability is greater than 5%, and the model should not be
rejected.
• Combining the two tests: if m = 2, 3, …, 10, the model can’t be rejected.
• Two-tail test (Kupiec, 1995), based on
𝑛𝑛−𝑚𝑚 𝑚𝑚
𝑚𝑚 𝑛𝑛−𝑚𝑚 𝑚𝑚 𝑚𝑚
−2 ln 1 − 𝑝𝑝 𝑝𝑝 + 2 ln 1 −
𝑛𝑛 𝑛𝑛
This model should be rejected if this quantity > 3.84
– Historical Simulation to Determine the Risk Measure
• How do we do historical simulation?
– Collect data on the daily movements in all market variables
– Each day’s percentage changes in all market variables (for
example asset returns) are considered as a scenario to simulate
the change in the portfolio value (from today to tomorrow)
» For example, to calculate 1-day 99% VaR with 500-day data,
we can evaluate the changes in the portfolio value for these
500 scenarios --- VaR can be estimated as the fifth-worst loss
(ES can be estimated as the average of the top four losses).
– Some Technicalities
» Suppose we use n days of historical data (today is day n)
» Let vi be the value of a variable on the day i
» The ith trial assumes that the value of the market variable
tomorrow (i.e., on day n+1) is
a) vn+1 = vn *(vi/vi-1) if the percentage change is matched
b) vn+1 = vn + (vi-vi-1) if the actual change is matched (mostly
for interest rate, credit spread, and possibly volatility).
• Stressed VaR and Stressed ES
– Instead of basing calculations on the movements in market variables over the
last n days, we can base calculations on movements during a period in the
past that would have been particularly bad for the current portfolio
– This produces measures known as “stressed VaR” and “stressed ES”
– The 2008 financial crisis is a good period to stress test an investment
portfolio.
– Aug. 2007 (a period of quant strategy “earthquakes” lasting for a week) may
be a good period to stress test the quant equity portfolio
• Accuracy of VaR
Suppose that x is the qth percentile of the loss distribution, Prob(loss > x)=q,
when it is estimated from n observations. The standard error of x is (Kendall
and Stuart, 1972)
1 𝑞𝑞(1 − 𝑞𝑞)
𝑓𝑓(𝑥𝑥) 𝑛𝑛
where f(x) is an estimate of the probability density of the loss at the qth percentile
Note: The standard error is proportional to 1/f(x) and 1/√n
– Example
• Supposed the 1-percentile loss from 500 observations is estimated
as $25 million
• We estimate that the actual empirical distribution f(x) can be
approximated with a normal distribution mean zero and standard
deviation of $10 million
• The 1-percentile of the approximating distribution is
NORMINV(0.99,0,10) = 23.26 and the value of f(x) is
NORMDIST(23.26,0,10,FALSE)=0.0027
[NORMINV and NORMDIST are all Excel functions]
• The estimate of the standard error is therefore
1 0.01 × 0.99
= 1.67
0.0027 500
• Volatility Scaling
– Use a volatility updating scheme to monitor the volatilities of all market
variables
– If the ratio of the current volatility for a market variable vs the volatility
on Day i, is β --- multiply the percentage change observed on Day i by
β
– The value of the market variable under ith scenario thus becomes
𝑣𝑣𝑖𝑖−1 + 𝑣𝑣𝑖𝑖−𝑣𝑣𝑖𝑖−1 𝜎𝜎𝑛𝑛+1 /𝜎𝜎𝑖𝑖 𝑣𝑣𝑛𝑛+1 −𝑣𝑣𝑛𝑛 𝑣𝑣𝑖𝑖 −𝑣𝑣𝑖𝑖−1
𝑣𝑣𝑛𝑛+1 = 𝑣𝑣𝑛𝑛 , so that = (𝜎𝜎𝑛𝑛+1 /𝜎𝜎𝑖𝑖 )
𝑣𝑣𝑖𝑖−1 𝑣𝑣𝑛𝑛 𝑣𝑣𝑖𝑖−1
– This scaling makes sense as the percentage change relative to volatility
is important (we often talk about the number of std moves). The stock
market volatility has generally declined for the past ten years due to an
increase in passive investing.
• A variation of volatility scaling
– Monitor variance of simulated losses on the portfolio using EWMA
– Use standard deviation of the simulated losses to scale --- If the current
standard deviation of losses is βP times the standard deviation of
simulated losses on Day i, multiply ith loss given by the standard
approach by βP
• Another way to obtain the accuracy of VaR --- The Bootstrap Method
to Determine Confidence Intervals
– Example: For a historical simulation based on 500 daily changes
• Sample 500,000 times with replacement from daily changes to
obtain 1000 sets of changes over 500 days
• Calculate VaR for each set and get a distribution of VaRs
• The 95% confidence interval is the range between the 2.5
percentage point and the 97.5% percentile point: Given the 1000-
set samples, this is between the 25th largest VaR and the 975th
largest VaR.
– Model Building Approach for Estimating Risks
• This approach uses variance-covariance for the risk measure (tail
risk is underestimated). Portfolio variance is used in Markowitz’s
pioneering work on portfolio theory.
• It is the main alternative to historical simulation for calculating
Portfolio VaR or ES.
• Assuming the probability distributions of the returns on the market
variables are multivariate normal distributions --- to define the
normal distribution only variance-covariance is needed
– the average values are often set to zero; for stock return, as the
magnitude of the average daily return is much smaller than the
daily volatility
• This setup requires linearity (no options in the portfolio; only long or
short positions in stocks, bonds, commodities, and other products).
• The simple assumptions make the calculation quite fast.
• Variance-covariance can be updated easily using EWMA or
GARCH model, which is a computationally efficient method.
– Example: Portfolio of n Stocks
The variance of Portfolio Return 𝜎𝜎𝑃𝑃 2 = ∑𝑛𝑛𝑖𝑖=1 ∑𝑛𝑛𝑗𝑗=1 𝜌𝜌𝑖𝑖𝑖𝑖 𝑤𝑤𝑖𝑖 𝑤𝑤𝑗𝑗 𝜎𝜎𝑖𝑖 𝜎𝜎𝑗𝑗
wi – the weight of ith asset in the portfolio
ϭi2 – variance of return on the ith asset
ρij – correlation between the ith asset and the jth asset. ρij ϭi ϭj is the
covariance.
– Monte Carlo Simulation Approach (works with non-normal
distributions)
To calculate VaR using MC simulation. The following steps are used
• Get the current value of the portfolio
• Sample once from the multivariate distributions of the ∆xi (the
change in the market variable)
• Use the ∆xi to determine new values of the market variables
• Revalue the portfolio with the new values (nonlinearity does not pose
a problem here)
• Calculate the change in portfolio value ∆P
• Repeat many times to build up a probability distribution for ∆P
• VaR is the value at the appropriate percentile of the distribution
– For example, with 1,000 trials the 1 percentile is the 10th worst
case.
– Aside: Sampling from the probability distribution
• Sampling a one-variable distribution using random numbers
uniformly distributed between 0 and 1.
• Example: how to sample probability distribution p(y)=e-y, for y≥0
using a random x, generated from a uniform distribution between 0
and 1?
– Use percentile-to-percentile mapping (to equate the cumulative
probability) to map y to x
𝑦𝑦 𝑥𝑥
∫0 𝑒𝑒 −𝑢𝑢 𝑑𝑑𝑑𝑑 = ∫0 𝑑𝑑𝑑𝑑 → 1 − 𝑒𝑒 −𝑦𝑦 = 𝑥𝑥;
This leads to y = - ln(1-x)
– This is equivalent to
p(y)dy =u(x)dx, so we have
e-ydy = dx, dx/dy = e-y x = 1 - e-y , or y = - ln(1-x)
– We can also write y=-ln(x), as x and 1-x follow the same
distribution.
– Aside: Sampling from a normal distribution
• Sampling two-variable distribution p(y1, y2) using two random
numbers x1 and x2 generated from a uniform distribution
• The mapping is defined using the Jacobian determinant
𝜕𝜕 𝑥𝑥1 , 𝑥𝑥2
𝑝𝑝 𝑦𝑦1 , 𝑦𝑦2 𝑑𝑑𝑦𝑦1 𝑑𝑑𝑦𝑦2 = 𝑢𝑢 𝑥𝑥1 , 𝑥𝑥2 𝑑𝑑𝑥𝑥1 𝑑𝑑𝑥𝑥2 = 𝑑𝑑𝑦𝑦1 𝑑𝑑𝑑𝑑2
𝜕𝜕(𝑦𝑦1 , 𝑦𝑦2 )
The Jacobian defines the two-variable distribution we want to sample.
This can be generalized to multi-variable cases.
• The Box-Muller method for sampling the normal distribution
The mapping is defined as follows
𝑦𝑦1 = −2𝑙𝑙𝑙𝑙𝑥𝑥1 cos 2𝜋𝜋𝑥𝑥2 ; 𝑦𝑦2 = −2𝑙𝑙𝑙𝑙𝑥𝑥1 sin 2𝜋𝜋𝑥𝑥2 . Or
1 1 𝑦𝑦
𝑥𝑥1 = 𝑒𝑒𝑒𝑒𝑒𝑒 − (𝑦𝑦12 + 𝑦𝑦22 ) ; 𝑥𝑥2 = 𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎 𝑦𝑦2
2 2𝜋𝜋 1

It is easy to show the Jacobian determinant is given by


𝜕𝜕 𝑥𝑥1 ,𝑥𝑥2 1 −𝑦𝑦12 /2 1 −𝑦𝑦22 /2
=− 𝑒𝑒 𝑒𝑒 --- The minus sign can be
𝜕𝜕(𝑦𝑦1 ,𝑦𝑦2 ) 2𝜋𝜋 2𝜋𝜋
easily removed by replacing one of x with 1-x

You might also like