BACKTESTING VAR
Modelling VaR → trying to make a
good guess !
Model Validation → Are those guesses
making sense
?
LO 4.a: Define backtesting and exceptions and explain the importance of backtesting VaR models.
Back testing Comparison
1 I ,
losses
tool for Actual
predicted US
model losses
by VaR .
Valiahorg
Model
The main goal of backtesting is to ensure that actual losses do not exceed expected losses at a given confidence level. The
number of actual observations that fall outside a given confidence level are called exceptions.
The number of exceptions falling outside of the VaR confidence level should not exceed one minus the confidence level. For
example, exceptions should occur less than 5% of the time if the confidence level is 95%.
Backtesting is extremely important for risk managers and regulators
-
The Basel Committee allows banks
to use internal VaR models to
measure their risk levels, and
backtesting provides a critical
evaluation technique to test the
adequacy of those internal VaR
models.
Banks with excessive exceptions (more than four exceptions in a sample size of 250) are
penalized with higher capital requirements.
Explain the significant difficulties in backtesting a VaR model.
VaR models are based on static portfolios, while actual portfolio compositions are constantly changing as relative prices
change and positions are bought and sold.
the trading portfolio evolves dynamically during the day. Thus the actual portfolio is " contaminated" by changes in its
composition. The actual return corresponds to the actual P&L, taking into account intraday trades and other profit items such
as fees, commissions, spreads, and net interest income. This contamination will be minimized if the horizon is relatively
short, which explains why backtesting usually is conducted on daily returns. Even so, intraday trading generally will increase
the volatility of revenues because positions tend to be cut down toward the end of the trading day. Counterbalancing this is
the effect of fee income, which generates steady profits that may not enter the VaR measure.
Since the VaR forecast really pertains to R*, backtesting ideally should be done with these hypothetical returns.
Actual returns do matter, though, because they entail real profits and losses and are scrutinized by bank
regulators. . Ideally, both actual and hypothetical returns should be used for backtesting because both sets of
numbers yield informative comparisons.
--
If the model passes
backtesting with hypothetical In contrast, if the model does not pass
but not actual returns, then the backtesting with hypothetical returns,
problem lies with intraday then the modeling methodology
trading should be reexamined.
Verify a model based on exceptions or failure rates
Failure rates define the percentage of times the VaR confidence level is exceeded in a given sample. If we use N to
represent the number of exceptions and T to represent the sample size, the failure rate is computed as N / T.
Under Basel rules, bank VaR models must use a 99% confidence level, which means a bank must report the VaR
amount at the 1% left tail level for a total of T days.
The probability of exception, p, equals one minus the confidence level (p = 1 − c).
Suppose daily loss exceeded a predetermined VaR level (at the 95% confidence level) on 24 days during a 252-day
period. Is this sample an unbiased sample?
2 Score = 24-12.6
= 3.2g
#
'
51×95%9252 i
I
Note that the confidence
level at which we choose
to reject or fail to reject a ← Reiect that VaR
model is not related to the
+196
confidence level at which
VaR was calculated
-
1.96
model is unbiased
Note that this test makes no assumption about the return distribution. The distribution could be normal, or skewed, or with heavy
tails, or time-varying. We simply count the number of exceptions. As a result, this approach is fully nonparametric.
The setup for this test is the classic testing framework for a sequence of success and failures, also called Bernoulli trials.
Under the null hypothesis that the model is correctly calibrated, the number of exceptions x follows a binomial probability
distribution:
TC PJ
" "
thx ) =
p (i -
We also know that x has expected value of E(x) = p T and variance V(x) = p(1 — p)T.
When T is large, we can use the central limit theorem and approximate the binomial distribution by the normal
distribution
Define and identify Type I and Type II errors.
Type I Type I
Rejecting an accurate model ( failing to reject an inaccurate model
tradeoff between Type I and Type II errors
The goal in backtesting is to create a VaR model with a low Type I error and include a test for a very low Type II error
rate. We can establish such ranges at different confidence levels using a binomial probability distribution based on the
size of the sample.
The binomial test is used to determine if the ←-
number of exceptions is acceptable at various y
confidence levels. .
Banks are required to use 250 days of data to I
be tested at the 99% confidence level. This
results in a failure rate, or p = 0.01, of only 2.5
exceptions in a 250-day time horizon.
Bank regulators impose a penalty in the form of
higher capital requirements if five or more -
-
-
→ 252cg 1%5×99%247=6.79 y
x .
exceptions are observed. Figure on the right
- . . .
- - - - -
252cg ly? 99%246 2.82-1
→ =
illustrates that we expect five or more exceptions x
-
-
×
.
-
10.8% of the time given a 99% confidence level.
Regulators will reject a correct model or commit
a Type I error in these cases at the far right tail.
← -
y
,
Figure on the right illustrates the far left tail of
the distribution, where we evaluate Type II
errors. For less than five exceptions, regulators I
will fail to reject an incorrect model at a 97%
confidence level (rather than a 99% confidence
level) 12.8% of the time. Note that if we lower
the confidence level to 95%, the probability of
committing a Type I error slightly increases,
while the probability of committing a Type II
error sharply decreases.
Unconditional coverage
t
refers to the fact that we are not concerned about independence of exception
observations or the timing of when the exceptions occur. We simply are
concerned with the number of total exceptions.
t ( LR )
Log likelihood ratio
determined by calculating
-
can be
t
model is correct
He Reject the hypothesis that
if
[Link]
3.84
"
Critical value
957
-
-
→ .
1.962
with one
squared
° O
Chi -
or
doff 95% CL
I
Normal dist value
-
LR →
prob level
p = .
Sample size
T =
Number of exceptions
N =
statistic for unconditional
( Roc = test
coverage .
Table below provides the nonrejection region for the number of failures (N) based on the probability level (p), confidence
level (c), and time period (T).
252×54--12.6=4 -
.
expected
value
'2 -
GI ([Link] )
① O
6 1-120
Tables above can also be used to illustrate how increasing the sample size allows us to reject the model more easily.
For example, at the 97.5% confidence, where T = 252, the test interval is 2 / 252 = 0.79%, 12 / 252 = 4.76%.
When T is increased to 1,000, the test interval shrinks to 16 / 1,000 = 1.6%, 36 / 1,000 = 3.6%.
Table above also illustrates that it is difficult to backtest VaR models constructed with higher levels of confidence, because
the number of exceptions is often not high enough to provide meaningful information.
Notice that at the 95% confidence level, the test interval for T = 252 is 6 / 252 = 0.024, 20 / 252 = 0.079.
With higher confidence levels (i.e., smaller values of p), the range of acceptable exceptions is smaller.
Thus, it becomes difficult to determine if the model is overstating risks (i.e., fewer than expected exceptions) or if the
number of exceptions is simply at the lower range of acceptable.
Banks will sometimes choose to use a higher value of p such as 5%, in order to validate the model with a sufficient
number of deviations.
'
LR =
-21N -4 -5%1252-9 (5%39)
,
[9/252]
-9
}
'"
+ 2 IN
{ [ I 9/252 )
-
= 1.1973
Examine
Suppose that a risk manager needs to backtest a daily VaR model that was constructed using a 95% confidence level over a
252-day period. If the sample revealed 15 exceptions, should we reject or fail to reject the null hypothesis that p is the true
probability of failure for this VaR model?
Hsing VaR to Measure potential Losses
There are two theories about choosing a holding period for VaR calculations
The first theory is that the holding The second theory is that the holding
period should correspond to the period should be chosen to match the
amount of time required to either period over which the portfolio is not
liquidate or hedge the portfolio. expected to change due to non-risk-
Thus, VaR would calculate related activity (e.g., trading). The two
possible losses before corrective theories are not that different. For
action could take effect. example, many banks use a daily VaR to
correspond with the daily profit and loss
measures.
Explain the need to consider conditional coverage in the backtesting framework.
So far in unconditional coverage, the timing of our exceptions was not considered.
Conditioning considers the time variation of the data. In addition to having a predictable number of exceptions, we also
anticipate the exceptions to be fairly equally distributed across time
A bunching of exceptions may indicate that market correlations have changed or that our trading positions have been
altered. In the event that exceptions are not independent, the risk manager should incorporate models that consider
time variation in risk.
We need some guide to determine if the bunching is random or caused by one of these changes. By including a measure of
the independence of exceptions, we can measure conditional coverage of the model.
Christofferson proposed extending the unconditional coverage test statistic (LRuc) to allow for potential time variation of the
data.
He developed a statistic to determine the serial independence of deviations using a log-likelihood ratio test (LRind).
The overall log-likelihood test statistic for conditional coverage (LRcc) is then computed as:
(Rec = [Link] t LR ind
Each individual component is independently distributed as chi-squared, and the sum is also distributed as chi-
squared.
At the 95% confidence level, we would reject the model if LRcc > 5.99 and we would reject the independence term
alone if LRind > 3.84.
If exceptions are determined to be serially dependent, then the VaR model needs to be revised to incorporate the
correlations that are evident in the current conditions.
Describe the Basel rules for backtesting.
In the backtesting process, we attempt to strike a balance between the probability of a Type I & II error
Regulators do not have access to every parameter input of
the model and must construct rules that are applicable
across institutions.
To mitigate the risk that banks willingly commit a Type II error
and use a faulty model, the Basel Committee designed the
Basel penalty zones presented in Table here.
The committee established a scale of the number of
exceptions and corresponding increases in the capital
multiplier, k.
Thus, banks are penalized for exceeding four exceptions per
year.
The multiplier is normally three but can be increased to as
much as four, based on the accuracy of the bank’s VaR
The current verification procedure consists of recording model. Increasing k significantly increases the amount of
daily exceptions of the 99 percent VaR over the last capital a bank must hold and lowers the bank’s performance
year. One would expect, on average, 1 percent of 250, measures, like return on equity.
or 2.5 instances of exceptions over the last year. The
Basel Committee has decided that up to four
exceptions are acceptable, which defines a " green
light" zone for the bank.
the yellow zone is quite broad (five to nine exceptions).
The penalty (raising the multiplier from three to four) is automatically required for banks with 10 or more exceptions.
However, the penalty for banks with five to nine exceptions is subject to supervisors’ discretions, based on what type of
model error caused the exceptions.
The Committee established four categories of causes for exceptions and guidance for supervisors for each category:
Model accuracy needs
The basic integrity of the
improvement. The exceptions
model is lacking.
occurred because the model
Exceptions occurred
does not accurately describe
because of incorrect data or
risks. The penalty should Intraday trading activity. The
errors in the model
apply. exceptions occurred due to
programming. The penalty
should apply. trading activity (VaR is based
on static portfolios). The
penalty should be considered.
Bad luck. The exceptions occurred
because market conditions (volatility
and correlations among financial
instruments) significantly varied from
an accepted norm. These exceptions
should be expected to occur at least
some of the time. No penalty
guidance is provided.
Although the yellow zone is broad, an accurate model could produce five or more exceptions 10.8% of the time at the 99%
confidence level. So even if a bank has an accurate model, it is subject to punishment 10.8% of the time (using the required
99% confidence level). However, regulators are more concerned about Type II errors, and the increased capital multiplier
penalty is enforced using the 97% confidence level. At this level, inaccurate models would not be rejected 12.8% of the time
(e.g., those with VaR calculated at the 97% confidence level rather than the required 99% confidence level). While this
seems to be only a slight difference, using a 99% confidence level would result in a 1.24 times greater level of required
capital, providing a powerful economic incentive for banks to use a lower confidence level.
--
Solution I solution 2
Industry analysts have suggested lowering the
required VaR confidence level to 95% and Another way to make variations
compensating by using a greater multiplier. This in the number of exceptions more
would result in a greater number of expected significant would be to use a
exceptions, and variances would be more longer backtesting period. This
statistically significant. The one-year exception approach may not be as practical
rate at the 95% level would be 13, and with more because the nature of markets,
than 17 exceptions, the probability of a Type I portfolios, and risk changes over
error would be 12.5% (close to the 10.8% time.
previously noted), but the probability of a Type II
error at this level would fall to 7.4% (compared to
12.8% at a 97.5% confidence level). Thus,
inaccurate models would fail to be rejected less
frequently.
The management of a financial institution reports that on a particular year, the daily revenue fell short of the downside 95%
VaR band on 26 occasions (days), or more than 5% of the time. Ten of these 26 occurrences fell within the May to July period.
Assuming 252 days in the year, test if this was a faulty model or bad luck using the binomial distribution.
model 8.5
A .
faulty ,
2 =
model 2=3.5
B .
Faulty ,
bad luck 2=1.82
C .
it is ,
luck
2=115
it is bad
D .
,
Vaibhav Neele , a risk analyst at a large multinational bank, is backtesting the VaR model of the bank. The model being tested is a
daily, 98% VaR model. If the backtest is conducted for one year at a 95% confidence level, what is the acceptable number of daily
losses that will lead Black to conclude that the model is calibrated correctly?
A. 25
B. 7
c. 2
D. 20
A model gives a VaR value of $5 million for a portfolio at a 99% confidence interval. A one-year backtest conducted at
the 90% confidence level reveals that losses exceeded $5 million on 12 occasions. The model is accepted as accurate.
Assuming 224 days in a year, which of these statements is most likely true?
2 error occurred
A .
Type
occurred
I error
B. Type
errors
occurred
C . Both
has been accepted correctly .
D. Model
Dash Sang and Vaibhav Pile are two junior risk analysts. They have recently been assigned to perform a 1-year backtest
of a 1 day 98% VaR model, assuming 225 days in the year.
During the next few days, they exchange a number of emails regarding the assignment:
Email 1 - Jones forecasts the number of expected exceptions for the model to be 4.5.
Email 2 - Atherton replies that according to the Basel Committee’s prescribed penalty zones, the yellow zone starts at six
exceptions and attracts a multiplier of four.
Email 3 - Jones states a type II error occurs when an accurate model is rejected.
The contents of which of the emails is/are not true ?
A . All three
B .
I only
C . 223
D .
I 22