Back testing VaR
“Disclosure of quantitative measures of market
risk, such as value-at-risk, is enlightening only
when accompanied by a thorough discussion of
how the risk measures were constructed.”
— Alan Greenspan, Bank supervision,
regulation, and risk, October 5 1996.
Backtesting
• It is a test of how well the current procedure for
calculating the measure would have worked in the
past.
• Back-testing involves looking at how often the loss in a
day would have exceeded the one-day 99% VaR when
the latter is calculated using the current procedure.
• Days when the actual loss exceeds VaR are referred
to as exceptions. If exceptions happen on about 1% of
the days, we can feel reasonably comfortable with the
current methodology for calculating VaR.
• If they happen on, say, 7% of days, the methodology is
suspect and it is likely that VaR is underestimated.
From a regulatory perspective, the capital calculated
using the current VaR estimation procedure is then too
low.
Statistical Test: Kupiec Test (1996)
Kupiec Unconditional Coverage (UC) test—also
called the Proportion-of-Failures (POF) test—
checks whether the observed exception rate of a
VaR model equals the model’s nominal tail
probability.
Step by step Computation
Symbol Meaning
Number of observations (e.g., 250 trading
T
days)
Number of VaR exceptions (days where
N
actual loss > VaR)
Expected tail probability (e.g., 1% for 99%
p
VaR)
Hypothesis
Example
Basel Traffic-Light Framework
In Basel backtesting, the penalty zone (or
“traffic-light” system) classifies the quality of a
bank’s Value at Risk (VaR) model by counting
the number of exceptions—days when the
actual trading loss exceeds the model’s 99% 1-
day VaR—over the most recent 250 trading
days.
Framework
No. of exceptions
Zone Model status Interpretation
(out of 250)
Exception frequency
Green 0–4 Model acceptable consistent with 99%
VaR.
Needs closer
Model under- or
supervision and
Yellow 5–9 over-estimating risk
possible
slightly
recalibration.
VaR severely
underestimates risk;
Red ≥ 10 Model unreliable new model or
methodology
required.
Likelihood Ratio (LR) Independence
Test (Christoffersen, 1998)
• It examines whether VaR exceptions occur
independently over time — i.e., whether the
model properly captures volatility clustering.
• Even if the number of exceptions (Kupiec UC test)
matches the expected rate, they might cluster in
turbulent markets.
• The LR independence test checks that exceptions
are independent across days — a good VaR
model should not produce runs of consecutive
breaches.
Setup & Notation
Hypothesis
Likelihood functions
LR Independence Statistic
Example
• Suppose over 250 days and you had 6
exceptions (days where 𝐼𝑡 = 1)
• Exceptions occurred on days: 23, 24, 101, 150,
151, 220.
• This produces clusters (like two consecutive
exceptions on 23–24 and 150–151).
Estimate probabilities
Likelihood under 𝐻0 independence
Likelihood under 𝐻1 1st-order Markov
LR statistic
Decision and Interpretation
Exceptions are clustered (note 𝜋11 ≈ 0.333 is far larger than 𝜋01
≈ 0.016. Thus, the VaR model is likely slow to adapt to volatility/regime shifts.
Thank You