Understanding Volatility in Finance
Understanding Volatility in Finance
• Definition of Volatility
– Suppose that Si is the value of a variable on the day i. The volatility per
day is the standard deviation of ln(Si /Si-1)
» Normally the days when markets are closed are ignored in volatility
calculations [volatility is most generated by trading itself]
– The volatility per year is 252 times the daily volatility
» The variance rate is the square of the volatility
– Volatility measures the average risk (not specifically for tail risks)
• Implied Volatility
– Of the variables needed to price an option, the one that cannot be
observed directly is the volatility
– We can imply volatilities from market prices -- One can get a better feel
for whether the option is over/undervalued from IV
– IV depends on its strike price and time to expiration, leading to an
implied volatility surface (it captures both the average and tail risks)
– Volatility smile/skew (due to fat tails of the return distribution)
• VIX Index: A Measure of the Implied Volatility of the S&P 500
90.00
80.00
70.00
60.00
50.00
40.00
30.00
20.00
10.00
0.00
This seems to fit the behavior of the returns on many market variables better
than the normal distribution
Cubic law for stock market returns, 𝛼𝛼 = 3
– Log-Log Test for Exchange Rate Data (v is the number of standard deviations
which the exchange rate moves)
0
ln[Prob(v > x)] = ln(K) –α ln(x) 0 0.5 1 1.5 2
-1
ln(x)
-2
ln(Prob(v>x)
-3
-4
-5
-6
-7
On a Side Note
For a normal distribution
Distribution of heights of men in China: 172 ± 7 cm (2.31m: 7σ deviation)
Prob( 𝑥𝑥 > 𝑥𝑥̅ + 3𝜎𝜎) : 0.27%
Prob(𝑥𝑥 > 𝑥𝑥̅ + 7𝜎𝜎) : 2.5x10-10
Prob(𝑥𝑥 > 𝑥𝑥̅ + 20𝜎𝜎) : 2.7x10-87 (black swan events)
not a good distribution for risk management
– We do have 20σ black swan events:
1987 Black Monday, 2015 Swiss Franc’s 30% move
Power-law tail distribution is a better choice.
Gutenberg-Richter law: P(>E) : E-b, valid for an enormous range of energies.
The magnitude of 4: 1000 tons of TNT (happens 6000 times a year)
The magnitude of 8: 32 billion tons (every ten years)
no typical-size earthquake can be defined!
Thinking in terms of a normal distribution ~ thinking in terms of a fixed
typical/maximum value (could it contribute to the Fukushima nuclear
disaster?)
• Can the stock market return be described by the normal
distribution? Date SP500 Return
1/2/2003 909.03
𝑟𝑟̅ = 0.00036, 𝜎𝜎 = 0.012
1/3/2003 908.59 -0.00048
Return of SP500
1/6/2003 929.01 0.022474
60
1/7/2003 922.93 -0.00654
40
1/8/2003 909.93 -0.01409
20 1/9/2003 927.58 0.019397
0 1/10/2003 927.57 -1.1E-05
-0.12 -0.07 -0.02 0.03 0.08 1/13/2003 926.26 -0.00141
-20
1/14/2003 931.66 0.00583
Probability Density Normal Distribution
1/15/2003 918.22 -0.01443
Return of SP500 1/16/2003 914.6 -0.00394
100
SP500
6000
0.1
-0.12 -0.07 -0.02 0.03 0.08
0.0001 5000
0.0000001
4000
1E-10
3000
1E-13
1E-16 2000
1E-19
1000
1E-22
0
Probability Density Normal Distribution 1/9/2002 28/5/2005 22/2/2008 18/11/2010 14/8/2013 10/5/2016 4/2/2019 31/10/2021
• The empirical cubic law for the tail distribution
𝑪𝑪
Right tail: 𝑷𝑷𝑷𝑷𝑷𝑷𝑷𝑷 𝒓𝒓 > 𝒙𝒙 = 𝒙𝒙𝟑𝟑𝟏𝟏
𝑪𝑪
Left tail: 𝑷𝑷𝑷𝑷𝑷𝑷𝑷𝑷 𝒓𝒓 < 𝒙𝒙 = (−𝒙𝒙)
𝟐𝟐
𝟑𝟑
Cumulative Probability
0.07
0.06
0.05
0.04
0.03
0.02
0.01
0
-0.12 -0.11 -0.1 -0.09 -0.08 -0.07 -0.06 -0.05 -0.04 -0.03 -0.02
-0.01
1E-06 0.1
1E-09
1E-12 0.01
1E-15
1E-18 0.001
1E-21
1E-24 0.0001
CPD CPD (Normal) CPD(Cubic Law) CPD CPD (Normal) CPD(Cubic Law)
𝑚𝑚
1
𝜎𝜎𝑛𝑛 2 = � 𝑢𝑢𝑛𝑛−𝑖𝑖 2
𝑚𝑚
𝑖𝑖=1
• Weighting Schemes
Compared to stock returns, volatility is easier to forecast
– Volatility clustering, Mean-reverting
It makes sense to give more weight to recent data (volatility clustering)
𝜎𝜎𝑛𝑛 2 = ∑𝑚𝑚 2 𝑚𝑚
𝑖𝑖=1 𝛼𝛼𝑖𝑖 𝑢𝑢𝑛𝑛−𝑖𝑖 , ∑𝑖𝑖=1 𝛼𝛼𝑖𝑖 = 1, 𝛼𝛼𝑖𝑖 < 𝛼𝛼𝑗𝑗 when i > j
An extension of this is to include a long-term average variance rate
𝜎𝜎𝑛𝑛 2 = 𝛾𝛾𝑉𝑉𝐿𝐿 + ∑𝑚𝑚 2
𝑖𝑖=1 𝛼𝛼𝑖𝑖 𝑢𝑢𝑛𝑛−𝑖𝑖 , 𝛾𝛾 + ∑𝑚𝑚𝑖𝑖=1 𝛼𝛼𝑖𝑖 = 1
• The Exponentially Weighted Moving Average Model (EWMA)
𝜎𝜎𝑛𝑛 2 = 𝜆𝜆𝜎𝜎𝑛𝑛−1 2 + (1 − 𝜆𝜆) 𝑢𝑢𝑛𝑛−1 2 (0<𝜆𝜆 < 1) --- online algorithm
This can be applied iteratively
𝜎𝜎𝑛𝑛 2 = 𝜆𝜆(𝜆𝜆𝜎𝜎𝑛𝑛−2 2 + 1 − 𝜆𝜆 𝑢𝑢𝑛𝑛−2 2 ) + (1 − 𝜆𝜆) 𝑢𝑢𝑛𝑛−1 2
= 𝜆𝜆2 (𝜆𝜆𝜎𝜎𝑛𝑛−3 2 + 1 − 𝜆𝜆 𝑢𝑢𝑛𝑛−3 2 ) + 1 − 𝜆𝜆 𝜆𝜆𝑢𝑢𝑛𝑛−2 2 + (1 − 𝜆𝜆) 𝑢𝑢𝑛𝑛−1 2
= (1 − 𝜆𝜆) ∑𝑚𝑚 𝑖𝑖=1 𝜆𝜆
𝑖𝑖−1 𝑢𝑢 2 𝑚𝑚
𝑛𝑛−𝑖𝑖 + 𝜆𝜆 𝜎𝜎𝑛𝑛−𝑚𝑚
2
For large m, the last term is very small. It is the weighted average with
𝛼𝛼𝑖𝑖 = 1 − 𝜆𝜆 𝜆𝜆𝑖𝑖−1 = 1 − 𝜆𝜆 𝑒𝑒 − 𝑖𝑖−1 𝑙𝑙𝑙𝑙𝑙𝑙
• EWMA (cont..)
𝜎𝜎𝑛𝑛 2 = 𝜆𝜆𝜎𝜎𝑛𝑛−1 2 + (1 − 𝜆𝜆) 𝑢𝑢𝑛𝑛−1 2 (0<𝜆𝜆 < 1)
» 𝜆𝜆 = 0.94 has been found to be a good choice across a wide
range of market variables
𝜎𝜎𝑛𝑛 2 = 0.94𝜎𝜎𝑛𝑛−1 2 + 0.06 𝑢𝑢𝑛𝑛−1 2
» Little data need to be stored
2
Day Price Return Return σ2(EWMA)
0 20
1 20.1 0.005 2.5E-05 2.5E-05
2 19.9 -0.01 9.901E-05 2.5E-05
3 20 0.00503 2.525E-05 2.94404E-05
4 20.5 0.025 0.000625 2.91891E-05
5 20.25 -0.0122 0.0001487 6.49378E-05
6 20.9 0.0321 0.0010303 6.99648E-05
7 20.9 0 0 0.000127587
8 20.9 0 0 0.000119932
9 20.6 -0.0144 0.000206 0.000112736
• The GARCH(1,1) Model
Combining EWMA and a long-run average
𝜎𝜎𝑛𝑛 2 = 𝛾𝛾𝑉𝑉𝐿𝐿 + 𝛼𝛼𝑢𝑢𝑛𝑛−1 2 + 𝛽𝛽𝜎𝜎𝑛𝑛−1 2 = 𝜔𝜔 + 𝛼𝛼𝑢𝑢𝑛𝑛−1 2 + 𝛽𝛽𝜎𝜎𝑛𝑛−1 2
𝛾𝛾 + 𝛼𝛼 + 𝛽𝛽 = 1, 𝛼𝛼 + 𝛽𝛽 < 1
– EWMA is a particular case with 𝛾𝛾=0, 𝛼𝛼=1-𝜆𝜆, 𝛽𝛽=𝜆𝜆
– 𝛽𝛽 describes the decay rate
– GARCH(1,1) model has the property that over time the variance
tends to get pulled back to a long-run average level
– The parameters can be estimated using maximum likelihood
methods
using data for the S&P index between July 18, 2005 and Aug. 13,
2010, one can estimate
𝛼𝛼=0.083394, 𝛽𝛽=0.910116, 𝜔𝜔=0.0000013465
VL =𝜔𝜔/(1-𝛼𝛼-𝛽𝛽)=0.0002075, this gives the long-term volatility of
1.4404% per day (~23% per annum)
• Using Garch(1,1) to Forecast Future Volatility
From
𝜎𝜎𝑛𝑛 2 = 1 − 𝛼𝛼 − 𝛽𝛽 𝑉𝑉𝐿𝐿 + 𝛼𝛼𝑢𝑢𝑛𝑛−1 2 + 𝛽𝛽𝜎𝜎𝑛𝑛−1 2
we have
𝜎𝜎𝑛𝑛 2 − 𝑉𝑉𝐿𝐿 = 𝛼𝛼(𝑢𝑢𝑛𝑛−1 2 − 𝑉𝑉𝐿𝐿 ) + 𝛽𝛽(𝜎𝜎𝑛𝑛−1 2 − 𝑉𝑉𝐿𝐿 )
On day n+t in the future
𝜎𝜎𝑛𝑛+𝑡𝑡 2 − 𝑉𝑉𝐿𝐿 = 𝛼𝛼(𝑢𝑢𝑛𝑛+𝑡𝑡−1 2 − 𝑉𝑉𝐿𝐿 ) + 𝛽𝛽(𝜎𝜎𝑛𝑛+𝑡𝑡−1 2 − 𝑉𝑉𝐿𝐿 )
(a) (b)
E(V2)
V1
(c)
• Monitoring Correlation Between Two Variables X and Y
Define xi=(Xi−Xi-1)/Xi-1 and yi=(Yi−Yi-1)/Yi-1
Also
varx,n: daily variance of X
vary,n: daily variance of Y
covn: covariance between x and y
These are calculated based on the data on day n-1 and before
Covn = E(xy)−E(x)E(y)
---- It is usually approximated as E(xy) when the
averages of the variables are small
For a simple estimate
𝑚𝑚
1
𝑐𝑐𝑐𝑐𝑐𝑐𝑛𝑛 = � 𝑥𝑥𝑛𝑛−𝑖𝑖 𝑦𝑦𝑛𝑛−𝑖𝑖
𝑚𝑚
𝑖𝑖=1
The correlation is
𝒄𝒄𝒄𝒄𝒄𝒄𝒏𝒏
𝒗𝒗𝒗𝒗𝒗𝒗𝒙𝒙,𝒏𝒏 𝒗𝒗𝒗𝒗𝒗𝒗𝒚𝒚,𝒏𝒏
• Estimating Covariance
EWMA: covn = 𝜆𝜆covn-1 + (1-𝜆𝜆)xn-1yn-1
GARCH(1,1): covn = 𝜔𝜔 + 𝛼𝛼xn-1yn-1 + 𝛽𝛽covn-1
• Positive Finite Definite Condition
A variance-covariance matrix Ω, is internally consistent if the positive
semi-definite condition
wTΩw ≥ 0 holds for all vectors w
wTΩw is the variance of the portfolio with weights given by w
This ensures that variances are nonnegative for all principal
components
– Example:
1 0 0.9
0 1 0.9 is not internally consistent (try w=[1,1,-1]T); this cannot
0.9 0.9 1
be a covariance matrix. (𝑤𝑤 𝑇𝑇 Ω𝑤𝑤 = −0.6)
When estimating variance/covariance, the same numerical procedure
should be used to ensure consistency (for example, use the same 𝜆𝜆 for
EWMA for variance and covariance)
– Estimation of covariance -- Factor Models
1. When there are N variables, Vi (i = 1, 2,..N), there are N(N−1)/2
covariance/correlations
2. In risk management, one never estimates all covariance directly
3. We can reduce the number of correlation parameters that have to be
estimated with a factor model and express covariance in terms of
factor covariance.
• A simple example: One-Factor Model (Gaussian variables)
– Correlations among variables following a normal distribution can be
defined easily. If {Ui } follows the standard normal distribution we can
set Ui = ai F + 1 − 𝑎𝑎𝑖𝑖2 Zi , where the common factor F and the
idiosyncratic component Zi have independent standard normal
distributions (𝐸𝐸 𝑍𝑍𝑖𝑖 𝑍𝑍𝑗𝑗 = 𝛿𝛿𝑖𝑖𝑖𝑖 , etc.)
Note var(Ui) = 1 because E(F2) =1, E(Zi2) =1, E(FZi) = 0
Correlation between Ui and Uj (i≠j), 𝐸𝐸 𝑈𝑈𝑖𝑖 𝑈𝑈𝑗𝑗 = 𝑎𝑎𝑖𝑖 𝑎𝑎𝑗𝑗 𝐸𝐸 𝐹𝐹 2 = 𝑎𝑎𝑖𝑖 𝑎𝑎𝑗𝑗
• What should we do if the variable does not follow the normal distribution?
– Gaussian Copula Model:
-0.2 0 0.2 0.4 0.6 0.8 1 1.2 -0.2 0 0.2 0.4 0.6 0.8 1 1.2
V1 V2
- -
-6 -4 -2 0 2 4 6 -6 -4 -2 0 2 4 6
U1 U2
N(x)
F(V1) = N(U1) : V1 U1
• Examples:
V1 V2
• V1 Mapping to U1 and V2 Mapping to U2
V1 Percentile U1 V2 Percentile U2
Gain Loss
-VaR VaR
-5.8
-9.8
z loss
• Changing the Time Horizon
– If losses in successive days are independent, normally distributed, and
have a mean of zero
T-day VaR = 1-day VaR x √T
T-day ES = 1-day ES x √T
• Aggregating VaRs
An approximate approach that seems to work well is
where VaRi is the VaR for the ith segment, VaRtotal is the total VaR, and ρij is
the coefficient of correlation between losses from the ith and jth segments
• Back-testing to check for the risk measure
– Back Test for VaR is relatively easy to do than ES.
– Back-testing a VaR involves looking at how often exceptions (loss >
VaR) occur --- Intuitively, for 99% VaR, 7% occurrence means the VaR
is likely to be under-estimated; on the other hand, 0.2% occurrence
means that VaR is likely to be over-estimated.
– A more quantitative statistical test
• Suppose that the theoretical probability of an exception is p (=1−X),
and the number of occurrences (out of n) is m with m/n > p.
– When can we reject the model for producing VaR being too
low? Consider the probability of m or more exceptions
𝑛𝑛 𝑚𝑚−1
𝑛𝑛! 𝑛𝑛!
� 𝑝𝑝𝑘𝑘 (1 − 𝑝𝑝)𝑛𝑛−𝑘𝑘 = 1 − � 𝑝𝑝𝑘𝑘 (1 − 𝑝𝑝)𝑛𝑛−𝑘𝑘
𝑘𝑘! 𝑛𝑛 − 𝑘𝑘 ! 𝑘𝑘! 𝑛𝑛 − 𝑘𝑘 !
𝑘𝑘=𝑚𝑚 𝑘𝑘=0
– A significant level often used is 5%. If the resulting probability is
less than 5%, the model can be rejected.
– Example: 600 days of data, consider 1 day 99%VaR, we can
calculate that if m>10, the probability is less than 5%, the model
should be rejected --- VaR is too low.
• Suppose that the number of occurrences (out of n) is m with m/n < p.
– When can we reject the model for producing VaR being high?
Consider the probability of m or fewer exceptions
𝑚𝑚
𝑛𝑛!
� 𝑝𝑝𝑘𝑘 (1 − 𝑝𝑝)𝑛𝑛−𝑘𝑘
𝑘𝑘! 𝑛𝑛 − 𝑘𝑘 !
𝑘𝑘=0
– Again, if the resulting probability is less than 5%, the model can be
rejected.
– The same example: 600 days of data, consider 1 day 99%VaR, we
can calculate that if m=1, the probability is 1.7%, less than 5%, the
model should be rejected --- VaR is too high. But m > 1 (m/n<p),
the probability is greater than 5%, and the model should not be
rejected.
• Combining the two tests: if m = 2, 3, …, 10, the model can’t be rejected.
• Two-tail test (Kupiec, 1995), based on
𝑛𝑛−𝑚𝑚 𝑚𝑚
𝑚𝑚 𝑛𝑛−𝑚𝑚 𝑚𝑚 𝑚𝑚
−2 ln 1 − 𝑝𝑝 𝑝𝑝 + 2 ln 1 −
𝑛𝑛 𝑛𝑛
This model should be rejected if this quantity > 3.84
– Historical Simulation to Determine the Risk Measure
• How do we do historical simulation?
– Collect data on the daily movements in all market variables
– Each day’s percentage changes in all market variables (for
example asset returns) are considered as a scenario to simulate
the change in the portfolio value (from today to tomorrow)
» For example, to calculate 1-day 99% VaR with 500-day data,
we can evaluate the changes in the portfolio value for these
500 scenarios --- VaR can be estimated as the fifth-worst loss
(ES can be estimated as the average of the top four losses).
– Some Technicalities
» Suppose we use n days of historical data (today is day n)
» Let vi be the value of a variable on the day i
» The ith trial assumes that the value of the market variable
tomorrow (i.e., on day n+1) is
a) vn+1 = vn *(vi/vi-1) if the percentage change is matched
b) vn+1 = vn + (vi-vi-1) if the actual change is matched (mostly
for interest rate, credit spread, and possibly volatility).
• Stressed VaR and Stressed ES
– Instead of basing calculations on the movements in market variables over the
last n days, we can base calculations on movements during a period in the
past that would have been particularly bad for the current portfolio
– This produces measures known as “stressed VaR” and “stressed ES”
– The 2008 financial crisis is a good period to stress test an investment
portfolio.
– Aug. 2007 (a period of quant strategy “earthquakes” lasting for a week) may
be a good period to stress test the quant equity portfolio
• Accuracy of VaR
Suppose that x is the qth percentile of the loss distribution, Prob(loss > x)=q,
when it is estimated from n observations. The standard error of x is (Kendall
and Stuart, 1972)
1 𝑞𝑞(1 − 𝑞𝑞)
𝑓𝑓(𝑥𝑥) 𝑛𝑛
where f(x) is an estimate of the probability density of the loss at the qth percentile
Note: The standard error is proportional to 1/f(x) and 1/√n
– Example
• Supposed the 1-percentile loss from 500 observations is estimated
as $25 million
• We estimate that the actual empirical distribution f(x) can be
approximated with a normal distribution mean zero and standard
deviation of $10 million
• The 1-percentile of the approximating distribution is
NORMINV(0.99,0,10) = 23.26 and the value of f(x) is
NORMDIST(23.26,0,10,FALSE)=0.0027
[NORMINV and NORMDIST are all Excel functions]
• The estimate of the standard error is therefore
1 0.01 × 0.99
= 1.67
0.0027 500
• Volatility Scaling
– Use a volatility updating scheme to monitor the volatilities of all market
variables
– If the ratio of the current volatility for a market variable vs the volatility
on Day i, is β --- multiply the percentage change observed on Day i by
β
– The value of the market variable under ith scenario thus becomes
𝑣𝑣𝑖𝑖−1 + 𝑣𝑣𝑖𝑖−𝑣𝑣𝑖𝑖−1 𝜎𝜎𝑛𝑛+1 /𝜎𝜎𝑖𝑖 𝑣𝑣𝑛𝑛+1 −𝑣𝑣𝑛𝑛 𝑣𝑣𝑖𝑖 −𝑣𝑣𝑖𝑖−1
𝑣𝑣𝑛𝑛+1 = 𝑣𝑣𝑛𝑛 , so that = (𝜎𝜎𝑛𝑛+1 /𝜎𝜎𝑖𝑖 )
𝑣𝑣𝑖𝑖−1 𝑣𝑣𝑛𝑛 𝑣𝑣𝑖𝑖−1
– This scaling makes sense as the percentage change relative to volatility
is important (we often talk about the number of std moves). The stock
market volatility has generally declined for the past ten years due to an
increase in passive investing.
• A variation of volatility scaling
– Monitor variance of simulated losses on the portfolio using EWMA
– Use standard deviation of the simulated losses to scale --- If the current
standard deviation of losses is βP times the standard deviation of
simulated losses on Day i, multiply ith loss given by the standard
approach by βP
• Another way to obtain the accuracy of VaR --- The Bootstrap Method
to Determine Confidence Intervals
– Example: For a historical simulation based on 500 daily changes
• Sample 500,000 times with replacement from daily changes to
obtain 1000 sets of changes over 500 days
• Calculate VaR for each set and get a distribution of VaRs
• The 95% confidence interval is the range between the 2.5
percentage point and the 97.5% percentile point: Given the 1000-
set samples, this is between the 25th largest VaR and the 975th
largest VaR.
– Model Building Approach for Estimating Risks
• This approach uses variance-covariance for the risk measure (tail
risk is underestimated). Portfolio variance is used in Markowitz’s
pioneering work on portfolio theory.
• It is the main alternative to historical simulation for calculating
Portfolio VaR or ES.
• Assuming the probability distributions of the returns on the market
variables are multivariate normal distributions --- to define the
normal distribution only variance-covariance is needed
– the average values are often set to zero; for stock return, as the
magnitude of the average daily return is much smaller than the
daily volatility
• This setup requires linearity (no options in the portfolio; only long or
short positions in stocks, bonds, commodities, and other products).
• The simple assumptions make the calculation quite fast.
• Variance-covariance can be updated easily using EWMA or
GARCH model, which is a computationally efficient method.
– Example: Portfolio of n Stocks
The variance of Portfolio Return 𝜎𝜎𝑃𝑃 2 = ∑𝑛𝑛𝑖𝑖=1 ∑𝑛𝑛𝑗𝑗=1 𝜌𝜌𝑖𝑖𝑖𝑖 𝑤𝑤𝑖𝑖 𝑤𝑤𝑗𝑗 𝜎𝜎𝑖𝑖 𝜎𝜎𝑗𝑗
wi – the weight of ith asset in the portfolio
ϭi2 – variance of return on the ith asset
ρij – correlation between the ith asset and the jth asset. ρij ϭi ϭj is the
covariance.
– Monte Carlo Simulation Approach (works with non-normal
distributions)
To calculate VaR using MC simulation. The following steps are used
• Get the current value of the portfolio
• Sample once from the multivariate distributions of the ∆xi (the
change in the market variable)
• Use the ∆xi to determine new values of the market variables
• Revalue the portfolio with the new values (nonlinearity does not pose
a problem here)
• Calculate the change in portfolio value ∆P
• Repeat many times to build up a probability distribution for ∆P
• VaR is the value at the appropriate percentile of the distribution
– For example, with 1,000 trials the 1 percentile is the 10th worst
case.
– Aside: Sampling from the probability distribution
• Sampling a one-variable distribution using random numbers
uniformly distributed between 0 and 1.
• Example: how to sample probability distribution p(y)=e-y, for y≥0
using a random x, generated from a uniform distribution between 0
and 1?
– Use percentile-to-percentile mapping (to equate the cumulative
probability) to map y to x
𝑦𝑦 𝑥𝑥
∫0 𝑒𝑒 −𝑢𝑢 𝑑𝑑𝑑𝑑 = ∫0 𝑑𝑑𝑑𝑑 → 1 − 𝑒𝑒 −𝑦𝑦 = 𝑥𝑥;
This leads to y = - ln(1-x)
– This is equivalent to
p(y)dy =u(x)dx, so we have
e-ydy = dx, dx/dy = e-y x = 1 - e-y , or y = - ln(1-x)
– We can also write y=-ln(x), as x and 1-x follow the same
distribution.
– Aside: Sampling from a normal distribution
• Sampling two-variable distribution p(y1, y2) using two random
numbers x1 and x2 generated from a uniform distribution
• The mapping is defined using the Jacobian determinant
𝜕𝜕 𝑥𝑥1 , 𝑥𝑥2
𝑝𝑝 𝑦𝑦1 , 𝑦𝑦2 𝑑𝑑𝑦𝑦1 𝑑𝑑𝑦𝑦2 = 𝑢𝑢 𝑥𝑥1 , 𝑥𝑥2 𝑑𝑑𝑥𝑥1 𝑑𝑑𝑥𝑥2 = 𝑑𝑑𝑦𝑦1 𝑑𝑑𝑑𝑑2
𝜕𝜕(𝑦𝑦1 , 𝑦𝑦2 )
The Jacobian defines the two-variable distribution we want to sample.
This can be generalized to multi-variable cases.
• The Box-Muller method for sampling the normal distribution
The mapping is defined as follows
𝑦𝑦1 = −2𝑙𝑙𝑙𝑙𝑥𝑥1 cos 2𝜋𝜋𝑥𝑥2 ; 𝑦𝑦2 = −2𝑙𝑙𝑙𝑙𝑥𝑥1 sin 2𝜋𝜋𝑥𝑥2 . Or
1 1 𝑦𝑦
𝑥𝑥1 = 𝑒𝑒𝑒𝑒𝑒𝑒 − (𝑦𝑦12 + 𝑦𝑦22 ) ; 𝑥𝑥2 = 𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎 𝑦𝑦2
2 2𝜋𝜋 1