Bayesian Machine Learning Overview
Bayesian Machine Learning Overview
Amirabbas Asadi
March 2022
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Outline
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Bayesian Machine Learning
2, 4, 6, 8, 10, ?
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Bayesian Machine Learning
2, 4, 6, 8, 10, ?
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Bayesian Machine Learning
2, 4, 6, 8, 10, ?
f (n) = 2n
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Bayesian Machine Learning
2, 4, 6, 8, 10, ?
f (n) = 2n
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Bayesian Machine Learning
2, 4, 6, 8, 10, ?
f (n) = 2n
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Bayesian Machine Learning
H1 : f (n) = 2n
H2 : f (n) = 0.0167n − 0.25n4 + 1.4167n3 − 3.75n2 + 6.5667n − 2
5
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Bayesian Machine Learning
H1 : f (n) = 2n
H2 : f (n) = 0.0167n − 0.25n4 + 1.4167n3 − 3.75n2 + 6.5667n − 2
5
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Bayesian Machine Learning
H1 : f (n) = 2n
H2 : f (n) = 0.0167n − 0.25n4 + 1.4167n3 − 3.75n2 + 6.5667n − 2
5
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Bayesian Machine Learning
H1 : f (n) = 2n
H2 : f (n) = 0.0167n − 0.25n4 + 1.4167n3 − 3.75n2 + 6.5667n − 2
5
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Bayesian Machine Learning
H1 : f (n) = 2n
H2 : f (n) = 0.0167n − 0.25n4 + 1.4167n3 − 3.75n2 + 6.5667n − 2
5
Then why do people choose the first one for the same data???
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Bayesian Machine Learning
H1 : f (n) = 2n
H2 : f (n) = 0.0167n − 0.25n4 + 1.4167n3 − 3.75n2 + 6.5667n − 2
5
Then why do people choose the first one for the same data???
But How can we quantify and take into account a prior belief?
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Bayesian Machine Learning
But How can we quantify and take into account a prior belief?
We can encode our prior belief p(H) as distribution over all hypotheses
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Bayesian Machine Learning
But How can we quantify and take into account a prior belief?
We can encode our prior belief p(H) as distribution over all hypotheses
p(D|H)
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Bayesian Machine Learning
But How can we quantify and take into account a prior belief?
We can encode our prior belief p(H) as distribution over all hypotheses
p(D|H)p(H)
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Bayesian Machine Learning
But How can we quantify and take into account a prior belief?
We can encode our prior belief p(H) as distribution over all hypotheses
p(D|H)p(H)
p(D|H)p(H)
p(H|D) =
p(D)
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Bayesian Machine Learning
But How can we quantify and take into account a prior belief?
We can encode our prior belief p(H) as distribution over all hypotheses
p(D|H)p(H)
p(D|H)p(H)
p(H|D) =
p(D)
Bayes Theorem
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Bayesian Machine Learning
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Bayesian Machine Learning
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Bayesian Machine Learning
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Bayesian Machine Learning
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Bayesian Machine Learning
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Bayesian Machine Learning
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Bayesian Machine Learning
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Bayesian Machine Learning
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Bayesian Machine Learning
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Bayesian Machine Learning
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Bayesian Machine Learning
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Bayesian Machine Learning
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Bayesian Machine Learning
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Bayesian Machine Learning
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Bayesian Machine Learning
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Bayesian Machine Learning
Frequentist Bayesian
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Bayesian Machine Learning
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Bayesian Machine Learning
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Bayesian Machine Learning
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Bayesian Machine Learning
Probabilistic Programming
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Probabilistic Graphical Models
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Probabilistic Graphical Models
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Probabilistic Graphical Models
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Probabilistic Graphical Models
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Probabilistic Graphical Models
B C D
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Probabilistic Graphical Models
B C D
p(A, B, C, D, E) =
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Probabilistic Graphical Models
B C D
p(A, B, C, D, E) = p(E|B, C)
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Probabilistic Graphical Models
B C D
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Probabilistic Graphical Models
B C D
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Probabilistic Graphical Models
B C D
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Probabilistic Graphical Models
B C D
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Probabilistic Graphical Models
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Probabilistic Graphical Models
x1 x2 x3
x4 x5 x6
x7 x8 x9
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Probabilistic Graphical Models
x1 x2 x3
x4 x5 x6
x7 x8 x9
1 ∏
p(x) = ψc (xc )
Z
c∈C
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Probabilistic Programming
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Probabilistic Programming
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Probabilistic Programming
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Probabilistic Programming
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Probabilistic Programming
Turing (Julia)
TensorFlow Probability
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Probabilistic Programming
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Probabilistic Programming
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Probabilistic Programming
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Probabilistic Programming
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Probabilistic Programming
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Coding Time
Coding Time!
Constructing Probabilistic Models in PyMC3
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Probabilistic Programming
[Link]()
[Link]()
[Link]()
[Link]()
[Link]()
[Link]()
[Link]()
[Link]()
[Link]()
[Link]()
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Probabilistic Programming
[Link]()
[Link]()
[Link]()
[Link]()
[Link]()
[Link]()
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Inference Problem
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Inference Problem
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Why inference is difficult?
p(x|z)p(z)
p(z|x) =
p(x)
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Why inference is difficult?
p(x|z)p(z)
p(z|x) =
p(x)
To obtain p(x) we have to marginalize all possible
hypotheses:
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Why inference is difficult?
p(x|z)p(z)
p(z|x) =
p(x)
To obtain p(x) we have to marginalize all possible
hypotheses: ∫
p(x) = p(x, z)dz
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Why inference is difficult?
p(x|z)p(z)
p(z|x) =
p(x)
To obtain p(x) we have to marginalize all possible
hypotheses: ∫
p(x) = p(x, z)dz
Now imagine what does p(x) look like if we have used something
like Neural Networks inside the model!
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Why inference is difficult?
p(x|z)p(z)
p(z|x) =
p(x)
To obtain p(x) we have to marginalize all possible
hypotheses: ∫
p(x) = p(x, z)dz
Now imagine what does p(x) look like if we have used something
like Neural Networks inside the model!
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Why inference is difficult?
p(x|z)p(z)
p(z|x) =
p(x)
To obtain p(x) we have to marginalize all possible
hypotheses: ∫
p(x) = p(x, z)dz
Now imagine what does p(x) look like if we have used something
like Neural Networks inside the model!
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
However, By carefully choosing likelihood and prior, Exact
Inference is possible.
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
However, By carefully choosing likelihood and prior, Exact
Inference is possible.
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Coding Time
Coding Time!
MAP Inference
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Approximate Inference Methods
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Approximate Inference Methods
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Approximate Inference Methods
Variational Inference
Expectation Propagation
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Markov Chain Monte Carlo
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Markov Chain Monte Carlo
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Markov Chain Monte Carlo
X0 , X1 , X2 , ...
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Markov Chain Monte Carlo
X0 , X1 , X2 , ...
Such a stochastic process is called Markov Chain
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Markov Chain Monte Carlo
X0 , X1 , X2 , ...
Such a stochastic process is called Markov Chain
Under some conditions after a time τ the Markov Chain will forget
it’s initial State and becomes stationary
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Markov Chain Monte Carlo
X0 , X1 , X2 , ...
Such a stochastic process is called Markov Chain
Under some conditions after a time τ the Markov Chain will forget
it’s initial State and becomes stationary
In other words the terms in the sequence
Xτ +1 , Xτ +2 , Xτ +3 , ...
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Markov Chain Monte Carlo
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Markov Chain Monte Carlo
Definition
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Markov Chain Monte Carlo
Definition
The idea is to define the transition probability P (x′ |x) such that it
satisfies the detailed balance equations for the target distribution
π(x)
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Metropolis-Hastings Algorithm
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Metropolis-Hastings Algorithm
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Metropolis-Hastings Algorithm
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Metropolis-Hastings Algorithm
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Metropolis-Hastings Algorithm
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Metropolis-Hastings Algorithm
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Metropolis-Hastings Algorithm
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Metropolis-Hastings Algorithm
′ ′
The funny fact is that for computing π(x ) g(x|x )
π(x) g(x′ |x) we don’t need to
know the normalization constant of π(x)!
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Metropolis-Hastings Algorithm
Algorithm Metropolis-Hastings
1: x0 is the initial State
2: Tmax is the maximum number of iterations
3: t ← 0
4: while t < Tmax do
5: x′ ← sample a new candidate from g(x′ |xt )
′ ) g(x|x′ )
6: α ← min(1, π(x π(x) g(x′ |x) )
7: u ← sample from a uniform distribution on [0, 1]
8: if u < α then
9: xt+1 ← x′
10: else
11: xt+1 ← xt
12: t←t+1
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Gibbs Sampling
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Gibbs Sampling
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Gibbs Sampling
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Gibbs Sampling
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Gibbs Sampling
xi+1
1 ∼ p(xi+1
1 |x2 , x3 )
i i
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Gibbs Sampling
xi+1
1 ∼ p(xi+1
1 |x2 , x3 )
i i
xi+1
2 ∼ p(xi+1
2 |x1 , x3 )
i+1 i
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Gibbs Sampling
xi+1
1 ∼ p(xi+1
1 |x2 , x3 )
i i
xi+1
2 ∼ p(xi+1
2 |x1 , x3 )
i+1 i
xi+1
3 ∼ p(xi+1
3 |x1 , x2 )
i+1 i+1
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Gibbs Sampling
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Gibbs Sampling
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Gibbs Sampling
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Advances in MCMC methods
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Advances in MCMC methods
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Advances in MCMC methods
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Advances in MCMC methods
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Advances in MCMC methods
Inference Compilation
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Coding Time
Coding Time!
Inference using MCMC in PyMC3
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Problems with MCMC methods
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Problems with MCMC methods
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Problems with MCMC methods
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Problems with MCMC methods
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Problems with MCMC methods
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Variational Inference
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Variational Inference
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Variational Inference
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Variational Inference
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Variational Inference
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Variational Inference
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Variational Inference
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Variational Inference
∫
log p(x) = log p(x, z)dz
∫
p(x, z)q(z; λ)
= log dz
q(z; λ)
p(x, z)
= log Eq(z;λ)
q(z; λ)
p(x, z)
≥ Eq(z;λ) log
q(z; λ)
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Variational Inference
∫
log p(x) = log p(x, z)dz
∫
p(x, z)q(z; λ)
= log dz
q(z; λ)
p(x, z)
= log Eq(z;λ)
q(z; λ)
p(x, z)
≥ Eq(z;λ) log
q(z; λ)
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Variational Inference
surprisingly we have
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Variational Inference
surprisingly we have
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Variational Inference
surprisingly we have
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Variational Inference
surprisingly we have
Wow!!!
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Coding Time
Coding Time!
Variational Inference in PyMC3
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Advances in Variational Inference
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Advances in Variational Inference
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Advances in Variational Inference
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Advances in Variational Inference
Stein Methods
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Summary
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Practical Examples
Practical Examples
A review on a few practical examples
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Advances in Bayesian ML
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Advances in Bayesian ML
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Advances in Bayesian ML
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Advances in Bayesian ML
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Advances in Bayesian ML
Normalizing Flows
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Advances in Bayesian ML
Normalizing Flows
Gaussian Processes
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Advances in Bayesian ML
Normalizing Flows
Gaussian Processes
Energy-based Models
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
Discussion
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
References I
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .
References II
Le, Tuan Anh, Atilim Gunes Baydin, and Frank Wood (2017).
“Inference compilation and universal probabilistic programming”.
In: Artificial Intelligence and Statistics. PMLR, pp. 1338–1348.
. . . . . . . . . . . . . . . . . . . .
. . . . . . . . . . . . . . . . . . . .