Generative Adversarial Network (GAN)
Two forms of Machine Learning
• Discriminative
• Generative
Language Model
Given a language, Language Model estimate the probability of a word
sequence occurring in the language.
𝐿: Target language, say English.
V: Set of vocabularies in the language L.
𝑆 = {𝑤1 , 𝑤2 , … , 𝑤𝑚 |𝑤𝑖 𝜖, 𝑉}: a word sequence
Goal: Estimate the Pr(𝑆)that S is a valid sentence in L
Language Model
Pr 𝑆 = Pr(𝑤1 , 𝑤2 , … , 𝑤𝑚 )
Apply chain rule.
Pr 𝑆 = Pr 𝑤1 , 𝑤2 , … , 𝑤𝑚
= Pr 𝑤𝑚 𝑤1 , 𝑤2 , … , 𝑤𝑚−1 . Pr(𝑤1 , 𝑤2 , … , 𝑤𝑚−1 )
=
Pr 𝑤𝑚 𝑤1 , 𝑤2 , … , 𝑤𝑚−1 . Pr 𝑤𝑚−1 𝑤1 , 𝑤2 , … , 𝑤𝑚−2 . Pr(𝑤1 , 𝑤2 , … , 𝑤𝑚−2 )
= Pr 𝑤𝑚 𝑤1 , 𝑤2 , … , 𝑤𝑚−1 . Pr 𝑤𝑚−1 𝑤1 , 𝑤2 , … , 𝑤𝑚−2 . . . Pr 𝑤2 𝑤1 Pr 𝑤1
Pr 𝑤1 , 𝑤2 , … , 𝑤𝑚 = ෑ Pr(𝑤𝑖 |𝑤1 𝑤2 . . 𝑤𝑖−1 )
𝑖
Language Model – Two Common Tasks
Sentence generation: Estimate Pr 𝑆 = Pr 𝑤1 , 𝑤2 , … , 𝑤𝑚
P(“it is a nice movie”) = P(it) × P(is|it) × P(a|it is) × P(nice|it is a) ×
P(movie|it is a nice)
Language Model – Two Common Tasks
Sentence generation: Estimate Pr 𝑆 = Pr 𝑤1 , 𝑤2 , … , 𝑤𝑚
Next word prediction: Estimate Pr 𝑤𝑚 𝑤1 , 𝑤2 , … , 𝑤𝑚−1
P(“it is a nice movie”) = P(it) × P(is|it) × P(a|it is) × P(nice|it is a) ×
P(movie|it is a nice)
Language Model – Challenge
Support of 𝑤1 , 𝑤2 , … , 𝑤𝑚 may be small for large m i.e., there may not
be enough data for such large sequence.
Language Model – Challenge
Restrict the window size
𝑃(𝑚𝑜𝑣𝑖𝑒|𝑖𝑡 𝑖𝑠 𝑎 𝑛𝑖𝑐𝑒) ≈ 𝑃(𝑚𝑜𝑣𝑖𝑒|𝑛𝑖𝑐𝑒)
Or,
𝑃(𝑚𝑜𝑣𝑖𝑒|𝑖𝑡 𝑖𝑠 𝑎 𝑛𝑖𝑐𝑒) ≈ 𝑃(𝑚𝑜𝑣𝑖𝑒|𝑎 𝑛𝑖𝑐𝑒)
Language Model – Challenge
Restrict the window size.
𝑃(𝑚𝑜𝑣𝑖𝑒|𝑖𝑡 𝑖𝑠 𝑎 𝑛𝑖𝑐𝑒) ≈ 𝑃(𝑚𝑜𝑣𝑖𝑒|𝑛𝑖𝑐𝑒)
Or,
𝑃(𝑚𝑜𝑣𝑖𝑒|𝑖𝑡 𝑖𝑠 𝑎 𝑛𝑖𝑐𝑒) ≈ 𝑃(𝑚𝑜𝑣𝑖𝑒|𝑎 𝑛𝑖𝑐𝑒)
So, restrict to a small window of size k.
𝑃(𝑤1 , 𝑤2 , … , 𝑤𝑚 )
= Pr 𝑤𝑚 𝑤𝑚−𝑘−1 , 𝑤𝑚−𝑘−2 , … , 𝑤𝑚−1 . Pr 𝑤𝑚−1 𝑤𝑚−𝑘 , 𝑤𝑚−𝑘−2 , … , 𝑤𝑚−2 . . . Pr 𝑤1 𝑤2 Pr 𝑤1
Language Model – n gram
Unigram (k=1): 𝑃(𝑚𝑜𝑣𝑖𝑒|𝑖𝑡 𝑖𝑠 𝑎 𝑛𝑖𝑐𝑒) ≈ 𝑃(𝑚𝑜𝑣𝑖𝑒)
Bigram (k=2): 𝑃 𝑚𝑜𝑣𝑖𝑒 𝑖𝑡 𝑖𝑠 𝑎 𝑛𝑖𝑐𝑒 ≈ 𝑃 𝑚𝑜𝑣𝑖𝑒 𝑛𝑖𝑐𝑒
Trigram(k=3):𝑃 𝑚𝑜𝑣𝑖𝑒 𝑖𝑡 𝑖𝑠 𝑎 𝑛𝑖𝑐𝑒 ≈ 𝑃 𝑚𝑜𝑣𝑖𝑒 𝑎 𝑛𝑖𝑐𝑒
n-gram (k=n)
𝑃 𝑤1 , 𝑤2 , … , 𝑤𝑚
= Pr 𝑤𝑚 𝑤𝑚−𝑘−1 , 𝑤𝑚−𝑘−2 , … , 𝑤𝑚−1 . Pr 𝑤𝑚−1 𝑤𝑚−𝑘 , 𝑤𝑚−𝑘−2 , … , 𝑤𝑚−2 . . . Pr 𝑤1 𝑤2 Pr 𝑤1
= ෑ Pr(𝑤𝑖 |𝑤𝑖−𝑘−1 𝑤𝑖−𝑘−2 . . 𝑤𝑖−1 )
𝑖
Language Generation
Consider a text corpus L in English, say, a collection of English sentences.
Let 𝑉 be the vocabulary set.
Let 𝑣 ⊆ 𝑉
The sentence generated from 𝑣 is defined by 𝑆 = argmax 𝑃(𝑆 ′ )
𝑠′
Give 𝒗 = {it, is, a, nice, movie} we may have various combinations such as
It is a nice movie, nice a is it movie, movie it a is nice, …..
𝑃 𝑖𝑡 𝑖𝑠 𝑎 𝑛𝑖𝑐𝑒 𝑚𝑜𝑣𝑖𝑒 > 𝑃(𝑛𝑖𝑐𝑒 𝑎 𝑖𝑠 𝑖𝑡 𝑚𝑜𝑣𝑖𝑒)
Language Model Could be built using
Statistical – we have just seen
Neural Network – Word2Vec
Generative Adversarial Network
Generative Adversarial Network
Counterfeiter Fraud Detector
Generative Adversarial Network
Counterfeiter Fraud Detector
Generator Discriminator
Generative Adversarial Network
generates fake samples as real as possible and tries to fool the Discriminator
Generator
tries to detect the fake samples generated by the Generator
Discriminator
Generative Adversarial Network
They are trained in an adversarial setting to master each
other’s task.
Generator
Discriminator
Magic of GANs….
● Images generated using StyleGAN- a
GAN variant
● These people don’t exist in real!!!!!!
● Image from Paper ‘A Style-Based
Generator Architecture for Generative
Adversarial Networks’
Karras, Tero et al. “A Style-Based Generator Architecture for Generative Adversarial Networks.” 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2019): 4396-4405.
GAN – both the generator and Discriminator are
neural network
Generator
Discriminator
GAN – both the generator and Discriminator are
neural network
𝐷
𝑥∼𝐷
Real Data Distribution
𝑍
𝑧∼𝑍
Random Sample
GAN – both the generator and Discriminator are
neural network
𝐷
𝑥∼𝐷
Real Data Distribution
𝐺(𝑧) 𝐷′
𝑍 G(z) ∼ 𝐷′
𝑧∼𝑍
𝐷 ≅ 𝐷′
Random Sample Generator
GAN – both the generator and Discriminator are
neural network
𝐷
Fake/Real
𝑥∼𝐷
Real Data Distribution
Discriminator
𝐺(𝑧) 𝐷′
𝑍 G(z) ∼ 𝐷′
𝑧∼𝑍
𝐷 ≅ 𝐷′
Random Sample Generator
GAN – both the generator and Discriminator are
neural network
𝐷
Real
𝑥∼𝐷
Real Data Distribution
Discriminator
GAN – both the generator and Discriminator are
neural network
𝐷
Fake
𝑥∼𝐷
Real Data Distribution
Discriminator
𝑍
𝑧∼𝑍
Random Sample Generator
GAN – both the generator and Discriminator are
neural network
𝐷
Fake
𝑥∼𝐷
Real Data Distribution
Discriminator
𝑍
𝑧∼𝑍
Random Sample Generator
GAN – both the generator and Discriminator are
neural network
𝐷
Fake
𝑥∼𝐷
Real Data Distribution
Discriminator
𝑍
𝑧∼𝑍
Random Sample Generator
GAN – both the generator and Discriminator are
neural network
𝐷
Fake
𝑥∼𝐷
Real Data Distribution
Discriminator
𝑍
𝑧∼𝑍
Random Sample Generator
GAN – both the generator and Discriminator are
neural network
𝐷
Real
𝑥∼𝐷
Real Data Distribution
Discriminator
𝑍
𝑧∼𝑍
Random Sample Generator
GAN – both the generator and Discriminator are
neural network
𝐷
Real
𝑥∼𝐷
Real Data Distribution
Discriminator
𝑍
𝑧∼𝑍
Random Sample Generator
GAN – both the generator and Discriminator are
neural network
𝐷
Fake
𝑥∼𝐷
Real Data Distribution
Discriminator
𝑍
𝑧∼𝑍
Random Sample Generator
GAN – both the generator and Discriminator are
neural network
𝐷
Fake
𝑥∼𝐷
Real Data Distribution
Discriminator
𝑍
𝑧∼𝑍
Random Sample Generator
GAN – A toy Example
GAN – A toy Example
GAN – A toy Example
GAN – A toy Example
GAN – A toy Example
GAN – A toy Example
GAN – A toy Example
GAN – A toy Example
GAN – A toy Example
GAN Framework
● Generator + Discriminator = GAN
● The latent vector belongs to some
random distribution
(Uniform/Gaussian)
● Both the generator and
discriminator network parameters
are updated during training
GAN – Loss Function
• Discriminator’s decision over real data should be accurate
• Maximize
• Discriminator’s decision over generated data should be considered
fake
• Maximize
• Generator is trained to increase the chances of D producing a high
probability for a fake sample
Variational AutoEncoder
Variational AutoEncoder
AutoEncoder Variational AutoEncoder
Variational AutoEncoder
x p 𝑃(𝑧|𝑥)
Variational AutoEncoder
x p p 𝑧 𝑥 ∼ 𝑞𝜃 𝑧 𝑥 , 𝑤ℎ𝑒𝑟𝑒 𝜃 =< 𝜇, 𝜎 >
Variational AutoEncoder
x p p 𝑧 𝑥 ∼ 𝑞𝜃 𝑧 𝑥 , 𝑤ℎ𝑒𝑟𝑒 𝜃 =< 𝜇, 𝜎 >
Using 𝒒𝜽 𝒛 𝒙 , generate a random sample z’
Variational AutoEncoder
p 𝑥𝑧
Variational AutoEncoder
p 𝑥𝑧
Variational AutoEncoder
p 𝑥𝑧
Variational AutoEncoder
p 𝑥𝑧
Variational AutoEncoder
p 𝑥𝑧
Variational AutoEncoder – Loss Function
p 𝑥𝑧
Reconstruction likelihood
Distance between learned distribution q and true prior
distribution p.
Variational AutoEncoder – Loss Function
p 𝑥𝑧
Reconstruction likelihood
Distance between learned distribution q and true prior
distribution p.
Variational AutoEncoder – Problem
p 𝑥𝑧
Z is randomly sample using < 𝜇, 𝜎 >. For a
random node, Backpropagation can not
flow through a random node
Variational AutoEncoder – Reparameterization
p 𝑥𝑧
Variational AutoEncoder
p 𝑥𝑧
Generator
Variational AutoEncoder – Reparameterization
p 𝑥𝑧