15

Introduction to Generative Models

Cantonese podcast title: 生成模型導論

Learning Objectives

  1. Derive the evidence lower bound (ELBO) for a latent-variable model and
  2. Write down the forward and reverse processes of a denoising diffusion
  3. State the GAN minimax objective and the saturation argument that
  4. Compare VAEs, GANs, and diffusion models on the axes of log-likelihood,
  5. Justify why the reparameterisation trick is the standard tool for
Introduction to Generative Models — visual guide
Lesson 15 — Diffusion forward and reverse chains A diagram of a denoising diffusion process showing six timesteps from x_0 to x_5: the forward chain adds Gaussian noise (top arrows), the reverse chain denoises with eps_theta (bottom arrows). The noise schedule alpha_bar_t monotonically decays. Diffusion · Forward & Reverse variance-preserving chain · noise prediction eps_θ(x_t, t) x₀ clean x₁ α̅≈0.9 x₂ α̅≈0.6 x₃ α̅≈0.3 x₄ α̅≈0.1 x_T 𝒩(0, I) +√β·ε +√β·ε +√β·ε +√β·ε +√β·ε FORWARD · q(xₜ | xₜ₋₁) μ_θ(xₜ, t) μ_θ(xₜ, t) μ_θ(xₜ, t) μ_θ(xₜ, t) μ_θ(xₜ, t) REVERSE · p_θ(xₜ₋₁ | xₜ) training loss (simplified) ℒ = ‖ ε − ε_θ(xₜ, t) ‖² variance-preserving Var(xₜ) = 1 ∀ t ⟹ α̅ₜ → 0

Assumes you know from ML-101

This lesson builds on ML-101 Lesson 14 (Neural Networks & Backprop) for the chain rule that lets us differentiate through sampling operations, and on ML-101 Lesson 9 (Model Evaluation) for the KL divergence as a measure of distributional fit. We will not re-derive the chain rule or re-define what a probability distribution is.

The reader is also expected to be familiar with ML-101 Lesson 13 (Unsupervised Learning) at the level of "k-means clusters, PCA reduces dimension, anomaly detection looks at low-density regions". We will use this vocabulary to introduce the generative version of each task — the problem of modelling p(x)p(x) rather than just summarising the data — and the surprise is that three seemingly different model families (VAEs, diffusion models, GANs) can be derived from a single underlying principle: the minimisation of a divergence between a model distribution and the data distribution.

Learning Objectives

  1. Derive the evidence lower bound (ELBO) for a latent-variable model and explain why maximising the ELBO is equivalent, up to a constant, to minimising the KL divergence between the model and the data distribution.
  2. Write down the forward and reverse processes of a denoising diffusion model, derive the simplified training objective that depends only on the noise prediction error, and identify the variance-preserving constraint that ties the forward and reverse noise schedules together.
  3. State the GAN minimax objective and the saturation argument that motivates the non-saturating loss used in practice, and explain why the two-player game has a unique Nash equilibrium at pG=pdatap_G = p_{\text{data}}.
  4. Compare VAEs, GANs, and diffusion models on the axes of log-likelihood, sample quality, mode coverage, and training stability, and predict from first principles which family is best suited to a given application.
  5. Justify why the reparameterisation trick is the standard tool for differentiating through a sampling operation, and identify the assumption whose violation makes the estimator biased.

The generative-modelling problem

A generative model specifies a probability distribution pmodelp_{\text{model}} over an observation space X\mathcal{X}, and the goal of generative modelling is to choose pmodelp_{\text{model}} so that, given samples {xi}i=1N\{x_i\}_{i=1}^N from an unknown data distribution pdatap_{\text{data}}, the model assigns high probability to those samples and generalises beyond them. The choice of loss function — what "fits well" means — determines the entire theory of generative modelling, and the three families in this lesson correspond to three different choices of loss.

The most direct loss is the negative log-likelihood

LNLL(θ)  =  −1N∑i=1Nlog⁡pmodel(xi∣θ),\mathcal{L}_{\text{NLL}}(\theta) \;=\; -\frac{1}{N}\sum_{i=1}^{N}\log p_{\text{model}}(x_i \mid \theta),

which is the maximum-likelihood estimator and is consistent for the true distribution in the limit of infinite data. Maximum likelihood is the loss implicit in VAEs and diffusion models, but with different parameterisations of pmodelp_{\text{model}}. GANs, by contrast, do not correspond to any explicit pmodelp_{\text{model}} — they learn a sampler, and their loss is a two-player adversarial game that converges to the data distribution at a Nash equilibrium.

The two methods that pick a likelihood are forced to confront the intractable partition function problem: many model families have

pmodel(x)  =  p~(x)Z(θ),p_{\text{model}}(x) \;=\; \frac{\tilde p(x)}{Z(\theta)},

with p~(x)\tilde p(x) easy to evaluate pointwise but Z(θ)=∫p~(x) dxZ(\theta) = \int \tilde p(x)\,dx hard to compute. The naïve Monte-Carlo estimator of log⁡Z\log Z has variance proportional to Z2Z^2, which is hopeless when the normalising constant is astronomically large (as it is for any high-dimensional distribution). The three families in this lesson are exactly three different tricks for getting around this problem: VAEs introduce a latent variable and bound the log-likelihood from below, diffusion models decompose the distribution into a chain of small conditional steps whose partition functions are tractable, and GANs sidestep the problem entirely by not modelling pp at all.

A 101-level misconception is that generative models "generate new data" — they generate samples from pmodelp_{\text{model}}, which is a different statement. A model can sample without being a good generative model (memorised training data, low-diversity outputs) and a model can be a good generative model without ever sampling (a likelihood-based model that is too slow to draw from).

Variational autoencoders: the ELBO

A VAE introduces a latent variable z∈Rmz \in \mathbb{R}^m and factorises the joint as pθ(x,z)=pθ(x∣z) p(z)p_\theta(x, z) = p_\theta(x \mid z)\,p(z) for a chosen prior p(z)p(z) (standard normal, in the canonical case). The marginal likelihood of an observation is

log⁡pθ(x)  =  log⁡∫pθ(x∣z) p(z) dz,\log p_\theta(x) \;=\; \log \int p_\theta(x \mid z)\,p(z)\,dz,

which is intractable for most choices of pθ(x∣z)p_\theta(x \mid z). The evidence lower bound (ELBO) is a tractable surrogate obtained by introducing a variational posterior qϕ(z∣x)q_\phi(z \mid x) and applying Jensen's inequality to the log:

log⁡pθ(x)  ≥  Eqϕ(z∣x) ⁣[log⁡pθ(x∣z)]  −  KL⁡ ⁣(qϕ(z∣x) ∥ p(z))  =:  ELBO(x;θ,ϕ).\log p_\theta(x) \;\ge\; \mathbb{E}_{q_\phi(z \mid x)}\!\left[\log p_\theta(x \mid z)\right] \;-\; \operatorname{KL}\!\bigl(q_\phi(z \mid x)\,\Vert\,p(z)\bigr) \;=:\; \text{ELBO}(x; \theta, \phi).

The bound is tight when qϕ(z∣x)=pθ(z∣x)q_\phi(z \mid x) = p_\theta(z \mid x) — i.e. when the variational posterior equals the true posterior — and the gap between log⁡pθ(x)\log p_\theta(x) and the ELBO is exactly KL⁡(qϕ(z∣x) ∥ pθ(z∣x))\operatorname{KL}(q_\phi(z \mid x)\,\Vert\,p_\theta(z \mid x)), which is the central quantity the VAE is minimising on the inference side.

Maximising the ELBO has two equivalent readings. From the inference viewpoint, the encoder qϕ(z∣x)q_\phi(z \mid x) is being pulled toward the true posterior pθ(z∣x)p_\theta(z \mid x) by the KL term, and the decoder pθ(x∣z)p_\theta(x \mid z) is being trained to reconstruct xx from samples of zz. From the model viewpoint, the decoder is being trained to match pdata(x)p_{\text{data}}(x) — under the bound, the marginal pθ(x)p_\theta(x) that the decoder defines is closer to pdatap_{\text{data}} to the extent that the ELBO is larger.

The reparameterisation trick is what makes the encoder gradient computable. For the canonical Gaussian encoder

qϕ(z∣x)  =  N ⁣(z; μϕ(x), diag⁡(σϕ2(x))),q_\phi(z \mid x) \;=\; \mathcal{N}\!\bigl(z;\, \mu_\phi(x),\,\operatorname{diag}(\sigma_\phi^2(x))\bigr),

we re-write the sample as z=μϕ(x)+σϕ(x)⊙ϵz = \mu_\phi(x) + \sigma_\phi(x) \odot \epsilon with ϵ∼N(0,I)\epsilon \sim \mathcal{N}(0, I), and the expectation in the ELBO becomes a deterministic function of ϕ\phi and a random ϵ\epsilon that does not depend on ϕ\phi. The gradient of the expectation is then

∇ϕ Eqϕ(z∣x)[ ⋅ ]  =  Eϵ∼N(0,I) ⁣[∇ϕ (expression in μϕ(x),σϕ(x),ϵ)],\nabla_\phi\,\mathbb{E}_{q_\phi(z \mid x)}[\,\cdot\,] \;=\; \mathbb{E}_{\epsilon \sim \mathcal{N}(0, I)}\!\left[\nabla_\phi\,\bigl(\text{expression in } \mu_\phi(x), \sigma_\phi(x), \epsilon\bigr)\right],

which is computable by backpropagation through the encoder network. The trick requires that the noise can be factored outside the parameters — for continuous distributions this is always possible; for discrete latents it is not, and discrete-latent VAEs need a different estimator (e.g. REINFORCE with a learned baseline). This is the assumption whose violation breaks the trick.

import torch
import torch.nn as nn
import torch.nn.functional as F

class VAE(nn.Module):
    def __init__(self, in_dim, latent_dim):
        super().__init__()
        self.enc = nn.Linear(in_dim, 2 * latent_dim)   # outputs mu and log-sigma
        self.dec = nn.Linear(latent_dim, in_dim)

    def encode(self, x):
        h = self.enc(x)
        mu, log_sigma = h.chunk(2, dim=-1)
        return mu, log_sigma

    def reparameterise(self, mu, log_sigma):
        # The noise epsilon is independent of phi; gradient flows through mu and log_sigma.
        eps = torch.randn_like(mu)
        return mu + eps * log_sigma.exp()

    def forward(self, x):
        mu, log_sigma = self.encode(x)
        z = self.reparameterise(mu, log_sigma)
        x_hat = self.dec(z)
        # Reconstruction term is per-sample MSE (or Bernoulli/Gaussian log-likelihood).
        recon = F.mse_loss(x_hat, x, reduction="sum")
        # Closed-form KL between q(z|x) and the standard normal prior.
        kl = -0.5 * torch.sum(1 + 2 * log_sigma - mu.pow(2) - (2 * log_sigma).exp())
        return recon, kl

A VAE's loss is recon + kl, and the KL term has a closed form only because the prior is Gaussian and the encoder is also Gaussian with diagonal covariance. For richer families (mixture priors, normalising-flow posteriors) the KL must be sampled.

Diffusion models: forward and reverse processes

A denoising diffusion model specifies a Markov chain that gradually corrupts the data distribution pdatap_{\text{data}} into a noise distribution p∞=N(0,I)p_{\infty} = \mathcal{N}(0, I) over TT steps. The forward process adds Gaussian noise according to a variance schedule {βt}t=1T\{\beta_t\}_{t=1}^T with βt∈(0,1)\beta_t \in (0, 1):

q(xt∣xt−1)  =  N ⁣(xt; 1−βt xt−1, βtI).q(x_t \mid x_{t-1}) \;=\; \mathcal{N}\!\bigl(x_t;\,\sqrt{1-\beta_t}\,x_{t-1},\,\beta_t I\bigr).

Because the noise is Gaussian at each step, the marginal at step tt conditioned on x0x_0 has a closed form

q(xt∣x0)  =  N ⁣(xt; αˉt x0, (1−αˉt)I),q(x_t \mid x_0) \;=\; \mathcal{N}\!\bigl(x_t;\,\sqrt{\bar\alpha_t}\,x_0,\,(1-\bar\alpha_t)I\bigr),

where αt=1−βt\alpha_t = 1 - \beta_t and αˉt=∏s=1tαs\bar\alpha_t = \prod_{s=1}^t \alpha_s — the cumulative product of noise-preservation factors. The variance-preserving constraint is that the marginal variance of xtx_t remains 11 for all tt, which requires βt\beta_t to satisfy αˉt→0\bar\alpha_t \to 0 as t→Tt \to T.

The reverse process is the generative direction: starting from xT∼N(0,I)x_T \sim \mathcal{N}(0, I) and applying the learned transition

pθ(xt−1∣xt)  =  N ⁣(xt−1; μθ(xt,t), Σθ(xt,t)),p_\theta(x_{t-1} \mid x_t) \;=\; \mathcal{N}\!\bigl(x_{t-1};\,\mu_\theta(x_t, t),\,\Sigma_\theta(x_t, t)\bigr),

with parameters θ\theta. The training objective simplifies dramatically when we parameterise μθ(xt,t)\mu_\theta(x_t, t) as

μθ(xt,t)  =  1αt ⁣(xt−βt1−αˉt ϵθ(xt,t)),\mu_\theta(x_t, t) \;=\; \frac{1}{\sqrt{\alpha_t}}\!\left(x_t - \frac{\beta_t}{\sqrt{1-\bar\alpha_t}}\,\epsilon_\theta(x_t, t)\right),

where ϵθ(xt,t)\epsilon_\theta(x_t, t) is a neural network that predicts the noise added at step tt. The training loss reduces to

Lsimple(θ)  =  Et,x0,ϵ ⁣[ ∥ϵ−ϵθ(αˉt x0+1−αˉt ϵ, t)∥2 ],\mathcal{L}_{\text{simple}}(\theta) \;=\; \mathbb{E}_{t, x_0, \epsilon}\!\left[\,\lVert \epsilon - \epsilon_\theta(\sqrt{\bar\alpha_t}\,x_0 + \sqrt{1-\bar\alpha_t}\,\epsilon,\,t)\rVert^2\,\right],

which is a mean-squared error between the true noise ϵ\epsilon and the network's prediction. The simplification is exact under the variance-preserving schedule; with a non-VP schedule, the loss acquires an LtL_t weighting whose optimal weighting is the signal-to-noise ratio αˉt/(1−αˉt)\bar\alpha_t / (1-\bar\alpha_t).

def diffusion_loss(model, x0, alphas_cumprod, device="cuda"):
    """
    Simplified DDPM training loss: predict the noise that was added.
    The schedule (alphas_cumprod) is fixed for the whole training run.
    """
    B = x0.size(0)
    t = torch.randint(0, len(alphas_cumprod), (B,), device=device)   # uniform timestep
    noise = torch.randn_like(x0)
    a_t = alphas_cumprod[t].view(-1, 1, 1, 1)                       # broadcast over (C, H, W)
    x_t = a_t.sqrt() * x0 + (1.0 - a_t).sqrt() * noise              # forward step
    pred_noise = model(x_t, t)                                       # reverse model
    return F.mse_loss(pred_noise, noise)

The diffusion forward process is the unique chain with Gaussian transitions whose marginal is always Gaussian in the limit and whose conditional pθ(xt−1∣xt)p_\theta(x_{t-1}\mid x_t) is tractable to fit. The reverse process's Markov property is the assumption whose violation breaks the chain — if the model needs to look at x0x_0 to denoise (i.e. if the optimal reverse transition depends on information not in xtx_t), the parameterisation above cannot represent it and the chain collapses.

A 101-level misconception is that diffusion models "generate by removing noise". They generate by running the reverse chain, which is a parameterised function of xtx_t and the timestep; the noise is just the source of randomness in the early steps and a residual term in the reparameterised loss.

GANs: the two-player game

A generative adversarial network defines a generator Gθ:z↦xG_\theta: z \mapsto x that maps noise samples z∼p(z)z \sim p(z) to data samples xx, and a discriminator Dϕ:x↦(0,1)D_\phi: x \mapsto (0, 1) that estimates the probability that a sample came from the data rather than the generator. The two networks play a minimax game with value function

min⁡θmax⁡ϕ  V(θ,ϕ)  =  Ex∼pdata[log⁡Dϕ(x)]  +  Ez∼p(z)[log⁡(1−Dϕ(Gθ(z)))].\min_\theta \max_\phi\; V(\theta, \phi) \;=\; \mathbb{E}_{x \sim p_{\text{data}}}[\log D_\phi(x)] \;+\; \mathbb{E}_{z \sim p(z)}[\log(1 - D_\phi(G_\theta(z)))].

The discriminator is trained to maximise this — assign high probability to real data, low to generated — and the generator is trained to minimise it — push its samples toward looking real. Under standard assumptions (both networks have infinite capacity and the optimisation is run until convergence), the game has a unique Nash equilibrium at pGθ=pdatap_{G_\theta} = p_{\text{data}}, and the discriminator at equilibrium outputs D∗(x)=1/2D^*(x) = 1/2 everywhere — i.e. it cannot distinguish real from generated.

The original GAN loss has a saturation problem in practice: when the discriminator confidently assigns low probability to generated samples, log⁡(1−Dϕ(Gθ(z)))\log(1 - D_\phi(G_\theta(z))) saturates to a small gradient and the generator stops learning. The fix is the non-saturating loss, where the generator is trained to maximise log⁡Dϕ(Gθ(z))\log D_\phi(G_\theta(z)) rather than minimise log⁡(1−Dϕ(Gθ(z)))\log(1 - D_\phi(G_\theta(z))). The two losses have the same fixed point but opposite signs of the gradient at saturation, so the non-saturating form keeps the generator's gradient large when the discriminator is winning. This is a 101-level misconception: the GAN "loss" is a sum of two distinct objectives for two distinct players, and they cannot be summed and treated as a single scalar to minimise.

The deeper question of what GANs actually optimise is answered by viewing the optimal discriminator D∗D^* as a function of the generator:

D∗(x)  =  pdata(x)pdata(x)+pGθ(x).D^*(x) \;=\; \frac{p_{\text{data}}(x)}{p_{\text{data}}(x) + p_{G_\theta}(x)}.

Substituting this optimal discriminator back into the value function gives

V(θ,D∗)  =  −2log⁡2+2JSD⁡ ⁣(pdata ∥ pGθ),V(\theta, D^*) \;=\; -2 \log 2 + 2 \operatorname{JSD}\!\bigl(p_{\text{data}}\,\Vert\,p_{G_\theta}\bigr),

where JSD⁡\operatorname{JSD} is the Jensen-Shannon divergence, bounded in [0,log⁡2][0, \log 2]. So the GAN game minimises JSD⁡(pdata,pGθ)\operatorname{JSD}(p_{\text{data}}, p_{G_\theta}) when the discriminator is optimal — which is why GAN training is unstable: the JSD is locally flat in regions where the two distributions do not overlap, and the discriminator can get stuck confidently assigning probability 0 to generated samples without giving the generator a useful gradient.

def gan_step(G, D, real_batch, z, opt_g, opt_d, device):
    """One non-saturating GAN update. The two optimisers step separately."""
    # --- train the discriminator ---
    D.train()
    opt_d.zero_grad()
    real_pred = D(real_batch.to(device))
    fake_pred = D(G(z.to(device)).detach())      # detach: do not flow grads into G here
    loss_d = (F.binary_cross_entropy(real_pred, torch.ones_like(real_pred)) +
              F.binary_cross_entropy(fake_pred, torch.zeros_like(fake_pred)))
    loss_d.backward(); opt_d.step()

    # --- train the generator with the non-saturating loss ---
    G.train()
    opt_g.zero_grad()
    fake_pred = D(G(z.to(device)))               # grads DO flow into G
    loss_g = F.binary_cross_entropy(fake_pred, torch.ones_like(fake_pred))
    loss_g.backward(); opt_g.step()

Comparing the three families

The three families are not interchangeable. They differ along four practical axes: log-likelihood, sample quality, mode coverage, and training stability.

FamilyLog-likelihoodSample qualityMode coverageTraining
VAEtractable lower boundgood but blurrycovers all modesstable
GANnone (implicit)sharp, photorealisticmode collapse riskunstable
Diffusiontractable lower boundstate-of-the-art FIDcovers all modesstable

The trade-off is not accidental. GANs win on sample quality because their training objective is a divergence rather than a likelihood, and the divergence is dominated by the modes that are easy to match — the "sharpness" of samples. They lose on mode coverage because the JSD is not sensitive to the absence of a mode; a generator that covers only a subset of the data can still drive the value function to zero. VAEs and diffusion models, by contrast, are tied to a likelihood and cannot ignore a mode whose contribution to the log-likelihood is non-zero.

The deeper reason for the stability difference is that VAE and diffusion losses are single-objective: a single scalar that gradients flow through, with no opposing player. GANs are two-objective: a minimax game whose dynamics can cycle, diverge, or get stuck. The non-saturating loss helps with the saturation pathology but does not fix the cycling pathology.

For practitioners: VAEs are the right default for tasks that need a likelihood (anomaly detection, density estimation, semi-supervised learning); diffusion models are the right default for high-fidelity image, audio, or video synthesis at moderate scale; GANs are the right choice when sample sharpness is the dominant metric and the training budget is large enough to tune the discriminator-generator balance.

Key Takeaways

  • The generative-modelling problem is to choose a model distribution pmodelp_{\text{model}} that fits the data. Maximum likelihood is the default loss, but most expressive models have an intractable partition function, and the three families in this lesson are three different tricks for getting around this intractability.
  • The ELBO bounds log⁡pθ(x)\log p_\theta(x) from below and equals it when the variational posterior matches the true posterior; maximising the ELBO is equivalent, up to a constant, to minimising the KL divergence between model and data distributions.
  • The reparameterisation trick factors the noise out of the encoder's parameters and is what makes the encoder gradient computable by backpropagation. The trick requires that the noise can be sampled outside the parameters — it is exact for continuous latents and approximate for discrete ones.
  • The diffusion forward process is a Markov chain with Gaussian transitions and a variance-preserving schedule; the training loss reduces to a noise-prediction MSE under the canonical parameterisation, and the reverse process is a neural network that predicts the noise given xtx_t and tt.
  • GANs are a two-player minimax game whose Nash equilibrium is pGθ=pdatap_{G_\theta} = p_{\text{data}}; the optimal discriminator at equilibrium outputs 1/21/2 everywhere, and the generator's value function with the optimal discriminator is the Jensen-Shannon divergence. The non-saturating loss is the practical fix for the saturation pathology.

Check your understanding

7 questions · 80% to complete the lesson

1 / 7

6 correct to pass

The evidence lower bound (ELBO) for a latent-variable model is tight when

0 of 7 answered

Pick a lesson to start the audio.