Skip to main content

📝 Diffusion Models

Description​

< What are they? >​

Diffusion models are generative models that learn to turn noise into data. During training, they learn how to reverse a gradual noising process. During generation, they start with random noise and repeatedly denoise it to create an image, audio clip, molecule, or other sample.

  • Example: a text-to-image system can start from a field of random pixels and gradually denoise it into an image matching “a red bicycle beside a lake.”

Key points​

< Forward diffusion: add noise >​

Training begins with a clean sample x0x_0. At a randomly chosen step tt, Gaussian noise ϵ\epsilon is added according to a noise schedule:

xt=αˉt x0+1−αˉt ϵ,ϵ∼N(0,I)x_t = \sqrt{\bar\alpha_t}\,x_0 + \sqrt{1-\bar\alpha_t}\,\epsilon, \qquad \epsilon\sim\mathcal N(0,I)

Here, xtx_t is a partially noised version of the original sample, and αˉt\bar\alpha_t controls how much signal remains. As tt increases, the sample approaches random noise.

< Reverse diffusion: learn to denoise >​

A neural network is trained to estimate the noise that was added:

ϵθ(xt,t,c)≈ϵ\epsilon_\theta(x_t,t,c)\approx\epsilon

cc is optional conditioning information, such as a text embedding or class label. At generation time, the model starts from noise xTx_T and applies the learned reverse process repeatedly:

xT→xT−1→⋯→x0x_T\rightarrow x_{T-1}\rightarrow\cdots\rightarrow x_0

For images, the denoising network is often a U-Net or Transformer. Some variants predict a denoised sample or a velocity instead of the noise; the core idea remains iterative denoising.

< Conditioning and guidance >​

Conditioning tells the model what to generate. A text-to-image model commonly supplies text embeddings to the denoiser through cross-attention. Classifier-free guidance strengthens the influence of that condition during sampling, often making an image match the prompt more closely, although too much guidance can reduce diversity or create artifacts.

< Latent diffusion >​

Pixel-space diffusion is expensive at high resolution. Latent diffusion first encodes an image into a smaller latent representation, runs diffusion there, and then decodes the result back into pixels. This substantially reduces compute while keeping the model controllable with inputs such as text, masks, or bounding boxes.

< Strengths and limitations >​

  • Strengths: high-quality, diverse samples; flexible conditioning; and generally stable training compared with adversarial training.
  • Limitation: sampling usually requires many sequential denoising steps, so it can be slower than a one-pass generator.
  • Practical trade-off: faster samplers and fewer steps reduce latency, but can change image quality or diversity.

Comparison​

< Diffusion models vs GANs >​

Diffusion modelGenerative Adversarial Network (GAN)
GenerationRepeatedly denoises a random sampleRuns a generator once from a latent vector
Training signalDenoising or score-matching objectiveGenerator and discriminator compete
Typical strengthHigh fidelity, diversity, and flexible conditioningFast sampling after training
Typical difficultySequential sampling can be slowAdversarial training can be unstable or suffer mode collapse

Neither is universally better. The right choice depends on output quality, diversity, controllability, latency, and training resources.

Video Tutorial​

  • Generative Adversarial Network (GAN) is another major family of generative models.
  • Autoencoders provide the compression and decoding stages used by latent diffusion.
  • Embeddings explains the vector representations commonly used to condition text-to-image diffusion models.
  • Transformer explains the attention architecture often used in modern denoisers and cross-attention conditioning.

Reference​