📝 Diffusion Models
Description
< What are they? >
Diffusion models are generative models that learn to turn noise into data. During training, they learn how to reverse a gradual noising process. During generation, they start with random noise and repeatedly denoise it to create an image, audio clip, molecule, or other sample.
- Example: a text-to-image system can start from a field of random pixels and gradually denoise it into an image matching “a red bicycle beside a lake.”
Key points
< Forward diffusion: add noise >
Training begins with a clean sample . At a randomly chosen step , Gaussian noise is added according to a noise schedule:
Here, is a partially noised version of the original sample, and controls how much signal remains. As increases, the sample approaches random noise.
< Reverse diffusion: learn to denoise >
A neural network is trained to estimate the noise that was added:
is optional conditioning information, such as a text embedding or class label. At generation time, the model starts from noise and applies the learned reverse process repeatedly:
For images, the denoising network is often a U-Net or Transformer. Some variants predict a denoised sample or a velocity instead of the noise; the core idea remains iterative denoising.
< Conditioning and guidance >
Conditioning tells the model what to generate. A text-to-image model commonly supplies text embeddings to the denoiser through cross-attention. Classifier-free guidance strengthens the influence of that condition during sampling, often making an image match the prompt more closely, although too much guidance can reduce diversity or create artifacts.
< Latent diffusion >
Pixel-space diffusion is expensive at high resolution. Latent diffusion first encodes an image into a smaller latent representation, runs diffusion there, and then decodes the result back into pixels. This substantially reduces compute while keeping the model controllable with inputs such as text, masks, or bounding boxes.
< Strengths and limitations >
- Strengths: high-quality, diverse samples; flexible conditioning; and generally stable training compared with adversarial training.
- Limitation: sampling usually requires many sequential denoising steps, so it can be slower than a one-pass generator.
- Practical trade-off: faster samplers and fewer steps reduce latency, but can change image quality or diversity.
Comparison
< Diffusion models vs GANs >
| Diffusion model | Generative Adversarial Network (GAN) | |
|---|---|---|
| Generation | Repeatedly denoises a random sample | Runs a generator once from a latent vector |
| Training signal | Denoising or score-matching objective | Generator and discriminator compete |
| Typical strength | High fidelity, diversity, and flexible conditioning | Fast sampling after training |
| Typical difficulty | Sequential sampling can be slow | Adversarial training can be unstable or suffer mode collapse |
Neither is universally better. The right choice depends on output quality, diversity, controllability, latency, and training resources.
Video Tutorial
Related ideas
- Generative Adversarial Network (GAN) is another major family of generative models.
- Autoencoders provide the compression and decoding stages used by latent diffusion.
- Embeddings explains the vector representations commonly used to condition text-to-image diffusion models.
- Transformer explains the attention architecture often used in modern denoisers and cross-attention conditioning.