📝 Generative Adversarial Network (GAN)
Description
< What is it? >
A Generative Adversarial Network (GAN) is a generative model with two neural networks trained against each other:
- The generator turns a random latent vector into a synthetic sample, such as an image
- The discriminator estimates whether a sample looks real or was generated
random noise z → generator G → synthetic image
↓
real image ─────────────────→ discriminator D → real or fake score
As the discriminator becomes better at spotting fake samples, the generator learns to make samples that are harder to distinguish from real data.
Key points
< Adversarial objective >
The original GAN objective is a two-player minimax game:
The discriminator tries to give real examples a high score and generated examples a low score. The generator tries to make high, so its generated samples appear real.
In practice, the generator often uses the non-saturating loss because it gives stronger gradients early in training:
< Alternating training >
GANs do not train one fixed model against fixed labels. Each iteration commonly alternates:
- Sample real data and random vectors
- Generate fake samples
- Update to distinguish real samples from fake ones
- Update while keeping 's parameters fixed, so its samples better fool
- Example: the generator is like a counterfeiter and the discriminator is like a detective. As each improves, the other must improve too.
< Latent space and generation >
The latent vector is sampled from a simple distribution, often a standard normal distribution. After training, changing produces different synthetic samples; smoothly changing can often smoothly change generated attributes such as pose, lighting, or style.
GANs can generate new samples without copying a training example exactly, but their quality and diversity depend heavily on the data, architecture, and training stability.
< Common challenges >
- Mode collapse: the generator produces only a narrow range of samples that reliably fool the discriminator
- Unstable balance: if one network becomes much stronger, the other can receive weak or unhelpful learning signals
- Hard evaluation: realistic-looking samples do not automatically mean the generator covers the full data distribution
Common stabilizers include careful learning-rate choices, normalization, data augmentation, spectral normalization, gradient penalties, and Wasserstein-style objectives. Exploding or vanishing gradients can also make adversarial training unstable.
< Typical uses >
GANs are used for image synthesis, image-to-image translation, super-resolution, inpainting, data augmentation, and synthetic-data research. They can generate samples quickly after training because generation usually needs one pass through the generator.
Comparison
< GAN compared with diffusion models >
| GAN | Diffusion model | |
|---|---|---|
| Core idea | Generator fools a discriminator | Model gradually removes noise from a sample |
| Sampling | Usually one generator pass, so often fast | Usually many denoising steps, so often slower |
| Training challenge | Adversarial instability and mode collapse | Long sampling and careful noise-schedule training |
| Common strength | Sharp, fast samples | Strong quality and diversity in many modern image generators |
Neither approach is automatically best. Compare output quality, diversity, latency, training cost, and the target application.
Related ideas
- Neural Networks introduces the building blocks used for and .
- Diffusion Models introduces another major generative-model family.
- Vanishing and Exploding Gradients explains one source of unstable training signals.