Skip to main content

📝 Generative Adversarial Network (GAN)

Description​

< What is it? >​

A Generative Adversarial Network (GAN) is a generative model with two neural networks trained against each other:

  • The generator GG turns a random latent vector zz into a synthetic sample, such as an image
  • The discriminator DD estimates whether a sample looks real or was generated
random noise z → generator G → synthetic image
↓
real image ─────────────────→ discriminator D → real or fake score

As the discriminator becomes better at spotting fake samples, the generator learns to make samples that are harder to distinguish from real data.

Key points​

< Adversarial objective >​

The original GAN objective is a two-player minimax game:

min⁡Gmax⁡DEx∼pdata[log⁡D(x)]+Ez∼p(z)[log⁡(1−D(G(z)))]\min_G\max_D \mathbb{E}_{x\sim p_{\text{data}}}[\log D(x)] + \mathbb{E}_{z\sim p(z)}[\log(1-D(G(z)))]

The discriminator tries to give real examples a high score and generated examples a low score. The generator tries to make D(G(z))D(G(z)) high, so its generated samples appear real.

In practice, the generator often uses the non-saturating loss because it gives stronger gradients early in training:

LG=−Ez∼p(z)[log⁡D(G(z))]L_G = -\mathbb{E}_{z\sim p(z)}[\log D(G(z))]

< Alternating training >​

GANs do not train one fixed model against fixed labels. Each iteration commonly alternates:

  1. Sample real data and random vectors zz
  2. Generate fake samples G(z)G(z)
  3. Update DD to distinguish real samples from fake ones
  4. Update GG while keeping DD's parameters fixed, so its samples better fool DD
  • Example: the generator is like a counterfeiter and the discriminator is like a detective. As each improves, the other must improve too.

< Latent space and generation >​

The latent vector zz is sampled from a simple distribution, often a standard normal distribution. After training, changing zz produces different synthetic samples; smoothly changing zz can often smoothly change generated attributes such as pose, lighting, or style.

GANs can generate new samples without copying a training example exactly, but their quality and diversity depend heavily on the data, architecture, and training stability.

< Common challenges >​

  • Mode collapse: the generator produces only a narrow range of samples that reliably fool the discriminator
  • Unstable balance: if one network becomes much stronger, the other can receive weak or unhelpful learning signals
  • Hard evaluation: realistic-looking samples do not automatically mean the generator covers the full data distribution

Common stabilizers include careful learning-rate choices, normalization, data augmentation, spectral normalization, gradient penalties, and Wasserstein-style objectives. Exploding or vanishing gradients can also make adversarial training unstable.

< Typical uses >​

GANs are used for image synthesis, image-to-image translation, super-resolution, inpainting, data augmentation, and synthetic-data research. They can generate samples quickly after training because generation usually needs one pass through the generator.

Comparison​

< GAN compared with diffusion models >​

GANDiffusion model
Core ideaGenerator fools a discriminatorModel gradually removes noise from a sample
SamplingUsually one generator pass, so often fastUsually many denoising steps, so often slower
Training challengeAdversarial instability and mode collapseLong sampling and careful noise-schedule training
Common strengthSharp, fast samplesStrong quality and diversity in many modern image generators

Neither approach is automatically best. Compare output quality, diversity, latency, training cost, and the target application.

Reference​