Skip to main content

πŸ“ Generative Pre-trained Transformer (GPT)

Description​

< What is it? >​

GPT stands for Generative Pre-trained Transformer. GPT-style language models use a decoder-only Transformer: they contain decoder blocks, token embeddings, and an output prediction head, but no Transformer encoder stack.

prompt tokens β†’ embeddings + positions β†’ decoder blocks with causal self-attention
β†’ next-token logits β†’ choose a token β†’ append it β†’ repeat

GPT is called generative because it predicts the next token, and pre-trained because it first learns this task from large collections of text before later fine-tuning or post-training.

Key points​

< What decoder-only means >​

The decoder blocks in GPT use causal self-attention: token tt can attend to earlier tokens, but not later ones. This lets the same model generate text from left to right.

GPT does not need an encoder because it is not translating an input sequence into a separate output sequence in the way an encoder–decoder model does. It also omits the cross-attention sublayer found in the original Transformer decoder, because there is no encoder output to attend to.

Original encoder–decoder Transformer decoder:
masked self-attention β†’ cross-attention to encoder output β†’ feed-forward network

GPT decoder-only block:
masked self-attention β†’ feed-forward network

< Causal attention and next-token prediction >​

During training, GPT models the probability of a token sequence as a product of next-token probabilities:

p(x1:T)=∏t=1Tp(xt∣x<t)p(x_{1:T}) = \prod_{t=1}^{T}p(x_t\mid x_{<t})

A causal mask prevents a token at position ii from seeing future position j>ij>i:

Mij={0j≀iβˆ’βˆžj>iM_{ij} = \begin{cases} 0 & j\leq i\\ -\infty & j>i \end{cases}

The negative-infinity entries become zero attention weight after softmax.

  • Example: when predicting the next token after Paris is the capital of, GPT can use all earlier tokens in that prefix but cannot use the future answer token while making its prediction.

< Pretraining and generation >​

For a training document, the document itself supplies the target next tokens. For example, the prefix The sky is can train the model to predict blue. This is self-supervised learning because no person needs to write a separate label.

At inference time, GPT repeatedly predicts one next-token distribution, selects a token, appends it to the context, and repeats. The selected token can be the most likely one or can be sampled for more variety.

< What GPT representations contain >​

GPT's initial token embeddings come from an embedding-table lookup. Each decoder layer then produces a context-dependent hidden state that can be used to predict the next token. Although GPT is not an encoder-only model, its hidden states are still useful contextual token representations.

Comparison​

For an encoder-only, decoder-only, and encoder-decoder model-family comparison, see Which models use an encoder or decoder?.

  • Transformer explains attention, causal masks, and the original encoder–decoder design.
  • Large Language Models puts GPT-style models in the broader LLM workflow.
  • BERT is the classic encoder-only Transformer model.
  • Embeddings explains GPT's initial token vectors and later contextual representations.
  • Softmax converts next-token logits into probabilities.
  • Post-Training explains how a base GPT-style model becomes an instruct model.

Reference​