π Generative Pre-trained Transformer (GPT)
Descriptionβ
< What is it? >β
GPT stands for Generative Pre-trained Transformer. GPT-style language models use a decoder-only Transformer: they contain decoder blocks, token embeddings, and an output prediction head, but no Transformer encoder stack.
prompt tokens β embeddings + positions β decoder blocks with causal self-attention
β next-token logits β choose a token β append it β repeat
GPT is called generative because it predicts the next token, and pre-trained because it first learns this task from large collections of text before later fine-tuning or post-training.
Key pointsβ
< What decoder-only means >β
The decoder blocks in GPT use causal self-attention: token can attend to earlier tokens, but not later ones. This lets the same model generate text from left to right.
GPT does not need an encoder because it is not translating an input sequence into a separate output sequence in the way an encoderβdecoder model does. It also omits the cross-attention sublayer found in the original Transformer decoder, because there is no encoder output to attend to.
Original encoderβdecoder Transformer decoder:
masked self-attention β cross-attention to encoder output β feed-forward network
GPT decoder-only block:
masked self-attention β feed-forward network
< Causal attention and next-token prediction >β
During training, GPT models the probability of a token sequence as a product of next-token probabilities:
A causal mask prevents a token at position from seeing future position :
The negative-infinity entries become zero attention weight after softmax.
- Example: when predicting the next token after
Paris is the capital of, GPT can use all earlier tokens in that prefix but cannot use the future answer token while making its prediction.
< Pretraining and generation >β
For a training document, the document itself supplies the target next tokens. For example, the prefix The sky is can train the model to predict blue. This is self-supervised learning because no person needs to write a separate label.
At inference time, GPT repeatedly predicts one next-token distribution, selects a token, appends it to the context, and repeats. The selected token can be the most likely one or can be sampled for more variety.
< What GPT representations contain >β
GPT's initial token embeddings come from an embedding-table lookup. Each decoder layer then produces a context-dependent hidden state that can be used to predict the next token. Although GPT is not an encoder-only model, its hidden states are still useful contextual token representations.
Comparisonβ
For an encoder-only, decoder-only, and encoder-decoder model-family comparison, see Which models use an encoder or decoder?.
Related ideasβ
- Transformer explains attention, causal masks, and the original encoderβdecoder design.
- Large Language Models puts GPT-style models in the broader LLM workflow.
- BERT is the classic encoder-only Transformer model.
- Embeddings explains GPT's initial token vectors and later contextual representations.
- Softmax converts next-token logits into probabilities.
- Post-Training explains how a base GPT-style model becomes an instruct model.
Referenceβ
- Improving Language Understanding by Generative Pre-Training
- Attention Is All You Need
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer (T5)
- BART: Denoising Sequence-to-Sequence Pre-training