๐ Transformer
Descriptionโ
< What is it? >โ
A Transformer is a neural-network architecture for sequences. Its attention mechanism lets each token combine information from other relevant tokens instead of processing tokens strictly one at a time. Transformers are the foundation of many large language models and Vision Transformers.
The original 2017 design is an encoderโdecoder model: the encoder turns a source sequence into contextual representations, while the decoder uses masked self-attention and cross-attention to generate a target sequence. The label means that each kind of block is repeated times. Many modern LLMs keep only the decoder stack.
- transformer arch
- transformer explain

Diagram source: transformer-model-in-PyTorch, based on the original Transformer architecture from Attention Is All You Need.
Key pointsโ
< What is attention? >โ
Attention is a learned way for one token to gather useful information from other tokens. For each token, the model assigns a non-negative weight to every possible source token; the weights sum to . It then uses those weights to mix the source tokens' value vectors into a new, contextual representation:
- Example: in โThe animal rested because it was tired,โ the token
itcan give a high weight toanimal. Information fromanimal's value vector then becomes part ofit's contextual representation.
In self-attention, the source tokens come from the same sequence. In cross-attention, they come from another sequence, such as source-language tokens in translation.
< Encoder layer and decoder layer >โ
The original Transformer repeats two different kinds of layers. For translation, the encoder reads the complete source sentence, while the decoder generates the target sentence one token at a time.
| Layer | Main components | What it produces |
|---|---|---|
| Encoder layer | Self-attention over all non-padding source tokens, then a feed-forward network | A contextual representation for every source token |
| Decoder layer | Masked self-attention over the target prefix, cross-attention over encoder outputs, then a feed-forward network | A representation used to predict the next target token |
For example, when translating โThe cat sleeps.โ into French, the encoder lets cat use the entire English sentence. While the decoder generates chat, its masked self-attention can use only earlier French output tokens, while cross-attention can retrieve information from English cat.
Each attention or feed-forward sublayer is wrapped in a residual connection and layer normalization. The residual path preserves the previous representation, while normalization helps the stack train stably. The blocks in the architecture diagram mean that these layers are repeated several times.
< Query, Key, and Value (Q, K, V) >โ
For a sequence representation , self-attention gives every token three learned roles. The same token-representation matrix is projected into , , and in parallel. The following sentence is a running example:
The animal rested because it was tired.
Lowercase , , and refer to one token's vectors. Uppercase , , and refer to the matrices formed from all tokens' vectors. While the model forms the contextual representation for it, the roles are:
| Symbol | Name | Intuition | Role in attention | Example |
|---|---|---|---|---|
| / | Query | What information is this token looking for? | It is compared with keys to decide where to attend | asks which token can help explain it |
| / | Key | What kind of information does this token offer? | It is matched against queries to produce attention scores | offers animal as a possible relevant token |
| / | Value | What information should this token contribute? | It is combined using the resulting attention weights | provides learned context that can be mixed into it |
- In the running sentence, the query for
itmay learn to look for a noun antecedent, while the key foranimalmay learn to advertise that it is one. If their score is high, softmax givesanimala high attention weight, so its value contributes strongly toit's new representation. - In โThe chef chopped the vegetables before serving them,โ the query for
themcan similarly attend tovegetables. - This is an intuition, not a hard-coded grammar rule: the model learns the vectors from data and can distribute attention across several tokens.
< Why use separate , , and ? >โ
Queries and keys serve the matching operation, while values carry the information that is passed onward. Separate query and key projections make attention directional:
The first score can be high when it needs information from animal, without requiring the reverse score to be high. Using for matching is possible when the dimensions fit, but it forces each value vector to serve both the lookup and content roles. Distinct projections give the model more flexibility to learn what to retrieve separately from what to pass to the next layer.
< From text to , , and >โ
The input starts as raw text. For the running example, an illustrative tokenization is:
The animal rested because it was tired.
โ tokenizer
[The, animal, rested, because, it, was, tired, .]
โ token embeddings + position information
xโ, xโ, xโ, xโ, xโ
, xโ, xโ, xโ
โ stack token vectors into X
โ three learned projections, calculated in parallel
Q, K, V
Actual tokenizers may split words into subwords, so the exact token list can differ.
< Token embeddings and positions >โ
The tokenizer turns each token into an integer token ID. For a token at position , the model retrieves its embedding from a learned embedding table :
animal โ token ID tแตข โ embedding-table lookup E[tแตข] โ token embedding eแตข
If the vocabulary has tokens, has shape . The lookup selects one row of . Mathematically, this is equivalent to multiplying a one-hot token-ID vector by :
The embedding table is a learned weight matrix, but implementations normally use a fast lookup rather than explicitly multiplying a one-hot vector. During training, backpropagation updates the embedding-table rows. Later attention and MLP layers transform these initial embeddings into contextual hidden-state representations.
Here, first Transformer block means layer 1 of the Transformer stack, not the first token. It receives the initial, non-contextual token representations. Before that block, each token embedding is combined with position information. In a simple additive-position view:
For the first Transformer block, is the matrix of token representations:
For one sequence, has shape . At this stage, there is one vector per token; the sequence has not been consolidated into one sentence embedding. The stack proceeds as follows:
raw text โ token embeddings + positions = Xโฝโฐโพ
โ Transformer block 1 โ Xโฝยนโพ
โ Transformer block 2 โ Xโฝยฒโพ
โ ...
In later Transformer blocks, is no longer the original token embeddings; it is the contextual hidden-state matrix output by the previous block.
Each attention head then applies three separate learned linear projections:
Here, , , and are separate learned weight matrices. The same matrix is used in all three multiplications, so the input sequence is not duplicated. For token it, the same token vector produces three learned vectors: , , and .
< Shapes in self-attention and cross-attention >โ
For one sequence and one attention head, omit the batch dimension and let have shape . The projection and output shapes are:
is the main hidden width of the Transformer: the number of learned numeric components in each token representation. It is not the number of tokens, layers, or vocabulary items. This width stays the same across Transformer blocks so residual connections can add their inputs and outputs. In standard multi-head attention, it is commonly divided across heads:
-
Simple example: after tokenizing
I like tea, suppose there are tokens and . The input matrix has shape , so every token has four learned numeric components:The rows represent
I,like, andtea. The numbers are invented for illustration; a real model learns them. Every later Transformer block still receives and returns three vectors with four components each.
Here, is the head dimension: the number of learned numeric components handled by one attention head. In the common equal-width design, , so must be divisible by the number of heads . Heads have independent learned projections; they may learn different or overlapping attention patterns, but no linguistic role is assigned to a particular head in advance.
- Eight-head example: if and , then . For each token, every head produces its own two-component query, key, and value vectors. For example, one head might form . The two numbers are learned latent coordinates, not two named aspects of language. Each head produces its own attention pattern over the complete sequence.
| Object | Shape | Role |
|---|---|---|
| Projects a token representation into a query | ||
| Projects a token representation into a key | ||
| Projects a token representation into a value | ||
| One query vector per token | ||
| One key vector per token | ||
| One value vector per token |
In self-attention, all three matrices have rows because they come from the same sequence of tokens. and must have the same vector width so their dot products work. may have a different width .
-
Concrete self-attention example: suppose there are tokens, , , and . Then is , and are , is , and are , and is :
The attention matrix gives one score for every query-token/key-token pair. After softmax, multiplying it by gives:
In cross-attention, the query sequence and the key/value sequence may have different numbers of tokens. For example, 3 decoder tokens can attend to 4 encoder tokens:
Then has shape , and has shape . Thus, and must have the same number of rowsโeach key needs a corresponding valueโwhile can have a different number of rows.
The , , and matrices are learned during pretraining with the rest of the model. A released checkpoint keeps them fixed during ordinary inference, while full fine-tuning can update them and adapter methods such as LoRA usually leave them frozen and learn small updates instead.
Lowercase , , and denote one token's vectors. Uppercase , , and stack the corresponding vectors for all tokens, which lets the model calculate attention for the complete sequence efficiently.
< Scaled dot-product attention >โ
An attention score is the unnormalized compatibility score between query and key . In scaled dot-product attention:
is the query and key width for one attention head: the number of learned numeric components in each vector. It is a feature width, not a token count or a number of heads. In the common equal-width design, ; for and , .
-
Why divide by : if the components have similar scale, a dot product tends to grow as its dimension grows. The division keeps scores from becoming so large that softmax becomes overly peaked and gradients become less useful.
-
Tiny numeric example: if and , each vector has three components, so . Their unscaled dot product compares corresponding components:
The scaled attention score is .
A higher score means key is a stronger match for query . Softmax converts the scores across all keys into normalized attention weights, which form a weighted sum of the values:
has shape : row contains how strongly token attends to every token. Some sources use โattention scoreโ for the normalized weights too; here, it means the raw value before softmax.
-
Concrete output example: In โThe animal rested because it was tired,โ suppose one head gives the query for
itthe following attention weights and two-dimensional value vectors. These are toy numbersโthe real vectors are learned and usually have many more dimensions.Key token Attention weight Value vector TheanimalrestedbecauseitwastiredThe weights sum to . The output for
itis their weighted mixture of value vectors:animalcontributes the most because it has the largest weight. The output is not readable English; it is a new contextual representation ofitthat carries information from the tokens it attended to.
< Multi-head attention >โ
Multi-head attention runs independent attention projections in parallel. Each head attends over the complete sequence using its own learned , , and , so heads can learn different or overlapping patterns. Their outputs are concatenated, then mixes them back into the model width.
For example, if and , each head commonly uses .
Diagram source: transformer-model-in-PyTorch, based on the original Transformer attention diagrams.
For head , the model computes attention using its own projections, then concatenates the results from all heads:
-
Concrete merge: suppose and . For token , each head returns a two-number attention output:
Concatenation places the outputs side by side; it does not sum or average them:
The output projection then mixes the 16 components into one updated representation for the same token:
Implementations commonly calculate all heads together, then rearrange and reshape the tensor:
(batch, heads, tokens, d_h)โ (batch, tokens, heads, d_h)โ (batch, tokens, heads ร d_h)This merge happens independently at every token position, so the result is one contextual hidden-state vector per token, not one embedding for the entire sentence. A sentence-level embedding, when needed, is created later by a task-specific method such as pooling token vectors or using a special
[CLS]token.
< Feed-forward network (FFN) >โ
After attention mixes information between tokens, a position-wise feed-forward network (FFN) transforms each token representation independently. The same FFN weights are applied at every sequence position:
The first linear layer usually expands the representation from to a larger hidden size , an activation adds non-linearity, and the second linear layer projects it back to . The original Transformer used ReLU; modern models often use GELU or SwiGLU. Because the FFN does not mix positions, attention handles token-to-token communication and the FFN refines what each token has learned.
< Positional encoding >โ
Attention alone is permutation-equivariant: without position information, it can tell which token representations are present but not their order. Positional encoding tells the model where each token occurs in the sequence.
In the original Transformer, a position vector is added to each token embedding before the first layer, as shown earlier:
- Example:
dog bites manandman bites dogcontain the same words but have different meanings. Position information lets the model distinguish their order. - Common choices: fixed sinusoidal encodings, learned position embeddings, and rotary position embeddings (RoPE). RoPE instead modifies the query and key vectors so their dot product can express relative position.
< Attention masks >โ
An attention mask tells a token which positions it is allowed to use. Position encoding says where a token is; a mask says which tokens it may attend to.
An attention mask is added to the raw scaled dot-product scores before softmax:
controls whether query token may attend to key token . Use for an allowed position and for a blocked position. After softmax, a blocked position has weight . In real implementations, is often represented by a very large negative number for numerical safety.
| Mask | What it blocks | Example |
|---|---|---|
| Causal / look-ahead mask | Future keys: | In โThe cat slept,โ the query at cat may attend to The and cat, but not the future token slept |
| Padding mask | Padding keys inserted to make a batch the same length | In a batch containing โBirds flyโ and โThe cat slept,โ the padding after Birds fly receives no attention weight |
A causal mask lets a decoder predict the next token without seeing the answer ahead of time. Padding masks prevent meaningless <PAD> tokens from affecting real token representations.
< KV cache during generation >โ
In an autoregressive decoder, a cache usually means a keyโvalue (KV) cache. When the model generates token , its new query must attend to the keys and values of all earlier tokens. Those earlier key and value vectors will not change, so the model stores them rather than recomputing the full prefix at every step.
For each layer , the cache after processing tokens through contains:
- Prefill: run the prompt through every layer once and fill the cache with its and vectors.
- Decode: for the newly generated token, compute one new , , and . Compare with the cached keys, mix the cached values, then append and to the cache.
- What it stores: numeric attention keys and values for every past token, head, and layerโcommonly shaped like
(batch, heads, past_length, head_dim). - Why it stores keys and values, not queries: every later query compares with each earlier key and needs the paired value for its weighted sum. An earlier query is used only to form that earlier token's attention output, so later tokens do not need it.
- What it does not normally store: Query vectors, model weights, raw text, or verified facts. The โvalueโ is an attention vector, not a next-token probability or the answer text.
The KV cache makes generation much faster, but it grows with sequence length and consumes memory. In an encoderโdecoder model, cross-attention can also cache the encoder's keys and values because they stay fixed while the decoder generates.
- KV Cache 1
- KV Cache 2
- KV Cache 3
- KV Cache 4




Comparisonโ
< Self-attention vs cross-attention >โ
| Attention type | Queries come from | Keys and values come from | Typical use | Example |
|---|---|---|---|---|
| Self-attention | A sequence | The same sequence | Understanding relationships within text, image patches, or another sequence | In โThe animal rested because it was tired,โ it can attend to animal |
| Cross-attention | One sequence | A different sequence | A decoder attending to encoder outputs, or text attending to image features | Translate โThe cat sleeps.โ to โLe chat dort.โ: while generating chat, the decoder can attend to source token cat |
In the self-attention example, every , , and vector comes from the English sentence. In the translation example, the decoder's query comes from its French-side representation, while the English encoder outputs provide the keys and values.
< Which models use an encoder or decoder? >โ
| Model family | Encoder stack | Decoder stack | How it uses context | Typical use |
|---|---|---|---|---|
| BERT | โ | โ | Bidirectional self-attention over the full input | Classification, search, token representations |
| GPT-style | โ | โ | Causal self-attention over earlier tokens only | Text generation, chat, code completion |
| T5 / BART | โ | โ | Encoder reads the input; causal decoder attends to earlier output tokens and encoder output | Translation, summarization, text-to-text tasks |
| Vision Transformer (ViT) | โ | โ | Self-attention over image-patch tokens | Image classification and vision tasks |
The words encoder and decoder describe Transformer components, not whether a model can create embeddings. A decoder-only GPT still creates useful hidden representations; it simply does so with left-to-right context rather than full bidirectional context.
Implementationโ
Favoritesโ
๐ค
Related ideasโ
- Embeddings explains how token IDs become the vectors supplied to Transformer blocks.
- Large Language Models use Transformer-based architectures.
- Vision Transformer applies Transformer encoders to image patches.
- Low-Rank Adaptation (LoRA) adapts Transformer projections with a small number of trainable parameters.
Video Tutorialโ
- Transformer Neural Networks, ChatGPT's foundation, Clearly Explained!!!
- Attention in transformers, step-by-step | Deep Learning Chapter 6
- Query, Key and Value Matrix for Attention Mechanisms in Large Language Models
- Why the name Query, Key and Value? Self-Attention in Transformers
- The KV Cache: Memory Usage in Transformers
Referenceโ
- Attention Is All You Need (2017)
- Transformer using PyTorch (geeksforgeeks.org)
- LLM Visualization (by Ben Bycroft)
- Transformers Explained Visually: Learn How LLM Transformer Models Work (by Aeree, Grace, Alex, Seongmin, and Alec)
- Neural Networks - series courses (3Blue1Brown | YouTube)
- transformer-model-in-PyTorch โ source of the locally copied diagrams
- Mastering LLama โ Understanding Residual Connection (Hugman Sangkeun Jung)
- Attention layers in Transformer (PyLessons: architecture, code, charts, and video)