Skip to main content

๐Ÿ“ Transformer

Descriptionโ€‹

< What is it? >โ€‹

A Transformer is a neural-network architecture for sequences. Its attention mechanism lets each token combine information from other relevant tokens instead of processing tokens strictly one at a time. Transformers are the foundation of many large language models and Vision Transformers.

The original 2017 design is an encoderโ€“decoder model: the encoder turns a source sequence into contextual representations, while the decoder uses masked self-attention and cross-attention to generate a target sequence. The Nร—N\times label means that each kind of block is repeated NN times. Many modern LLMs keep only the decoder stack.

Diagram source: transformer-model-in-PyTorch, based on the original Transformer architecture from Attention Is All You Need.

Key pointsโ€‹

< What is attention? >โ€‹

Attention is a learned way for one token to gather useful information from other tokens. For each token, the model assigns a non-negative weight to every possible source token; the weights sum to 11. It then uses those weights to mix the source tokens' value vectors into a new, contextual representation:

oi=โˆ‘j=1TAijvjo_i = \sum_{j=1}^{T}A_{ij}v_j
  • Example: in โ€œThe animal rested because it was tired,โ€ the token it can give a high weight to animal. Information from animal's value vector then becomes part of it's contextual representation.

In self-attention, the source tokens come from the same sequence. In cross-attention, they come from another sequence, such as source-language tokens in translation.

< Encoder layer and decoder layer >โ€‹

The original Transformer repeats two different kinds of layers. For translation, the encoder reads the complete source sentence, while the decoder generates the target sentence one token at a time.

LayerMain componentsWhat it produces
Encoder layerSelf-attention over all non-padding source tokens, then a feed-forward networkA contextual representation for every source token
Decoder layerMasked self-attention over the target prefix, cross-attention over encoder outputs, then a feed-forward networkA representation used to predict the next target token

For example, when translating โ€œThe cat sleeps.โ€ into French, the encoder lets cat use the entire English sentence. While the decoder generates chat, its masked self-attention can use only earlier French output tokens, while cross-attention can retrieve information from English cat.

Each attention or feed-forward sublayer is wrapped in a residual connection and layer normalization. The residual path preserves the previous representation, while normalization helps the stack train stably. The Nร—N\times blocks in the architecture diagram mean that these layers are repeated several times.

< Query, Key, and Value (Q, K, V) >โ€‹

For a sequence representation XX, self-attention gives every token three learned roles. The same token-representation matrix is projected into QQ, KK, and VV in parallel. The following sentence is a running example:

The animal rested because it was tired.

Lowercase qiq_i, kik_i, and viv_i refer to one token's vectors. Uppercase QQ, KK, and VV refer to the matrices formed from all tokens' vectors. While the model forms the contextual representation for it, the roles are:

SymbolNameIntuitionRole in attentionExample
qiq_i / QQQueryWhat information is this token looking for?It is compared with keys to decide where to attendqitq_{\textit{it}} asks which token can help explain it
kik_i / KKKeyWhat kind of information does this token offer?It is matched against queries to produce attention scoreskanimalk_{\textit{animal}} offers animal as a possible relevant token
viv_i / VVValueWhat information should this token contribute?It is combined using the resulting attention weightsvanimalv_{\textit{animal}} provides learned context that can be mixed into it
  • In the running sentence, the query for it may learn to look for a noun antecedent, while the key for animal may learn to advertise that it is one. If their score is high, softmax gives animal a high attention weight, so its value contributes strongly to it's new representation.
  • In โ€œThe chef chopped the vegetables before serving them,โ€ the query for them can similarly attend to vegetables.
  • This is an intuition, not a hard-coded grammar rule: the model learns the vectors from data and can distribute attention across several tokens.
Query, key, and value roles in attention

< Why use separate QQ, KK, and VV? >โ€‹

Queries and keys serve the matching operation, while values carry the information that is passed onward. Separate query and key projections make attention directional:

qitkanimalโŠคโ‰ qanimalkitโŠคq_{\textit{it}}k_{\textit{animal}}^\top \ne q_{\textit{animal}}k_{\textit{it}}^\top

The first score can be high when it needs information from animal, without requiring the reverse score to be high. Using QVโŠคQV^\top for matching is possible when the dimensions fit, but it forces each value vector to serve both the lookup and content roles. Distinct projections give the model more flexibility to learn what to retrieve separately from what to pass to the next layer.

< From text to QQ, KK, and VV >โ€‹

The input starts as raw text. For the running example, an illustrative tokenization is:

The animal rested because it was tired.
โ†“ tokenizer
[The, animal, rested, because, it, was, tired, .]
โ†“ token embeddings + position information
xโ‚, xโ‚‚, xโ‚ƒ, xโ‚„, xโ‚…, xโ‚†, xโ‚‡, xโ‚ˆ
โ†“ stack token vectors into X
โ†“ three learned projections, calculated in parallel
Q, K, V

Actual tokenizers may split words into subwords, so the exact token list can differ.

< Token embeddings and positions >โ€‹

The tokenizer turns each token into an integer token ID. For a token at position ii, the model retrieves its embedding from a learned embedding table EE:

animal โ†’ token ID tแตข โ†’ embedding-table lookup E[tแตข] โ†’ token embedding eแตข
ei=E[ti]e_i = E[t_i]

If the vocabulary has โˆฃVโˆฃ|\mathcal{V}| tokens, EE has shape (โˆฃVโˆฃ,dmodel)(|\mathcal{V}|,d_{\mathrm{model}}). The lookup selects one row of EE. Mathematically, this is equivalent to multiplying a one-hot token-ID vector by EE:

ei=onehotโก(ti)โŠคEe_i = \operatorname{onehot}(t_i)^\top E

The embedding table is a learned weight matrix, but implementations normally use a fast lookup rather than explicitly multiplying a one-hot vector. During training, backpropagation updates the embedding-table rows. Later attention and MLP layers transform these initial embeddings into contextual hidden-state representations.

Here, first Transformer block means layer 1 of the Transformer stack, not the first token. It receives the initial, non-contextual token representations. Before that block, each token embedding eie_i is combined with position information. In a simple additive-position view:

xi=ei+pix_i = e_i+p_i

For the first Transformer block, XX is the matrix of token representations:

X=[x1x2โ‹ฎxT]X = \begin{bmatrix} x_1\\ x_2\\ \vdots\\ x_T \end{bmatrix}

For one sequence, XX has shape (T,dmodel)(T,d_{\mathrm{model}}). At this stage, there is one vector per token; the sequence has not been consolidated into one sentence embedding. The stack proceeds as follows:

raw text โ†’ token embeddings + positions = Xโฝโฐโพ
โ†’ Transformer block 1 โ†’ Xโฝยนโพ
โ†’ Transformer block 2 โ†’ Xโฝยฒโพ
โ†’ ...

In later Transformer blocks, XX is no longer the original token embeddings; it is the contextual hidden-state matrix output by the previous block.

Each attention head then applies three separate learned linear projections:

Q=XWQ,K=XWK,V=XWVQ=XW_Q, \qquad K=XW_K, \qquad V=XW_V

Here, WQW_Q, WKW_K, and WVW_V are separate learned weight matrices. The same matrix XX is used in all three multiplications, so the input sequence is not duplicated. For token it, the same token vector xitx_{\textit{it}} produces three learned vectors: qit=xitWQq_{\textit{it}}=x_{\textit{it}}W_Q, kit=xitWKk_{\textit{it}}=x_{\textit{it}}W_K, and vit=xitWVv_{\textit{it}}=x_{\textit{it}}W_V.

< Shapes in self-attention and cross-attention >โ€‹

For one sequence and one attention head, omit the batch dimension and let XX have shape (T,dmodel)(T,d_{\mathrm{model}}). The projection and output shapes are:

dmodeld_{\mathrm{model}} is the main hidden width of the Transformer: the number of learned numeric components in each token representation. It is not the number of tokens, layers, or vocabulary items. This width stays the same across Transformer blocks so residual connections can add their inputs and outputs. In standard multi-head attention, it is commonly divided across hh heads:

  • Simple example: after tokenizing I like tea, suppose there are T=3T=3 tokens and dmodel=4d_{\mathrm{model}}=4. The input matrix has shape (3,4)(3,4), so every token has four learned numeric components:

    X=[0.2โˆ’0.10.70.4โˆ’0.30.50.1โˆ’0.20.80.0โˆ’0.40.6]โˆˆR3ร—4X = \begin{bmatrix} 0.2 & -0.1 & 0.7 & 0.4\\ -0.3 & 0.5 & 0.1 & -0.2\\ 0.8 & 0.0 & -0.4 & 0.6 \end{bmatrix} \in\mathbb{R}^{3\times4}

    The rows represent I, like, and tea. The numbers are invented for illustration; a real model learns them. Every later Transformer block still receives and returns three vectors with four components each.

dmodel=hโ‹…dhd_{\mathrm{model}} = h\cdot d_h

Here, dhd_h is the head dimension: the number of learned numeric components handled by one attention head. In the common equal-width design, dh=dk=dv=dmodel/hd_h=d_k=d_v=d_{\mathrm{model}}/h, so dmodeld_{\mathrm{model}} must be divisible by the number of heads hh. Heads have independent learned projections; they may learn different or overlapping attention patterns, but no linguistic role is assigned to a particular head in advance.

  • Eight-head example: if dmodel=16d_{\mathrm{model}}=16 and h=8h=8, then dh=2d_h=2. For each token, every head produces its own two-component query, key, and value vectors. For example, one head might form qit=[0.6,โˆ’0.2]q_{\textit{it}}=[0.6,-0.2]. The two numbers are learned latent coordinates, not two named aspects of language. Each head produces its own attention pattern over the complete sequence.
ObjectShapeRole
WQW_Q(dmodel,dk)(d_{\mathrm{model}},d_k)Projects a token representation into a query
WKW_K(dmodel,dk)(d_{\mathrm{model}},d_k)Projects a token representation into a key
WVW_V(dmodel,dv)(d_{\mathrm{model}},d_v)Projects a token representation into a value
QQ(T,dk)(T,d_k)One query vector per token
KK(T,dk)(T,d_k)One key vector per token
VV(T,dv)(T,d_v)One value vector per token

In self-attention, all three matrices have TT rows because they come from the same sequence of TT tokens. QQ and KK must have the same vector width dkd_k so their dot products work. VV may have a different width dvd_v.

  • Concrete self-attention example: suppose there are T=3T=3 tokens, dmodel=12d_{\mathrm{model}}=12, dk=7d_k=7, and dv=5d_v=5. Then XX is (3,12)(3,12), WQW_Q and WKW_K are (12,7)(12,7), WVW_V is (12,5)(12,5), QQ and KK are (3,7)(3,7), and VV is (3,5)(3,5):

    QKโŠค=(3,7)(7,3)=(3,3)QK^\top = (3,7)(7,3) = (3,3)

    The (3,3)(3,3) attention matrix gives one score for every query-token/key-token pair. After softmax, multiplying it by VV gives:

    AV=(3,3)(3,5)=(3,5)AV = (3,3)(3,5) = (3,5)

In cross-attention, the query sequence and the key/value sequence may have different numbers of tokens. For example, 3 decoder tokens can attend to 4 encoder tokens:

Q:(3,7),K:(4,7),V:(4,5)Q:(3,7), \qquad K:(4,7), \qquad V:(4,5)

Then QKโŠคQK^\top has shape (3,4)(3,4), and AVAV has shape (3,5)(3,5). Thus, KK and VV must have the same number of rowsโ€”each key needs a corresponding valueโ€”while QQ can have a different number of rows.

The WQW_Q, WKW_K, and WVW_V matrices are learned during pretraining with the rest of the model. A released checkpoint keeps them fixed during ordinary inference, while full fine-tuning can update them and adapter methods such as LoRA usually leave them frozen and learn small updates instead.

Lowercase qiq_i, kik_i, and viv_i denote one token's vectors. Uppercase QQ, KK, and VV stack the corresponding vectors for all tokens, which lets the model calculate attention for the complete sequence efficiently.

< Scaled dot-product attention >โ€‹

An attention score is the unnormalized compatibility score between query qiq_i and key kjk_j. In scaled dot-product attention:

sij=qikjโŠคdks_{ij} = \frac{q_i k_j^\top}{\sqrt{d_k}}

dkd_k is the query and key width for one attention head: the number of learned numeric components in each vector. It is a feature width, not a token count or a number of heads. In the common equal-width design, dk=dh=dmodel/hd_k=d_h=d_{\mathrm{model}}/h; for dmodel=512d_{\mathrm{model}}=512 and h=8h=8, dk=64d_k=64.

  • Why divide by dk\sqrt{d_k}: if the components have similar scale, a dot product tends to grow as its dimension grows. The division keeps scores from becoming so large that softmax becomes overly peaked and gradients become less useful.

  • Tiny numeric example: if q=[2,โˆ’1,0.5]q=[2,-1,0.5] and k=[1,3,4]k=[1,3,4], each vector has three components, so dk=3d_k=3. Their unscaled dot product compares corresponding components:

    qkโŠค=2(1)+(โˆ’1)(3)+0.5(4)=1qk^\top = 2(1)+(-1)(3)+0.5(4) = 1

    The scaled attention score is 1/3โ‰ˆ0.5771/\sqrt{3}\approx0.577.

A higher score means key jj is a stronger match for query ii. Softmax converts the scores across all keys into normalized attention weights, which form a weighted sum of the values:

Attentionโก(Q,K,V)=softmaxโก(QKโŠคdk)V\operatorname{Attention}(Q,K,V) = \operatorname{softmax}\left(\frac{QK^\top}{\sqrt{d_k}}\right)V

QKโŠคQK^\top has shape (T,T)(T,T): row ii contains how strongly token ii attends to every token. Some sources use โ€œattention scoreโ€ for the normalized weights too; here, it means the raw value before softmax.

  • Concrete output example: In โ€œThe animal rested because it was tired,โ€ suppose one head gives the query for it the following attention weights and two-dimensional value vectors. These are toy numbersโ€”the real vectors are learned and usually have many more dimensions.

    Key token jjAttention weight Ait,jA_{\textit{it},j}Value vector vjv_j
    The0.030.03[0.2,โ€…โ€Š0.6][0.2,\;0.6]
    animal0.620.62[0.8,โ€…โ€Š0.3][0.8,\;0.3]
    rested0.080.08[0.2,โ€…โ€Š0.9][0.2,\;0.9]
    because0.050.05[0.4,โ€…โ€Š1.0][0.4,\;1.0]
    it0.120.12[0.4,โ€…โ€Š0.5][0.4,\;0.5]
    was0.040.04[0.1,โ€…โ€Š0.6][0.1,\;0.6]
    tired0.060.06[0.5,โ€…โ€Š1.0][0.5,\;1.0]

    The weights sum to 11. The output for it is their weighted mixture of value vectors:

    oit=โˆ‘jAit,jvj=[0.62,โ€…โ€Š0.47]o_{\textit{it}} = \sum_j A_{\textit{it},j}v_j = [0.62,\;0.47]

    animal contributes the most because it has the largest weight. The output is not readable English; it is a new contextual representation of it that carries information from the tokens it attended to.

< Multi-head attention >โ€‹

Multi-head attention runs hh independent attention projections in parallel. Each head attends over the complete sequence using its own learned WQW_Q, WKW_K, and WVW_V, so heads can learn different or overlapping patterns. Their outputs are concatenated, then WOW_O mixes them back into the model width.

For example, if dmodel=512d_{\mathrm{model}}=512 and h=8h=8, each head commonly uses dh=dk=dv=64d_h=d_k=d_v=64.

Scaled dot-product attention on the left; multiple attention heads combined on the right. Click to expand.

Diagram source: transformer-model-in-PyTorch, based on the original Transformer attention diagrams.

For head rr, the model computes attention using its own projections, then concatenates the results from all hh heads:

headโกr=Attentionโก(Qr,Kr,Vr)\operatorname{head}_r = \operatorname{Attention}(Q_r,K_r,V_r) MultiHeadโก(Q,K,V)=Concatโก(headโก1,โ€ฆ,headโกh)WO\operatorname{MultiHead}(Q,K,V) = \operatorname{Concat}(\operatorname{head}_1,\ldots,\operatorname{head}_h)W_O
  • Concrete merge: suppose h=8h=8 and dh=2d_h=2. For token ii, each head returns a two-number attention output:

    headโกi,1=[a,b],headโกi,2=[c,d],โ€ฆ,headโกi,8=[o,p]\operatorname{head}_{i,1}=[a,b], \qquad \operatorname{head}_{i,2}=[c,d], \qquad \ldots, \qquad \operatorname{head}_{i,8}=[o,p]

    Concatenation places the outputs side by side; it does not sum or average them:

    ci=[a,b,c,d,โ€ฆ,o,p]โˆˆR16c_i = [a,b,c,d,\ldots,o,p] \in\mathbb{R}^{16}

    The output projection then mixes the 16 components into one updated representation for the same token:

    zi=ciWOโˆˆR16z_i = c_iW_O \in\mathbb{R}^{16}

    Implementations commonly calculate all heads together, then rearrange and reshape the tensor:

    (batch, heads, tokens, d_h)
    โ†’ (batch, tokens, heads, d_h)
    โ†’ (batch, tokens, heads ร— d_h)

    This merge happens independently at every token position, so the result is one contextual hidden-state vector per token, not one embedding for the entire sentence. A sentence-level embedding, when needed, is created later by a task-specific method such as pooling token vectors or using a special [CLS] token.

< Feed-forward network (FFN) >โ€‹

After attention mixes information between tokens, a position-wise feed-forward network (FFN) transforms each token representation independently. The same FFN weights are applied at every sequence position:

FFNโก(x)=W2โ€‰ฯ•(W1x+b1)+b2\operatorname{FFN}(x) = W_2\,\phi(W_1x+b_1)+b_2

The first linear layer usually expands the representation from dmodeld_{\mathrm{model}} to a larger hidden size dffd_{\mathrm{ff}}, an activation ฯ•\phi adds non-linearity, and the second linear layer projects it back to dmodeld_{\mathrm{model}}. The original Transformer used ReLU; modern models often use GELU or SwiGLU. Because the FFN does not mix positions, attention handles token-to-token communication and the FFN refines what each token has learned.

< Positional encoding >โ€‹

Attention alone is permutation-equivariant: without position information, it can tell which token representations are present but not their order. Positional encoding tells the model where each token occurs in the sequence.

In the original Transformer, a position vector pip_i is added to each token embedding eie_i before the first layer, as shown earlier:

xi=ei+pix_i=e_i+p_i
  • Example: dog bites man and man bites dog contain the same words but have different meanings. Position information lets the model distinguish their order.
  • Common choices: fixed sinusoidal encodings, learned position embeddings, and rotary position embeddings (RoPE). RoPE instead modifies the query and key vectors so their dot product can express relative position.

< Attention masks >โ€‹

An attention mask tells a token which positions it is allowed to use. Position encoding says where a token is; a mask says which tokens it may attend to.

An attention mask MM is added to the raw scaled dot-product scores before softmax:

Attentionโก(Q,K,V;M)=softmaxโก(QKโŠคdk+M)V\operatorname{Attention}(Q,K,V;M) = \operatorname{softmax}\left( \frac{QK^\top}{\sqrt{d_k}} + M \right)V

MijM_{ij} controls whether query token ii may attend to key token jj. Use 00 for an allowed position and โˆ’โˆž-\infty for a blocked position. After softmax, a blocked position has weight 00. In real implementations, โˆ’โˆž-\infty is often represented by a very large negative number for numerical safety.

MaskWhat it blocksExample
Causal / look-ahead maskFuture keys: j>ij>iIn โ€œThe cat slept,โ€ the query at cat may attend to The and cat, but not the future token slept
Padding maskPadding keys inserted to make a batch the same lengthIn a batch containing โ€œBirds flyโ€ and โ€œThe cat slept,โ€ the padding after Birds fly receives no attention weight

A causal mask lets a decoder predict the next token without seeing the answer ahead of time. Padding masks prevent meaningless <PAD> tokens from affecting real token representations.

< KV cache during generation >โ€‹

In an autoregressive decoder, a cache usually means a keyโ€“value (KV) cache. When the model generates token tt, its new query must attend to the keys and values of all earlier tokens. Those earlier key and value vectors will not change, so the model stores them rather than recomputing the full prefix at every step.

For each layer โ„“\ell, the cache after processing tokens 11 through tโˆ’1t-1 contains:

Kcache(โ„“)=[k1(โ„“)โ‹ฎktโˆ’1(โ„“)],Vcache(โ„“)=[v1(โ„“)โ‹ฎvtโˆ’1(โ„“)]K_{\mathrm{cache}}^{(\ell)} = \begin{bmatrix} k_1^{(\ell)}\\ \vdots\\ k_{t-1}^{(\ell)} \end{bmatrix}, \qquad V_{\mathrm{cache}}^{(\ell)} = \begin{bmatrix} v_1^{(\ell)}\\ \vdots\\ v_{t-1}^{(\ell)} \end{bmatrix}
  • Prefill: run the prompt through every layer once and fill the cache with its KK and VV vectors.
  • Decode: for the newly generated token, compute one new qtq_t, ktk_t, and vtv_t. Compare qtq_t with the cached keys, mix the cached values, then append ktk_t and vtv_t to the cache.
  • What it stores: numeric attention keys and values for every past token, head, and layerโ€”commonly shaped like (batch, heads, past_length, head_dim).
  • Why it stores keys and values, not queries: every later query compares with each earlier key and needs the paired value for its weighted sum. An earlier query is used only to form that earlier token's attention output, so later tokens do not need it.
  • What it does not normally store: Query vectors, model weights, raw text, or verified facts. The โ€œvalueโ€ is an attention VV vector, not a next-token probability or the answer text.

The KV cache makes generation much faster, but it grows with sequence length and consumes memory. In an encoderโ€“decoder model, cross-attention can also cache the encoder's keys and values because they stay fixed while the decoder generates.

KV-cache prefill and decode flow

Comparisonโ€‹

< Self-attention vs cross-attention >โ€‹

Attention typeQueries come fromKeys and values come fromTypical useExample
Self-attentionA sequenceThe same sequenceUnderstanding relationships within text, image patches, or another sequenceIn โ€œThe animal rested because it was tired,โ€ it can attend to animal
Cross-attentionOne sequenceA different sequenceA decoder attending to encoder outputs, or text attending to image featuresTranslate โ€œThe cat sleeps.โ€ to โ€œLe chat dort.โ€: while generating chat, the decoder can attend to source token cat

In the self-attention example, every QQ, KK, and VV vector comes from the English sentence. In the translation example, the decoder's query comes from its French-side representation, while the English encoder outputs provide the keys and values.

< Which models use an encoder or decoder? >โ€‹

Model familyEncoder stackDecoder stackHow it uses contextTypical use
BERTโœ…โŒBidirectional self-attention over the full inputClassification, search, token representations
GPT-styleโŒโœ…Causal self-attention over earlier tokens onlyText generation, chat, code completion
T5 / BARTโœ…โœ…Encoder reads the input; causal decoder attends to earlier output tokens and encoder outputTranslation, summarization, text-to-text tasks
Vision Transformer (ViT)โœ…โŒSelf-attention over image-patch tokensImage classification and vision tasks

The words encoder and decoder describe Transformer components, not whether a model can create embeddings. A decoder-only GPT still creates useful hidden representations; it simply does so with left-to-right context rather than full bidirectional context.

Implementationโ€‹

Favoritesโ€‹

๐Ÿค–

Video Tutorialโ€‹

Referenceโ€‹