Skip to main content

📝 Vision Transformer (ViT)

Description​

< What is it? >​

A Vision Transformer (ViT) applies a Transformer encoder to images. It divides an image into fixed-size patches, turns each patch into a token embedding, adds position information, and lets self-attention combine information from across the image.

image → patches → patch embeddings + positions → Transformer encoder blocks → task head

Unlike a CNN, which starts with small local filters, a standard ViT can let every patch attend to every other patch from its first attention layer.

Key points​

< What is a patch? >​

A patch is a small local region of an input (usually an image). The idea: instead of looking at the whole image at once, you look at small chunks. A standard ViT splits an image into a grid of fixed-size, usually non-overlapping patches; for example, a 224 × 224 image split into 16 × 16 patches produces a 14 × 14 grid of 196 patches.

Each patch contains its pixel values from every input channel and becomes one token for the Transformer. A patch is input data, whereas a CNN kernel is a learned set of weights that slides over input data. See CNN patch vs ViT patch for the full comparison.

< Patch embeddings >​

For an input image of shape (N, C, H, W), a patch size of P × P creates a sequence of

L=HP⋅WPL = \frac{H}{P} \cdot \frac{W}{P}

patch tokens, assuming H and W are divisible by P. Each raw patch contains P × P × C values and is linearly projected into an embedding of dimension D.

< Position embeddings >​

Self-attention does not inherently know whether a token came from the top-left or bottom-right of an image. Position embeddings tell the model where each patch belongs, so spatial arrangement is preserved.

< Self-attention >​

Within each Transformer block, a patch can attend to other patches and gather relevant information. This gives a standard ViT a global receptive field from the first block, which can help with relationships between distant image regions.

The tradeoff is cost: full self-attention grows roughly quadratically with the number of patch tokens. Small patches or high-resolution images therefore make attention more expensive.

< Outputs >​

For image classification, the original ViT prepends a learned classification token and uses its final representation for the prediction. For tasks such as detection or segmentation, models can use the patch-token representations directly, often with task-specific or hierarchical components.

< Strengths and limitations >​

  • Strengths: a simple, scalable architecture; global interactions between patches; strong results with suitable pretraining and transfer learning
  • Limitations: fewer built-in locality and translation biases than CNNs; full attention can be costly; training from scratch often needs more data, augmentation, or pretraining

< A patch-count example >​

An RGB image of size 224 × 224 with 16 × 16 patches becomes 14 × 14 = 196 patch tokens. Each raw patch has 16 × 16 × 3 = 768 values. A classification ViT typically adds one class token, so the encoder receives 197 tokens.

< A Transformer encoder block >​

tokens → LayerNorm → multi-head self-attention → residual connection
→ LayerNorm → MLP → residual connection

Stacking these blocks lets each patch representation become increasingly contextual.

Comparison​

< CNN compared with ViT >​

CNNs use shared local kernels and build global context over layers. ViTs use patch tokens and self-attention to exchange global information directly. Neither is always better: compare pretrained models on the target data, accuracy, latency, memory, and deployment hardware.

< Why ViTs often need more data >​

With a small labeled dataset, a ViT trained from scratch often needs more data to match a comparable CNN. CNNs build in useful image assumptions: local connectivity, shared kernels, and translation equivariance. Those biases help a CNN learn visual patterns such as edges and textures from fewer examples.

A standard ViT has weaker built-in image-specific assumptions, so it must learn more of the local visual structure and spatial relationships from data. Large-scale pretraining, strong augmentation, knowledge distillation, and hybrid or hierarchical designs can greatly reduce this gap, so it is not an absolute limitation.

  • Transformer introduces self-attention for sequences.
  • CNN explains kernels, feature maps, and spatial downsampling.
  • Activation Functions explains the non-linear components used in Transformer MLP blocks.

Video Tutorial​

Reference​