📝 Vision Transformer (ViT)
Description
< What is it? >
A Vision Transformer (ViT) applies a Transformer encoder to images. It divides an image into fixed-size patches, turns each patch into a token embedding, adds position information, and lets self-attention combine information from across the image.
image → patches → patch embeddings + positions → Transformer encoder blocks → task head
Unlike a CNN, which starts with small local filters, a standard ViT can let every patch attend to every other patch from its first attention layer.
Key points
< What is a patch? >
A patch is a small local region of an input (usually an image). The idea: instead of looking at the whole image at once, you look at small chunks. A standard ViT splits an image into a grid of fixed-size, usually non-overlapping patches; for example, a 224 × 224 image split into 16 × 16 patches produces a 14 × 14 grid of 196 patches.
Each patch contains its pixel values from every input channel and becomes one token for the Transformer. A patch is input data, whereas a CNN kernel is a learned set of weights that slides over input data. See CNN patch vs ViT patch for the full comparison.
< Patch embeddings >
For an input image of shape (N, C, H, W), a patch size of P × P creates a sequence of
patch tokens, assuming H and W are divisible by P. Each raw patch contains P × P × C values and is linearly projected into an embedding of dimension D.
< Position embeddings >
Self-attention does not inherently know whether a token came from the top-left or bottom-right of an image. Position embeddings tell the model where each patch belongs, so spatial arrangement is preserved.
< Self-attention >
Within each Transformer block, a patch can attend to other patches and gather relevant information. This gives a standard ViT a global receptive field from the first block, which can help with relationships between distant image regions.
The tradeoff is cost: full self-attention grows roughly quadratically with the number of patch tokens. Small patches or high-resolution images therefore make attention more expensive.
< Outputs >
For image classification, the original ViT prepends a learned classification token and uses its final representation for the prediction. For tasks such as detection or segmentation, models can use the patch-token representations directly, often with task-specific or hierarchical components.
< Strengths and limitations >
- Strengths: a simple, scalable architecture; global interactions between patches; strong results with suitable pretraining and transfer learning
- Limitations: fewer built-in locality and translation biases than CNNs; full attention can be costly; training from scratch often needs more data, augmentation, or pretraining
< A patch-count example >
An RGB image of size 224 × 224 with 16 × 16 patches becomes 14 × 14 = 196 patch tokens. Each raw patch has 16 × 16 × 3 = 768 values. A classification ViT typically adds one class token, so the encoder receives 197 tokens.
< A Transformer encoder block >
tokens → LayerNorm → multi-head self-attention → residual connection
→ LayerNorm → MLP → residual connection
Stacking these blocks lets each patch representation become increasingly contextual.
Comparison
< CNN compared with ViT >
CNNs use shared local kernels and build global context over layers. ViTs use patch tokens and self-attention to exchange global information directly. Neither is always better: compare pretrained models on the target data, accuracy, latency, memory, and deployment hardware.
< Why ViTs often need more data >
With a small labeled dataset, a ViT trained from scratch often needs more data to match a comparable CNN. CNNs build in useful image assumptions: local connectivity, shared kernels, and translation equivariance. Those biases help a CNN learn visual patterns such as edges and textures from fewer examples.
A standard ViT has weaker built-in image-specific assumptions, so it must learn more of the local visual structure and spatial relationships from data. Large-scale pretraining, strong augmentation, knowledge distillation, and hybrid or hierarchical designs can greatly reduce this gap, so it is not an absolute limitation.
Related ideas
- Transformer introduces self-attention for sequences.
- CNN explains kernels, feature maps, and spatial downsampling.
- Activation Functions explains the non-linear components used in Transformer MLP blocks.