Skip to main content

πŸ“ Vision Language Models (VLM)

Description​

< What is it? >​

A Vision Language Model (VLM) learns relationships between visual information and language. Depending on its architecture and training, it can match images with descriptions, retrieve images using text, or generate answers about an image.

Two common designs are embedding models, such as CLIP, which compare image and text representations, and generative models, such as LLaVA, which generate text conditioned on images and a prompt. In conversational applications, β€œVLM” often refers to the generative design.

  • Example: given a photograph of a red bicycle and the question β€œWhat color is the bicycle?”, a generative VLM can produce β€œRed.” An embedding VLM can instead score how well the photograph matches candidate descriptions such as β€œa red bicycle” and β€œa blue car.”
TaskInputOutput
Image–text retrievalA text query and candidate imagesSimilarity scores used to rank images
Image captioningAn image and a captioning promptA description of the image
Visual Question Answering (VQA)An image and a questionA text answer
Document or chart understandingA document image or chart and a questionExtracted information or an explanation

The supported tasks depend on the model and its training data. For example, answering questions about photographs does not by itself establish reliable chart reading or document extraction.

Key points​

< How a generative VLM processes an image >​

A common design connects a vision encoder to a language model through a learned connector:

Image β†’ Vision encoder β†’ Connector β†’ Visual tokens ─┐
β”œβ†’ Language model β†’ Text answer
Question β†’ Tokenizer β†’ Text embeddings β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
  1. Vision encoder: converts the image into visual features. A Vision Transformer (ViT) typically processes image patches as a sequence.
  2. Connector: adapts the visual features for the language model. It may project each feature vector to a different width or use learned queries to select and combine information.
  3. Language model: generates answer tokens using the visual information and the text prompt.

The connection varies by architecture. LLaVA connects a vision encoder to a language model, while BLIP-2 uses a Querying Transformer (Q-Former) between pretrained vision and language components. See the LLaVA paper and BLIP-2 paper.

< Visual tokens and embedding dimensions >​

A visual token is a numerical feature vector carrying information about the image. It is not necessarily a recognized word or an object label. The number of visual tokens and the dimension of each token describe different parts of the representation.

  • Example: suppose an encoder produces 196 visual tokens, each containing 768 numbers. A projector can transform each token into a 4,096-dimensional vector for a language model. For a batch of BB images, the shape changes from (B,196,768)(B, 196, 768) to (B,196,4096)(B, 196, 4096). These are illustrative dimensions; actual shapes depend on the model and image processing.

An MLP can serve as this projector. Its output width determines the number of components in each projected vector. Training makes those components useful to the language model; matching the vector width alone does not align the modalities. Google's Machine Learning Crash Course: Obtaining embeddings explains how task training shapes learned representations.

< How a generative VLM is trained >​

Training often begins with pretrained vision and language components, followed by stages that teach them to work together. Which components are frozen or updated depends on the training recipe.

StageTraining dataPrediction targetPurpose
Vision–language alignmentImages paired with descriptionsDescription tokens in a generative training recipeTeach the connector to supply useful visual information
Visual instruction tuningImages, instructions or questions, and reference responsesResponse tokensTeach the model to follow instructions about visual inputs

For answer generation, a typical objective is next-token cross-entropy: predict each reference answer token given the image, prompt, and preceding answer tokens. The labels are text tokens, and the visual representations receive learning signals through this objective. The LLaVA implementation describes feature alignment followed by visual instruction tuning.

  • Example: a training example contains a bicycle image, the question β€œWhat color is the bicycle?”, and the reference answer β€œRed.” Training increases the probability of the reference answer tokens when the model receives that image and question.

< How an embedding VLM is trained >​

CLIP uses separate image and text encoders to produce vectors in a shared embedding space. Contrastive training increases the similarity of matched image–text pairs relative to mismatched pairs in the batch. The supervision identifies which image belongs with which text; it does not prescribe the coordinates of their embeddings. See the CLIP paper.

At inference, image and text vectors can be compared using cosine similarity. This supports retrieval and classification using text descriptions of candidate classes. CLIP itself does not generate a conversational answer.

Comparison​

< CLIP, LLaVA, and BLIP-2 >​

DesignConnection between vision and languageTypical outputRepresentative use
CLIPSeparately encoded image and text vectors aligned through contrastive trainingEmbeddings and similarity scoresRetrieve images using a text query
LLaVAA learned projector connects a vision encoder to a language modelGenerated textAnswer an instruction about an image
BLIP-2A Q-Former bridges a frozen image encoder and a frozen language model during its proposed pretrainingGenerated text with the language modelCaption an image or answer a visual question

These describe representative architecture choices, rather than a quality ranking. See the original papers for CLIP, LLaVA, and BLIP-2.

< VLM, LLM, multimodal model, and foundation model >​

TermWhat it describesRelationship to a VLM
Text-only Large Language Model (LLM)A model that processes and generates textCan provide the language component of a generative VLM
Multimodal modelA model that works with multiple kinds of dataVLMs connect the visual and language modalities
Foundation modelA broadly pretrained model that can be adapted to many downstream tasksA broadly pretrained VLM can also be a foundation model

These terms describe overlapping properties. β€œVision language” identifies the modalities; β€œfoundation” describes broad pretraining and adaptability. See the foundation models report.

Implementation​

< A small vision-to-language connector >​

LLaVA 1.5 supports a two-layer MLP connector with a GELU activation. The following PyTorch example demonstrates the shape transformation using smaller, illustrative dimensions. It creates an untrained connector, not a complete VLM. See the LLaVA implementation.

import torch
from torch import nn

# 2 images, 196 visual tokens per image, 64 numbers per token.
visual_features = torch.zeros(2, 196, 64)

connector = nn.Sequential(
nn.Linear(64, 128),
nn.GELU(),
nn.Linear(128, 128),
)

visual_tokens = connector(visual_features)
print(tuple(visual_tokens.shape))

Expected output:

(2, 196, 128)

The connector transforms the last dimension independently at each token position. Integrating it with a language model also requires the model's expected token placement, attention handling, and training.

< Asking a generative VLM about an image >​

The application supplies an image and a text prompt. The model's processor prepares the image and formats the conversation; the model then generates answer tokens. The following is pseudocode; exact APIs and input formats depend on the checkpoint.

model, processor = load_pretrained_vlm("chosen-checkpoint")
image = load_image("bicycle.jpg")

inputs = processor.prepare(
image=image,
prompt="What color is the bicycle?",
)
answer_tokens = model.generate(inputs, max_new_tokens=32)
answer = processor.decode_answer(answer_tokens)

Use the processor and conversation template supplied for the checkpoint. Hugging Face's image-text-to-text guide provides a concrete inference workflow.

Troubleshoot​

< Common problems >​

SymptomWhat to check
The answer is generic or ignores the imageConfirm the image reaches the processor and that the checkpoint's image markers and chat template are used
Small text or chart labels are misreadInspect the processed image resolution and cropping; evaluate on representative documents or charts
The model mentions an object that is absentCheck the answer against the image; confident language does not establish visual evidence
Long image prompts exhaust memoryCheck image resolution, number of images, visual token count, and generated sequence length against the model's supported limits

Object hallucination means describing objects that are not present in the image. It is a documented failure mode of generative VLMs; the POPE paper studies its evaluation. Validate performance on the actual task, including incorrect answers, rather than judging only fluent examples.

Q & A​

< Does a VLM first convert the image into text? >​

A generative VLM can pass learned visual features directly to its language component. Optical Character Recognition (OCR) can be added to a document-processing pipeline, but an intermediate text transcript is not required by this architecture.

< Can every VLM generate images or understand video? >​

Supported outputs and inputs depend on the architecture and training. An image-to-text VLM generates text. Image generation requires an appropriate generation component, while video support requires handling multiple frames and learning to use their temporal information. Check the specific model's supported modalities.

Reference​