π Vision Language Models (VLM)
Descriptionβ
< What is it? >β
A Vision Language Model (VLM) learns relationships between visual information and language. Depending on its architecture and training, it can match images with descriptions, retrieve images using text, or generate answers about an image.
Two common designs are embedding models, such as CLIP, which compare image and text representations, and generative models, such as LLaVA, which generate text conditioned on images and a prompt. In conversational applications, βVLMβ often refers to the generative design.
- Example: given a photograph of a red bicycle and the question βWhat color is the bicycle?β, a generative VLM can produce βRed.β An embedding VLM can instead score how well the photograph matches candidate descriptions such as βa red bicycleβ and βa blue car.β
| Task | Input | Output |
|---|---|---|
| Imageβtext retrieval | A text query and candidate images | Similarity scores used to rank images |
| Image captioning | An image and a captioning prompt | A description of the image |
| Visual Question Answering (VQA) | An image and a question | A text answer |
| Document or chart understanding | A document image or chart and a question | Extracted information or an explanation |
The supported tasks depend on the model and its training data. For example, answering questions about photographs does not by itself establish reliable chart reading or document extraction.
Key pointsβ
< How a generative VLM processes an image >β
A common design connects a vision encoder to a language model through a learned connector:
Image β Vision encoder β Connector β Visual tokens ββ
ββ Language model β Text answer
Question β Tokenizer β Text embeddings ββββββββββββββ
- Vision encoder: converts the image into visual features. A Vision Transformer (ViT) typically processes image patches as a sequence.
- Connector: adapts the visual features for the language model. It may project each feature vector to a different width or use learned queries to select and combine information.
- Language model: generates answer tokens using the visual information and the text prompt.
The connection varies by architecture. LLaVA connects a vision encoder to a language model, while BLIP-2 uses a Querying Transformer (Q-Former) between pretrained vision and language components. See the LLaVA paper and BLIP-2 paper.
< Visual tokens and embedding dimensions >β
A visual token is a numerical feature vector carrying information about the image. It is not necessarily a recognized word or an object label. The number of visual tokens and the dimension of each token describe different parts of the representation.
- Example: suppose an encoder produces 196 visual tokens, each containing 768 numbers. A projector can transform each token into a 4,096-dimensional vector for a language model. For a batch of images, the shape changes from to . These are illustrative dimensions; actual shapes depend on the model and image processing.
An MLP can serve as this projector. Its output width determines the number of components in each projected vector. Training makes those components useful to the language model; matching the vector width alone does not align the modalities. Google's Machine Learning Crash Course: Obtaining embeddings explains how task training shapes learned representations.
< How a generative VLM is trained >β
Training often begins with pretrained vision and language components, followed by stages that teach them to work together. Which components are frozen or updated depends on the training recipe.
| Stage | Training data | Prediction target | Purpose |
|---|---|---|---|
| Visionβlanguage alignment | Images paired with descriptions | Description tokens in a generative training recipe | Teach the connector to supply useful visual information |
| Visual instruction tuning | Images, instructions or questions, and reference responses | Response tokens | Teach the model to follow instructions about visual inputs |
For answer generation, a typical objective is next-token cross-entropy: predict each reference answer token given the image, prompt, and preceding answer tokens. The labels are text tokens, and the visual representations receive learning signals through this objective. The LLaVA implementation describes feature alignment followed by visual instruction tuning.
- Example: a training example contains a bicycle image, the question βWhat color is the bicycle?β, and the reference answer βRed.β Training increases the probability of the reference answer tokens when the model receives that image and question.
< How an embedding VLM is trained >β
CLIP uses separate image and text encoders to produce vectors in a shared embedding space. Contrastive training increases the similarity of matched imageβtext pairs relative to mismatched pairs in the batch. The supervision identifies which image belongs with which text; it does not prescribe the coordinates of their embeddings. See the CLIP paper.
At inference, image and text vectors can be compared using cosine similarity. This supports retrieval and classification using text descriptions of candidate classes. CLIP itself does not generate a conversational answer.
Comparisonβ
< CLIP, LLaVA, and BLIP-2 >β
| Design | Connection between vision and language | Typical output | Representative use |
|---|---|---|---|
| CLIP | Separately encoded image and text vectors aligned through contrastive training | Embeddings and similarity scores | Retrieve images using a text query |
| LLaVA | A learned projector connects a vision encoder to a language model | Generated text | Answer an instruction about an image |
| BLIP-2 | A Q-Former bridges a frozen image encoder and a frozen language model during its proposed pretraining | Generated text with the language model | Caption an image or answer a visual question |
These describe representative architecture choices, rather than a quality ranking. See the original papers for CLIP, LLaVA, and BLIP-2.
< VLM, LLM, multimodal model, and foundation model >β
| Term | What it describes | Relationship to a VLM |
|---|---|---|
| Text-only Large Language Model (LLM) | A model that processes and generates text | Can provide the language component of a generative VLM |
| Multimodal model | A model that works with multiple kinds of data | VLMs connect the visual and language modalities |
| Foundation model | A broadly pretrained model that can be adapted to many downstream tasks | A broadly pretrained VLM can also be a foundation model |
These terms describe overlapping properties. βVision languageβ identifies the modalities; βfoundationβ describes broad pretraining and adaptability. See the foundation models report.
Implementationβ
< A small vision-to-language connector >β
LLaVA 1.5 supports a two-layer MLP connector with a GELU activation. The following PyTorch example demonstrates the shape transformation using smaller, illustrative dimensions. It creates an untrained connector, not a complete VLM. See the LLaVA implementation.
import torch
from torch import nn
# 2 images, 196 visual tokens per image, 64 numbers per token.
visual_features = torch.zeros(2, 196, 64)
connector = nn.Sequential(
nn.Linear(64, 128),
nn.GELU(),
nn.Linear(128, 128),
)
visual_tokens = connector(visual_features)
print(tuple(visual_tokens.shape))
Expected output:
(2, 196, 128)
The connector transforms the last dimension independently at each token position. Integrating it with a language model also requires the model's expected token placement, attention handling, and training.
< Asking a generative VLM about an image >β
The application supplies an image and a text prompt. The model's processor prepares the image and formats the conversation; the model then generates answer tokens. The following is pseudocode; exact APIs and input formats depend on the checkpoint.
model, processor = load_pretrained_vlm("chosen-checkpoint")
image = load_image("bicycle.jpg")
inputs = processor.prepare(
image=image,
prompt="What color is the bicycle?",
)
answer_tokens = model.generate(inputs, max_new_tokens=32)
answer = processor.decode_answer(answer_tokens)
Use the processor and conversation template supplied for the checkpoint. Hugging Face's image-text-to-text guide provides a concrete inference workflow.
Troubleshootβ
< Common problems >β
| Symptom | What to check |
|---|---|
| The answer is generic or ignores the image | Confirm the image reaches the processor and that the checkpoint's image markers and chat template are used |
| Small text or chart labels are misread | Inspect the processed image resolution and cropping; evaluate on representative documents or charts |
| The model mentions an object that is absent | Check the answer against the image; confident language does not establish visual evidence |
| Long image prompts exhaust memory | Check image resolution, number of images, visual token count, and generated sequence length against the model's supported limits |
Object hallucination means describing objects that are not present in the image. It is a documented failure mode of generative VLMs; the POPE paper studies its evaluation. Validate performance on the actual task, including incorrect answers, rather than judging only fluent examples.
Q & Aβ
< Does a VLM first convert the image into text? >β
A generative VLM can pass learned visual features directly to its language component. Optical Character Recognition (OCR) can be added to a document-processing pipeline, but an intermediate text transcript is not required by this architecture.
< Can every VLM generate images or understand video? >β
Supported outputs and inputs depend on the architecture and training. An image-to-text VLM generates text. Image generation requires an appropriate generation component, while video support requires handling multiple frames and learning to use their temporal information. Check the specific model's supported modalities.
Related ideasβ
- Multimodal Models covers models and systems that connect different data modalities.
- Vision Transformer (ViT) explains image patches and visual features.
- Large Language Models (LLMs) explains the language generation component.
- Embeddings covers learned vector representations.
- Multilayer Perceptron (MLP) explains a model used for visionβlanguage projection.
Referenceβ
- Google Machine Learning Crash Course: Obtaining embeddings
- CLIP: Learning Transferable Visual Models From Natural Language Supervision
- LLaVA: Visual Instruction Tuning
- LLaVA implementation
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- On the Opportunities and Risks of Foundation Models
- Evaluating Object Hallucination in Large Vision-Language Models
- Hugging Face Transformers: Image-text-to-text