Skip to main content

πŸ“ Multimodal Models

Description​

< What is it? >​

A modality is a kind of data, such as text, images, audio, or video. Multimodal models work across multiple modalities, for example by answering a question about an image or relating speech to text. The supported input and output types depend on the model; a multimodal model does not necessarily handle every modality. See Google's machine-learning glossary.

  • Example: a voice assistant can turn spoken audio into text, generate a response, and turn that response into speech. This can be a pipeline of separate models, with each component responsible for one stage.

This section groups models and tasks that connect modalities: Vision Language Models (VLM), Speech-to-Text (STT), and Text-to-Speech (TTS). STT and TTS are specific audio–text conversion tasks; they can be components of a broader multimodal system.

Key points​

< Representing different modalities >​

Neural models convert raw inputs into numerical representations. Text may be tokenized and embedded, images may be divided into patches, and audio may be processed as waveforms or time–frequency features. Training teaches the model how those representations relate to the target task.

An embedding is a learned vector representation. Google's Machine Learning Crash Course explains how embeddings can be learned within a task model. Sharing a vector size alone does not establish meaningful relationships between modalities.

< A voice assistant pipeline >​

User's speech
|
v
Speech-to-Text (STT)
|
v
Transcribed text
|
v
Language model or application
|
v
Response text
|
v
Text-to-Speech (TTS)
|
v
Spoken response

STT transcribes the request. The language model or application interprets it and decides how to respond. TTS speaks the response. Hugging Face's voice assistant tutorial demonstrates this modular design.

  • Example: β€œSet a timer” becomes text through STT. The application creates the timer and supplies β€œYour timer is ready,” which TTS reads aloud.

< Model and system boundaries >​

A multimodal model is a learned component; a multimodal system includes the models, input processing, application logic, and output handling needed for the experience. A voice interface can use a text-only language model between separate speech models. Its audio input and output do not require every component to process audio.

Comparison​

< Speech-to-Text and Text-to-Speech >​

SystemFull nameInput β†’ OutputCommon uses
STTSpeech-to-TextSpoken audio β†’ Written textVoice typing, captions, recording transcripts
TTSText-to-SpeechWritten text β†’ Spoken audioReading aloud, narration, spoken assistant responses

STT is also called Automatic Speech Recognition (ASR). TTS is also called speech synthesis. They connect the same two modalities in opposite directions, using different prediction tasks.

Reference​