π Multimodal Models
Descriptionβ
< What is it? >β
A modality is a kind of data, such as text, images, audio, or video. Multimodal models work across multiple modalities, for example by answering a question about an image or relating speech to text. The supported input and output types depend on the model; a multimodal model does not necessarily handle every modality. See Google's machine-learning glossary.
- Example: a voice assistant can turn spoken audio into text, generate a response, and turn that response into speech. This can be a pipeline of separate models, with each component responsible for one stage.
This section groups models and tasks that connect modalities: Vision Language Models (VLM), Speech-to-Text (STT), and Text-to-Speech (TTS). STT and TTS are specific audioβtext conversion tasks; they can be components of a broader multimodal system.
Key pointsβ
< Representing different modalities >β
Neural models convert raw inputs into numerical representations. Text may be tokenized and embedded, images may be divided into patches, and audio may be processed as waveforms or timeβfrequency features. Training teaches the model how those representations relate to the target task.
An embedding is a learned vector representation. Google's Machine Learning Crash Course explains how embeddings can be learned within a task model. Sharing a vector size alone does not establish meaningful relationships between modalities.
< A voice assistant pipeline >β
User's speech
|
v
Speech-to-Text (STT)
|
v
Transcribed text
|
v
Language model or application
|
v
Response text
|
v
Text-to-Speech (TTS)
|
v
Spoken response
STT transcribes the request. The language model or application interprets it and decides how to respond. TTS speaks the response. Hugging Face's voice assistant tutorial demonstrates this modular design.
- Example: βSet a timerβ becomes text through STT. The application creates the timer and supplies βYour timer is ready,β which TTS reads aloud.
< Model and system boundaries >β
A multimodal model is a learned component; a multimodal system includes the models, input processing, application logic, and output handling needed for the experience. A voice interface can use a text-only language model between separate speech models. Its audio input and output do not require every component to process audio.
Comparisonβ
< Speech-to-Text and Text-to-Speech >β
| System | Full name | Input β Output | Common uses |
|---|---|---|---|
| STT | Speech-to-Text | Spoken audio β Written text | Voice typing, captions, recording transcripts |
| TTS | Text-to-Speech | Written text β Spoken audio | Reading aloud, narration, spoken assistant responses |
STT is also called Automatic Speech Recognition (ASR). TTS is also called speech synthesis. They connect the same two modalities in opposite directions, using different prediction tasks.
Related ideasβ
- Vision Language Models (VLM) connects visual inputs and language.
- Speech-to-Text (STT) converts speech into a transcript.
- Text-to-Speech (TTS) generates spoken audio from text.
- Large Language Models (LLMs) can generate the text response in a voice assistant.
- Embeddings explains learned numerical representations.
- Transformer describes an architecture used in many text, image, and audio models.