π Speech-to-Text (STT)
Descriptionβ
< What is it? >β
Speech-to-Text (STT) converts spoken audio into written text. It is also called Automatic Speech Recognition (ASR). It is used for voice typing, captions, meeting transcripts, and the input stage of a voice assistant.
- Example: you say βSet a timer,β and the STT system outputs the text
Set a timer. A separate application interprets the request and creates the timer.
Spoken audio β Speech recognition model β Written transcript
Key pointsβ
< From audio to words >β
The system receives an audio waveform, prepares it in the format expected by the model, and predicts text tokens. Some models process waveform samples; others use features such as a spectrogram, which describes how frequency content changes over time. Decoding converts model predictions into the transcript.
The input sample rate must match the model's requirements. Changing sample-rate metadata alone does not resample a recording. Hugging Face's audio data introduction explains waveform and sampling concepts.
< Models and training labels >β
Supervised speech recognition uses audio paired with its transcript. The words are the target; many training methods do not require manually assigned word-level timestamps.
Two common approaches are:
- Connectionist Temporal Classification (CTC): predicts token distributions across audio time steps and trains by summing over valid alignments with the transcript.
- Sequence-to-sequence (Seq2Seq): encodes the audio and uses a decoder to generate transcript tokens, conditioned on the audio and earlier tokens.
For model examples and their differences, see Hugging Face's ASR model guide.
< Recorded audio and streaming audio >β
A system can transcribe a complete recording or process audio as it arrives. Streaming recognition can provide partial transcripts that change as more context becomes available. This is useful for live captions and voice interaction. Support for streaming depends on the model and service. Google's STT overview describes recognition request modes.
< Measuring transcription quality >β
Word Error Rate (WER) measures the edits needed to turn a prediction into a reference transcript:
Here, is substitutions, is deletions, is insertions, and is the number of words in the reference. Lower is better. WER can exceed 100% when there are many insertions. Evaluation also depends on consistent handling of punctuation, case, and numbers. See the ASR evaluation guide.
- Example: if a five-word reference requires one word substitution and no other edits, WER is .
Comparisonβ
< Transcription, translation, and speaker diarization >β
| Task | What it answers | Output |
|---|---|---|
| Speech transcription | What words were spoken? | Text in the spoken language |
| Speech translation | What was said in another language? | Translated text or speech, depending on the system |
| Speaker diarization | Which speaker spoke when? | Time segments grouped by speaker |
These tasks can be combined, but transcription alone does not identify speakers or translate languages. Diarization usually assigns speaker labels rather than identifying people by name.
Implementationβ
< Using STT in a voice application >β
- Capture or load audio and check its encoding, sample rate, and channels.
- Prepare the audio for a recognizer that supports the desired language and operating mode.
- Run recognition and collect the transcript, plus timestamps if available.
- Send the text to application logic or a language model.
- Evaluate with recordings representative of the actual microphones, noise, and speakers.
For a working application walkthrough, see Hugging Face's voice assistant tutorial.
Related ideasβ
- Multimodal Models places audioβtext conversion in a wider system.
- Text-to-Speech (TTS) converts the response text into spoken audio.
- Large Language Models (LLMs) can interpret transcripts and generate responses.
- Transformer explains a common foundation for speech encoders and decoders.