Skip to main content

πŸ“ Speech-to-Text (STT)

Description​

< What is it? >​

Speech-to-Text (STT) converts spoken audio into written text. It is also called Automatic Speech Recognition (ASR). It is used for voice typing, captions, meeting transcripts, and the input stage of a voice assistant.

  • Example: you say β€œSet a timer,” and the STT system outputs the text Set a timer. A separate application interprets the request and creates the timer.
Spoken audio β†’ Speech recognition model β†’ Written transcript

Key points​

< From audio to words >​

The system receives an audio waveform, prepares it in the format expected by the model, and predicts text tokens. Some models process waveform samples; others use features such as a spectrogram, which describes how frequency content changes over time. Decoding converts model predictions into the transcript.

The input sample rate must match the model's requirements. Changing sample-rate metadata alone does not resample a recording. Hugging Face's audio data introduction explains waveform and sampling concepts.

< Models and training labels >​

Supervised speech recognition uses audio paired with its transcript. The words are the target; many training methods do not require manually assigned word-level timestamps.

Two common approaches are:

  • Connectionist Temporal Classification (CTC): predicts token distributions across audio time steps and trains by summing over valid alignments with the transcript.
  • Sequence-to-sequence (Seq2Seq): encodes the audio and uses a decoder to generate transcript tokens, conditioned on the audio and earlier tokens.

For model examples and their differences, see Hugging Face's ASR model guide.

< Recorded audio and streaming audio >​

A system can transcribe a complete recording or process audio as it arrives. Streaming recognition can provide partial transcripts that change as more context becomes available. This is useful for live captions and voice interaction. Support for streaming depends on the model and service. Google's STT overview describes recognition request modes.

< Measuring transcription quality >​

Word Error Rate (WER) measures the edits needed to turn a prediction into a reference transcript:

WER=S+D+IN\mathrm{WER}=\frac{S+D+I}{N}

Here, SS is substitutions, DD is deletions, II is insertions, and NN is the number of words in the reference. Lower is better. WER can exceed 100% when there are many insertions. Evaluation also depends on consistent handling of punctuation, case, and numbers. See the ASR evaluation guide.

  • Example: if a five-word reference requires one word substitution and no other edits, WER is 1/5=20%1/5=20\%.

Comparison​

< Transcription, translation, and speaker diarization >​

TaskWhat it answersOutput
Speech transcriptionWhat words were spoken?Text in the spoken language
Speech translationWhat was said in another language?Translated text or speech, depending on the system
Speaker diarizationWhich speaker spoke when?Time segments grouped by speaker

These tasks can be combined, but transcription alone does not identify speakers or translate languages. Diarization usually assigns speaker labels rather than identifying people by name.

Implementation​

< Using STT in a voice application >​

  1. Capture or load audio and check its encoding, sample rate, and channels.
  2. Prepare the audio for a recognizer that supports the desired language and operating mode.
  3. Run recognition and collect the transcript, plus timestamps if available.
  4. Send the text to application logic or a language model.
  5. Evaluate with recordings representative of the actual microphones, noise, and speakers.

For a working application walkthrough, see Hugging Face's voice assistant tutorial.

Reference​