Skip to main content

πŸ“ Text-to-Speech (TTS)

Description​

< What is it? >​

Text-to-Speech (TTS) generates spoken audio from written text. It is also called speech synthesis. It is used for reading text aloud, audiobook narration, accessibility, and voice-assistant responses.

  • Example: the application supplies Your timer is ready, and TTS produces an audio recording speaking those words.
Written text β†’ Speech synthesis model β†’ Spoken audio

TTS supplies the voice for an existing text response. The application or language model decides what that response says.

Key points​

< From text to a waveform >​

A TTS system prepares the text, predicts a speech representation, and generates playable audio. Text preparation may resolve numbers and abbreviations and convert words into tokens or phonemes, which represent speech sounds.

One common neural design predicts a spectrogram and then uses a vocoder to convert it into a waveform:

Text β†’ Text processing β†’ Acoustic model β†’ Spectrogram β†’ Vocoder β†’ Audio waveform

This is one architecture rather than a requirement for every TTS model. Hugging Face's audio architecture guide illustrates spectrogram output and vocoders.

< Training inputs and targets >​

A common supervised setup pairs text with a recording of someone speaking it. The text conditions generation, while the target comes from the recording, such as acoustic features or encoded audio tokens, depending on the model. Speaker information may also condition a model trained on multiple voices.

The audio target contains more than the words: pronunciation, timing, pitch, and pauses all affect the spoken result. A sentence can have several valid spoken realizations. See the audio sequence-to-sequence guide.

< Voice, pronunciation, and prosody >​

Prosody includes rhythm, stress, intonation, and pauses. These features help speech sound natural and convey the intended phrasing. Some TTS systems offer controls for voice, speaking rate, pronunciation, or pauses.

Speech Synthesis Markup Language (SSML) is supported by some services for specifying aspects of spoken output. Available controls depend on the service and voice. Google's TTS basics explains text and SSML inputs.

  • Example: the same text, β€œReally?”, can sound like a neutral question or an expression of surprise depending on pitch and emphasis.

< Evaluating generated speech >​

Evaluate whether the speech is intelligible, pronounces the intended words correctly, and has natural pacing. For interactive use, also measure the delay before playback starts and whether generation keeps up with playback. A low training loss alone does not establish that listeners will find a voice natural.

Implementation​

< Using TTS in a voice application >​

  1. Obtain the response text from the application or language model.
  2. Choose a supported language and voice, and prepare pronunciation or pause controls if needed.
  3. Run synthesis to generate audio in a supported format.
  4. Play the audio using its actual sample rate and encoding.
  5. Check names, numbers, abbreviations, and longer passages with listening tests.

Hugging Face's TTS pipeline tutorial demonstrates generating audio from text.

Reference​