text to speech vs voice cloning
Both turn text into speech. The difference is whose voice does the talking.
- TTS uses catalog voices; cloning recreates a specific voice
- Cloning needs a voice sample and consent
- Both synthesize from text the same way
- Cloned voices work on select models
definition
Text-to-speech and voice cloning both turn text into spoken audio. The difference is whose voice speaks. TTS uses a voice from a catalog. Voice cloning first recreates a specific person’s voice from a sample, then synthesizes text in that voice. The synthesis step is identical; the input voice is what changes.
how-they-differ
With TTS, you pick a ready-made voice by ID and synthesize. With cloning, you first create a voice from an audio sample, with the speaker’s consent, and then reference that new voice by ID exactly like a catalog voice. Cloning adds a one-time creation step; after that, using it is the same.
when-to-use-each
Reach for a catalog voice when any suitable natural voice works: most narration, IVR prompts, notifications, and accessibility reading. Reach for cloning when the identity of the voice is the point: a branded narrator, an author reading their own audiobook, a recurring character, or a signature voice across a channel.
a-constraint-to-know
Cloned voices are supported on select synthesis models rather than the entire catalog. Plan for that when choosing a model. See the voice cloning page for how creation and consent work.
Frequently asked questions
What is the difference between text-to-speech and voice cloning?
When should I use voice cloning instead of TTS?
Do cloned voices work everywhere TTS does?
Start building
Both turn text into speech. The difference is whose voice does the talking.