what is voice cloning
Voice cloning recreates a specific voice from a sample. Here is the plain definition.
- Recreates a specific voice from an audio sample
- Then synthesizes any text in that voice
- Consent is required to create a voice
- Zero-shot from a short sample or fine-tuned from hours
definition
Voice cloning is technology that recreates a specific person’s voice from an audio sample, then synthesizes any text in that voice. Where stock text-to-speech picks a voice from a catalog, cloning reproduces a particular voice. Creating one requires the speaker’s consent, which is captured as a record with their full name and email.
how-it-works
You supply an audio sample and a consent record, and the system builds a voice you use by ID. There are two approaches. Zero-shot cloning works from a short sample, 10 to 30 seconds of clean speech, with good quality and self-serve. Fine-tuned cloning trains on hours of audio for the best quality. After creation, you synthesize speech in the voice like any catalog voice.
why-consent-matters
A voice is personal, so cloning it needs authorization. Responsible cloning makes consent a requirement, not an afterthought. On the Speechify Build API the consent record is mandatory on every create call, so a voice is only ever cloned with permission. See the ethics page.
what-its-used-for
Cloning powers signature narrators for audiobooks and podcasts, dubbing that keeps the original voice across languages, creator tools, game characters, and accessibility. See what voice cloning is used for and how to clone a voice.
Frequently asked questions
What is voice cloning?
How does voice cloning work?
Is voice cloning legal and safe?
Start building
Voice cloning recreates a specific voice from a sample. Here is the plain definition.