what is voice cloning

Voice cloning recreates a specific voice from a sample. Here is the plain definition.

  • Recreates a specific voice from an audio sample
  • Then synthesizes any text in that voice
  • Consent is required to create a voice
  • Zero-shot from a short sample or fine-tuned from hours

definition

Voice cloning is technology that recreates a specific person’s voice from an audio sample, then synthesizes any text in that voice. Where stock text-to-speech picks a voice from a catalog, cloning reproduces a particular voice. Creating one requires the speaker’s consent, which is captured as a record with their full name and email.

how-it-works

You supply an audio sample and a consent record, and the system builds a voice you use by ID. There are two approaches. Zero-shot cloning works from a short sample, 10 to 30 seconds of clean speech, with good quality and self-serve. Fine-tuned cloning trains on hours of audio for the best quality. After creation, you synthesize speech in the voice like any catalog voice.

A voice is personal, so cloning it needs authorization. Responsible cloning makes consent a requirement, not an afterthought. On the Speechify Build API the consent record is mandatory on every create call, so a voice is only ever cloned with permission. See the ethics page.

what-its-used-for

Cloning powers signature narrators for audiobooks and podcasts, dubbing that keeps the original voice across languages, creator tools, game characters, and accessibility. See what voice cloning is used for and how to clone a voice.

FAQ

Frequently asked questions

What is voice cloning?
Voice cloning is technology that recreates a specific person's voice from an audio sample, then synthesizes any text in that voice. Unlike stock text-to-speech, which uses catalog voices, cloning reproduces a particular voice. Creating one requires the speaker's consent, captured as a record with their full name and email.
How does voice cloning work?
You provide an audio sample and a consent record, and the system builds a voice you can synthesize with by ID. Zero-shot cloning works from a short 10-to-30-second sample with good quality; fine-tuned cloning trains on hours of audio for the best quality. After that, you generate speech in the voice like any other.
Is voice cloning legal and safe?
Responsible voice cloning requires consent. On the Speechify Build API, creating a voice requires a consent record with the speaker's full name and email, and there is no path that skips it, so voices are created only with authorization.

Start building

Voice cloning recreates a specific voice from a sample. Here is the plain definition.