fine-tuned voice cloning
Train on hours of audio for the highest-fidelity cloned voice.
- Trained on hours of audio
- Best quality tier
- Arranged with sales
- Consent required
definition
Fine-tuned voice cloning trains on hours of a speaker’s audio to reach the best quality tier. Where zero-shot clones directly from a short sample, fine-tuning invests a larger dataset and a training run for higher fidelity. It is arranged with sales and, like all cloning, requires consent.
why-more-audio
Hours of audio give the model the speaker’s full range: their intonation across many sentences, their cadence, their handling of emphasis and pauses. Capturing that range is what lifts fidelity from good to best, and it needs more than a short sample can provide.
how-its-set-up
Because it involves collecting hours of material and a fine-tune, this tier is arranged with sales rather than a self-serve call. Once built, the voice is used by ID on the normal speech endpoints like any other. See professional cloning.
consent
A consent record is required, captured as part of the engagement. A fine-tuned clone is built on a speaker’s identity, so authorization is mandatory. Compare with zero-shot cloning.
Frequently asked questions
What is fine-tuned voice cloning?
When should I choose fine-tuned over zero-shot?
How do I get fine-tuned cloning?
Start building
Train on hours of audio for the highest-fidelity cloned voice.