multilingual text to speech

Speak to a global audience with natural voices in 30+ languages.

  • 30+ languages supported
  • Natural multilingual voices
  • Localize by re-synthesizing text
  • One API for every language

definition

Multilingual text-to-speech synthesizes natural speech in 30+ languages from one API. It lets a product reach a global audience without integrating a separate voice service per language: you translate the text, pick a voice, and synthesize.

localization-as-synthesis

Because everything is text-driven, localizing content is a synthesis job. An audiobook, a course library, a set of IVR prompts, or in-app notifications can each be produced in another language by re-synthesizing the translated text. Nothing needs re-recording.

one-integration

The same endpoint serves every supported language, so you integrate once and select a language per request. That keeps a multi-market product simple instead of stitching together per-language vendors.

natural-voices

The multilingual voices are built to sound natural in their language, not like an accent laid over English. See the language support guide for the full list of languages and voices.

FAQ

Frequently asked questions

How many languages does the text-to-speech support?
The API synthesizes speech in 30+ languages with natural multilingual voices. You localize content by re-synthesizing the translated text through the same API, so adding a language is a synthesis job rather than a separate integration.
How do I localize existing content?
Translate the text, then re-synthesize it with a voice for the target language. Because the content is text-driven, an audiobook, course, or set of prompts can be localized without re-recording anything.
Is it one API for all languages?
Yes. The same endpoint handles every supported language, so you do not integrate a separate service per language. You select a voice and synthesize.

Start building

Speak to a global audience with natural voices in 30+ languages.