SpeechifyAI Build

Every building block for speech. One API.

The developer API for voice AI: text to speech and voice cloning on the same speech stack that powers Speechify, exposed as building blocks you can compose.

500K characters free every month. Commercial use included, no card.

  • Text to speech
  • Voice cloning
  • <100ms first byte
  • One key, one bill

Synthesis request

Production REST endpoint

POST /v1/audio/speech

bash
curl -X POST https://api.speechify.ai/v1/audio/speech \
  -H "Authorization: Bearer $SPEECHIFY_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "input": "Ship production voice in a few lines of code.",
    "model": "simba-3.2",
    "voice_id": "sabrina",
    "audio_format": "mp3"
  }'
Model
simba-3.2
Voice
sabrina
Audio format
mp3
Explore speech

Hear it before you write a line.

Hear prepared samples of expressive speech, voice cloning, and multilingual synthesis. Then open the playground to try your own text.

Simba 3.2 · Emotion control

Same words.
Different feeling.

Choose a delivery. Hear how the same voice changes its tone and rhythm.

Neutral.

Beatrice · Delivery sample

The quick brown fox jumped over the lazy dog.

Ready to listen

Prepared Speechify samples. Choose a sample, then press play.Try your own text
Building blocks

The complete speech layer, kept composable.

Generate a voice or create one. Both capabilities share one API, one key, and one bill, so the first prototype and the production integration stay on the same path.

01Generate

Text to Speech

Generate lifelike speech from text, with streaming, SSML, and emotion on the Simba models.

  • Realtime streaming
  • SSML and emotion
  • Speech marks
02Create

Voice Cloning

Create a voice from a short sample, get a voice ID, and synthesize with it on the same endpoints.

  • Consent verified
  • One voice ID
  • Same synthesis API
Benchmark

Independently evaluated, ready to build on.

The speech block runs on Simba 3.2, a production text-to-speech model evaluated on the independent, blind-listening Artificial Analysis leaderboard. The quality evaluation is independent, not self-reported.

streaming text-to-speech
Native
per 1M additional characters on $10/month Starter
$10
self-serve languages
6
included on paid plans
Cloning
Model system

Choose for the experience, not the integration.

Both Simba models use the same endpoint. Select a model and a compatible voice without rebuilding the rest of your speech stack.

Recommended for English

Simba 3.2

Streaming-native speech for real-time playback.

Built for real-time playback, with finer-grained emotion control and SSML prosody, served from the full English voice roster with your cloned voices included. English only, and the recommendation for any English build.

Delivery
Streaming-native
Playback
Real-time
Languages
English
Voice cloning
Self-serve

Alternative

Simba 3.0

Multilingual streaming

Streaming-native across 6 languages and 7 locales. Set the language on simba-3.0 and one model covers the supported set, with zero-shot voice cloning.

  • 6 languages
  • 7 locales
  • Streaming-native
  • Voice cloning
Model docs
Start building on Simba

Start with Simba 3.2 for English. Choose Simba 3.0 for its supported multilingual locale set.

Speech system

The details that turn audio into a product.

Streaming, identity, expression, synchronization, and output formats sit behind the same synthesis request. Add only the controls your experience needs.

01

Delivery

Streaming for realtime

Chunked audio that starts playing before the sentence finishes generating. Up to 20,000 characters a request.

02

Identity

Voice cloning

Create a voice from a short sample, get a voice ID, and synthesize with it on Simba 3.2 or Simba 3.0, on the same endpoints. Consent required.

03

Direction

SSML control

Full SSML (prosody, break, emphasis, sub) for exact control over delivery and timing.

04

Expression

13 emotion styles

Shift a single voice across narration, dialogue, and emotional registers with speechify:style.

05

Synchronization

Word-level speech marks

Timestamps for every word, returned alongside the audio, for synced captions and read-along highlighting.

06

Output

Five audio formats

wav, mp3, ogg, aac, and pcm, plus u-law from the stream endpoint for telephony.

Quickstart

From API key to production in three steps.

The production path is the quickstart path. Start with one request, then add realtime delivery and synchronization when the product calls for them.

  1. 01

    Create one key

    Open a workspace and start with 500K characters each month. Keep the key on your server.

    SPEECHIFY_API_KEY

  2. 02

    Send one request

    Choose a model and voice, send text, and receive audio from the same endpoint you will use in production.

    POST /v1/audio/speech

  3. 03

    Move to production

    Switch to streaming where latency matters, add speech marks when the interface needs timing, and keep the same key and billing model.

    stream · observe · scale

Use cases

One speech layer, shaped around the experience.

Use the same models for a realtime reply, a long-form library, a recognizable brand voice, or localized audio. The surrounding product is yours to compose.

01

Conversational products

Stream a responsive voice into assistants, tutors, games, and interactive characters.

02

Narration at scale

Turn articles, lessons, books, and product content into consistent long-form audio.

03

A recognizable voice

Create a consent-verified voice identity and use it across every synthesized output.

04

Localized experiences

Ship spoken content across languages while keeping one integration and one account.

Pricing

One balance. The real unit on the bill.

Text to speech is billed per input character, with no credit conversion or token math. Your plan funds one balance across the speech capabilities you use.

per 1M additional chars on $499/month Scale
$6
per 1M additional chars on $99/month Pro
$8
per 1M additional chars on $10/month Starter
$10
free every month
500K

Measured in input characters

No separate model surcharge

Commercial use on every plan

Start building free

No credit card required.

Enterprise

Built for teams at scale.

SOC 2 Type II, SSO, a signed DPA, and volume pricing, plus a Forward Deployed Engineer who builds alongside heavy-traffic and regulated teams. Most ship a first use case in two to three weeks.

FAQ

Common questions.

What is SpeechifyAI Build?
The developer API for voice AI, organized as building blocks. Text to speech and voice cloning are available today: generate lifelike speech, stream long-form audio, and clone a voice from a short sample, all on one API, one key, and one bill.
What can I build with it today?
Anything that needs synthesized speech: narration, agents, accessibility, captions, IVR, dubbing. Text to speech and voice cloning are live now on one key and one bill.
Which model should I use?
Use Simba 3.2 for English or Simba 3.0 for 6 languages across 7 locales. Both are streaming-native and support voice cloning on paid plans. Check each voice's model and language compatibility. Broader language coverage is available on request.
How much does it cost?
Paid plans have a monthly subscription and an included USD usage balance. Additional text-to-speech usage is $10 per 1M characters on Starter, $8 on Pro, and $6 on Scale. The Free plan includes 500K characters a month with commercial use. Voice cloning requires a paid plan. See Pricing for the complete monthly costs.

Start building with voice.

Your first audio file is one request away. 500K characters free every month, commercial use included.

No credit card required.

Privacy preferences

Choose what we may store on this device. You can change this at any time from the footer.

Strictly necessary

Sign-in, security, load balancing, and remembering your privacy choices. These cannot be switched off.

Always on

Analytics

How the site is used in aggregate - which pages get read, where people get stuck - so we can improve it.

Marketing

Measures which campaigns bring people here, and lets us show relevant ads on other platforms.