Artificial Analysis TTS leaderboard: Elo explained, Simba 3.2 ranked

On the Artificial Analysis Speech Arena on 23 September 2026, Simba 3.2 was in the top five of 92 models at an Elo of 1,237, and the lowest-priced model in the top ten. How the leaderboard's Elo works, and what two independent latency benchmarks measured.

Updated 5 min read
The Simba 3.2 sculpture, a satin silver wave crest, beside the headline Top five of 92, the lowest price in the top ten, for the Artificial Analysis Speech Arena provider-voice board read on 23 September 2026, with an Elo of 1,237 (plus or minus 14) and a rank band of 4 to 7.
The Simba 3.2 sculpture, a satin silver wave crest, beside the headline Top five of 92, the lowest price in the top ten, for the Artificial Analysis Speech Arena provider-voice board read on 23 September 2026, with an Elo of 1,237 (plus or minus 14) and a rank band of 4 to 7.

On 23 September 2026, Simba 3.2 was in the top five of the 92 models on the Artificial Analysis Speech Arena, at an Elo of 1,237 (±14) with a rank band of fourth to seventh, and it was the lowest-priced model in the top ten (our reviewed snapshot of the full board). Nothing on the board was both cheaper and rated higher. On 24 September 2026, two independent benchmarks timed it from their own runners to the first audible sound: Coval measured a median of 106 ms over the 24 hours to that day, and Voice Arena read 123 ms p50 on its US English board, the lowest of the 14 models it gives a latency reading (Coval snapshot, Voice Arena snapshot).

This post is about the leaderboard itself: how the Elo is scored and what the standing means. For the whole top ten with prices, read our comparison of the best TTS APIs in 2026. For what changed in Simba 3.2 and how to migrate, read the Simba 3.2 announcement. To build with the model, start from the Text-to-Speech API.

Recorded in July 2026. The board has moved since; the figures on this page are its 23 September reading.

How the Artificial Analysis Elo works

Artificial Analysis runs a blind speech arena. Listeners hear two clips of the same text from two models, without knowing which model made which, and pick the one they prefer. Every vote moves both models’ ratings, the same Elo method chess uses and the Chatbot Arena made standard for language models. It is not our benchmark and not our numbers.

A gap between two ratings is a head-to-head preference. Cartesia’s Sonic 3.6 held the top score on 23 September 2026 at 1,273, 36 points above Simba 3.2, which means blind listeners pick it about 55 times in 100, at about seven times the price on Artificial Analysis’s normalized figures (23 Sep snapshot).

Read the bands before you read the gaps. Artificial Analysis publishes a confidence interval and a rank band beside every score, and on 23 September 2026 Simba 3.2’s band of fourth to seventh overlapped five models: Google’s Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, Alibaba’s Qwen-Audio-3.0-TTS-Plus, Inworld’s Realtime TTS-2 and VUI Labs’ Luna TTS (23 Sep snapshot). Those five cost 1.7 to 4.2 times as much on the same normalized prices, and the board cannot place any of them above Simba 3.2 with confidence.

Artificial Analysis’s normalized price for Simba 3.2, $6.60 per million characters, is its own estimate and assumes 80% use of a plan (23 Sep snapshot). Our list rates are $10 per 1M characters on Starter, $8 on Pro and $6 on Scale, on the pricing page.

Voice Arena, a second blind board

Voice Arena ranks text to speech the same way, with native-speaker panels choosing between two clips of the same text, blind. It runs six languages, a balanced voice slate per model rather than whichever default sounds best, and sentences written for the places TTS actually ships, with a methodology built with advice from Prof. Shinji Watanabe at Carnegie Mellon. It also times models to their first audio. On its US English board on 24 September 2026, Simba 3.2’s quality rank was 7 of 20 (range 6 to 8), and its median time to first audio, 123 ms, was the lowest of the 14 models with a latency reading (Voice Arena snapshot).

How fast it starts talking

Three numbers, measured three ways.

  • Our own first byte. On our production US East streaming path, Simba 3.2 returned its first audio bytes in 56 ms at the median and 102 ms at the 90th percentile, read on 15 September 2026. That is time from our servers, not time to audible audio from your region.
  • Coval. Its open benchmark measured a median time to first audio of 106 ms over the 24 hours to 24 September 2026, from its own runner, counting the network and any silence before the first audible sample (Coval snapshot).
  • Voice Arena. It read 123 ms p50 on its US English board the same day, from its own runner (Voice Arena snapshot).

How both benchmarks measured it is in our write-up of the two readings, and how to measure it yourself is in our guide to choosing a low-latency TTS API.

Voice cloning on Simba 3.2

Clone a voice you have permission to clone from a 10-30 second sample. The speaker reads a one-time phrase aloud as the consent record, and the clone exists as soon as that checks out. On plans that include cloning, Starter and above, it is self-serve through the API or the Console with no per-voice review, and the clone works on simba-3.2 straight away. Simba 3.2 is English-only, so a clone speaks English on it; for German, Spanish, French, Italian or Brazilian Portuguese, use simba-3.0.

Full details are in the voice cloning docs.

Moving from another provider

Switching is real work: voices to re-map, SSML to port, latency to re-check under your own load.

  • The REST API. Every call is one HTTP request, and the REST API reference shows exactly what goes on the wire.
  • Forward-deployed engineers. For teams with volume, our forward-deployed engineers work with you on voice mapping, prosody parity, load testing and cutover.

Hear it yourself

Do not take our word for it; that is the point of an independent board. Check the Artificial Analysis provider-voice leaderboard and vote on Voice Arena, then try Simba 3.2 with a free API key on your own text.

FAQ

Common questions

Where does Simba 3.2 rank on the Artificial Analysis TTS leaderboard?
In the top five. On the Artificial Analysis Speech Arena provider-voice board on 23 September 2026, Simba 3.2 scored 1,237 (±14), fifth of 92 models with a rank band of fourth to seventh, and it was the lowest-priced model in the top ten. Nothing on the board was both cheaper and rated higher. Cartesia's Sonic 3.6 held the top score, 1,273.
How is the Artificial Analysis TTS Elo calculated?
From blind votes. Listeners hear two clips of the same text from two models without knowing which is which and pick the one they prefer, and each vote moves both models' Elo ratings. The gap between two ratings is a head-to-head preference: 36 points means listeners pick the higher model about 55 times in 100. Artificial Analysis publishes a confidence interval and a rank band beside each score, and on 23 September 2026 Simba 3.2's band was fourth to seventh.
How fast is Simba 3.2?
Two independent benchmarks timed it on 24 September 2026. Coval measured a median time to first audio of 106 ms over the 24 hours to that day, and Voice Arena read 123 ms p50 on its US English board, the lowest of the 14 models it gives a latency reading. Our own production first byte on the US East streaming path was 56 ms at the median and 102 ms at the 90th percentile, read 15 September 2026.

Privacy preferences

Choose what we may store on this device. You can change this at any time from the footer.

Strictly necessary

Sign-in, security, load balancing, and remembering your privacy choices. These cannot be switched off.

Always on

Analytics

How the site is used in aggregate - which pages get read, where people get stuck - so we can improve it.

Marketing

Measures which campaigns bring people here, and lets us show relevant ads on other platforms.