Artificial Analysis TTS leaderboard: Elo explained, Simba 3.2 ranked
On the Artificial Analysis Speech Arena on 23 September 2026, Simba 3.2 was in the top five of 92 models at an Elo of 1,237, and the lowest-priced model in the top ten. How the leaderboard's Elo works, and what two independent latency benchmarks measured.

On 23 September 2026, Simba 3.2 was in the top five of the 92 models on the Artificial Analysis Speech Arena, at an Elo of 1,237 (±14) with a rank band of fourth to seventh, and it was the lowest-priced model in the top ten (our reviewed snapshot of the full board). Nothing on the board was both cheaper and rated higher. On 24 September 2026, two independent benchmarks timed it from their own runners to the first audible sound: Coval measured a median of 106 ms over the 24 hours to that day, and Voice Arena read 123 ms p50 on its US English board, the lowest of the 14 models it gives a latency reading (Coval snapshot, Voice Arena snapshot).
This post is about the leaderboard itself: how the Elo is scored and what the standing means. For the whole top ten with prices, read our comparison of the best TTS APIs in 2026. For what changed in Simba 3.2 and how to migrate, read the Simba 3.2 announcement. To build with the model, start from the Text-to-Speech API.
How the Artificial Analysis Elo works
Artificial Analysis runs a blind speech arena. Listeners hear two clips of the same text from two models, without knowing which model made which, and pick the one they prefer. Every vote moves both models’ ratings, the same Elo method chess uses and the Chatbot Arena made standard for language models. It is not our benchmark and not our numbers.
A gap between two ratings is a head-to-head preference. Cartesia’s Sonic 3.6 held the top score on 23 September 2026 at 1,273, 36 points above Simba 3.2, which means blind listeners pick it about 55 times in 100, at about seven times the price on Artificial Analysis’s normalized figures (23 Sep snapshot).
Read the bands before you read the gaps. Artificial Analysis publishes a confidence interval and a rank band beside every score, and on 23 September 2026 Simba 3.2’s band of fourth to seventh overlapped five models: Google’s Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, Alibaba’s Qwen-Audio-3.0-TTS-Plus, Inworld’s Realtime TTS-2 and VUI Labs’ Luna TTS (23 Sep snapshot). Those five cost 1.7 to 4.2 times as much on the same normalized prices, and the board cannot place any of them above Simba 3.2 with confidence.
Artificial Analysis’s normalized price for Simba 3.2, $6.60 per million characters, is its own estimate and assumes 80% use of a plan (23 Sep snapshot). Our list rates are $10 per 1M characters on Starter, $8 on Pro and $6 on Scale, on the pricing page.
Voice Arena, a second blind board
Voice Arena ranks text to speech the same way, with native-speaker panels choosing between two clips of the same text, blind. It runs six languages, a balanced voice slate per model rather than whichever default sounds best, and sentences written for the places TTS actually ships, with a methodology built with advice from Prof. Shinji Watanabe at Carnegie Mellon. It also times models to their first audio. On its US English board on 24 September 2026, Simba 3.2’s quality rank was 7 of 20 (range 6 to 8), and its median time to first audio, 123 ms, was the lowest of the 14 models with a latency reading (Voice Arena snapshot).
How fast it starts talking
Three numbers, measured three ways.
- Our own first byte. On our production US East streaming path, Simba 3.2 returned its first audio bytes in 56 ms at the median and 102 ms at the 90th percentile, read on 15 September 2026. That is time from our servers, not time to audible audio from your region.
- Coval. Its open benchmark measured a median time to first audio of 106 ms over the 24 hours to 24 September 2026, from its own runner, counting the network and any silence before the first audible sample (Coval snapshot).
- Voice Arena. It read 123 ms p50 on its US English board the same day, from its own runner (Voice Arena snapshot).
How both benchmarks measured it is in our write-up of the two readings, and how to measure it yourself is in our guide to choosing a low-latency TTS API.
Voice cloning on Simba 3.2
Clone a voice you have permission to clone from a 10-30 second sample.
The speaker reads a one-time phrase aloud as the consent record, and the clone exists as soon as that checks out.
On plans that include cloning, Starter and above, it is self-serve through the API or the Console with no per-voice review, and the clone works on simba-3.2 straight away.
Simba 3.2 is English-only, so a clone speaks English on it; for German, Spanish, French, Italian or Brazilian Portuguese, use simba-3.0.
Full details are in the voice cloning docs.
Moving from another provider
Switching is real work: voices to re-map, SSML to port, latency to re-check under your own load.
- The REST API. Every call is one HTTP request, and the REST API reference shows exactly what goes on the wire.
- Forward-deployed engineers. For teams with volume, our forward-deployed engineers work with you on voice mapping, prosody parity, load testing and cutover.
Hear it yourself
Do not take our word for it; that is the point of an independent board. Check the Artificial Analysis provider-voice leaderboard and vote on Voice Arena, then try Simba 3.2 with a free API key on your own text.
Common questions
Where does Simba 3.2 rank on the Artificial Analysis TTS leaderboard?
How is the Artificial Analysis TTS Elo calculated?
How fast is Simba 3.2?
This post is narrated by Harper on Simba 3.2 through our text to speech API. How the player works: add read-aloud to your docs and turn speech marks into captions.