TTS emotion control

Match the delivery to the moment with selectable emotional styles.

  • 13 selectable emotional styles
  • Warm, energetic, calm, and more
  • Set per request
  • Pairs with SSML for fine control

definition

Emotion control sets the emotional tone of synthesized speech. Instead of one neutral read, you choose from 13 styles so the delivery fits the content: warm for a welcome, energetic for an ad, calm for a tutorial. It is the difference between speech that sounds present and speech that sounds like a screen reader.

the-styles

There are 13 selectable styles spanning warm, energetic, calm, and more. You set one per synthesis request, so a single script can be voiced differently for different contexts by changing the style rather than the text.

pairing-with-ssml

Emotion sets the broad tone; SSML handles the fine detail. Combine them and you get both: an energetic style for the segment, with SSML pauses and emphasis placed exactly where the line needs them.

where-it-matters

Podcasts, ads, game dialogue, and video voiceover all live or die on delivery. Emotion control gives those use cases the expressiveness that flat TTS lacks. See the emotion guide for the full style list.

FAQ

Frequently asked questions

Can AI text-to-speech convey emotion?
Yes. Emotion control sets the emotional tone of the synthesized speech across 13 styles such as warm, energetic, and calm. You pick a style per request so the delivery fits the moment, an upbeat ad or a measured tutorial, rather than a single neutral read.
How many emotional styles are there?
There are 13 selectable styles. You choose one per synthesis request to shape the overall tone, and combine it with SSML for finer control of pacing and emphasis within that tone.
Does emotion control work with SSML?
Yes. Emotion control sets the broad tone while SSML handles fine-grained pacing, pauses, and emphasis. Using both gives you tone plus rhythm in the same request.

Start building

Match the delivery to the moment with selectable emotional styles.