TTS emotion control
Match the delivery to the moment with selectable emotional styles.
- 13 selectable emotional styles
- Warm, energetic, calm, and more
- Set per request
- Pairs with SSML for fine control
definition
Emotion control sets the emotional tone of synthesized speech. Instead of one neutral read, you choose from 13 styles so the delivery fits the content: warm for a welcome, energetic for an ad, calm for a tutorial. It is the difference between speech that sounds present and speech that sounds like a screen reader.
the-styles
There are 13 selectable styles spanning warm, energetic, calm, and more. You set one per synthesis request, so a single script can be voiced differently for different contexts by changing the style rather than the text.
pairing-with-ssml
Emotion sets the broad tone; SSML handles the fine detail. Combine them and you get both: an energetic style for the segment, with SSML pauses and emphasis placed exactly where the line needs them.
where-it-matters
Podcasts, ads, game dialogue, and video voiceover all live or die on delivery. Emotion control gives those use cases the expressiveness that flat TTS lacks. See the emotion guide for the full style list.
Frequently asked questions
Can AI text-to-speech convey emotion?
How many emotional styles are there?
Does emotion control work with SSML?
Start building
Match the delivery to the moment with selectable emotional styles.