Expressive Voices
There is no emotion parameter, so range comes from the voice and from how the line is written. Fourteen voices carry the character tag, and those have the widest delivery available.

VOICES FOR THIS USE CASE
Range comes from casting and writing
There is no emotion parameter, so expression cannot be dialled in. It comes from the voice you cast and from how the line is written. That is a real constraint worth stating rather than working around, because it changes where the effort goes.
Consistency across a scene
Requests cap at 250 characters, so an extended exchange spans several calls with nothing holding a performance steady across them. Generate a full scene rather than a line and listen across the joins before building a production pipeline around it.
The character pool
Fourteen voices carry the character and animation tag, which is the group with the widest delivery available. Julia and Kiara carry the most range, Albus and Blofeld the most weight. Beyond those fourteen, the catalog offers nothing that describes expressiveness.
Generating at runtime
For games and interactive work, the WebSocket transport streams audio as it renders, so expressive dialogue can be produced during play. That avoids pre-rendering every branch, which in a game with meaningful choice is where asset counts usually get out of hand.
SPECIFICATION
Sample rate
44.1 kHz native, resampled to 8 kHz for telephony
Latency
~200 ms to first byte (p50, warm region)
Output formats
ulaw, alaw, PCM 16-bit, WAV, MP3
Streaming transports
WebSocket, HTTP chunked transfer
Speed range
0.5× – 2.0×, set per request
Explore Voice Similar to
Expressive Voices
FAQs










