Conversational Voices
One hundred and ninety five voices carry the conversational tag, 84 percent of the catalog. Streaming is what makes them feel live: audio begins playing before the sentence has finished rendering.

VOICES FOR THIS USE CASE
The largest tag in the catalog
One hundred and ninety five of 232 voices carry the conversational tag, which is 84 percent of the catalog. That breadth is also its weakness as a filter: almost everything qualifies, so narrowing a shortlist usually means using accent, language or age rather than this tag.
Where the rest of the latency lives
Synthesis is usually the smallest component of a conversational turn. Speech recognition and the language model generating the reply typically account for more. Measuring the whole turn before optimising the voice layer avoids work that improves a number nobody was waiting on.
What conversational means technically
In a real time application it means responding inside the pause a person expects, which requires streaming rather than batch generation. First audio arrives in around 200 milliseconds on the SSE and WebSocket transports. The synchronous endpoint waits for the full clip, and that delay is audible in dialogue.
Casting for extended interaction
A voice heard once in a video and a voice heard across a twenty turn conversation are judged differently. Character that seems appealing in a sample can become tiring over a session. Auditioning a realistic exchange rather than a single line is the more useful test.
SPECIFICATION
Sample rate
44.1 kHz native, resampled to 8 kHz for telephony
Latency
~200 ms to first byte (p50, warm region)
Output formats
ulaw, alaw, PCM 16-bit, WAV, MP3
Streaming transports
WebSocket, HTTP chunked transfer
Speed range
0.5× – 2.0×, set per request
Explore Voice Similar to
Conversational Voices
FAQs










