A synthetic voice is speech produced by a model rather than recorded by a person. These six are a cross-section of the catalog: different accents, different languages, all generated the moment you press play.

VOICES FOR THIS USE CASE
What synthetic means here
Every voice in the catalog is generated by a model at request time. Nothing is a stitched recording or a sampled clip, which is why the same voice can say text it has never encountered in a language it was not recorded in. That is also why output is consistent between runs in a way a human read never is.
What the model produces
44.1 kHz native, mono, with first audio in around 200 milliseconds and generation running at 3.3 times real time. Output is available as PCM, WAV, MP3, ulaw or alaw. Requests cap at 250 characters, so longer text is assembled from several calls.
Synthetic does not mean robotic
The two get conflated, and they are different things. Robotic speech is a deliberate effect, produced with pitch shifting and vocoding. There is no effects layer here and no pitch control, so these voices sound like people rather than machines. If you specifically want a machine-sounding voice, this is not the right tool.
Where synthetic beats recorded
Anywhere the script changes often, anywhere the same content ships in several languages, and anywhere audio has to be produced at a volume no booking schedule could absorb. A correction costs one line rather than a session, and 31 languages are available from one integration.
SPECIFICATION
Sample rate
44.1 kHz native, resampled to 8 kHz for telephony
Latency
~200 ms to first byte (p50, warm region)
Output formats
ulaw, alaw, PCM 16-bit, WAV, MP3
Streaming transports
WebSocket, HTTP chunked transfer
Speed range
0.5× – 2.0×, set per request
Explore Voice Similar to
Synthetic Voices
FAQs








