Voice Library
Every voice in the catalog, playable right here without signing in. 244 of them across 22 languages and 27 accents. Filter by accent, gender or use case, hear the one you want, then take its ID straight to the API.
Use Case

Course audio has to sound the same in module ten as in module one. These voices hold consistent across sessions, handle technical vocabulary through pronunciation dictionaries, and cover nine Indic languages alongside English.

Explainer scripts get rewritten after the first cut. Generating line by line means a late change costs one line rather than a re-record, and audio arrives in about a second at 44.1 kHz.

IVR voices need to survive a compressed phone line, not just sound good in a browser. These output ulaw and alaw directly, generate in roughly 200 milliseconds, and handle Hindi and English in the same prompt.

Voice agents live or die on the pause before a reply. Synthesis starts in about 200 milliseconds and runs at 3.3 times real time, which leaves the budget where it usually belongs: your model.

Support calls reach people who are already inconvenienced, often on a poor line. These voices favour clarity over character, hold steady across renders, and move between Hindi and English the way callers actually do.

Web chat is unusual because the visitor reads and listens at once. These voices suit an unhurried delivery, begin playing in about 200 milliseconds over WebSocket, and cover twelve languages from one integration.

Pre-rendering every branch means shipping every branch. WebSocket streaming lets dialogue generate during play instead, so a conversation tree costs what players actually hear rather than what they might.

A novel is thousands of requests stitched together, and the joins are where narration falls apart. These voices hold consistent across a full book and cover nine Indic languages that most vendors do not.

Announcements are heard in reverberant halls and repeated thousands of times. Pronunciation dictionaries fix station and place names permanently, ulaw and alaw feed public address hardware directly, and one voice carries a full multilingual chain.

Podcast production is segment based already, which suits generation in short calls. Intros and ad reads generate once and cache, and pairing contrasting voices produces a two-hander without booking two people.

Experienced screen reader users often run well above normal speed. Speed adjusts from 0.5 to 2.0, and nine Indic languages are covered, where assistive audio is thin across the whole industry.

Guided audio lives on pacing rather than voice character. Speed goes down to 0.5, pauses come from how you write the script, and nine Indic languages are available for regional wellness content.

Short form rewards pace, and speed adjusts up to 2.0. Scripts change after the edit, so generating line by line means retiming a section costs one line rather than the whole track.

Commercial reads are short, fast, and worth testing in variants. A thirty second spot is one or two requests, which makes producing four versions cheaper than booking one session.

Generation runs at 3.3 times real time, so audio renders faster than it plays and the buffer stays ahead. First audio arrives in about 200 milliseconds over SSE or WebSocket.

At volume the cost driver is duplication, not generation. Hashing text and skipping what already exists removes most of it, since catalog and template work repeats the same phrases constantly.

Course audio has to sound the same in module ten as in module one. These voices hold consistent across sessions, handle technical vocabulary through pronunciation dictionaries, and cover nine Indic languages alongside English.

Explainer scripts get rewritten after the first cut. Generating line by line means a late change costs one line rather than a re-record, and audio arrives in about a second at 44.1 kHz.

IVR voices need to survive a compressed phone line, not just sound good in a browser. These output ulaw and alaw directly, generate in roughly 200 milliseconds, and handle Hindi and English in the same prompt.

Voice agents live or die on the pause before a reply. Synthesis starts in about 200 milliseconds and runs at 3.3 times real time, which leaves the budget where it usually belongs: your model.

Support calls reach people who are already inconvenienced, often on a poor line. These voices favour clarity over character, hold steady across renders, and move between Hindi and English the way callers actually do.

Web chat is unusual because the visitor reads and listens at once. These voices suit an unhurried delivery, begin playing in about 200 milliseconds over WebSocket, and cover twelve languages from one integration.

Pre-rendering every branch means shipping every branch. WebSocket streaming lets dialogue generate during play instead, so a conversation tree costs what players actually hear rather than what they might.

A novel is thousands of requests stitched together, and the joins are where narration falls apart. These voices hold consistent across a full book and cover nine Indic languages that most vendors do not.

Course audio has to sound the same in module ten as in module one. These voices hold consistent across sessions, handle technical vocabulary through pronunciation dictionaries, and cover nine Indic languages alongside English.

Explainer scripts get rewritten after the first cut. Generating line by line means a late change costs one line rather than a re-record, and audio arrives in about a second at 44.1 kHz.

IVR voices need to survive a compressed phone line, not just sound good in a browser. These output ulaw and alaw directly, generate in roughly 200 milliseconds, and handle Hindi and English in the same prompt.

Voice agents live or die on the pause before a reply. Synthesis starts in about 200 milliseconds and runs at 3.3 times real time, which leaves the budget where it usually belongs: your model.

Support calls reach people who are already inconvenienced, often on a poor line. These voices favour clarity over character, hold steady across renders, and move between Hindi and English the way callers actually do.

Web chat is unusual because the visitor reads and listens at once. These voices suit an unhurried delivery, begin playing in about 200 milliseconds over WebSocket, and cover twelve languages from one integration.
ACCENT

One hundred and four voices carry an Indian accent, more than twice any other group. They handle English words inside Devanagari sentences without changing register, which is how Indian English is actually spoken.

Seventy two voices cover eleven Indic languages plus English, so one voice serves a Hindi caller and a Tamil caller without re-casting. Hindi and English can alternate inside a single sentence.

Seven British voices, listed in full rather than selected. Isla, Julia and Poppy are female, Alistair, Edward and Noah male. Enough for most projects, small enough to audition in five minutes.

Forty three American voices, the deepest English set in the catalog. Enough range to cast for tone across a campaign rather than reusing the same voice because nothing else fits.

Five Australian voices, and all five are here. Chloe, Nyah and Sienna are female, Cooper and Flynn male. A small set, but enough to cast a project without repeating a voice.

Three Canadian voices: Erica, Alec and William. That is the whole group. If you need more range in a North American register, the forty three American voices sound close.
STYLE

One hundred and one female voices, spanning six accent groups and all twelve recommended languages. At that scale accent and language usually constrain casting more than gender does.

One hundred and twenty male voices across every accent group and all twelve recommended languages. Most carry an Indic recommendation, which is unusual in catalogs that treat Indian languages as an afterthought.

Worth saying plainly: nothing in the catalog records pitch, so these six were chosen by ear rather than filtered. There is no pitch control either, so depth is a casting decision.

Corporate narration needs consistent terminology more than it needs a particular tone. Pronunciation dictionaries fix brand and product names once, so every module in a library matches the first one.

One hundred and ninety five voices carry the conversational tag, 84 percent of the catalog. Streaming is what makes them feel live: audio begins playing before the sentence has finished rendering.

There is no tone field in the catalog, so these were chosen by listening rather than filtered. Warmth is a property of the voice, not a setting, which makes casting the decision that counts.

Calm comes from pace more than from voice. Speed runs down to 0.5, and pauses come from punctuation rather than a parameter, so the script does as much work as the casting.

Speed is the reliable lever here, not casting. Settings between 1.2 and 1.4 lift almost any voice in the catalog, though longer sentences start to blur past roughly 1.5.

Slowing down reads as more authoritative than speeding up. A setting near 0.9 does more than any voice choice, and short declarative sentences carry more weight than qualified ones.

One hundred and forty six voices are tagged young, 63 percent of the catalog, across every accent and all twelve recommended languages. Unlike tone, age is a real tagged field rather than an inference.

Eighty six voices are tagged mature, and twenty also carry the narrative tag. That overlap is the usual starting point for audiobooks and documentary work, where consistency shows over hours rather than seconds.

There is no emotion parameter, so range comes from the voice and from how the line is written. Fourteen voices carry the character tag, and those have the widest delivery available.

Only two voices carry the meditative tag, so this set is drawn wider and chosen by ear. Pace matters more than casting: around 0.7 suits most soothing material.

For Indian customer facing work the deciding factor is not tone but language. These voices switch between Hindi and English mid sentence, which reads as considerably more natural than either alone.

For compliance and legal reads the goal is that exact wording lands, not that the voice sounds grave. A speed near 0.9 improves intelligibility more than any casting choice.

Only six voices in the catalog show cheerful or positive markers, the thinnest tone signal available. Speed slightly above 1.0 lifts almost any voice, which is the more dependable route.

One hundred and one female voices, spanning six accent groups and all twelve recommended languages. At that scale accent and language usually constrain casting more than gender does.

One hundred and twenty male voices across every accent group and all twelve recommended languages. Most carry an Indic recommendation, which is unusual in catalogs that treat Indian languages as an afterthought.

Worth saying plainly: nothing in the catalog records pitch, so these six were chosen by ear rather than filtered. There is no pitch control either, so depth is a casting decision.

Corporate narration needs consistent terminology more than it needs a particular tone. Pronunciation dictionaries fix brand and product names once, so every module in a library matches the first one.

One hundred and ninety five voices carry the conversational tag, 84 percent of the catalog. Streaming is what makes them feel live: audio begins playing before the sentence has finished rendering.

There is no tone field in the catalog, so these were chosen by listening rather than filtered. Warmth is a property of the voice, not a setting, which makes casting the decision that counts.

Calm comes from pace more than from voice. Speed runs down to 0.5, and pauses come from punctuation rather than a parameter, so the script does as much work as the casting.

Speed is the reliable lever here, not casting. Settings between 1.2 and 1.4 lift almost any voice in the catalog, though longer sentences start to blur past roughly 1.5.

One hundred and one female voices, spanning six accent groups and all twelve recommended languages. At that scale accent and language usually constrain casting more than gender does.

One hundred and twenty male voices across every accent group and all twelve recommended languages. Most carry an Indic recommendation, which is unusual in catalogs that treat Indian languages as an afterthought.

Worth saying plainly: nothing in the catalog records pitch, so these six were chosen by ear rather than filtered. There is no pitch control either, so depth is a casting decision.

Corporate narration needs consistent terminology more than it needs a particular tone. Pronunciation dictionaries fix brand and product names once, so every module in a library matches the first one.

One hundred and ninety five voices carry the conversational tag, 84 percent of the catalog. Streaming is what makes them feel live: audio begins playing before the sentence has finished rendering.

There is no tone field in the catalog, so these were chosen by listening rather than filtered. Warmth is a property of the voice, not a setting, which makes casting the decision that counts.
Build the future of voice agent orchestration
311 California Street, Suite 320
San Francisco, CA 94104
Documentation
Resources
Initiatives
Build the future of voice agent orchestration
311 California Street, Suite 320
San Francisco, CA 94104
Build the future of voice agent orchestration
311 California Street, Suite 320
San Francisco, CA 94104
Documentation
Resources
Initiatives
