American English AI Voices
Forty three American voices, the deepest English set in the catalog. Enough range to cast for tone across a campaign rather than reusing the same voice because nothing else fits.

VOICES FOR THIS USE CASE
The deepest English set
Forty three voices carry the American tag, making this the accent group with the most room to cast for tone rather than availability. Across a campaign you can vary voices without repetition, which the smaller accent groups do not allow.
Neutrality is audience dependent
American reads as the default international register for many audiences, and one voice in the catalog is explicitly tagged neutral. For Indian audiences that reverses: an Indian accent reads as neutral and American reads as foreign. Choose against the audience rather than a general notion of neutrality.
Language pairing and what it covers
Set language to en. These voices sit in the European family, which accepts ten European languages in total, though only English and Spanish are recommended within it. Anything beyond those two should be tested before you depend on it.
Consistent output at volume
Latency and audio quality are identical across every voice in the catalog, so performance is never a casting consideration. Output is 44.1 kHz mono natively, with PCM, MP3 and WAV available, plus ulaw and alaw where the audio is destined for a phone path.
SPECIFICATION
Sample rate
44.1 kHz native, resampled to 8 kHz for telephony
Latency
~200 ms to first byte (p50, warm region)
Output formats
ulaw, alaw, PCM 16-bit, WAV, MP3
Streaming transports
WebSocket, HTTP chunked transfer
Speed range
0.5× – 2.0×, set per request
Explore Voice Similar to
American English AI Voices
FAQs










