FOR EXPLAINER & PRODUCT DEMO VIDEOS
Explainer scripts get rewritten after the first cut. Generating line by line means a late change costs one line rather than a re-record, and audio arrives in about a second at 44.1 kHz.

VOICES FOR THIS USE CASE
Short scripts, tight timing
Explainer scripts are written to picture, so the audio has to fit a cut rather than the other way round. Requests cap at 250 characters, which maps roughly to a single spoken sentence and encourages writing in beats. Generating line by line makes it easier to retime a section without re-rendering the whole track.
Matching a brand voice
There is no tone parameter, so a voice cannot be nudged toward a brand sound. The two routes are casting from the catalog, where 232 voices span six accent groups, or cloning a voice you already use in campaigns. Cloning requires consent from the person whose voice it is.
Iterating without a booth
Most explainer work involves several script revisions after the first edit. Regenerating a line takes about a second and costs the characters in that line, so late changes stop being expensive. That tends to change how teams write, since the cost of trying a different phrasing drops close to zero.
Formats for an editing timeline
Audio is 44.1 kHz mono natively. Take WAV rather than MP3 into an editing timeline so your export applies compression once rather than twice. MP3 is fine for review copies and client approvals, where file size matters more than the final quality of the master.
SPECIFICATION
Sample rate
44.1 kHz native, resampled to 8 kHz for telephony
Latency
~200 ms to first byte (p50, warm region)
Output formats
ulaw, alaw, PCM 16-bit, WAV, MP3
Streaming transports
WebSocket, HTTP chunked transfer
Speed range
0.5× – 2.0×, set per request
Explore Voice Similar to
FOR EXPLAINER & PRODUCT DEMO VIDEOS
FAQs










