Benchmarks

Voice AI Benchmarks: Speech-to-Text, Text-to-Speech & Speech-to-Speech

Real-world and public-benchmark results for Pulse, Lightning and Hydra, with full methodology and reproducible scripts.

Last updated

19 Aug 2026

Every change is logged in the Benchmark history.

Section 01

Three numbers, three levels of proof

Every figure on this page carries a badge naming who measured it. Only one of these three is independently verifiable, and we say so.

Pulse Pro · STT

5.42%

Average ESB word error rate

Tied 2nd of 86 systems on the public Open ASR Leaderboard, across 8 datasets.

Lightning v3.1 · TTS

~200ms

Time to first byte

At 40 concurrent requests over WebSocket streaming. Real-time factor 3.3× on one L40S.

Hydra · S2S

1,624ms

Tool voice-to-voice latency

Mean, fastest of 9 production models on AIEWF. 864 ms median without tool calls.

Hydra · S2S

1,624ms

Tool voice-to-voice latency

Mean, fastest of 9 production models on AIEWF. 864 ms median without tool calls.

Providers compared on this page

Deepgram

Cartesia

OpenAI

AssemblyAI

Murf

Speechmatics

WellSaid

Soniox

Gladia

NVIDIA

ElevenLabs

Sarvam

Section 02

Latency comparison

Plotted on a shared millisecond scale so the real spread is visible. Averages hide the calls that feel broken, so percentiles appear wherever they exist.

Speech-to-text, time to first transcript

Sorted fastest first. Solid marker is a point measurement, capsule is a P50 to P90 span.

Speechmatics Flow publishes only a sub-1 second figure, too coarse to plot honestly.

Under 150 ms

150 to 300 ms

Over 300 ms

Smallest measured

Speko

Text-to-speech, time to first audio

Sorted fastest first. Vendor ranges are drawn as capsules across their whole span, not collapsed to a midpoint.

Murf and Cartesia are faster than Lightning on these figures. Cartesia's is a 90th percentile and WellSaid's is per 30 characters, so the set is not measured like for like.

Under 200 ms

200 to 300 ms

Over 400 ms

Smallest measured

Our own numbers disagree. Internal sources give Pulse time to first transcript as both 64 ms and 70 ms, and Deepgram Nova 3 as both 71 ms and 100 ms. Lightning appears as 100 ms end-to-end on our comparison pages and ~200 ms time to first byte in the docs, and end-to-end cannot be faster than its own first byte. The conservative figure is plotted for us and the favourable one for competitors until a single run settles it.

Why percentiles, not means

A voice agent is judged on its worst turns. A model with a 200 ms mean and a 2 second P99 feels broken more often than one with a 300 ms mean and a 400 ms P99, and a mean cannot tell you which one you have. We publish percentiles where we have them and mark the gap where we do not.

Section 03

How we benchmark

The parameters stated up front rather than hidden behind a disclosure. Public benchmarks can be overfitted and gamed, which is an argument for showing your method, not for hiding behind one number.

Datasets

ESB suite, 9 datasets, English. FLEURS for multilingual, 26 languages. WildASR for degraded recording conditions. Internal English and Hindi perturbation suites. EmergentTTS, 1,088 samples drawn from the 1,645-sample public set. AIEWF, 300 tasks over 10 runs of 30 turns. Plus 44 real Hindi-English collections calls.

Accuracy metric

Word error rate under the Whisper normalizer, with character error rate alongside. This choice is material: a third-party board reported Pulse at 8.0% where inverse text normalisation scored spelled-out numbers as errors, against 5.1% otherwise. On numbers-heavy call audio the normaliser changes the answer.

Latency metric

Time to first transcript runs from end of speech to complete transcript. Time to first byte is the first audio byte received. Voice-to-voice is end of user speech to start of agent audio. Reported as P50, P90, P95 and P99 wherever the run produced them.

Ground truth

Manual human transcription for the call benchmark. Public datasets use their own published references, unmodified.

Hardware

Throughput on a single L40S, 48 GB. Streaming tests at 16 kHz linear16. Competitor models called through their public production APIs, not self-hosted.

Refresh cadence

Quarterly, and immediately when any model in the comparison set ships a new version. Every table carries its own measurement date.

Reproduce it

The evaluation walkthrough and scripts are public, and every result below can be regenerated against your own audio.

Section 04

Speech-to-text

Switch regimes to see the same models on clean benchmark audio and on real call recordings.

Clean benchmarks

Real call audio

ESB suite word error rate by domain. Lower is better, green marks the best model in each row.

Domain

Audiobook, clean

Audiobook, noisy

Crowdsourced

Parliament

TED talks

Podcasts

Financial

Earnings calls

Meetings

Overall

Dataset

LibriSpeech Clean

LibriSpeech Other

Common Voice

VoxPopuli

TED-LIUM

GigaSpeech

SPGISpeech

Earnings22

AMI

Aggregate

Pulse

2.46

5.31

10.89

7.16

4.07

10.43

2.86

12.25

10.58

7.33

AssemblyAI

1.65

2.86

6.73

7.28

2.95

9.12

1.74

11.52

14.60

6.49

Deepgram

3.20

6.60

14.22

9.55

3.59

10.05

2.99

15.79

17.04

9.23

ElevenLabs

1.97

4.45

9.83

7.91

3.16

9.66

4.40

12.20

12.23

7.31

Smallest's Pulse achievement. Pulse wins parliamentary speech and meetings, the two hardest multi-speaker conditions. Note also that Pulse streams at 7.33 aggregate here, while the 5.42% headline is Pulse Pro on the pre-recorded path. They are different models. While Assembly AI leads 7 of 9 domains and the aggregate, and ElevenLabs Scribe edges Pulse overall by 0.02.

FLEURS streaming word error rate by language, Pulse against Deepgram Nova 3

LANGUAGE

Italian

Spanish

English

Portuguese

Hindi

Germen

French

Marathi

Gujarati

PULSE

4.41

5.99

6.03

8.32

8.30

9.50

10.71

15.68

20.05

Deepgram NOVA 3

6.99

7.55

11.21

11.46

15.46

10.15

12.07

not supported

not supported

RATIO

1.6x

1.3x

1.9x

1.4x

1.19

1.1x

1.1x

Marathi and Gujarati have no Deepgram column because the model does not cover them. Coverage is itself a result.

Smallest measured

WildASR robustness, word error rate under degraded recording conditions

CONDITION

Clean Audio

Accented speech

Phone codec

Background noise

Reverberation

Far-field mic

Clipping

Overall

Pulse

5.98

5.82

7.19

8.90

9.06

13.38

14.03

9.63

AssemblyAI

3.33

2.80

3.45

4.04

23.50

26.07

6.59

12.52

Deepgram

11.62

7.31

9.13

15.04

27.27

62.99

43.35

28.17

ElevenLabs

4.24

4.01

4.98

6.30

6.48

7.38

11.20

6.74

ElevenLabs Scribe is the most robust model here overall. The useful finding is about failure modes: AssemblyAI is the most accurate on clean audio yet degrades to 26.07 far-field and 23.50 on reverberation, and Deepgram reaches 62.99 far-field. Pulse loses most single conditions but never collapses on any of them, which is what matters when you cannot control how your callers are recorded.

Smallest measured

Pulse: 21 streaming languages, 26 pre-recorded.
Pulse Pro: English only, HTTP only, streaming worker on roadmap. Throughput 250-300x real time on 1x L40S.
Pricing: our own materials quote $0.003, $0.004, $0.005 and $0.008 per minute for real-time Pulse. Not published here until reconciled.

Section 05

Public benchmarks vs your actual audio

Our Hindi word error rate is 8.30% on FLEURS and 29.1% on real collections calls. Both are measured. Neither is wrong.

Same model, same language, two kinds of audio

Word error rate. Lower is better.

Read-speech benchmarks measure clean, scripted audio recorded in good conditions. Production calls are code-switched, noisy, spontaneous and compressed by telephony codecs. The gap is roughly 3.5× for us and 3.5× for Sarvam, so it is a property of the audio rather than of any one model.


We publish both because a single number would mislead you. Whichever figure a vendor quotes, it is probably the easier one, and the only number that predicts your production accuracy is the one you measure on your own recordings.

Section 06

Real Hindi-English call audio

Five providers on 44 real collections calls, scored against manual human transcripts. We win two of six categories. The other four are published unchanged.

Green marks the category winner. Recall scores run 0 to 1, higher is better. July 2026.

Dataset

SMALLEST AI

PULSE

DEEPGRAM

NOVA 3

SERVAM AI

Saaras

ELEVENLABS

Scribe, KEYTERMS

ELEVENLABS

Scribe, standard

WER

0.291

0.325

0.376

0.460

0.554

LOAN TERMS

0.872

0.919

0.910

0.902

0.815

PROMISE TO PAY

0.557

0.509

0.502

0.432

0.381

NUMBERS

0.387

0.282

0.381

0.410

0.389

AMOUNT

0.334

0.280

0.330

0.349

0.313

DATES

0.531

0.510

0.522

0.611

0.485

What to take from this

Pulse gives the most accurate and most consistent transcript. Deepgram is strongest on loan terminology. ElevenLabs Keyterms captures financial entities best but produces a noticeably worse transcript around them. Always choose Keyterms over standard ElevenLabs; it wins every metric at the same speed and cost.

An industry-wide gap

Across all five providers, amounts and numbers recall stayed low and none cleared 0.35. That is an unsolved problem for the whole category rather than a gap in any one product, and it means any amount-critical workflow needs a verification step whichever API you choose.

Scope and limits. 44 recordings, July 2026, Hindi-English collections and loan-servicing calls from a single domain. Ground truth is manual human transcription. The dataset carries no noise-condition labels and no language-mix breakdown, so nothing here supports a conclusion about specific noise levels or code-switching ratios. At n=44, small gaps between adjacent models may not be significant.

Section 07

Entity-level recall

Word error rate treats every word as equal. In a collections call, one missed rupee figure matters more than ten missed filler words. No public leaderboard measures this, so we defined it.

Loan terms

Did the model capture loan vocabulary such as EMI, tenure and interest.

Promise to pay

Did it capture commitment phrases such as "I will pay by" or "will transfer today".

Numbers

Did it correctly transcribe spoken numeric values.

Amounts

Did it capture spoken rupee figures.

Scoring

All scored 0 to 1 and averaged over 44 recordings. These are Smallest-defined metrics. The term lists and matching logic still need publishing before an outside team can reproduce them.

Promise-to-pay recall

Higher is better, on the full 0 to 1 scale so the true size of the gaps is visible.

Amounts recall

The dashed line marks 0.35. No provider clears it.

Section 08

Text-to-speech

Blind head-to-head listener preference on the EmergentTTS benchmark, then the same models rated dimension by dimension.

Lightning win rate against each competitor

1,088 samples. The solid line is 50%, where neither model is preferred.

GPT-4o-mini TTS beats Lightning on this benchmark. That row stays in. The 1,088 samples are a subset of the 1,645-sample public EmergentTTS set.

Smallest measured

Naturalness

Metric

Overall

Naturalness

Intonation

Prosody

Lightning Pro

3.16

2.55

3.06

2.81

GPT-4o-mini

3.13

2.41

3.06

2.73

EL Turbo v2.5

3.16

2.52

3.07

2.73

Sonic-3

3.20

2.57

3.12

2.83

EXPRESSIVENESS

Overall

Paralinguistics

Emotions

3.55

3.64

3.47

3.45

3.60

3.30

3.44

3.59

3.28

3.38

3.56

2.83

MOS V2

Mean MOS

UTMOS, predicted

WV-MOS, predicted

4.22

3.77

5.05

4.16

3.76

4.55

3.98

3.37

4.60

3.76

2.77

4.76

DELIVERY

Natural pace

Pause placement

Breathing naturalness

Pronunciation style

Boundary consistency

4.72

4.66

3.82

4.98

4.96

4.57

4.54

3.06

4.96

4.94

4.51

4.49

3.14

4.95

4.93

4.01

4.28

2.79

4.96

4.93

Performance Lightning's advantage is expressiveness, delivery and breathing, where the gap is large: 3.82 against 2.79 on breathing naturalness. This also contradicts the head-to-head chart above, where Lightning was preferred over Sonic-3 68.3% of the time. Blind preference and dimensional rating measure different things, and both are reported rather than only the flattering one. UTMOS and WV-MOS are model-predicted, not human ratings.

Synthesis accuracy, transcribed back with Whisper and scored against the input text

Metric

Lightning v3.1

Lightning v3.1 Pro

wer

1.57%

1.36%

cer

0.67%

0.40%

hallucination

0.03%

0.00%

15 languages including Hindi, Tamil, Kannada, Telugu, Malayalam, Marathi and Gujarati, with automatic language detection and mid-sentence switching.
Time to first byte ~200 ms at 40 concurrent streams.
Pricing: three different figures appear across our own materials for the same product, $0.05, $0.145 and $0.175 per 10K characters. Not published here until reconciled.

Section 09: Beta

Speech-to-speech

Voice-to-voice latency across 9 production models on the AIEWF benchmark: 300 tasks, 10 runs of 30 turns

Tool-call voice-to-voice latency, mean

Lower is better. The dashed line marks 2 seconds, where a turn starts to feel slow.

Non-tool voice-to-voice median is 864 ms, tied fastest of the set. Percentile distributions are not yet published.

Smallest measured

Task performance, AIEWF pass rate and Audio MultiChallenge

MODEL

Hydra

Ultravox v0.7

GPT Realtime 1.5

GPT Realtime

GPT Realtime 2

WER

95.9

97.7%

93.3%

86.7%

not run

Audio multichallenge

35.52%

not run

34.73%

not run

37.61%

Section 10

Methodology and changelog

Dataset inventory, provenance policy, and the open items we have not resolved.

Dataset inventory

BENCHMARK

ESB suite

FLURS

WildASR

EmergentTTS

AIEWF

Audio MultiChallenge

Hindi-English calls

Perturbation suites

Scope

9 datasets, English

26 languages, read speech

7 degradation conditions

1,088 of 1645 samples

300 tasks, 10x30 turns

Multi-turn audio reasoning

44 collections recordings

English and Hindi

used in

Section 04

Section 04, 05

Section 04

Section 08

Section 09

Section 09

Section 05, 06, 08

Section 04

publicly available

Yes

Yes

Yes

Yes

Yes

Yes

No, Internal

No, Internal

Section 11

FAQ

Why do your two Hindi numbers differ so much?

What is word error rate?

What is the difference between TTFB, TTFT and voice-to-voice?

Are these numbers independent?

Which model should I use for noisy call audio?

Do you support Hindi and code-switching?