Fast to the First Word, Right on the Hardest Ones: Pulse's Benchmark on Speko

Listen to the article
2:00

TABLE OF CONTENT

Agent Workflows

AI-Powered Solutions

Revolutionizing Industries

Automate your Contact Centers with Us

Experience fast latency, strong security, and unlimited speech generation.

Most of what you hear about a speech-to-text model's performance comes from the company that built it. That's exactly why it matters when an independent voice AI gateway runs its own numbers and Speko just ran rounds on Pulse, Smallest AI's streaming + batch STT.

Speko is the team behind the Speko voice AI gateway: the same gateway that now runs Pulse in production, routable to any voice agent on the platform with a single config change. When a gateway that has to be right about latency and accuracy for its own paying customers benchmarks a model, the results carry a different kind of weight. Here's what they found.

Round one: the two clocks that actually matter

Speko's first look at Pulse ran under the title "Fast to the First Word, Fast to the Last One." Their framing is worth repeating, because it cuts against how most STT leaderboards work: word error rate on clean, pre-recorded audio is the standard number, and it's the wrong one for a live voice agent. A caller doesn't experience your STT's accuracy — they experience its latency. The dead air after they stop talking. The beat too long before the agent starts responding. That gap is where a conversation stops feeling like a conversation.

So Speko watched two clocks instead of one:


Metric

Pulse

First partial token

~64 ms

End-of-turn (with end-of-speech finalization)

~180 ms

WER — English (FLEURS, n=50)

5.1%

Most streaming STT models are good at one of these clocks. Speko's conclusion was that Pulse is the rare model that's genuinely fast on both. First partial tokens land in about 64 milliseconds — the transcript is already scrolling while the caller is still mid-sentence, which matters enormously for a downstream LLM that needs a head start to begin reasoning. And on the end of the turn, paired with end-of-speech finalization, Pulse commits its final transcript in around 180 milliseconds after the caller stops talking — not the one-to-three second silence-timer lag you get from a lot of production STT stacks. In a live call, that's the difference between a reply landing on the beat and an agent that feels like it's catching up.


Speko didn't stop at speed. They were explicit that speed only counts if the words are right — and Pulse held 5.1% WER on English, which they placed in the competitive band with the established names, at a price point they called hard to argue with.

Their conclusion: Pulse is now live in the Speko gateway as a routable STT, and it's one of the first models they'd point builders at for agents where responsiveness is the product.

Round two: tested against five real voice agents actually hear

Four days later, Speko went further — and this is the round that really matters, because it's not a solo review. It's a head-to-head across six streaming STT models, all run through the same gateway, on the conditions a real phone call actually presents: accented speech, telephone-band audio, background noise, medical vocabulary, and spontaneous, overlapping speech. Not audiobook audio in a quiet room.

The result, from their post "What a Voice Agent Actually Hears":


Model

Avg. WER across 5 conditions

Smallest Pulse

9.9%

Qwen3-ASR

9.9%

ElevenLabs Scribe v2

12.9%

Cartesia Ink-2

13.1%

Deepgram Nova-3

13.5%

Soniox stt-rt-v5

17.2%


Pulse tied for the top spot on the leaderboard, with Qwen3-ASR — a much larger general-purpose multilingual model — matching it at the same score. Speko's own framing made a point of calling out that the two leaders "aren't the ones with the loudest launch posts," and that a competitor which markets itself as the #1 streaming WER model landed fourth once tested on the full mess of real production audio rather than a curated accent set.


What's most telling is where Pulse held up. Speko's breakdown found that telephony and background noise barely move the needle for modern streaming models — the real difficulty is content. Medical vocabulary and spontaneous, overlapping speech are where the field pulls apart, and every model's error rate roughly doubles on spontaneous speech. On that hardest condition, Speko specifically called out Pulse as having the flattest curve in the field — the smallest degradation when the audio stopped being tidy.

That's the number that matters most for anyone building a real agent: not how a model performs on its best day, but how little it falls apart on a normal one.

Why this is the review that counts

Anyone can publish their own benchmark. What makes Speko's numbers worth building around is the rigor behind them — including the part most benchmarks would quietly bury. When two other models initially looked broken in their harness, Speko didn't publish the bad numbers and move on. They went back, called each vendor's API directly, found the bug was in their own gateway's handling of streaming finals, fixed it, and republished corrected results. That's the kind of self-scrutiny that makes a "tied for #1" claim actually mean something.

Pulse earned its place on that leaderboard against real production noise, real accents, real medical terms, and real people talking over themselves — scored by a team with every incentive to be skeptical of vendor claims, including ours.

Pulse is live in the Speko gateway today. If you're building a voice agent where every millisecond of turn-taking latency is felt by a real caller, that's a good place to try it — or come talk to us directly at smallest.ai.

Frequently asked questions

Frequently asked questions

How accurate is Pulse compared to other STT models?

How fast is Pulse?

Why does an independent benchmark from Speko matter more than a vendor's own numbers?

What makes Pulse well-suited for real voice agents?