Smallest AI vs Speechmatics: Which Speech-to-Text API Delivers?

Compare Pulse STT by Smallest AI vs Speechmatics on WER, latency, emotion recognition, pricing, and deployment

Smallest AI vs Speechmatics: Which Speech-to-Text API Delivers?

Compare Pulse STT by Smallest AI vs Speechmatics on WER, latency, emotion recognition, pricing, and deployment

For real-time speech-to-text, Smallest AI's Pulse leads on latency with ~64ms time-to-first-token and ranks #2 on the Open ASR Leaderboard at 5.42% average WER. Speechmatics is strong on global accents and on-premise deployment, but Pulse is the better fit when sub-100ms streaming and voice-agent responsiveness are the priority.

Tap the mic, upload a file, or pick a sample below.
Our engine

Pulse Pro

Model: Pulse Pro · smallest.ai

Response time · transcription only

Speechmatics

Model: Ursa 2

Response time · transcription only

Why teams switch to Pulse Speech to Text

Here's what that difference looks like in production.

Why teams switch to Pulse Speech to Text

Here's what that difference looks like in production.

Why teams switch to Pulse Speech to Text

Here's what that difference looks like in production.

Pulse Speech-to-Text

Real-time transcription with industry-leading accuracy and ~64ms time-to-first-token, built for live voice agents.

30+ Languages

Accurate transcription across 30+ languages, with native English performance ranked #2 on the Open ASR Leaderboard.

Enterprise-Grade Compliance

SOC 2 Type II, GDPR, ISO 27001, and HIPAA-ready — with a Business Associate Agreement available for healthcare deployments.

Pulse vs Speechmatics

A factual, model-level comparison on the metrics that matter in production.

Pulse vs Speechmatics

A factual, model-level comparison on the metrics that matter in production.

Pulse vs Speechmatics

A factual, model-level comparison on the metrics that matter in production.

Features

Features
Pulse
Ursa 2
Streaming latency (TTFT)
~64 ms
Sub-1s real-time
Streaming approach
Token-level, low-latency
Larger-chunk streaming
Languages
35 (21 streaming + 26 batch)
55+ languages
Deployment
Managed API + on-prem
Cloud, on-prem, on-device
Compliance
SOC 2, HIPAA, GDPR
Enterprise, data residency
Pricing (real-time)
~$0.008/min
~$0.0537/min (Flow)

Benchmarks

Benchmark
Domain
Dataset
Pulse
AssemblyAI
Deepgram Nova 3
ElevenLabs Scribe
Audiobook (clean)
LibriSpeech Clean
2.46
1.65
3.20
1.97
Audiobook (noisy)
LibriSpeech Other
5.31
2.86
6.60
4.45
Crowdsourced
Common Voice
10.89
6.73
14.22
9.83
Parliament
VoxPopuli
7.16
7.28
9.55
7.91
TED talks
TED-LIUM
4.07
2.95
3.59
3.16
Podcasts
GigaSpeech
10.43
9.12
10.05
9.66
Financial
SPGISpeech
2.86
1.74
2.99
4.40
Earnings calls
Earnings22
12.25
11.52
15.79
12.20
Meetings
AMI
10.58
14.60
17.04
12.23
Overall
Aggregate
7.33
6.49
9.23
7.31

WER % on the Hugging Face ESB benchmark (9 English datasets, streaming). Lower is better. Pulse is competitive on aggregate and leads on meeting/noisy audio; AssemblyAI leads on clean read speech. Source: HF ESB, internal evaluation.

Certified & Compliant

Guarding your data with enterprise security

Certified & Compliant

Guarding your data with enterprise security

Proactive Defense

Anticipating threats before they emerge, thanks to our advanced monitoring.

Proactive Defense

Anticipating threats before they emerge, thanks to our advanced monitoring.

Proactive Defense

Anticipating threats before they emerge, thanks to our advanced monitoring.

Frequently
asked questions

Can Pulse STT run on-premise?

Why are teams looking for a Speechmatics alternative?

How does Pulse STT's accuracy compare to Speechmatics?

What does switching from Speechmatics to Pulse STT involve?

Does Speechmatics offer emotion recognition?

Build the future of voice agent orchestration

Build the future of voice agent orchestration

Build the future of voice agent orchestration