Smallest AI vs Gladia: Which Speech-to-Text API Delivers?

Compare Pulse STT by Smallest AI vs Gladia Solaria-1 on latency, WER, language support, on-premise deployment, and pricing for real-time voice agents.

Smallest AI vs Gladia: Which Speech-to-Text API Delivers?

Compare Pulse STT by Smallest AI vs Gladia Solaria-1 on latency, WER, language support, on-premise deployment, and pricing for real-time voice agents.

Smallest AI's Pulse is a real-time speech-to-text API delivering ~64ms time-to-first-token and 5.42% average WER (#2 on the Open ASR Leaderboard) across 30+ languages. Gladia bundles transcription with audio-intelligence add-ons, but Pulse is the stronger choice when raw accuracy and low-latency streaming for voice agents come first.

Tap the mic, upload a file, or pick a sample below.
Our engine

Pulse Pro

Model: Pulse Pro · smallest.ai

Response time · transcription only

Gladia

Model: Solaria

Response time · transcription only

Why teams switch to Pulse Speech to Text

Here's what that difference looks like in production.

Why teams switch to Pulse Speech to Text

Here's what that difference looks like in production.

Why teams switch to Pulse Speech to Text

Here's what that difference looks like in production.

Pulse Speech-to-Text

Real-time transcription with industry-leading accuracy and ~64ms time-to-first-token, built for live voice agents.

30+ Languages

Accurate transcription across 30+ languages, with native English performance ranked #2 on the Open ASR Leaderboard.

Enterprise-Grade Compliance

SOC 2 Type II, GDPR, ISO 27001, and HIPAA-ready — with a Business Associate Agreement available for healthcare deployments.

Pulse vs Solaria

A factual, model-level comparison on the metrics that matter in production.

Pulse vs Solaria

A factual, model-level comparison on the metrics that matter in production.

Pulse vs Solaria

A factual, model-level comparison on the metrics that matter in production.

Features

Features
Pulse
Solaria
Time to First Transcript
sub 70 ms
270ms
Emotion Recognition
Yes
No
Compliance
SOC 2, HIPAA, GDPR
SOC 2
Pricing
$0.30/hr
$0.55/hr

Benchmarks

Benchmark
Domain
Dataset
Pulse
AssemblyAI
Deepgram Nova 3
ElevenLabs Scribe
Audiobook (clean)
LibriSpeech Clean
2.46
1.65
3.20
1.97
Audiobook (noisy)
LibriSpeech Other
5.31
2.86
6.60
4.45
Crowdsourced
Common Voice
10.89
6.73
14.22
9.83
Parliament
VoxPopuli
7.16
7.28
9.55
7.91
TED talks
TED-LIUM
4.07
2.95
3.59
3.16
Podcasts
GigaSpeech
10.43
9.12
10.05
9.66
Financial
SPGISpeech
2.86
1.74
2.99
4.40
Earnings calls
Earnings22
12.25
11.52
15.79
12.20
Meetings
AMI
10.58
14.60
17.04
12.23
Overall
Aggregate
7.33
6.49
9.23
7.31

WER % on the Hugging Face ESB benchmark (9 English datasets, streaming). Lower is better. Pulse is competitive on aggregate and leads on meeting/noisy audio; AssemblyAI leads on clean read speech. Source: HF ESB, internal evaluation.

Certified & Compliant

Guarding your data with enterprise security

Certified & Compliant

Guarding your data with enterprise security

Proactive Defense

Anticipating threats before they emerge, thanks to our advanced monitoring.

Proactive Defense

Anticipating threats before they emerge, thanks to our advanced monitoring.

Proactive Defense

Anticipating threats before they emerge, thanks to our advanced monitoring.

Frequently
asked questions

Can Pulse STT run on-premise?

Is Gladia good for real-time voice agents?

How does Pulse STT's latency compare to Gladia Solaria-1?

What does switching from Gladia to Pulse STT involve?

Build the future of voice agent orchestration

Build the future of voice agent orchestration

Build the future of voice agent orchestration