Announcing our Series A Funding

Announcing our Series A Funding

Text-to-speech API pricing and latency compared

Listen to the article
2:00

Summarize with AI

Automate your Contact Centers with Us

Experience fast latency, strong security, and unlimited speech generation.

Text-to-Speech API Pricing and Latency Compared
Text-to-Speech API Pricing and Latency Compared

Text-to-speech API pricing and latency compared across 6 providers, with cost per million characters, TTFB benchmarks, and fit for real-time vs batch.

Text-to-speech API pricing almost never comes in a clean, comparable unit. One vendor bills per character, another quotes per thousand tokens, and a third hides usage inside seat-based subscriptions. Add latency that can swing from very low to over 800ms, and the wrong pick shows up fast: either your bill balloons, or your product sounds great but feels slow.

This comparison looks at six providers: Smallest.ai, ElevenLabs, Deepgram, OpenAI, Cartesia, and AssemblyAI. I focus on pricing model and cost per million characters, Time to First Byte (TTFB) latency, voice quality, streaming support, and where each option fits in real-time vs. asynchronous workloads.

How the TTS Market Has Structured Itself in 2026

By 2026, the TTS API market has settled into three recognizable lanes: expressive models aimed at offline content creation, low-latency models built for real-time voice agents, and low-cost, high-volume models meant for bulk processing. If you can place a provider into one of those lanes, you can usually cut your shortlist before you even open a pricing tab.

Most text-to-speech API pricing still maps back to characters sent to the service, with three common patterns: pay-as-you-go per character, subscriptions with a character allowance, and token-based pricing. For real-time conversational work, the practical line in the sand is latency: a TTFB under 300ms is generally considered a practical target for responsive voice interactions. That single constraint eliminates a surprising number of otherwise solid-sounding options for voice agents.

Smallest.ai Lightning: Built for Sub-100ms Workloads


Smallest.ai's Lightning model is designed around low-latency speech synthesis for real-time applications. Smallest.ai's fastest text-to-speech stack is engineered to deliver low latency with a TTFB under 100ms on streaming requests, which makes it suitable for real-time voice agent deployments. Developers reach Lightning through the Waves API, and streaming is on by default rather than bolted on as an afterthought. 

Pricing stays simple: character-based, with pay-as-you-go plus volume discounts. The free tier is enough to prototype an end-to-end voice agent flow without immediately running into a paywall. In production, the pricing structure is designed for usage-based scaling and volume-based deployments, and you can see the details on the pricing plans page. Voice cloning is exposed through the API, and the broader platform includes Pulse (speech-to-text), Hydra (speech-to-speech), and Atoms (voice and text agent platform) for teams that want one vendor for the whole voice pipeline. The product lineup is laid out on Smallest.ai's Text-to-Speech solution page.

Where Lightning stands out:

  • TTFB consistently under 100ms for streaming, well below the 300ms conversational threshold. 

  • Character-based pricing with no hidden per-request fees.

  • Voice cloning accessible through the same API without a separate product tier

  • Full voice stack available: STT, TTS, speech-to-speech, and agent platform under one roof

ElevenLabs


ElevenLabs is commonly evaluated for expressive voice generation. Its multilingual v2 and Turbo v2.5 models are commonly used for expressive voice generation and content production workflows, and its voice library includes a large range of voices across different use cases. If you're producing pre-rendered audio for audiobooks, marketing, or creator workflows, voice quality and expressiveness are often prioritized over latency.

ElevenLabs primarily uses a subscription-based pricing structure. The free tier has a low character limit per month. Paid plans offer increasing character allowances, with pay-as-you-go available but not the center of gravity. For high-volume API usage, the cost curve can climb quickly.

If you care most about voice quality for asynchronous content and can treat budget as a secondary constraint, ElevenLabs offers a feature set oriented toward expressive voice generation and content production. If you're building a real-time voice agent at scale, the cost-per-character and latency profile start to look like ongoing friction. The ElevenLabs vs. Smallest.ai comparison breaks down those exact trade-offs across quality, latency, and pricing.

Deepgram Aura


Deepgram's Aura TTS is offered on a pay-as-you-go basis per character, which makes forecasting spend more predictable. This pricing model makes spending relatively predictable as usage scales. One characteristic of the platform is that teams already using Deepgram STT can keep STT and TTS within the same vendor ecosystem.

Aura supports real-time speech applications and streaming use cases. Aura focuses on a more limited set of voices and voice customization options. If your team is already deep in Deepgram's STT ecosystem, Aura is a straightforward option. If latency is the primary evaluation criterion, teams often compare Aura alongside providers that focus specifically on low-latency speech generation.

OpenAI TTS


OpenAI sells two TTS models with different quality and pricing tiers. The standard model is priced per million characters, while the HD model is priced higher for audibly better quality. The higher rate adds up quickly once you're generating real volume. The appeal is operational simplicity: if you're already building on the OpenAI API for language models, TTS slots into the same auth, billing, and tooling without much extra work.

OpenAI TTS is generally evaluated differently from streaming-first voice systems because real-time conversational latency is not its primary focus. For batch narration, offline rendering, or low-frequency voice output where speed isn't the bottleneck, the HD model can make sense. For anything that needs sub-300ms response in an actual back-and-forth, it may be less suitable for applications that prioritize real-time conversational responsiveness.

Cartesia


Cartesia is built on the same premise as Smallest.ai: if you're serving live voice, latency isn't a nice-to-have. Cartesia uses a state-space model architecture meant to reduce TTFB for streaming, and its Sonic model is positioned squarely at real-time voice agents. It offers character-based pricing with a free tier for development and pay-as-you-go rates for production.

The real difference between Cartesia and Smallest.ai isn't the latency thesis, it's scope. Cartesia stays focused on TTS. Smallest.ai wraps TTS into a broader voice stack that includes STT, speech-to-speech, and an agent platform. If you want TTS only and you're benchmarking the low-latency category, Cartesia is another option commonly evaluated in the low-latency TTS category. If you're building an end-to-end voice product, platform breadth becomes a practical consideration, not a marketing bullet.

AssemblyAI


AssemblyAI is best known for speech-to-text, especially since Universal-2 is widely used for speech-to-text workloads. AssemblyAI is primarily known for speech-to-text products, with TTS representing a smaller part of its offering. If your roadmap is STT-first and you want a basic TTS option from the same vendor, it's worth a look. If your workload is TTS-first, it doesn't stand out against the rest of this field.


Head-to-Head: Pricing and Latency at a Glance

Provider

Pricing Model

Cost per 1M Chars

Latency Profile

Voice Cloning

Common Use Cases

Smallest.ai Lightning

Pay-as-you-go with volume tiers

Usage-based pricing (see pricing page)

Very low latency streaming

Yes, via API

Real-time voice agents, full voice stack

ElevenLabs

Subscription-first; PAYG available

Subscription-based pricing

Low latency streaming

Yes (premium tiers)

Expressive content, audiobooks

Deepgram Aura

Pay-as-you-go

Pay-as-you-go per character

Low-latency streaming

Limited

Unified STT+TTS platform users

OpenAI TTS-1

Pay-as-you-go

Per-character pricing

Realtime voice support

No

Low-frequency, OpenAI-integrated apps

OpenAI TTS-1-HD

Pay-as-you-go

Per-character pricing (premium)

Realtime voice support

No

High-quality batch narration

Cartesia Sonic

Pay-as-you-go with tiers

Usage-based pricing (see pricing page)

Very low latency streaming

Yes

Real-time agents, TTS-only stack

AssemblyAI

Pay-as-you-go

See pricing page

Not optimized for low-latency TTS

No

STT-primary teams adding basic TTS

The clearest through-line is architectural: providers designed around streaming-first architectures are the ones that treated streaming as the default from day one. Services that started life as batch TTS and later layered on streaming (OpenAI, ElevenLabs standard) tend to carry extra overhead into TTFB.

Which Provider Fits Which Use Case

If you're building IVR, voice agents, or anything where a person is literally waiting for the system to speak back, start with latency and be ruthless. A provider with TTFB above 300ms will introduce a pause users can hear, even if the voice itself is pristine.

For content production, podcasts, e-learning, audiobooks, you can trade speed for output quality because the audio is rendered offline. In that world, ElevenLabs is commonly evaluated for expressive voice generation and content production workflows. OpenAI's HD model can work if you're already committed to the OpenAI ecosystem and your volume stays moderate. If you're trying to put structure around what "natural" actually means beyond vibes, human-like text-to-speech quality offers a practical way to evaluate providers without getting stuck in spec-sheet theater.

If you're shipping streaming audio and want the engineering context, UX, latency budgets, and what the bill looks like when you scale, the streaming TTS guide for developers pulls those threads together. 

Choosing a TTS API

The right text-to-speech API depends on more than voice quality alone. Teams typically evaluate latency requirements, pricing structure, streaming support, voice customization, deployment complexity, and whether they need a broader voice stack beyond speech synthesis.

For real-time conversational applications, latency often becomes the deciding factor. Voice agents, phone automation, and interactive experiences benefit from streaming-first architectures that can begin generating audio quickly and integrate cleanly with speech recognition and orchestration layers. If you're evaluating providers specifically for real-time workloads, the Fastest Text-to-Speech APIs for Low-Latency Voice Agents guide can provide additional context. 

Another consideration is platform scope. Some providers focus primarily on text-to-speech, while others offer additional capabilities such as speech-to-text, speech-to-speech, voice cloning, and agent infrastructure. The more components involved in a voice application, the more important integration overhead, operational complexity, and vendor management become.

This comparison keeps surfacing the same failure mode: teams discover latency constraints, pricing limitations, or missing capabilities only after a product is already built around a provider. Evaluating those requirements early helps avoid expensive migrations later.

For teams building real-time voice products, Smallest.ai combines Text-to-Speech API, speech-to-text, speech-to-speech, voice cloning, and agent infrastructure within a single platform with clear Pricing Plans. The Lightning API is designed for low-latency speech generation, while the broader platform reduces the need to stitch together multiple voice vendors as requirements expand. 

Frequently asked questions

Frequently asked questions

What pricing models show up most often in text-to-speech APIs?

What TTS latency is realistic for a real-time voice agent?

Does voice cloning change how TTS providers price their APIs?

Is the cheapest TTS API automatically the best choice at high volume?

What should I prioritize in a TTS API when building a full voice product?

Summarize with AI