Pulse™

Speech to text that keeps up with the conversation

Turn live speech, calls, audio and video into structured text with 64ms model latency, 38+ languages, speaker labels and speech intelligence.

Click anywhere to start transcribing

Experience pulse speech to text

Pulse™

Speech to text that keeps up with the conversation

Turn live speech, calls, audio and video into structured text with 64ms model latency, 38+ languages, speaker labels and speech intelligence.

Click anywhere to start transcribing

Experience pulse speech to text

Pulse™

Speech to text that keeps up with the conversation

Turn live speech, calls, audio and video into structured text with 64ms model latency, 38+ languages, speaker labels and speech intelligence.

Click anywhere to start transcribing

Experience pulse speech to text

What is Speech-to-text?

Speech to text converts spoken audio into written words that can be searched, analyzed, displayed as captions or used by an application.

Built for a useful next step, not just a response.

Built for a useful next step, not just a response.

Transcription that keeps a finger on the Pulse™

Go beyond text with automated speaker labeling, real-time sentiment analysis, and language identification for global workloads

Transcription that keeps a finger on the Pulse™

Go beyond text with automated speaker labeling, real-time sentiment analysis, and language identification for global workloads

World’s Most Advanced Speech Intelligence

Go beyond text with automated speaker labeling, real-time sentiment analysis, and language identification for global workloads

Emotion Recognition

Detects user emotions to make conversations more empathetic

Industry-Leading acccuracy and speed

Pulse STT outperforms the competition with the lowest Word Error Rates across 30+ languages and sub-70ms latency for seamless, real-time conversations.

Context Switching

Pulse adapts to healthcare, legal, finance, media, customer support, and enterprise.

Context Switching

Pulse adapts to healthcare, legal, finance, media, customer support, and enterprise.

Context Switching

Pulse adapts to healthcare, legal, finance, media, customer support, and enterprise.

Transcribe in 38+ languages

Delivering exceptional accuracy across accents, dialects, and recording conditions.

Transcribe in 38+ languages

Delivering exceptional accuracy across accents, dialects, and recording conditions.

Transcribe in 38+ languages

Delivering exceptional accuracy across accents, dialects, and recording conditions.

Speaker diarization for clarity

Identify transitions between speakers and accurately label each contribution

Speaker diarization for clarity

Identify transitions between speakers and accurately label each contribution

Speaker diarization for clarity

Identify transitions between speakers and accurately label each contribution

Follow every capability to its source

Review implementation details and evidence for accuracy, intelligence, language handling, speakers, redaction, and model choice.

Accuracy and latency

See the evaluation methodology, latency measurements, and accuracy results.

Emotion detection

Go beyond words by returning emotional context with the transcript.

Language detection

Automatically identify the language in incoming audio before downstream processing.

Speaker diarization

Separate and label speakers in multi-speaker audio.

Sensitive-data redaction

Handle sensitive information in transcripts before it moves into downstream workflows.

Pulse

Review the model card for current real-time use cases, capabilities, and release details.

Pulse Pro

Review the model card for current high-accuracy recorded-audio positioning.

Our models

Speed or accuracy? you don't have to guess.

Our models

Speed or accuracy? you don't have to guess.

Our models

World’s Most Advanced Speech Intelligence

Pulse

Best for live audio, voice agents, and streaming transcription.

Live streaming & recorded

64ms ultra-low latency

38+ languages

~$0.005/min

Newly launched

Pulse Pro

Built for the highest transcription accuracy on recorded audio

Pre - recorded audio only

Max accuracy in transcription

Available only in English

~$0.004/min

Calculate your costs based on your usage needs

Select model

Select a model above to see pricing

Recommended plan

Pay as you go

Select a feature

No monthly commitments

Scale with your usage

Access to all features

Start building

Calculate your costs based on your usage needs

Select model

Select a model above to see pricing

Recommended plan

Pay as you go

Select a feature

No monthly commitments

Scale with your usage

Access to all features

Start building

Calculate your costs based on your usage needs

Select model

Select a model above to see pricing

Recommended plan

Pay as you go

Select a feature

No monthly commitments

Scale with your usage

Access to all features

Start building

Turn conversations into useful data

Build voice agents, captions, meeting notes, subtitles, call analytics, documentation, search and accessibility tools.

API

SDK

Visual tools

Start with a working request and API key setup in the Speech-to-Text quickstart

Need a full implementation path? Follow the Speech-to-Text API integration guide for Python, Node, and streaming.

01const res = await fetch(02 "https://api.smallest.ai/waves/v1/lightning-v3.1/get_speech",03 {04 method: "POST",05 headers: {06 Authorization: "Bearer YOUR_API_KEY",07 "Content-Type": "application/json",08 },09 body: JSON.stringify({10 text: "Modern problems require modern solutions.",11 voice_id: "magnus",12 sample_rate: 44100,13 output_format: "wav",14 }),15 },16);17 18writeFileSync("output.wav", Buffer.from(await res.arrayBuffer()));19console.log("Saved to output.wav");
01const res = await fetch(02 "https://api.smallest.ai/waves/v1/lightning-v3.1/get_speech",03 {04 method: "POST",05 headers: {06 Authorization: "Bearer YOUR_API_KEY",07 "Content-Type": "application/json",08 },09 body: JSON.stringify({10 text: "Modern problems require modern solutions.",11 voice_id: "magnus",12 sample_rate: 44100,13 output_format: "wav",14 }),15 },16);17 18writeFileSync("output.wav", Buffer.from(await res.arrayBuffer()));19console.log("Saved to output.wav");

Audio in. Structured text out.

Build real-time pipelines with Node and Python SDKs, without managing speech infrastructure or middleware.

  1. Node SDK

    Real-time pipelines

  2. Python SDK

    Backend workflows

  3. Live streams

    Voice agents and calls

  4. Recorded audio

    Media and archives

Speech intelligence for the next workflow

Follow the product or solution path that matches what you need to put into production.

Voice AI

Stream transcripts into conversational systems that need fast turn-taking.

AI notetakers

Combine transcription, speaker separation, and timestamps for searchable meeting records.

Call centers

Transcribe conversations for QA, analytics, and automation.

Planning a larger rollout? See speech to text for call centers at scale

Pay for the minutes you use

Choose the right model, scale with usage and estimate costs through the on-page calculator.

Choose a model

Pulse or Pulse Pro

Set your usage

Minutes, not seats

Estimate cost

Before production

Scale as needed

No fixed workflow

Certified & Compliant

We listen and we don't tell.

Certified & Compliant

We listen and we don't tell.

Proactive Defense

Anticipating threats before they emerge, thanks to our advanced monitoring.

Proactive Defense

Anticipating threats before they emerge, thanks to our advanced monitoring.

Proactive Defense

Anticipating threats before they emerge, thanks to our advanced monitoring.

Frequently
asked questions

Can I try speech to text online?

Does Pulse work in real time?

Can it identify different speakers?

How many languages are supported?

Can I use Pulse for voice agents?