Announcing our Series A Funding

Announcing our Series A Funding

Meet Hydra.

Our full duplex speech to speech model

A new generation of speech-to-speech AI that

remembers, reasons, and responds like a human.

Introduction to Hydra

Hydra is Smallest AI's speech-native model for the most demanding real-time conversations. Built to reason continuously across long-running context, overlapping speech, interruptions, and tool-driven workflows while delivering fluid, low-latency responses.

Why voice agents stall

Most production voice agents use a cascaded pipeline. Every response passes through each stage. Latency accumulates across transcription, inference, tool execution, and speech generation. Turn detection adds another failure point around interruptions, pauses, and background speech.

Conventional cascade approach

STT

LLM

TTS

Example conversation: using standard cascaded approach

Context decays over long calls

Summaries preserve the broad state of a conversation, but can lose details such as:

• Names and identifiers

• Dates and amounts

• User corrections

• Confirmation state

• Values returned by earlier tools

• Conditions introduced much earlier

System overview

Hydra replaces the conventional ASR → language model → text-to-speech cascade with a speech-native pipeline that remains in a continuous latent space. A unified conversational tokenizer converts incoming audio into ~50 ms latent frames that preserve semantics, prosody, timing, speaker characteristics, and emotional cues for the model to reason over.

As each causal audio frame arrives, Hydra continuously updates its internal conversational state and begins forming possible response trajectories before the speaker finishes. The latent conversational transformer refines those trajectories while listening, and a conditional decoder produces speech or text directly from the resulting state—supporting lower-latency turn-taking, natural backchannels, overlap, and fluid interruption handling without an intermediate transcript.

What Hydra handles

Hydra is optimized around three requirements of deployed voice systems: overlap-aware interaction, durable conversational state, and reliable tool execution.

In-session tool use

Hydra is evaluated on the complete tool-use loop:

• Selection: Choose the correct tool from a catalogue containing hard distractors.

• Arguments: Extract the expected values from speech.

• Deferred reuse: Reuse values returned by previous calls.

• Precision and recall: Avoid both unnecessary calls and missed required calls.

• Guardrails: Reject forbidden or premature operations.

• Clarification: Ask for missing information instead of guessing.

• Recovery: Handle tool errors and empty results without fabricating an answer.

• Spoken accuracy: State the retrieved facts correctly in the final response.

These capabilities allow Hydra to manage multi-step workflows without requiring the caller to repeat information between operations.

Tool use accuracy across 60 turns

HydraSmallest.ai
0%
Grok RealtimeSpaceXAI
0%
GPT RealtimeOpen AI
0%
Gemini 2.5Google DeepMind
0%

ComplexFuncBenchmark (Single-turn tool calling)

HydraSmallest.ai
0%
Claude 3.5 SonnetAnthropic
0%
GPT 4oOpenAI
0%
GLM 4 LongZhipuAI
0%
GPT Realtime 1.5OpenAI
0%

Long-session state

Hydra tracks information across extended conversations and makes it available when required later. This includes:

• Names and identifiers

• Dates, times, and amounts

• Earlier corrections

• Accepted and rejected actions

• User preferences

• Values returned by tools

The goal is exact retention across turns. A value introduced early in a call should remain available for a decision or tool call much later.

Overlapping speech

Hydra is trained and evaluated on interaction states such as:

• Direct user interruptions

• Backchannels

• Speech directed at another person

• Background speech

• Mid-sentence corrections

• Input received during a response

The model must determine whether to stop, continue, yield, or resume without discarding the current state.

Benchmarks

We evaluate Hydra across interaction quality, speech reasoning, tool use, long-context retention, and end-to-end task completion.

AIEWF

This test evaluates tool use, instruction following, knowledge-base grounding, turn-taking, and latency across 300 tasks.

Model

Hydra

GPT Realtime

GPT Realtime 1.5

Ultravox v0.7

Tool V2V mean

1624ms

2199ms

2251ms

2406ms

Non-tool

V2V median

864ms

1536ms

1152ms

864ms

Pass Rate

95.9%

86.7%

93.3%

97.7%

τ-Voice

This test evaluates voice agents on retail,

airline, and telecom support tasks.

Model

Grok Voice Fast 1.0

Hydra

GPT Realtime 1.5

GPT Realtime 1.0

Gemini Live 2.5 Flash

Average pass^1

38.3

35.6

35.3

30.4

25.8

Average Latency

1.15s

1.3ms

1.39s

1.53s

1.43

Audio MultiChallenge

This test measures reasoning and instruction following directly from audio.

Model

GPT Realtime 2

GPT Realtime 1.5

Hydra

Gemini 3.1 Flash

Qwen 3 Omni 30B

Average Pass Rate

37.61%

34.73

32.52%

26.77%

24.24%

AIEWF

This test evaluates tool use, instruction following, knowledge-base grounding, turn-taking, and latency across 300 tasks.

Model

Hydra

GPT Realtime

GPT Realtime 1.5

Ultravox v0.7

Tool V2V mean

1624ms

2199ms

2251ms

2406ms

Non-tool

V2V median

864ms

1536ms

1152ms

864ms

Pass Rate

95.9%

86.7%

93.3%

97.7%

τ-Voice

This test evaluates voice agents on retail,

airline, and telecom support tasks.

Model

Grok Voice Fast 1.0

Hydra

GPT Realtime 1.5

GPT Realtime 1.0

Gemini Live 2.5 Flash

Average pass^1

38.3

35.6

35.3

30.4

25.8

Average Latency

1.15s

1.3ms

1.39s

1.53s

1.43

Audio MultiChallenge

This test measures reasoning and instruction following directly from audio.

Model

GPT Realtime 2

GPT Realtime 1.5

Hydra

Gemini 3.1 Flash

Qwen 3 Omni 30B

Average Pass Rate

37.61%

34.73

32.52%

26.77%

24.24%

AIEWF

This test evaluates tool use, instruction following, knowledge-base grounding, turn-taking, and latency across 300 tasks.

Model

Hydra

GPT Realtime

GPT Realtime 1.5

Ultravox v0.7

Tool V2V mean

1624ms

2199ms

2251ms

2406ms

Non-tool

V2V median

864ms

1536ms

1152ms

864ms

Pass Rate

95.9%

86.7%

93.3%

97.7%

τ-Voice

This test evaluates voice agents on retail,

airline, and telecom support tasks.


Model

Grok Voice Fast 1.0

Hydra

GPT Realtime 1.5

GPT Realtime 1.0

Gemini Live 2.5 Flash

Average pass^1

38.3

35.6

35.3

30.4

25.8

Average Latency

1.15s

1.3ms

1.39s

1.53s

1.43

Audio MultiChallenge

This test measures reasoning and instruction following directly from audio.


Model

GPT Realtime 2

GPT Realtime 1.5

Hydra

Gemini 3.1 Flash

Qwen 3 Omni 30B

Average Pass Rate

37.61%

34.73

32.52%

26.77%

24.24%

Hydra in real conversations

Benchmarks isolate individual capabilities. The recordings below show how those capabilities combine inside complete conversations, when callers interrupt, revise information, wait for tools, or refer back to something said much earlier.

Handle an interruption

The caller interrupts a response with a correction. Hydra yields, incorporates the corrected value, and resumes.

Patient intake with delayed recall

A patient provides an allergy early in the call. Hydra recalls it later before completing the scheduling workflow.

Complete and order return

Hydra retrieves an order, checks eligibility, asks for confirmation, and initiates the return.

Handle an interruption

The caller interrupts a response with a correction. Hydra yields, incorporates the corrected value, and resumes.

Patient intake with delayed recall

A patient provides an allergy early in the call. Hydra recalls it later before completing the scheduling workflow.

Complete and order return

Hydra retrieves an order, checks eligibility, asks for confirmation, and initiates the return.

Handle an interruption

The caller interrupts a response with a correction. Hydra yields, incorporates the corrected value, and resumes.

Patient intake with delayed recall

A patient provides an allergy early in the call. Hydra recalls it later before completing the scheduling workflow.

Complete and order return

Hydra retrieves an order, checks eligibility, asks for confirmation, and initiates the return.

Limitations

The evaluation results show that Hydra can complete long, tool-heavy voice workflows, but they also reveal where its behavior is not yet consistent enough.

Uneven domain performance

Tau Voice results vary significantly by domain. Hydra ranks third in telecom, fifth in airline, and tenth in retail. These gaps require domain-specific error analysis across tool selection, scenario coverage, and conversation policy.

Overlap classification

Hydra performs better on direct interruptions and backchannels than on speech directed at another person or background speakers. These cases require the model to determine both what was said and whether it belongs to the active conversation.

Complex word handling

Hydra can be inconsistent with rare names, domain-specific terms, acronyms, and alphanumeric identifiers. Providing spelling or contextual guidance can help, but there is not yet a reliable way to enforce an exact pronunciation. This can limit customization for brands, medication names, locations, and other specialized vocabulary.

Uneven domain performance

Tau Voice results vary significantly by domain. Hydra ranks third in telecom, fifth in airline, and tenth in retail. These gaps require domain-specific error analysis across tool selection, scenario coverage, and conversation policy.

Overlap classification

Hydra performs better on direct interruptions and backchannels than on speech directed at another person or background speakers. These cases require the model to determine both what was said and whether it belongs to the active conversation.

Complex word handling

Hydra can be inconsistent with rare names, domain-specific terms, acronyms, and alphanumeric identifiers. Providing spelling or contextual guidance can help, but there is not yet a reliable way to enforce an exact pronunciation. This can limit customization for brands, medication names, locations, and other specialized vocabulary.

Uneven domain performance

Tau Voice results vary significantly by domain. Hydra ranks third in telecom, fifth in airline, and tenth in retail. These gaps require domain-specific error analysis across tool selection, scenario coverage, and conversation policy.

Overlap classification

Hydra performs better on direct interruptions and backchannels than on speech directed at another person or background speakers. These cases require the model to determine both what was said and whether it belongs to the active conversation.

Complex word handling

Hydra can be inconsistent with rare names, domain-specific terms, acronyms, and alphanumeric identifiers. Providing spelling or contextual guidance can help, but there is not yet a reliable way to enforce an exact pronunciation. This can limit customization for brands, medication names, locations, and other specialized vocabulary.

Request access

Hydra is currently available through a limited beta. Book a call with our team, to discuss your use case and next steps. He will guide you through evaluation, integration, and onboarding.

Hydra is currently in beta

beta access only

beta access only

beta access only

beta access only

beta access only