Meet Hydra,
our full duplex model.
A new generation of speech-to-speech AI that
remembers, reasons, and responds like a human.
We fixed the most annoying thing about older AI voice models: the waiting. Hydra lets you talk to it like an actual human conversation.
No taking turns. You can talk while it's talking, cut in mid-sentence, change your mind halfway through and it just keeps up, the way a person would. It's not sitting there waiting for you to finish before it starts thinking.
It also doesn't forget. Long calls stay long calls Hydra remembers everything said from the start, so you're never repeating yourself or explaining something you already mentioned ten minutes ago. And when the conversation needs to actually get something done checking a booking, pulling up an order, looking something up Hydra can go do that right in the middle of talking to you, without going quiet or making you wait around for it.
In the end it doesn't feel like using a voice assistant. It feels like talking to someone who's actually paying attention. Currently in beta, and already showing up in conversations where that kind of attention actually matters.
Prefer to wait? Join the waitlist
The cost of the cascading approach.
Context leaks
The real issue is that no single model ever holds the whole conversation.
The speech-to-text model only hears the sentence you just said. The LLM only sees the bit of text handed to it for that turn. Nobody's tracking the full conversation from start to finish. So the longer a call runs, the more gets trimmed or dropped just to keep things moving. Things you said early on disappear fast.
When speech moves between models
Speech to text
Large language model
Text to speech
what leaks out at each handoff
tone, emphasis
what leaks out at each handoff
Earlier details
No real memory
The real issue is that no single model ever holds the whole conversation.
The speech-to-text model only hears the sentence you just said. The LLM only sees the bit of text handed to it for that turn. Nobody's tracking the full conversation from start to finish. So the longer a call runs, the more gets trimmed or dropped just to keep things moving. Things you said early on disappear fast.
A conversation, left to right over time
start of call
Early details faded and dropped
1 hour in
Stalls when it matters
Every turn of the conversation has to travel through all three models before you hear anything back, your voice into text, text into a reply, reply back into speech. That adds up to real latency, and it happens on every single turn, not just once. The result is that pause before it answers, or a full freeze if any one step runs slow.
Play
Conventional cascade approach
STT
LLM
TTS
Example conversation: using standard cascaded approach
Our approach
Hydra is built on two changes that solve the problems above at the source, not around them.
Hydra isn't three models passed in a relay, it's one model that hears and speaks at the same time, over a single connection. There's no handoff from a listening step to a thinking step to a speaking step, because there's no separate steps to begin with. It's all happening in the same place, continuously, the whole time you're talking.
One continuous model, not three stitched together
Older voice AI works by passing you through separate models in sequence — one to listen, one to think, one to speak. Each handoff is a place where something can get lost.
Hydra is built as one continuous model instead, meaning there's no handoff for anything to fall through. It stays present for the whole conversation, start to finish, so context isn't passed along and quietly dropped along the way. It's remembered because it was never split apart to begin with.
New approach
(SALM)
ASR + LLM, merged · audio in
Conv-Text to speech
Speaks, and listens back
One speech-augmented model hears and reasons directly; a conversational TTS speaks and listens back.
Asynchronous approach, listening and speaking, at once
Older systems wait for you to finish talking before they start working, and then work through their steps one at a time.
Hydra is built on a synchronous architecture, meaning it processes what you're saying while you're still saying it. Listening and responding happen in the same moment, not one after another. There's no delay stacking up between steps, because there are no separate steps, just one continuous, real-time exchange.
Play
Prefer to wait? Join the waitlist
Meet Hydra
Watch it actually do the work
Most voice assistants can only talk. Hydra can act. This is tool calling reaching into real apps mid-conversation to actually get something done. Ask it to book a flight, add a note to Notion, or put an itinerary together, and it handles the whole thing without app-switching, without "I can't help with that," and without dropping the thread while it works.
Where Hydra fits
Most voice assistants can only talk. Hydra can act. This is tool calling reaching into real apps mid-conversation to actually get something done. Ask it to book a flight, add a note to Notion, or put an itinerary together, and it handles the whole thing without app-switching, without "I can't help with that," and without dropping the thread while it works.
Healthcare
Financial services
Customer support
Ecommerce
Alfred
Calm, highly capable assistant
who offers thoughtful advice.
Hi, I'm calling about my blood test results from last week
"Of course — I've got your file here. Dr. Bennett noted they'd be ready this Tuesday. Want me to book your follow-up now?"
Pause