Pocket captures conversations wherever they happen. That means its transcription layer has to deal with something much harder than clean meeting-room audio.
Transcript
Diversity
Speed
After 3 months of trialling 10+ speech-to-text models, Pocket realised Smallest's Pulse was the only models capable of handling real world usecases.
Products like Pocket are defined by how reliable they can be in everyday environments.
For a speech-to-text model to handle the everyday, they need to:
Ability to handle big chunks of audio, and generate transcripts super fast
Differentiate real conversations from noise of everyday environments
Supports true far-field audio capability
Understand conversational nuances like emotion, dialects & accents
Follow a conversation even when a human switches languages
Big things take small time
Every meeting ends with "Hey, can you share the moments of the meeting". If the transcription, summarization and everything else takes longer than 1 minutes we are too slow for a fast moving environment.
Pulse can handle a large amount of audio data and generate accurate transcripts at record speed.
Let there be noise
Pulse not only leads in Word Error Rates (WER) across all other models but it’s real strength is to perform when it’s noisier than the carnival in Rio.
Hello from the other side.
Rooms have their own roles. You walk into a room and start talking without thinking about room acoustics, reflections of your audio and the inverse-square law. But your note-taker still cares about what you’re saying. Pulse works phenomenally well in these rooms. Pulse cares about every word you say.
It’s not what you say. It’s how you say it.
Every human intrinsically picks an emotion for every word, says the same word in multiple different ways and even pronounce it differently at different times of the day. A good model would know how to differentiate.
Et tu, bilingual?
Great conversations don’t have a script, they flow however they want to, from one language to another. Pulse not only support 70+ languages but it also understand when you code switch from “Ma ma mia” to “Another one!” while you enjoy that spicy margarita.















