Announcing our Series A Funding

Announcing our Series A Funding

Congrats, Pocket -
The AI notetaker to reach $100M ARR in
7 months, powered by Smallest's Pulse STT

Congrats, Pocket.
The AI notetaker to reach $100M ARR in 7 months, powered by Smallest's Pulse

Pocket is a physical notetaker that captures every conversation. As of today, Pocket has sold over 274,098 note taking devices. They are loved by a variety of users from sales operators to therapists equally.


We played a small part in Pocket reaching this milestone.
Our speech to text model, Pulse powers Pocket to generate transcripts.

Pocket is a physical notetaker that lives on the back of 274,098 phones.


It started as an honest attempt to capture every conversation, anywhere they happened. We played a small (pun intended) part reaching this milestone.

Congrats to

Congrats to

Akshay,

Akshay,

Gabriel,

Gabriel,

Mihir,

Mihir,

and the team.

and the team.

Capture your mind in motion.

Take action without lifting a finger.

Capture your mind in motion.

Take action without lifting a finger.

Pocket is on Pulse.

Pocket is on Pulse.

Pocket captures conversations wherever they happen. That means its transcription layer has to deal with something much harder than clean meeting-room audio.

Transcript

2,070,360,621 minutes of conversations per year

2,070,360,621 minutes of conversations per year

2,070,360,621 minutes of conversations per year

Diversity

23 Languages including Spanish & Mandarin in 27 Accents & more..

23 Languages including Spanish & Mandarin in 27 Accents & more..

23 Languages including Spanish & Mandarin in 27 Accents & more..

Speed

an average 2-hour meeting takes just 23 seconds to transcribe

an average 2-hour meeting takes just 23 seconds to transcribe

an average 2-hour meeting takes just 23 seconds to transcribe

After 3 months of trialling 10+ speech-to-text models, Pocket realised Smallest's Pulse was the only models capable of handling real world usecases.

What does built-for-real-world mean?

What does built-for-real-world mean?

Products like Pocket are defined by how reliable they can be in everyday environments.

For a speech-to-text model to handle the everyday, they need to:


  1. Ability to handle big chunks of audio, and generate transcripts super fast

  2. Differentiate real conversations from noise of everyday environments

  3. Supports true far-field audio capability

  4. Understand conversational nuances like emotion, dialects & accents

  5. Follow a conversation even when a human switches languages

Big things take small time

Every meeting ends with "Hey, can you share the moments of the meeting". If the transcription, summarization and everything else takes longer than 1 minutes we are too slow for a fast moving environment.


Pulse can handle a large amount of audio data and generate accurate transcripts at record speed.

Let there be noise

Pulse not only leads in Word Error Rates (WER) across all other models but it’s real strength is to perform when it’s noisier than the carnival in Rio.

Hello from the other side.

Rooms have their own roles. You walk into a room and start talking without thinking about room acoustics, reflections of your audio and the inverse-square law. But your note-taker still cares about what you’re saying. Pulse works phenomenally well in these rooms. Pulse cares about every word you say.

It’s not what you say. It’s how you say it.

Every human intrinsically picks an emotion for every word, says the same word in multiple different ways and even pronounce it differently at different times of the day. A good model would know how to differentiate.

Et tu, bilingual?

Great conversations don’t have a script, they flow however they want to, from one language to another. Pulse not only support 70+ languages but it also understand when you code switch from “Ma ma mia” to “Another one!” while you enjoy that spicy margarita.

Benchmarks

Smallest pulse accuracy is measured on real world datasets.

While Pulse is great in all these use cases there are obvious reasons why we need to be on the top of our game when it comes to benchmarks.

Benchmarks

All Extraordinary, also needs to be Ordinary.

Word Error Rate

Lowest error rates in the industry as compared to Mistral, Deepgram, Assembly or Gemini models

For every 100 words you speak,
we might make an error with 2 of them.

COST Factor

3x cheaper than Elevenlabs with zero compromise in accuracy

Best in accuracy,
Best in speed to transcript.
While being 3x cheaper than Elevenlabs

The real world is the benchmark.

Pocket wasn’t built for perfect microphones, quiet rooms or carefully pronounced sentences. It was built to capture conversations wherever life happens.


That’s the same world we built Pulse for.


274,098 Pockets later, tens of millions of minutes have passed through Pulse, across crowded rooms, distant voices, accents, dialects and languages switching mid-sentence.

Pocket is proving that physical AI can become an everyday habit.


We’re proud to be the part underneath it that listens.

Fin.