
Dictionary

Dictionary

Dictionary
A
AI Agents
Learn what AI agents are, how they work, their types, and how they differ from agentic AI. A clear technical glossary entry on autonomous AI systems.
AI Song Generator
Learn what an AI song generator is, how it creates music from text or audio inputs, and key concepts behind vocal synthesis and neural composition.
AI Voice
AI voice uses deep learning to generate natural-sounding synthetic speech from text. Learn how AI voice generators, cloning, and changers work.
Artificial Intelligence
Learn what artificial intelligence is, how it works, its core types, and real-world examples. A clear, technical glossary definition of AI fundamentals.
Audio Codec
Learn what an audio codec is, how lossy and lossless compression work, and how codec ICs and software power voice AI, streaming, and telephony.
Automatic Language Detection
Learn what automatic language detection is, how it identifies spoken or written language in real time, and why accuracy varies by input length and context.
A
AI Agents
Learn what AI agents are, how they work, their types, and how they differ from agentic AI. A clear technical glossary entry on autonomous AI systems.
AI Song Generator
Learn what an AI song generator is, how it creates music from text or audio inputs, and key concepts behind vocal synthesis and neural composition.
AI Voice
AI voice uses deep learning to generate natural-sounding synthetic speech from text. Learn how AI voice generators, cloning, and changers work.
Artificial Intelligence
Learn what artificial intelligence is, how it works, its core types, and real-world examples. A clear, technical glossary definition of AI fundamentals.
Audio Codec
Learn what an audio codec is, how lossy and lossless compression work, and how codec ICs and software power voice AI, streaming, and telephony.
Automatic Language Detection
Learn what automatic language detection is, how it identifies spoken or written language in real time, and why accuracy varies by input length and context.
C
Code-Switching
Code-switching is the alternation between languages within a conversation or utterance. Learn how it works in linguistics, voice AI, and speech recognition.
Concurrency
Learn what concurrency means in voice AI and speech APIs, how it differs from parallelism, and why concurrent session limits matter for scalable platforms.
C
Code-Switching
Code-switching is the alternation between languages within a conversation or utterance. Learn how it works in linguistics, voice AI, and speech recognition.
Concurrency
Learn what concurrency means in voice AI and speech APIs, how it differs from parallelism, and why concurrent session limits matter for scalable platforms.
E
End-of-Utterance Detection
Learn how end-of-utterance detection identifies when a speaker finishes talking using acoustic, linguistic, and semantic cues for natural voice AI.
Endpointing
Endpointing detects when a speaker finishes an utterance so speech recognition can finalize results. Learn how VAD, silence thresholds, and tuning work.
Entropy
Entropy measures uncertainty in a probability distribution. Learn how entropy works in AI, information theory, and speech technology with formulas and examples.
E
End-of-Utterance Detection
Learn how end-of-utterance detection identifies when a speaker finishes talking using acoustic, linguistic, and semantic cues for natural voice AI.
Endpointing
Endpointing detects when a speaker finishes an utterance so speech recognition can finalize results. Learn how VAD, silence thresholds, and tuning work.
Entropy
Entropy measures uncertainty in a probability distribution. Learn how entropy works in AI, information theory, and speech technology with formulas and examples.
M
Machine Learning
Learn what machine learning is, how it works, its main types and algorithms, and how ML powers modern speech and voice AI systems.
Mel Spectrogram
Learn what a mel spectrogram is, how it works, and why it is essential for speech recognition, TTS, and voice AI. Includes formula, Python tips, and FAQs.
Model Latency
Model latency is the time between sending input to an ML model and receiving output. Learn how it affects voice AI and how to reduce it.
MOS Score (Mean Opinion Score)
Learn what MOS Score (Mean Opinion Score) means, how it measures voice quality on a 1-to-5 scale, and what counts as a good MOS for VoIP and speech tech.
M
Machine Learning
Learn what machine learning is, how it works, its main types and algorithms, and how ML powers modern speech and voice AI systems.
Mel Spectrogram
Learn what a mel spectrogram is, how it works, and why it is essential for speech recognition, TTS, and voice AI. Includes formula, Python tips, and FAQs.
Model Latency
Model latency is the time between sending input to an ML model and receiving output. Learn how it affects voice AI and how to reduce it.
MOS Score (Mean Opinion Score)
Learn what MOS Score (Mean Opinion Score) means, how it measures voice quality on a 1-to-5 scale, and what counts as a good MOS for VoIP and speech tech.
P
Partial Transcript
A partial transcript is an interim speech recognition result delivered in real time while audio is still being processed. Learn how partials work.
Phoneme
A phoneme is the smallest unit of sound that distinguishes meaning in a language. Learn how phonemes work in linguistics and speech technology.
Prosody
Prosody is the rhythm, stress, intonation, and timing of speech. Learn how prosody works in linguistics, voice AI, and text-to-speech systems.
P
Partial Transcript
A partial transcript is an interim speech recognition result delivered in real time while audio is still being processed. Learn how partials work.
Phoneme
A phoneme is the smallest unit of sound that distinguishes meaning in a language. Learn how phonemes work in linguistics and speech technology.
Prosody
Prosody is the rhythm, stress, intonation, and timing of speech. Learn how prosody works in linguistics, voice AI, and text-to-speech systems.
S
Sample Rate
Learn what sample rate means in audio and speech technology, how it relates to bit depth, and why rates like 16 kHz and 48 kHz matter for voice AI.
Speaker Diarization
Speaker diarization identifies who spoke when in an audio stream. Learn how it works, key methods, real-time use cases, and open-source tools.
Speech to Text
Learn what speech to text is, how ASR technology converts spoken language to written text, and explore key tools, models, and use cases.
Streaming Transcription
Streaming transcription converts speech to text in real time as audio is captured. Learn how it works, key concepts, and top use cases.
S
Sample Rate
Learn what sample rate means in audio and speech technology, how it relates to bit depth, and why rates like 16 kHz and 48 kHz matter for voice AI.
Speaker Diarization
Speaker diarization identifies who spoke when in an audio stream. Learn how it works, key methods, real-time use cases, and open-source tools.
Speech to Text
Learn what speech to text is, how ASR technology converts spoken language to written text, and explore key tools, models, and use cases.
Streaming Transcription
Streaming transcription converts speech to text in real time as audio is captured. Learn how it works, key concepts, and top use cases.
T
Text to Speech
Learn what text to speech (TTS) is, how it works, key synthesis methods, and common use cases. Explore AI and free TTS options for your projects.
Text-to-Speech (TTS)
Learn what TTS (text-to-speech) is, how speech synthesis works, common use cases, and how to choose the best TTS solution for your application.
Transformers
Learn what transformers are in AI: the self-attention architecture behind modern speech recognition, language models, and text-to-speech systems.
Turn Detection
Turn detection determines when a user starts or finishes speaking in voice AI conversations, enabling natural, well-timed agent responses.
T
Text to Speech
Learn what text to speech (TTS) is, how it works, key synthesis methods, and common use cases. Explore AI and free TTS options for your projects.
Text-to-Speech (TTS)
Learn what TTS (text-to-speech) is, how speech synthesis works, common use cases, and how to choose the best TTS solution for your application.
Transformers
Learn what transformers are in AI: the self-attention architecture behind modern speech recognition, language models, and text-to-speech systems.
Turn Detection
Turn detection determines when a user starts or finishes speaking in voice AI conversations, enabling natural, well-timed agent responses.
V
Vocoder
Learn what a vocoder is, how it works, and its role in speech synthesis and music. Covers neural vocoders, vocoder plugins, and key concepts.
Voice Activity Detection (VAD)
Learn what voice activity detection (VAD) is, how it works, and how to implement it in Python with models like Silero VAD and WebRTC VAD.
Voice Changer
Learn what a voice changer is, how real-time and AI voice changers work, common use cases, and key factors like latency, audio quality, and ethical use.
V
Vocoder
Learn what a vocoder is, how it works, and its role in speech synthesis and music. Covers neural vocoders, vocoder plugins, and key concepts.
Voice Activity Detection (VAD)
Learn what voice activity detection (VAD) is, how it works, and how to implement it in Python with models like Silero VAD and WebRTC VAD.
Voice Changer
Learn what a voice changer is, how real-time and AI voice changers work, common use cases, and key factors like latency, audio quality, and ethical use.
W
Word Error Rate (WER)
Learn what word error rate (WER) is, how to calculate it with the WER formula, and what counts as a good WER for speech recognition systems.
Word-Level Timestamps
Learn what word-level timestamps are, how ASR systems like Whisper generate them, and their uses in subtitles, audio editing, and translation alignment.
W
Word Error Rate (WER)
Learn what word error rate (WER) is, how to calculate it with the WER formula, and what counts as a good WER for speech recognition systems.
Word-Level Timestamps
Learn what word-level timestamps are, how ASR systems like Whisper generate them, and their uses in subtitles, audio editing, and translation alignment.
Build the future of voice agent orchestration
311 California Street, Suite 320
San Francisco, CA 94104
Documentation
Resources
Initiatives
Build the future of voice agent orchestration
311 California Street, Suite 320
San Francisco, CA 94104
Build the future of voice agent orchestration
311 California Street, Suite 320
San Francisco, CA 94104
Documentation
Resources
Initiatives
