Voice Activity Detection
The detection of speech versus non-speech regions in an audio stream to drive segmentation, endpointing and downstream ASR workload.
Voice Activity Detection (VAD) identifies speech and non-speech regions in an audio stream. Implementations range from energy-based rules to neural classifiers.
Why It Matters for ASR
VAD does more than remove silence. Speech start/end decisions determine segment length and therefore affect:
- available linguistic context,
- inference duration,
- queue load,
- user-visible latency.
Hangover and Gap Merging
A short end-of-speech hangover can reduce clipped word endings but increases endpoint latency. Merging short gaps reduces segment count, while aggressive merging can create unnecessarily long chunks.
Boundary
VAD does not identify the speaker or transcript content. Noise, music, breathing and channel conditions can affect the speech/non-speech decision, so thresholds should be evaluated on the target acoustic domain.
Applied context: Developing a VAD Library, Whisper Architecture.
Related article: Audio Segmentation
Direct source: The primary paper or official specification for Voice Activity Detection is linked here for verification.