Automatic Speech Recognition
The automated conversion of spoken audio into text, typically combining acoustic modeling, language constraints, decoding, and segmentation logic.
Boundaries
ASR transcribes speech; speaker diarization attributes time regions to speakers. The two tasks can be combined but are not interchangeable.
Related Concepts
- Word Error Rate
- Voice Activity Detection
- Speaker Diarization
- Real-Time Factor