Forced Alignment
The estimation of time boundaries for known transcript units by aligning provided text with the corresponding audio signal.
Speech Processing Context
Forced alignment estimates time boundaries for a transcript that is already known. Acoustic-model posteriors, CTC scores, or specialized alignment algorithms can produce word or phoneme start and end times for corpus preparation, subtitles, and phonetic analysis.
Modeling Boundary
Forced alignment is not free-form speech recognition. If the supplied transcript is wrong or strongly mismatched to the audio, the resulting timestamps can also be unreliable.