RNN-T

Turkish equivalent: Tekrarlayan sinir ağı dönüştürücüsüDomain: Speech Processing

Recurrent Neural Network Transducer, an end-to-end streaming ASR architecture that jointly models acoustic input and output-token history through a transducer objective.

Speech Processing Context

RNN-T, or the Recurrent Neural Network Transducer, combines an acoustic encoder with a prediction network and a joint network so output-token dependencies can be modeled while supporting streaming decoding. Its blank emissions and beam-search state management are important to online recognition latency.

Modeling Boundary

The name does not require every modern transducer encoder to be recurrent. Transformer or conformer encoders can preserve the transducer prediction-and-joint formulation while changing the acoustic encoder architecture.

Related technical publications

Publications whose title or summary directly references this concept.

Automatic Speech Recognition

Automatic speech recognition from isolated-word, word-spotting and continuous-speech tasks through DTW, HMM/WFST and n-gram systems to CTC, RNN-T, Transformer and Conformer models.