RNN-T
Recurrent Neural Network Transducer, an end-to-end streaming ASR architecture that jointly models acoustic input and output-token history through a transducer objective.
Speech Processing Context
RNN-T, or the Recurrent Neural Network Transducer, combines an acoustic encoder with a prediction network and a joint network so output-token dependencies can be modeled while supporting streaming decoding. Its blank emissions and beam-search state management are important to online recognition latency.
Modeling Boundary
The name does not require every modern transducer encoder to be recurrent. Transformer or conformer encoders can preserve the transducer prediction-and-joint formulation while changing the acoustic encoder architecture.