US 12,394,409 B2
Separating acoustic and linguistic information in neural transducer models for end-to-end speech recognition
Gakuto Kurata, Tokyo (JP)
Assigned to INTERNATIONAL BUSINESS MACHINES CORPORATION, Armonk, NY (US)
Filed by INTERNATIONAL BUSINESS MACHINES CORPORATION, Armonk, NY (US)
Filed on Sep. 17, 2021, as Appl. No. 17/478,064.
Prior Publication US 2023/0104244 A1, Apr. 6, 2023
Int. Cl. G10L 15/16 (2006.01); G10L 15/06 (2013.01)
CPC G10L 15/16 (2013.01) [G10L 15/063 (2013.01)] 18 Claims
OG exemplary drawing
 
1. A computer-implemented method for training a Recurrent Neural Network Transducer (RNN-T), the method comprising:
training, by inputting a set of audio data, a first RNN-T which comprises a common encoder, a forward prediction network, and a first joint network combining outputs of both the common encoder and the forward prediction network, wherein the forward prediction network predicts label sequences forward, wherein training the first RNN-T includes forming a first output probability lattice from the set of audio data and an output of the first RNN-T with the set of audio data along an x-axis and the output of the first RNN-T along a y-axis and computing a RNN-T loss on the first output probability matrix in an upper right direction; and
training, by inputting the set of audio data, a second RNN-T which comprises the common encoder, a backward prediction network, and a second joint network combining outputs of both the common encoder and the backward prediction network, wherein the backward prediction network predicts label sequences backward, and wherein the trained first RNN-T is used for inference.