Bi-directional recurrent neural networks for speech recognition
Mike Schuster · 1996
While many possible network architectures have been used to estimate conditional probabilities of class membership, recurrent neural networks (RNNs) have been most successful for speech recognition. In the past optimal results were achieved by merging the outputs of two RNNs trained in each time direction. Merging outputs of different experts to form one resulting opinion is theoretically difficult - a direct combination of the experts during training would be desirable. This paper presents a bi-directional neural network (BRNN) structure which can be trained in both time directions simultanously and hence avoids the difficult merging process. 1. INTRODUCTION Almost all current large vocabulary speech recognition systems are based on Hidden Markov Models (HMMs) with parametric observation density distributions. As the kernels for the distributions usually Gaussians are chosen. The parameters of the distributions can easily be estimated with maximum likelihood methods. For optimal resu...