An efficient multistage Rover method for Automatic Speech recognition
Haihua Xu, Zhu Jie, Guanyong Wu · 2009
In this paper, we implemented a multistage recognizer output voting error reduction (ROVER) method for better automatic speech recognition (ASR). The first stage ROVER is conducted by combining three recognizers, which are respectively trained with maximum likelihood estimation (MLE), minimum phone error (MPE) and recently proposed boosted maximum mutual information (BMMI) criteria. After that the second stage ROVER is performed on two groups of recognizers, which are separately adapted with maximum a posteriori (MAP) and maximum likelihood linear regression (MLLR) methods based on results from the first stage ROVER. It is found MAP adapted recognizers based ROVER does not lead to word error rate (WER) reduction while MLLR adapted recognizers based ROVER is still effective on recognition accuracy improvement. Finally, the third stage ROVER is constructed by combining adapted and unadapted recognizers, and the best WER is obtained by combining MLLR adapted recognizers with BMMI trained recognizers that are not adapted, which achieved 6.0% and 3.0% relative WER reduction, as opposed to the best result from the one-best decoding method and the single stage ROVER method accordingly.