Stable-time prediction during incremental speech recognition
Xiao Chen, Bo Xu · 2016
Incremental speech recognition (ISR) is the key technology to increase the efficiency of human-computer interaction and obtain the good user experience. However, ISR's results are not stable. The earlier methods decide whether to output the current best partial result by stable-time. These methods can increase stability, but introduce a lag which lowers the advantage of user experience. In this paper, a new method based on the stable-time prediction is proposed. It predicts the probable stable-time of the current best partial result in the future, using the acoustic score information of N-best paths of successive frames. So it can determine whether to output the current best partial result in advance. It can reduce lags and improve performance. Results indicate that, the proposed method outperforms the baseline. At the lag of 0.2s, the proposed method results in an absolute improvement of 1.2% and achieves a stability of 92.2%. And at the other lags, the proposed method also results in a similar improvement.