Random-forests-based phonetic decision trees for conversational speech recognition
Jian Xue, Yunxin Zhao · IEEE International Conference on Acoustics Speech and Signal Processing · 2008
In this paper we present a novel technique of constructing phonetic decision trees (PDTs) for acoustic modeling in conversational speech recognition. We use random forests (RF) to train a set of PDTs for each phone-state unit and obtain multiple acoustic models accordingly, and we extend the PDT-based state tying to RF-based state-tying. We combine acoustic scores at the model level in decoding search. Several methods are investigated to estimate the weight parameters for model combination, including maximum likelihood estimation of the weights from training data, as well as using confidence scores of P-value or relative entropy to obtain the weights dynamically from online data. Experimental results on a telemedicine automatic captioning task demonstrate that the proposed RF-PDT technique leads to significant improvements in word recognition accuracy.