Comparative Study on Smoothing Algorithms for Domain-Specific Chinese Language Models

Yonghong Yan · Computer Engineering and Applications Journal · 2006

It is important to build a powerful language model by using limited corpora in the field of speech recognition for a specific domain.To deal with this problem,two methods concerning how to process new words with high frequencies in a specific domain are presented.One way is to add the new words to the dictionary directly and then give them a high weight in the procedure of training.The other is to work out a new dictionary according to the new words.And based on some comparative experiments,these two methods and various smoothing algorithms are studied in detail.At last,it can be concluded that the performance of language model is affected by the smoothing algorithm greatly,and the Witten-Bell interpolation method could improve the recognition rate to 88.4%,which is 18.18% higher than the general language model.

Read the paper · More papers on PaperTik