A smoothing algorithm for the task adaptation Chinese Trigram model
Jiang Minghu, Baozong Yuan, Lin Biqin, Tang Xiaofang · 2002
The paper mainly solves two problems. A Chinese trigram model of task adaptation ability is set up. A zerogram to trigram probability statistics information base of the 1994 "People Daily" are built, making use of the successful experience of using HMM in speech recognition, and the adopted Baum-Welch algorithm for optimisation of the weights. Each weight stands for the correlation statistic reliability of these models. The probability statistics matrix smoothing algorithm of the parameter space is carried out, in order to offset the matrix sparse data of statistic probability. The "People Daily" corpus statistic results are regarded as the preliminary statistic results. When changing the application domain, the recognition accuracy rate of the preliminary statistic results decline, and we adopt "PC World" as the corpus of the changing domain and carry out successive training, then a second smoothing of the preliminary statistic results, and the successive statistic results are looked upon as the final results. A trigram model of task adaptation is obtained. The experimental results show this language model reduces the workload of successive training, and can effectively reduce the perplexity of language models in the task changing domain. It has a higher language adaptation ability in the task changing domain.