A Re-estimation Method for Stochastic Language Modeling from Ambiguous Observations
Mikio Yamamoto · 1996
This paper describes a reestimation method for stochastic language models such as the N-gram model and the Hidden Markov Model(HMM) from ambiguous observations. It is applied to model estination for a tagger from an untagged corpus. We make extensions to a previous algorithm that reestimates the N-gram model from an untagged segmented language (e.g., English) text as training data. The new method can estimate not only the N-gram model, but also the HMM from untagged, unsegmented language (e.g., Japanese) text. Credit factors for training data to improve the reliability of the estimated models are also introduced. In experiments, the extended algorithm could estimate the HMM as well as the N-gram model from an untagged, unsegmented Japanese corpus and the credit factor was effective in improving model accuracy. The use of credit factors is a useful approach to estimating a reliable stochastic language model from untagged corpora which are noisy by nature.