A comparison of some methods for training hidden Markov models on sequences with missing observations

Alexander A. Popov, Tatyana A. Gultyaeva, Vadim E. Uvarov · 2016

The three approaches to the problem of hidden Markov models training on sequences with missing observations are discussed: marginalization of missing observations, gluing of available parts of the sequence and training on the multisequence formed from the available parts of the sequence. The training performance of the three approaches is evaluated for various numbers of gaps in training sequences. The results were compared to the standard imputation method based on the mode (the most frequent value) of nearest observations. The marginalization approach showed the best training accuracy. The multisequence approach demonstrated a very poor performance hence it is considered inapplicable. Both the marginalization method and the gluing method performed better than the mode imputation method. The dependence of training accuracy on the position of gaps in training sequence was investigated. It has been found that the biggest decrease in training accuracy is achieved when the gap is situated at the beginning or in the middle of the sequence while the lowest decrease is observed when it is situated at the end.

Read the paper · More papers on PaperTik