Hidden Markov models with finite state supervision

Eric Sven Ristad · 1999

In this chapter we provide a supervised training paradigm for hidden Markov models (HMMs). Unlike popular ad-hoc approaches, our paradigm is completely general, need not make any simplifying assumptions about independence, and can take better advantage of the information contained in the training corpus. 1 Introduction A central part of modern research on spoken/written language recognition is the development of a comprehensive training corpus, that is used to estimate the statistical parameters of a recognition system. The training corpus typically consists of a segmented linguistic signal along with an aligned transcript. The transcript represents the language user's intention; it is a sequence of symbols drawn from a finite transcript alphabet. For speech recognition, the transcript alphabet might be as small as the set of all phonemes in a chosen language, or it might be as large as the set of the 100,000 most frequent words in that language. For handwriting recognition, the transc...

Read the paper · More papers on PaperTik