Improved HMM/SVM methods for automatic phoneme segmentation
Jen-Wei Kuo, Hung-Yi Lo, Hsin‐Min Wang · 2007
This paper presents improved HMM/SVM methods for a two-stage phoneme segmentation framework, which tries to imitate the human phoneme segmentation process. The first stage per-forms hidden Markov model (HMM) forced alignment accord-ing to the minimum boundary error (MBE) criterion. The objec-tive is to align a phoneme sequence of a speech utterance with its acoustic signal counterpart based on MBE-trained HMMs and explicit phoneme duration models. The second stage uses the support vector machine (SVM) method to refine the hy-pothesized phoneme boundaries derived by HMM-based forced alignment. The efficacy of the proposed framework has been validated on two speech databases: the TIMIT English database