A unified way in incorporating segmental feature and segmental model into HMM
Jun Yi Derek He, Henri Leich · 2002
There are two major approaches to speech recognition: frame-based and segment-based approach. The frame-based approach, e.g. HMM, assumes a statistical independence and an identical distribution of the observation in each state. In addition it incorporates weak duration constraints. The segment-based approach is computational expensive and rough modelling easily occurs if not much 'templates' are stored. This paper presents a new framework to incorporate the segmental feature and the segmental model in a unified way into frame-based HMM to exploit the advantage of both methods. In the modified Viterbi algorithm, frame-based information prunes out the most probable path at each segment level to which the segmental model can be applied with dramatically reduced computational load; at the same time, the segmental score refines the score obtained by the frame-based model at each level. In this way, the best path found in the end, by the Viterbi algorithm, is optimal.