Sparse Forward-Backward for Fast Training of Conditional Random Fields
Charles A. Sutton, Chris Pal, Andrew McCallum · ScholarWorks@UMassAmherst (University of Massachusetts Amherst) · 2006
Complex tasks in speech and language processing often include random variables with large state spaces, both in speech tasks that involve pre-dicting words and phonemes, and in joint processing of pipelined sys-tems, in which the state space can be the labeling of an entire sequence. In large state spaces, however, discriminative training can be expen-sive, because it often requires many calls to forward-backward. Beam search is a standard heuristic for controlling complexity during Viterbi decoding, but during forward-backward, standard beam heuristics can be dangerous, as they can make training unstable. We introduce sparse forward-backward, a variational perspective on beam methods that uses an approximating mixture of Kronecker delta functions. This motivates a novel minimum-divergence beam criterion based on minimizing KL di-vergence between the respective marginal distributions. Our beam selec-tion approach is not only more efficient for Viterbi decoding, but also more stable within sparse forward-backward training. For a standard text-to-speech problem, we reduce CRF training time fourfold—from over a day to six hours—with no loss in accuracy. 1