Fast, Piecewise Training for Discriminative Finite-state and Parsing Models

Charles A. Sutton, Andrew McCallum · ScholarWorks@UMassAmherst (University of Massachusetts Amherst) · 2005

Discriminitive models for sequences and trees—such as linear-chain conditional random fields (CRFs) and max-margin parsing—have shown great promise be-cause they combine the ability to incorpo-rate arbitrary input features and the ben-efits of principled global inference over their structured outputs. However, since parameter estimation in these models in-volves repeatedly performing this global inference, training can be very slow. We present piecewise training, a new train-ing method that combines the speed of lo-cal training with the accuracy of global training by incorporating a limited amount of global information derived from pre-vious errors of the model. On named-entity and part-of-speech data, we show that our new method not only trains in less than one-fifth the time of a CRF and yields improved accuracy over the MEMM, but surprisingly also provides a statistically-significant gain in accuracy over the CRF. Also, we present prelimi-nary results showing a potential applica-tion to efficient training of discriminative parsers. 1

Read the paper · More papers on PaperTik