Accelerated Training of Maximum Margin Markov Models for Sequence Labeling: A Case Study of NP Chunking
Xiaofeng Yu, Wai Pang Lam · 2010
We present the first known empirical results on sequence labeling based on maximum margin Markov networks (M 3 N), which incorporate both kernel methods to efficiently deal with high-dimensional feature spaces, and probabilistic graphical models to capture correlations in structured data. We provide an efficient algorithm, the stochastic gradient descent (SGD), to speedup the training procedure of M 3 N. Using official dataset for noun phrase (NP) chunking as a case study, the resulting optimizer converges to the same quality of solution over an order of magnitude faster than the structured sequential minimal optimization (structured SMO). Our model compares favorably with current state-of-the-art sequence labeling approaches. More importantly, our model can be easily applied to other sequence labeling tasks. 1