Augmented statistical models: exploiting generative models in discriminative classifiers
M. Layton, Mark Gales · 2005
In recent years, many algorithms have been proposed for discriminative classification of data. Popular examples are support vector machines (SVMs) [1] and conditional random fields (CRFs) [2]. These techniques make extensive use of fixed-dimensional mappings from the observation-space to a (often high-dimensional) feature-space. Unfortunately, for applications with variable-length sequences of observations – text processing, speech recognition and computational biology – it is not clear how these mappings should be defined. For variable-length sequences, it is usual to estimate class-conditional latent-variable generative models, such as Gaussian mixture models (GMMs) and hidden Markov models (HMMs). Bayes ’ rule is then used to calculate the posterior probability of the class labels. This allows missing data and variable-length sequences to be handled in a simple yet robust manner. However, for many tasks, the independence and conditional-independence assumptions associated with standard latent-variable models are not correct and may degrade classification performance. In [3], Jaakkola and Haussler proposed the Fisher score-space as a powerful method of