Data-dependent kernels in svm classification of speech patterns
Nathan D. Smith, Mahesan Niranjan · 2000
Support Vector Machines (SVMs) have recently proved to be powerful pattern classification tools with a strong connection to statistical learning theory. One of the hurdles to using SVMs in speech recognition, and a crucial aspect of SVM design in general, is the choice of the kernel function for non-separable data, and the setting of its parameters. This is often based on experience or a potentially costly search. This paper gives some experimental justification for the Fisher kernels proposed in [4]; kernels are obtained and their extra regularisation and use of labelled and un- labelled data discussed. Fisher kernels are derived from generarive probability models of the data, and are a firststep to implementing kernels for variable length sequences. 1.