Minimum Bayes error feature selection

George Saon, Mukund Padmanabhan · 2000

We consider the problem of designing a linear transformation 2 IR pn , of rank p n, which projects the features of a classier x 2 IR n onto y = x 2 IR p such as to achieve minimum Bayes error (or probability of misclassication). Two avenues will be explored: the rst is to maximize the -average divergence between the class densities and the second is to minimize the union Bhattacharyya bound in the range of . While both approaches yield similar performance in practice, they outperform standard LDA features and show a 10% relative improvement in the word error rate over state-of-the-art cepstral features on a large vocabulary telephony speech recognition task. 1 Introduction Modern speech recognition systems use cepstral features characterizing the short-term spectrum of the speech signal for classifying frames into phonetic classes. These features are augmented with dynamic information from the adjacent frames to capture transient spectral events in the signal. What ...

Read the paper · More papers on PaperTik