FEATURE SELECTION VS. FEATURE TRANSFORMATION IN REDUCING DIMENSIONALITY FOR SPEAKER RECOGNITION
Maider Zamalloa, Luis Javier Rodríguez-Fuentes, Mikel Peñagarikano, Germán Bordel, Juan Pedro Uribe · 2008
Mel-Frequency Cepstral Coefficients and their derivatives are commonly used as acoustic features for speaker recog-nition. Reducing the dimensionality of the feature set leads to more robust estimates of the model parame-ters, and speeds up the classification task, which is cru-cial for real-time speaker recognition applications run-ning on low-resource devices. In this paper, a feature selection procedure based on genetic algorithms (GA) is compared to two well-known dimensionality reduc-tion techniques based on linear transforms, namely Prin-cipal Component Analysis (PCA) and Linear Discrimi-nant Analysis (LDA). Evaluation is carried out for two speech databases, containing laboratory read speech and telephone spontaneous speech, and applying a state-of-the-art speaker recognition system. Results with GA-based feature selection suggest that dynamic features are less discriminant than static ones, since the low-size op-timal subsets found by the GA did not include dynamic features. GA-based feature selection outperformed PCA and LDA when dealing with clean speech, but not for telephone speech, probably due to some noise compen-sation implicit in linear transforms, which cannot be ac-complished just by selecting a subset of features. 1.