Discriminant sub-space projection of spectro-temporal speech features based on maximizing mutual information
Martin Heckmann, Claudius Gläser · 2011
We previously developed noise robust Hierarchical Spectro-Temporal (HIST) speech features. The learning of the features was performed in an unsupervised way with unlabeled speech data. In a final stage we deployed Principal Component Anal-ysis (PCA) to reduce the feature dimensions and to diagonalize them. In this paper we investigate if a discriminant projection can further increase the performance. We maximize the mu-tual information between the features and the phoneme cate-gories using a procedure known as Maximizing Renyi’s Mu-tual Information (MRMI) and also compare it to Linear Dis-criminant Analysis (LDA). Based on recognition tests in clean and in noise, i. e. in matching and mismatching conditions, we show that the discriminant projections increases recogni-tion scores compared to PCA in matching conditions. How-ever, this improvement does not transfer to the mismatching, i. e. noisy, conditions. We discuss measures to alleviate this problem. Overall MRMI performs better than LDA. Index Terms: Spectro-temporal, discriminant, mutual informa-tion, robust speech recognition, auditory