Audio context recognition in variable mobile environments from short segments using speaker and language recognizers
Tomi Kinnunen, Rahim Saeidi, Jussi Leppänen, Jukka P. P. Saarinen · 2012
The problem of context recognition from mobile audio data is con-sidered. We consider ten different audio contexts (such as car, bus, office and outdoors) prevalent in daily life situations. We choose mel-frequency cepstral coefficient (MFCC) parametrization and present an extensive comparison of six different classifiers: k-nearest neighbor (kNN), vector quantization (VQ), Gaussian mixture model trained with both maximum likelihood (GMM-ML) and max-imum mutual information (GMM-MMI) criteria, GMM supervector support vector machine (GMM-SVM) and, finally, SVM with gener-alized linear discriminant sequence (GLDS-SVM). After all param-eter optimizations, GMM-MMI and and VQ classifiers perform the best with 52.01 %, and 50.34 % context identification rates, respec-tively, using 3-second data records. Our analysis reveals further that none of the six classifiers is superior to each other when class-, user-or phone-specific accuracies are considered. Index Terms — Audio context recognition, speaker and lan-guage recognition, short duration, mobile environment 1.