Speaker's gender classification and segmentation using spectral and cepstral feature averaging

Marko Kos, Damjan Vlaj, Zdravko Kačič · International Conference on Systems, Signals and Image Processing · 2011

This paper presents speaker gender classification and segmentation. Such classification is frequently used in broadcast news domain. Because pitch is a feature that is difficult to calculate reliably in noisy environment, and because telephone speech is present in broadcast material, we focused on using general acoustic features for gender discrimination task. We also averaged the feature values to emphasize general speaker's properties to discard short-time properties of speech production. Test show that Average Mel-Frequency Cepstral Coefficients (AMFCC) perform best. The AMFCC features are very convenient gender discriminator for automatic speech recognition system where MFCC features are used, as they perform better than classic MFCC features and only one additional calculation step is needed.

Read the paper · More papers on PaperTik