Exploration of excitation source information for shouted and normal speech classification
Shikha Baghel, S. R. Mahadeva Prasanna, Prithwijit Guha · The Journal of the Acoustical Society of America · 2020
Discrimination between shouted and normal speech is an essential prerequisite for many speech processing applications. Existing works have established that excitation source information plays a significant role in shouted speech production. In speech processing literature, various features have been proposed to model different aspects of the excitation source. The principal contribution of this work is to explore three such features, Discrete Cosine Transform of Integrated Linear Prediction Residual (DCT-ILPR), Mel-Power Difference of Spectrum in Sub-bands (MPDSS), and Residual Mel-Frequency Cepstral Coefficient (RMFCC), for shouted and normal speech classification. The DCT-ILPR feature represents the shape of the glottal cycle, MPDSS estimates the periodicity of the excitation source spectrum, and RMFCC characterizes smoothed spectral information of the excitation source. The authors have also contributed a dataset containing shouted and normal speech. This work is evaluated on three datasets and benchmarked against three baseline methods. Deep neural networks are used to study the classification performance of individual features and their combinations. The generalization performance of features (and combinations) is also investigated. Fusion of excitation source features with Mel-Frequency Cepstral Coefficients (MFCC) provides the best performance compared to other combinations. Noise analysis shows that adding excitation features with MFCC+ΔΔ provides a more robust classification system.