Evaluating audio features for speech/non-speech discrimination

Herman Redelinghuys, Zenghui Wang · 2022

In this paper, the suitability of audio features for application in speech-music discrimination was evaluated to select a feature set that produces high mean accuracy in the classification algorithm, while also reducing the total feature space. The first four standardized moments of twelve audio features were evaluated namely the mean, variance, skewness and kurtosis of the Root Mean Square value, Short Time Energy Ratio, Zero Crossing Rate, Spectral Rolloff, Spectral Flux, Spectral Centroid, Energy Entropy, Spectral Entropy, the first 13 Mel Frequency Cepstral Coefficients (MFCC), Percentage Low Energy Frames, Modified Low Energy Ratio and 4 Hz Modulation Energy. The 4 Hz modulation Energy feature was computed by two different methods, firstly as a by-product of the MFCC feature and secondly using the Hilbert transform for envelope detection. This resulted in an 88-dimensional feature space. It was demonstrated that with a thorough feature selection process a higher mean accuracy and 50% reduction in dimensionality was achieved.

Read the paper · More papers on PaperTik