EMD inspired Spectral Peak Sequence-based Features for Speech-Music Classification

Kamlesh Kishore, Gayadhar Pradhan, Arvind Kumar · 2024

Speech and music segments differ in pattern in the spectrogram. However, the differences are more in higher frequencies as compared to lower frequencies as speech signals are band-limited to 4 KHz. This motivated us to extract spectral peak sequences (SPS) from low-order Intrinsic Mode Functions(IMFs) decomposed using Empirical Mode Decomposition. In the first stage, IMF-SPS features are extracted from the first five IMFs to capture the high-frequency components. Further, these IMFS-SPS features are used to derive statistical features like imfSPS-Mean, imfSPS-StandardDeviation, and imfSPS-Gradient. These features are then used to train three different classifiers - Support Vector Machines, Decision Trees, and Gaussian Naive Bayes Classifier. The results are validated on three datasets (Slaney, GTZAN, and BITM) and compared with previous benchmarks. We achieved the best classification accuracy of 98.3%, 91.4%, and 98.3% with SVM, GTZAN, and BITM datasets respectively.

Read the paper · More papers on PaperTik