Multitaper Spectrogram for Classification of Speech and Music With Pretrained Audio Neural Networks

G.B Rakshith, Krish Narendra, Sanjeev Gurugopinath · 2021

In this paper, we demonstrate the viability of multitaper (MT) features for classification of s peech and music with pretrained audio neural networks (PANN). Among several well-known features for audio tagging, log-mel is widely-used. Therefore, log-mel has been used to train and establish a near-perfect accurate PANN for audio tagging. For the classification problem at hand, we study the performance of MT numerator group delay (MT-NGD) and MT magnitude (MT-Mag) spectral features and compare it with the log-mel feature. Our experimental results on the MARSYAS speech and music database shows that the accuracy of the PANN converges faster as opposed to other features, when trained with MT-NGD spectrogram. Further, the multitaper representations are observed to be robust to the presence of noise in both speech and music.

Read the paper · More papers on PaperTik