DNN Audio Classification Based on Extracted Spectral Attributes
Pei‐Chen Lo, Chuanyi Liu, Tsung-Hsien Chou · 2022
Recent advances in multimedia systems provide remarkable audio-visual experiences to various fields including entertainment, education, communication, industrial design, etc. To facilitate the audio-visual experience, audio quality enhancement becomes important. However, methods and techniques for improving audio quality highly depend on such audio attributes like human voices, music of different genres, or audio of various programs. This study is devoted to the development of an effective method for real-time audio classification based on deep learning scheme. Three classes of interest include classical music, non-classical music and news. Subband-power distribution (SPD) is a one-dimensional feature based on the audio power in frequency domain, which effectively reflects the spectral attributes of various audio content and allows us to implement DNN (deep neural network) audio classifier in real time. This study develops different DNN models according to various input designs, original SPD of different frequency resolutions and SPD pre-processed by principal component analysis (PCA). Overall accuracy Acc and prediction accuracy of each class using confusion matrix (CFM) will be evaluated to compare the performance. According to our results, the DNN audio classifier implemented with the input SPD pre-processed by PCA not only achieves better performance but remarkably reduces the memory capacity and computational time.