RD-CNN: A Compact and Efficient Convolutional Neural Net for Sound Classification
Radu Dogaru, Ioana Dogaru · 2020
Classification of signals (sounds, biomedical, etc.) occurs in various circumstances, often requiring low-complexity implementations for various resources-constrained platforms such as mobile or embedded systems within the IoT framework. Herein we describe preliminary results for a novel recognition system where a compact and fast transform, the reaction-diffusion transform (RDT), is used to generate spectral images. Such images are then processed into a novel type of compact convolution neural network (called NL-CNN) where nonlinear convolution is emulated. The ESC-50 environmental sound database with 50 sound categories was considered for parameter tuning and performance evaluation. Despite its low complexity, a reasonable well accuracy is demonstrated (up to 76.5%) close to the human accuracy reported on the same dataset and better in both accuracy and complexity when compared to similar approaches reported in the literature and based on mel-cepstrum. These results provide a basis for a novel and effective method for time-series recognition in general.