Using Convolutional Neural Networks to Classify Audio Signal in Noisy Sound Scenes
M.V. Gubin · 2018 Global Smart Industry Conference (GloSIC) · 2018
The issue of source separation of audio signals from a mixture of sounds is an important problem that arises in various industrial applications. For example, it can be mentioned that the issue of sound-based fault detection and diagnosis in industrial equipment or a multi-speaker recognition task in a cocktail party problem. The last one is of great importance for the creation of new generation hearing aids that isolate and amplify a certain speech signal in noisy scenes. In this paper, I propose a new approach to this problem solving, based on an ensemble of artificial neural networks. The solution of the problem is divided into two stages. At the first stage, the ensemble of convolutional neural networks determines the presence or absence of the speech signal in a noisy environment, using a set of speech signal samples, prepared in advance. At the second stage, another ensemble of neural networks filters the speech signal, determined at the first stage and cuts the rest of the signals as noise. The ensemble of convolutional neural networks, used at the first stage, consists of neural networks, each of which includes three convolutional layers and one fully connected layer. The analysis of the sound scene is performed on the basis of its spectrogram, obtained by using the fast Fourier transform. This neural network is implemented in Python by the use of the TensorFlow and Keras software libraries. Here are the results of computational experiments on using the designed and trained neural network for analyzing and filtering an audio stream, which contains several superimposed male and female voices with music in the background. The performed experiments confirm the efficiency of the proposed approach.