Deep Learning Based Modeling of Audio Signal Enhancement and Segmentation
Tongyu Liu, Daofan Xiong · 2024
In this paper, a convolutional neural network (CNN)-based audio segmentation and enhancement model is proposed to solve the segmentation and enhancement tasks in audio signal processing. First, a one-dimensional convolutional neural network model for audio segmentation is designed, which is able to slice a long audio signal into multiple short segments and recognize the presence of content of interest among them. Second, another one-dimensional convolutional neural network model for audio enhancement is proposed, which can improve the quality and clarity of audio signals and reduce the effect of noise interference. For the loss function, binary cross entropy and mean square error are used and the Adam optimizer is used for training. Experimental results show that the proposed model achieves significant results in audio segmentation and enhancement tasks. In addition, to further enhance the audio quality, filtering operations, especially low-pass filters, are employed to reduce high-frequency noise and improve the signal-to-noise ratio. Experiments have demonstrated that the power spectral density of the filtered signal is significantly reduced in the high frequency region, indicating that the filtering operation successfully reduces the effect of high frequency noise, smoothes the signal, and significantly improves the signal-to-noise ratio. These results provide effective methods and ideas for further research in the field of audio signal processing.