Fusion of L2 Regularisation and Hybrid Sampling Methods for Multi-Scale SincNet Audio Recognition
Zhun Li, Yan Zhang, Wei Li · 2024
SincNet is a nonlinear neural network architecture based on the Sinc function, designed for processing data in the field of digital signal processing, particularly speech and audio data. This paper proposes an improved deep learning model based on the fundamental SincNet neural network structure, incorporating L2 regularization and a hybrid sampling strategy, named L2-MMS-SincNet. Next, L2-MS-SincNet utilizes a selective iterative convolution structure to optimize the model parameters through multiple iterations. In each iteration, according to the current classification results, the model selectively adjusts the weight of the samples, focuses on the misclassified samples, and gradually improves classification accuracy. Experiments conducted on multiple datasets demonstrate that L2-MS-SincNet effectively addresses data imbalance, areduces model overfitting, and significantly enhancing classification performance. In addition, by combining L2 regularization with a hybrid sampling strategy, L2-MS-SincNet reduces training and prediction time while enhancing the robustness and generalization capability of the model.