Performance Optimization and Model Research of Machine Learning in Speech Recognition System
Wenwen Shang · 2024
This paper proposes an improved method based on machine learning, which combines the deep neural network (DNN) architecture and speech enhancement technology to significantly improve the recognition accuracy and robustness of the system. In traditional speech recognition models, complex noise environments often lead to inaccurate feature extraction and poor model training effects, thereby reducing the practical application value of the system. To solve the above problems, this paper introduces adaptive noise reduction algorithms and time-frequency domain enhancement technologies in speech preprocessing to optimize the quality of input signals. In addition, through the improved DNN model structure, the ability to model the spatiotemporal features of speech signals is enhanced. At the same time, the dynamic learning rate strategy and gradient optimization algorithm are adopted to improve the convergence efficiency and generalization performance of the model. Experiments verify the superior performance of the improved model in multi-noise scenarios. Compared with traditional methods, the speech recognition accuracy of the improved model is improved by 12.5%, indicating that it has stronger robustness under noise interference. The processing delay of the system is reduced by an average of 15.3%, which greatly improves the real-time performance.