Research on Audio Recognition and Optimization Processing based on Deep Learning
Chuxin Hang, Mandan Zhuang, Tongyuan Bai, Kang Sun · 2022 3rd International Conference on Electronic Communication and Artificial Intelligence (IWECAI) · 2022
In order to study the audio recognition technology, this paper used the waveform and spectrum of 7434 audio data to know the different kinds of audio data in the part of time and frequncy. Then, MFCC and Chroma features were extracted, and construct CNN model to classify the audio features. Finally, it is concluded that the efficiency of MFCC feature matrix is much higher than that of Chroma feature matrix, and the recognition accuracy of CNN model is 87.72%. To improve the accuracy between audio categories, this paper uses Nonnegative Matrix Factorization (NMF) to optimize audio data. Based on the optimized audio CNN model, the accuracy is improved to 90.88%, and the distinction between processed and unprocessed audio data is effectively improved.