A Novel Speech Recognition Model with Bayesian Optimization
Yunfei Zhang, Meiling Xu · 2020
In order to solve the problem of the speech recognition, this paper proposes a novel framework to process speech signals by using the deep fully convolutional neural network (DFCNN) as the acoustic model and the Transformer as the linguistic model. DFCNN can learn more about speech signals through a deep network structure, while Transformer's self-attention mechanism makes it an efficient linguistic model. In order to better extract the characteristic information of speech signal, we transform the speech signal into a two-dimensional language spectrum by frame division and Fourier transform. In this way, the language spectrum can be processed by using convolutional neural network, which has promising ability to process two-dimensional data. The output of the acoustic model is passed into the linguistic model, and a complete sentence is output. At the same time, considering that there are many hyper parameters in the model, in order to adjust these hyper parameters efficiently, we employ the Bayesian optimization algorithm to optimize the hyper parameters of the acoustic model and the linguistic model, so as to further improve the accuracy of speech recognition. Experimental results on speech recognition library THCHS30 show that compared with other models, the optimized model has higher recognition rate for audio and stronger generalization ability.