Research on Silent Speech Reconstruction Based on the BigVGAN Network

Hanlin Qi, Deli Fu, Weiping Hu · 2024

To make the audio of silent electromyographic facial motion speech reconstruction sound smoother, more natural, and realistic, this paper combines the Transformer network with the BigVGAN network and designs a silent speech reconstruction process based on the BigVGAN network. The model extracts features using the ResNet network, then aligns the features through the DTW network before feeding them into the Transformer network to convert the electromyographic signals into an Mel-frequency spectrogram, which is subsequently input into the BigVGAN network for speech reconstruction. The final result is the voiced speech signal corresponding to the silent motion.In addition, this paper also explores the use of the WaveNet decoder and HiFi-GAN decoder for speech reconstruction from the obtained Mel-frequency spectrograms. The experimental results show that the BigVGAN reconstruction network used in this paper achieves a phoneme recognition rate of 90.88%, with a word error rate of 26.1% calculated through the ASR model. Compared to structures such as the WaveNet decoder and HiFi-GAN decoder, the BigVGAN reconstruction network effectively improves recognition accuracy.

Read the paper · More papers on PaperTik