Throat microphone speech recognition using wav2vec 2.0 and feature mapping
Kohta Masuda, Jun Ogata, Masafumi Nishida, Masafumi Nishimura · 2022 IEEE 11th Global Conference on Consumer Electronics (GCCE) · 2022
Throat microphones can record the voice which and simultaneously suppress the impact of external noise. This work aims to improve speech recognition performance using throat microphones within a high-noise environment. However, as there is no large database on throat microphone speech, training data are insufficient. This study proposes a method to realize the improved throat microphone speech recognition by utilizing self-supervised learning models such as wav2vec 2.0. However, because the volume of throat microphone speech data available for training of the model is rather limited, linguistic information is not particularly well trained. Therefore, we apply feature mapping to a large Japanese speech corpus to generate a quantity of pseudo-throat microphone speech features. It was confirmed that a significant improvement in the recognition rate could be achieved by utilizing the generated data for fine tuning of the wav2vec 2.0 model.