Speech Recognition Using Double Data Augmentation Strategy

Yan Zhou, Jing Sun · 2022 IEEE 10th Joint International Information Technology and Artificial Intelligence Conference (ITAIC) · 2022

Nowadays, deep learning-based voice recognition technology has achieved incredible advancements, significantly increasing the accuracy of speech recognition. However, this type of technology requires a huge amount of training data to prevent the model from becoming overfit. To address this issue, this article suggests a dual data augmentation strategy for voice recognition. To begin, the data set is enhanced using the vocal tract length perturbation (VTLP) technique, and then the data set is enhanced again using noise perturbation technology based on genetic algorithms. To accomplish the objective of speech recognition, the DeepSpeech2 model is trained as the acoustic model. The experimental results indicate that the proposed strategy reduces the character error rate (CER) by 3.61 percent when compared to the baseline and by about 0.92 percent when compared to vtlp.

Read the paper · More papers on PaperTik