DataShift: A Cross-Modal Data Augmentation Method for Speech Recognition and Machine Translation

Haodong Cheng, Yuhang Guo · 2022 4th International Conference on Natural Language Processing (ICNLP) · 2022

Data augmentation has been successful in the tasks of different modalities such as speech and text. In this paper, we present a cross-modal data augmentation method, DataShift, to improve the performance of automatic speech recognition (ASR) and machine translation (MT) by randomly shifting values of the feature sequence along the time or frequency dimensions respectively. Experimental results show that our data augmentation method can improve the performance by 4% of word error rate (WER) and 0.36 BLEU score on average on the ASR and MT datasets separately.

Read the paper · More papers on PaperTik