Investigation Of Data Augmentation Techniques For Bi-LSTM Based Direct Speech To Speech Translation
Lalaram Arya, Ayush Agarwal, S. R. Mahadeva Prasanna · 2023
Direct speech-to-speech translation (DS2ST) system translates the speech in the source language directly to the speech in the target language. It has been shown in the literature that deep learning systems trained using parallel datasets have given a good translation of the speech. However, getting large datasets to train deep learning networks for DS2ST tasks extensively is not easy. Also, the parallel data might not capture the variabilities like session, gender, speaker, and domain variation that might be present in the real-world dataset. In this work, we explore the data-augmentation techniques such that the pool of the training data can be increased and the DS2ST task can be generalized for all the variations. This work uses noise injection, speed perturbation, pitch perturbation, and vocal tract modification-based data-augmentation approaches as an initial attempt. From the experimental results, it has been found that these augmentation approaches improve the performance of the DS2ST system when compared with the clean/original data. Mel-cepstral distortion (MCD) and intelligibility score (IS) are used as metrics to compare the translated speech with the target language speech. Among the augmentation approaches explored, speed perturbation provides the best improvement of 6.125 in terms of MCD. Vocal tract modification improves the performance of the speaker variability in the dataset. This study shows the robustness of the DS2ST system trained on augmented data.