Generative Adversarial Network-Based Voice Synthesis from Spectrograms for Low-Resource Speech Recognition in Mismatched Conditions

Puneet Bawa, Virender Kadyan, Gunjan Chhabra · 2024

The use of Generative Adversarial Networks (GANs) has been increasing in speech recognition tasks but there has been significant hurdle due to limited availability. The use of GAN have shown promise in speech synthesis tasks, yet their application in low-resource speech systems faces a significant hurdle owing to limited data availability. The progress of effective Automatic Speech Recognition (ASR) systems faces multiple challenges due to a limited range of options and scarcity, resulting in decreased adaptability and efficiency. This article proposes an innovative approach for integrating Generative Adversarial Networks (GANs) to create speech for both adults and children. Experiments have been conducted on using Mel-Spectrograms for synthetic augmentation to address the problem of limited data availability, particularly for low-resource languages and children. The experiments were conducted under both matched and mismatched conditions. The results demonstrate a noteworthy decrease in the Word Error Rate (WER), showcasing the potential of the GAN-based Vocoder model. This leads to an overall Relative Improvement (RI) of $\mathbf{1 2. 7 4 \%}$ and $\mathbf{1 3. 9 5 \%}$ for the adult and children ASR system, respectively. The research has yielded useful insights on the advancement of ASR systems, particularly in relation to the potential benefits of using GAN-based augmentation in real-world scenarios.

Read the paper · More papers on PaperTik