A hybrid voice cloning for inclusive education in low-resource environments
Muhammad Mohtad Younus, Arshad Iqbal, Esha e Noor Durrani, Naveed Ahmad, Mohamad Ibrahim Ladan · Frontiers in Computer Science · 2025
Introduction Voice cloning can personalize speech technologies but typically requires large datasets and compute, limiting use in low-resource educational settings. Methods We propose a hybrid pipeline combining a GE2E-trained speaker encoder, a Tacotron-based text-to-spectrogram synthesizer, and a modified WaveRNN vocoder with gated GRUs and skip connections. The system targets few-shot adaptation (5–10 s of target speech) and near real-time synthesis on modest hardware. Results On LibriSpeech, VCTK, and noisy YouTube/local corpora, the system achieves MCD ≈ 4.8–5.1 and improves MOS over baselines (e.g., LibriSpeech: 4.55 vs. 4.33; YouTube: 3.82 vs. 3.10), with EER < 12% on an external ASV, indicating strong speaker similarity. Discussion Results show data-efficient, robust voice cloning suitable for inclusive education, with practical considerations for deployment (compute, noise) and responsible use (consent, watermarking, detection). The approach supports assistive and multilingual classroom scenarios in low-resource contexts.