Voice Cloning for Low‐Resource Languages

Vishnu Radhakrishnan, A. Aadharsh Aadhithya, Jayanth Mohan, M. Visweswaran, G. Jyothish Lal, B. Premjith · 2024

With the emergence of artificial intelligence (AI)-powered personalized assistive tools, the surge of futuristic AI agents, and democratization of AI, techniques like voice cloning helps in blurring the line between man and machine. Although there are existing methods for voice synthesis, the task of voice cloning is challenging because the model needs to adapt to an unseen speaker with very less data. Voice cloning is a relatively new task that has not received much attention until recently. While traditional text-to-speech (TTS) systems tries to aid man-machine interaction, voice cloning takes it a step further by enabling to replicate the voice of near or dear ones. However, it is practically difficult to gather large datasets for voice cloning in domestic environments. Apart from the major limitation of data unavailability, designing a compact, mobile, and efficient model for cloning voices with only a few samples of data remains an unaddressed problem. While voice cloning models continue to improve, it remains challenging to incorporate region-specific accents and indigenous low-resource languages into machine-generated audio outputs that accurately differentiate a human voice from a synthesized one. This is primarily because speech synthesis and cloning require a large amount of data, which is not available for low-resource languages. This work examines some preliminary results for voice cloning in Tamil. Despite being in the early stages, our results show promise for the development of a successful voice cloning system for Tamil. We believe that this work will serve as an invitation for the Tamil speech processing community to explore this exciting area of research further. With the potential to revolutionize the field of speech synthesis, voice cloning in Tamil could have significant implications for the development of speech-based applications and assistive technologies. The present work addresses the above-listed problems, specifically for Tamil.

Read the paper · More papers on PaperTik