Gen AI driven multilingual audio dubbing and synthesis system for cross-language video platforms
Rishabh Kannojia, Anuj Kumar Singh, I. N. S. Sharma, Shivani Gupta · Results in Engineering · 2025
The demand for accessible, multilingual video content has grown significantly with the global rise of streaming platforms, social media, and online learning. The paper aims to design a Gen AI-powered multilingual audio dubbing and synthesis system for cross-language video websites. In particular, it aims to develop a Chrome extension that can provide real-time dubbing of YouTube videos in several languages. The system uses speech recognition, machine translation, and text-to-speech synthesis to provide multilingual access. The system architecture, implementation, challenges, and enhancements, as well as experimental results demonstrating its utility, are discussed in the study. It further presents comparative analysis, performance metrics, and user feedback to evaluate the solution's efficiency. This paper presents the creation of a Chrome extension used for real-time dubbing of YouTube videos in various languages. The extension uses speech recognition, machine translation, and text-to-speech synthesis to create a seamless and efficient user experience. We explain system architecture, implementation, challenges, and areas for improvement. Experimental results show the success of the proposed system in facilitating multilingual accessibility. Besides, comparative study, performance measures, and user reviews are provided to assess the efficiency of the system. • Utilizes Generative AI for seamless speech-to-text, machine translation, and text-to-speech synthesis to automate multilingual dubbing. • Maintains speaker identity, tone, and emotional expression across languages using advanced voice cloning. • Supports a wide range of global and regional languages, including low-resource Indian languages. • Integrates AI-powered lip-syncing and optional visual alignment to synchronized dubbed video content.