Generative Adversarial Networks based Viable Solution on Dubbing Videos With Lips Synchronization

Abhijit Thikekar, Riya Menon, Saurav Telge, Gauravi Tolamatti, R. L. Priya · 2022 6th International Conference on Computing Methodologies and Communication (ICCMC) · 2022

Based on a survey in 2020, out of the total number of people using the internet, the percentage of users who watch video content on any device is on average 85-90%. Out of this, regional languages make up to 55% of television consumption and 30% of streaming video consumption, highlighting the fact that these regional language videos are confined to an area and cannot be viewed and understood by a significant number of global audiences. Also, live video conferencing has seen a boom in recent times, leading to the need for live translation features which would allow any person to attend any live event. The proposed solution aims at dubbing a video file. It includes giving a video and an audio source file as an input to the system, of which the audio source file will be translated into the targeted language. In order to retain the voice tones and annotations of the original audio, voice style transfer will also be implemented which combined with translation will produce the final audio file. The video source file will then be passed through a Generative Adversarial Networks (GANs) which will transform the facial expressions of characters to achieve lip sync in accordance with the output audio file generated previously. This would convert the video from one language to another while still providing a seamless experience in viewing the content. It will facilitate people with methods to overcome various language barriers. The generated video and audio files will be compiled to make a single file and together produce the final output of the proposed system.

Read the paper · More papers on PaperTik