Multi-language Audio-Visual Content Generation based on Generative Adversarial Networks

Aaditya Shivprakash Barve, Yashodhan Ghule, Pratham Madhani, Purva Potdukhe, Krunal Pawar · 2023

In this paper, we introduce an approach for producing audio-visual content in multiple languages using only a facial image, a manuscript/audio, and a target language as inputs. The proposed technique allows the sound portion to be delivered as a previously recorded file for Audio to Audio conversion using Neural Machine Translators, or as a piece of text from which the audio can be generated by utilizing a pipeline to convert Text to Audio depending on the provided manuscript. Unlike previous approaches, we focus on creating more expressive, dynamic, and captivating audio-visual content by incorporating voice cloning, facial expressions, and head movements. We propose a method that operates by training advanced techniques such as LipGAN to generate lip-synced videos, LSTM-MLP to produce corresponding facial expressions and head movements by learning from filtered content specific to individual speakers, and CycleGAN to transfer voices in the desired language while preserving natural features as much as possible. Towards the end, we present a thorough evaluation of the proposed solution, both qualitatively and quantitatively, and compares it to other state of the art existing methods.

Read the paper · More papers on PaperTik