Synchronized Speech and Video Synthesis
Aaditya Shivprakash Barve, Pratham Madhani, Yashodhan Ghule, Purva Potdukhe, Krunal Pawar, Nilesh Bhandare · 2023
The paper proposes a method to generate expressive talking head videos artificially in any context by synthesizing and syncing speech and video, with the only provided inputs of a facial image and an audio. The method allows the audio either to be provided directly in a pre-recorded format, or to be synthesized using a Text-to-Speech architecture from the textual input words provided. In contrast to previous methods, emphasis has been on generating more expressive, lively and appealing talking head videos through enhanced facial expressions, head dynamics and naturally enhanced audio. The proposed method works by training Long Short Term Memory (LSTM) and Multi-Layer Perceptron (MLP) networks to disentangle content and speaker identity from the input audio for generating facial landmarks dependent on these representations, in contrast to directly mapping audio with the pixels. The disentangled audio content enhances the lip movements and facial regions motions, whereas the speaker information enhances expressions and head dynamics. The method also proposes the prediction of facial landmarks, allowing enhanced, faster and more appropriate head dynamics. The paper also presents extensive qualitative and quantitative evaluation of the proposed methodology, and demonstrations of some generated talking head videos in comparison to previous methods.