GNNAE-AVSS: Graph Neural Network Based Autoencoders for Audio-Visual Speech Synthesis

Subhayu Ghosh, Nanda Dulal Jana · 2024

The rapidly growing field of audio-visual speech synthesis (AVSS) involves converting the speech of one speaker into the audio-visual stream of another speaker while preserving linguistic content. AVSS comprises two sequential components: voice conversion (VC), which modifies vocal features from the source speaker to the target speaker, and audio-visual synthesis (AVS), responsible for generating the audio-visual stream of the converted output from VC for the target speaker. Despite the notable advancements in deep learning (DL) technologies, the application of DL models in AVSS has been relatively unexplored in existing literature. This study proposes the use of graph neural network (GNN)-based autoencoders for VC and AVS, respectively, with the aim of enhancing AVSS capabilities. GNNs are specifically designed to operate on graph-structured data, enabling them to effectively capture relationships and dependencies among both audio and visual features. Moreover, trained GNNs have the ability to make predictions for new nodes or edges based on the learned relationships within the graph, showcasing effectiveness in both non-parallel VC and video synthesis tasks associated with AVSS. The proposed framework is named as GNNAE-AVSS, and it is trained and tested on the VoxCeleb2 and LRS3-TED datasets. Both the subjective and objective evaluations of the generated samples demonstrate the superiority of the GNNAE-AVSS framework over the state-of-the-art (SOTA) approaches.

Read the paper · More papers on PaperTik