TH-Mamba: Spatial-Temporal Correlation Learning for Mamba-Based Talking Head Generation

Xin Xu, Zhixi Yu, Kui Jiang, Xiaocheng Feng, Chia‐Wen Lin · IEEE Transactions on Circuits and Systems for Video Technology · 2025

Talking head generation aims to synthesize high-quality and lip-synchronized talking head videos from the given portrait images and audio. However, previous methods directly learn the alignment between lip movements and the driven audio, barely focusing on the fidelity and continuity of the generated videos, suffering from visual distortions and jitter. To deal with this issue, we propose to promote the consistency of audio and image by exploring their spatiotemporal relations, and construct a Mamba-based spatiotemporal fusion scheme. Specifically, we devise an Intra-frame Mamba module to characterize facial features from the source image, which encourages the content consistence between the generated frame and the current source frame. Meanwhile, an Inter-frame Mamba module is designed to excavate the complementary information across sequential frames, which provides clues for better motion simulation. The aggregated spatiotemporal representation with audio features are then aligned with a deformation network to alleviate visual distortions and jitter. In addition, we investigate the practical composite constraints on the structure, details, and motion aspects, involving the keypoint constraint, multi-scale content constraint, and displacement constraint to promote the training stability and model performance. With the above strategies, we construct a novel Talking Head Mamba network, termed as TH-Mamba for high-quality talking head generation. Extensive experiments on the HDTF and Mead-Neutral datasets verify the superiority of our proposed TH-Mamba, which significantly outperforms the current state-of-the-art method by 0.78dB and 1.55dB in PSNR, respectively. The demo is available at https://github.com/YZX-codesky/TH-Mamba.

Read the paper · More papers on PaperTik