NewTalker: Exploring frequency domain for speech‐driven 3D facial animation with Mamba
Weiran Niu, Zan Wang, Yi Li, Tangtang Lou · IET Image Processing · 2025
Abstract In the current field of speech‐driven 3D facial animation, transformer‐based methods are limited in practical applications due to their high computational complexity. A new model—NewTalker—is proposed, which has core modules consisting of the residual bidirectional Mamba (RBM) and the time–frequency domain Kolmogorov–Arnold networks (TFK). The RBM module incorporates the philosophy of Mamba, enhancing the model's predictive ability for sequence data by utilizing both past and future contextual information, thereby reducing the computational complexity. The TFK module integrates the temporal and frequency domain information of audio data through Kolmogorov–Arnold networks, allowing the model to generate 3D facial animations smoothly while learning more detailed features. Extensive experiments and user studies have shown that the proposed NewTalker significantly surpasses current mainstream algorithms in terms of animation quality and inference speed, achieving the state‐of‐the‐art level in this domain.