xADA: Controllable and Expressive Audio-Driven Animation
Sarah L. Taylor, Salvador Medina, Jonathan Windle, Erica Alcusa Sáez, Iain A. Matthews · 2025
Fig. 1. 𝑥ADA generates high fidelity animation of the face, head, and tongue from speech audio, with accurate blinking and expression.𝑥ADA is fully automatic, and can accurately animate a diverse range of speech and non-verbal sounds.We introduce 𝑥ADA, a generative model for creating expressive and realistic animation of the face, tongue, and head directly from speech audio.Our approach leverages the pretrained Whisper audio encoder to extract rich speech features which are decoded into face and head animation using a series of gated recurrent unit (GRU) networks.The generated animation maps directly onto MetaHuman compatible rig controls enabling seamless integration into industry-standard content creation pipelines.𝑥ADA operates fully automatically, with an option for users to override the detected emotion and/or blink timings.𝑥ADA generalizes across languages, and voice styles, and can animate non-verbal sounds.Quantitative evaluation and a user study demonstrate that 𝑥ADA produces state-of-the-art animation with high realism, frequently indistinguishable from ground truth performance.Additionally, we outline a comprehensive data capture protocol designed to collect an extensive range of speech and non-verbal sounds for training animation models.