EavaNet: Enhancing Emotional Facial Expressions in 3D Avatars through Speech-Driven Animation

Seyun Um, Yong-Ju Lee, WooSeok Ko, Yuan Hua Zhou, Sangyoun Lee, Hong-Goo Kang · 2024

Speech-driven 3D facial animation models are essential for creating human-like avatars that synchronize lip movements realistically with speech. Despite these advancements, it is still difficult to effectively convey a wide range of emotional facial expressions that align with the voice. This issue arises due to the lack of clearly labeled datasets for individual emotions as well as insufficient inputs to adequately describe these emotions. To overcome this challenge, we propose a re-categorization process that reduces the data into four emotional groups: angry, sadness, happy, and neutral. We use the re-categorized datasets to estimate style embeddings, which serve to distinctly express emotions and control their intensity. Additionally, we tackle the challenge of slow inference speed in autoregressive models by introducing EavaNet, a non-autoregressive model utilizing gated activation units (GAUs) and bidirectional long short-term memory (BLSTM) modules for efficient prediction of 3D face mesh vertices. Our proposed model outperforms previous state-of-the-art models in terms of emotional expressiveness and lip synchronization accuracy in both subjective and objective evaluations.1

Read the paper · More papers on PaperTik