Multiple Facial Reaction Generation Using Gaussian Mixture of Models and Multimodal Bottleneck Transformer
Dang-Khanh Nguyen, Prabesh Paudel, Seungwon Kim, Ji-eun Shin, Soo-Hyung Kim, Hyung-Jeong Yang · 2024
Facial reaction generation has gained prominence in recent years. However, while there has been extensive research on synthesizing facial expressions from the perspective of the speaker, the generation of reactions from the listener's standpoint remains relatively unexplored. Predicting the facial reactions of the listener in a conversational setting presents a challenge due to the diverse range of reactions that can be elicited by the behavior of a single speaker. In this study, we introduce a Multimodal Transformer-based Variational Autoencoder designed to learn the distribution of listener facial reactions based on speaker audiovisual cues. Our proposed approach incorporates the Multimodal Bottleneck Token mechanism to capture interactions between acoustic and visual speaker features and utilizes the Variational Au-to encoder framework to generate latent representations of multiple listener reactions. Additionally, we employ Gaussian Mixture Models to enhance the generative capabilities of the Autoencoder. Experimental results demonstrate that our method surpasses baseline models and previous approaches on the REACT24 benchmark dataset.