Exploring Compression Strategies for Blendshape-Based Avatar Facial Animation: Subjective and Objective Analysis

Anthony Trioux, Wei Zhang, Giuseppe Valenzise, Fuzheng Yang · 2025

Blendshapes (BS) have been widely adopted as a key method for generating facial animations on avatars due to their ease of manipulation, flexibility in capturing diverse expressions, and compatibility with real-time rendering. Despite their importance, current frameworks, including the recent MPEG Avatar Representation Format one, lack efficient methods for blendshape compression. This paper addresses this gap by proposing and analyzing a pioneering compression scheme for blendshape animation parameters. The method explores three key dimensions of compression: the number of transmitted BS, their quantization, as well as their frequency of transmission. The use of a linear interpolation for reduced blendshape transmission rates allow to mitigate flickering introduced by strong quantization, enhancing the viewing experience. Subjective evaluations demonstrate that the proposed approach achieves substantial data transmission savings while maintaining acceptable visual quality. Furthermore, an in-depth comparison of classical objective metrics against Mean Opinion Scores (MOS) reveals their limitations in accurately capturing perceived quality after blendshape compression, with Pearson and Spearman correlation scores reaching at most ~ 0.6. Among the evaluated metrics, Detail Loss Metric (DLM), Video Multimethod Assessment Fusion (VMAF), and Mean Peak Signal-to-Noise Ratio in the BS domain (PSNR-BS) exhibit the highest correlation with MOS. This study provides a comprehensive benchmark for blendshape compression and has driven the creation of an Exploratory Experiment (EE) within the ongoing MPEG avatar-related efforts, highlighting the study’s relevance to standardization activities.

Read the paper · More papers on PaperTik