Trans-APL: transformer model for audio and prior landmark fusion for talking landmark generation
Xuan-Nam Cao, Quoc-Huy Trinh, Minh–Triet Tran · 2025
The generation of talking landmarks from audio is pivotal for advancing talking head generation. This challenge poses a significant concern in landmark generation from audio and holds potential applications in various domains, including virtual assistants, education, and entertainment. However, existing audio-based methods exhibit limitations, such as inconsistencies in generated landmark frames and a lack of emotion features from the speech. In this research, we propose Trans-APL, a foundational approach that addresses these limitations by integrating a fusion of landmark and audio information. Additionally, we introduce the Conv-Attention module, a combination of the Convolution layer with the Attention mechanism designed to capture features from both low and high frequencies. This capability aids in capturing emotion information from the audio. Our extensive experimental and visualization results demonstrate the enhancement of talking landmark generation by adapting our methods. Consequently, our approach yields competitive results compared to state-of-the-art methods and significantly contributes to the advancement of realistic talking head motion.