AMCIU: An Adaptive Multimodal Complementary Intent Understanding Method
Meng Lv, Zhiquan Feng, Xiaohui Yang, Qingbei Guo, Xiaoyan Wang, Guixia Zhang, Qun Wang · International Journal of Human-Computer Interaction · 2025
With the aging of society, assistive companion robots have become an important solution to address the daily needs of the elderly. However, the elderly may suffer from slow thinking and memory loss, which makes it difficult for robots to accurately capture executable intentions during human-robot interaction. To address this challenge, this paper proposes an Adaptive Multimodal Complementary Intent Understanding Approach (AMCIU). First, a self-attentive mechanism is utilized to extract multimodal features of gestures, speech and images. Second, knowledge graph is utilized for the first time to obtain complementary information between image and audio modalities. Finally, robust intent integration is achieved through a multimodal hybrid expert intent fusion technique. In comparison with multiple state-of-the-art multimodal learning methods, AMCIU shows significant performance improvement in the building block tower scenario, where the intention understanding accuracy reaches 96.11%. This paper provides a novel research idea for natural human-robot interaction.