Generative AI-Driven Multimodal Interaction System Integrating Voice and Motion Recognition

DaeSung Jang, JongChan Kim · International Journal on Advanced Science Engineering and Information Technology · 2025

This research proposes a two-way interactive algorithm based on voice and motion recognition using generative AI technology to overcome the limitations of existing systems limited to simple command recognition. Current voice and motion recognition technologies are essential in enabling interaction between smart devices and users to enhance user experience. Still, they are mainly limited to recognizing and executing prescribed commands, which do not meet the diverse and complex needs of users. To solve these problems, this research aims to develop a technology that fuses and integrates voice and motion data based on advanced learning and prediction capabilities of generative AI, provides customized data optimized for each user's personality and situation in real-time, and enables more natural and efficient interactions. The main research content includes developing data analysis and processing algorithms that can integrally process multiple input channels, designing generative AI-based models for providing customized data to users, and implementing a two-way interactive system that maintains a natural conversation flow. In particular, the research is intended to combine generative AI language models with computer vision technology to comprehensively analyze user voice and motion data, enabling smart devices to understand and respond to user intent accurately. These technologies can potentially revolutionize the user experience in various areas, including smart homes, healthcare, education, and more. This study's results are expected to significantly contribute to the development of next-generation smart device interaction systems that could improve both efficiency and engagement of interactions.

Read the paper · More papers on PaperTik