LLM-Powered Embodied Intelligence for Socially-Aware Robot Navigation in Human-Robot Interaction

Xuqing Liu, Ahmed Farid, Tatsuya Amano, Hamada Rizk, Hirozumi Yamaguchi · 2025

This doctoral research proposes a framework for developing sociallyaware robot navigation systems by integrating the cognitive capabilities of Large Language Models (LLMs) with the demands of real-world Human-Robot Interaction (HRI).Our work follows a four-stage plan that systematically addresses the challenges of applying LLMs to time-sensitive, safety-critical tasks.This paper details the completion of the first two stages, wherein we developed and evaluated a foundational navigation model.Our system features a meticulously designed multimodal fusion pipeline that integrates LiDAR and camera data, processed by a YOLO model and a Hungarian algorithm for semantic association, providing rich, contextual input to the LLM.Through knowledge distillation and fine-tuning on data from a custom simulator, our model demonstrates robust spatial reasoning and superior performance in low-frequency decision-making scenarios compared to traditional reinforcement learning methods.We successfully validated this foundational model and identified its inference latency as a key challenge.These results establish a solid basis for our future work.This includes developing a "brain-cerebellum" hybrid architecture for real-time performance and exploring multi-robot social compliance.This research contributes to HRI by creating more predictable and trustworthy robots, and to the LLM field by investigating the symbol grounding problem through embodied intelligence.

Read the paper · More papers on PaperTik