Alleviating Action Hallucination for LLM-based Embodied Agents via Inner and Outer Alignment
Kanxue Li, Qi Zheng, Yibing Zhan, Chong Zhang, Tianle Zhang, Xu Lin, Chongchong Qi, Lusong Li, Dapeng Tao · 2024
Large language models (LLMs) have demonstrated impressive potential in empowering embodied agents, fortifying them with task planning and reasoning capabilities closely akin to humans. However, there remain significant challenges in aligning LLM-based embodied agent actions with the executable and safe action space to reduce hallucination. In this paper, we propose a flexible and resource-efficient framework for aligning the action space of LLM-based embodied agents. The framework employs a parameter-efficient fine-tuning method for inner alignment and a retrieval-based generation approach for outer alignment. Specifically, when the inner aligned model generates an action, the outer alignment employs ROUGE to calculate the similarity between the action and all actions within a safe and valid action space, ultimately selecting the action with the highest similarity as the output. For situations with multiple alternative actions, the outer alignment introduces a policy model, which could be either open-source small LLMs or commercial LLMs, to determine the optimal action based on the current context of the agent's task execution. The retrieval-based outer alignment ensures all actions align with an executable action space, alleviating action hallucination for LLM-based agents and significantly improving the controllability and interpretability of their decision-making process. Through extensive experimentation with the LlaMA-2, Bloomz, and OPT models on the ALFWorld benchmark, we validate the effectiveness and adaptability of our framework.