Generation Modeling of Robot Action Sequences Using Image Captioning and Large Language Models
Weihao CAI, Yoshiki Mori, Nobutaka Shimada · The Proceedings of JSME annual Conference on Robotics and Mechatronics (Robomec) · 2024
In this paper, we describe a method for generating robot action sequences using a large-scale model, which allows achieving appropriate robot behaviors while reducing the learning costs. We utilize fine-tuned large-scale models to build the action generation modeling. The task for validating action generation involves placing cubes into drawers. Action sequences are generated based on the current state to achieve the goal, and new action sequences are generated in case of action failure or external interference.