VisionCube: 3D-Aware Vision-Language Model for Multi-Step Spatial Reasoning

Feiyang Wang, Nan Luo, Wangyu Wu · 2025

Solving a Rubik's Cube requires precise spatial reasoning, sequential planning, and adaptive decision-making. Traditional solvers depend on hand-crafted heuristics or symbolic planning, which limit generalizability across diverse cube states. In this work, we introduce VisionCube, a 3D-aware vision-language model that combines multi-view spatial reasoning with multimodal embodied planning to tackle Rubik's Cube manipulation. VisionCube integrates three core components: (1) Cube3D, which reconstructs structured 3D representations from multi-view images; (2) a Dual-loop VisionCoT module for hierarchical task decomposition and refinement; and (3) a Memory Stream to support long-horizon reasoning and adaptive control. We also propose CubeCoT, a new benchmark dataset containing Rubik's Cube tasks at low, medium, and high difficulty levels. Experimental results show that VisionCube achieves 100%, 100%, and 90% accuracy on low-, medium-, and high-level tasks respectively, significantly outperforming GPT-4V, Qwen2-VL, and CubeRobot baselines. We further deploy VisionCube in a RoboGuide virtual robotic environment, demonstrating accurate real-time execution of cube manipulation actions and robust spatial generalization.

Read the paper · More papers on PaperTik