High-Fidelity Contrastive Language-State Pre-Training for Embodied Agent State Representation

Fuxian Huang, Qi Zhang, Haoran Zhang, Tianyi Zhang, Ming Zhou, Jinouwen Zhang, Shaopeng Zhai · 2025

With the rapid development of AI, multimodal learning has become crucial, especially with multimodal large language models and embodied agent. However, the represen-tation of the state modality still lags behind other modalities like images, videos, and language. To this end, we propose a High-Fidelity Contrastive Language-State Pre-training method (CLSP), which can accurately encode state information into rep-resentations for both embodied agent and multimodallarge language models. Extensive experiments demonstrate the superior precision and generalization capabilities of our representation, achieving outstanding results in text-state retrieval, navigation tasks, and multimodal large language model understanding.

Read the paper · More papers on PaperTik