A dual‐dimension collaborative enhancement framework to boost language model spatial semantic understanding
Chenyang Li, Maoyuan Zhang · Annals of the New York Academy of Sciences · 2025
Spatial expressions, which describe the spatial relationships between objects, are a common phenomenon in natural language. Accurately understanding the semantics of spatial expressions in text requires not only linguistic knowledge but also spatial cognition to construct spatial scenarios, and world knowledge for spatial reasoning. The task of spatial semantic understanding demands strong logical reasoning capabilities, posing a significant challenge for small language models with limited parameters. To overcome the performance bottleneck of small language models in this task, this study proposes a cognition-data collaborative enhancement framework. By injecting chain-of-thought, the logical reasoning ability of a large language model (LLM) is decomposed into transferable cognitive units. Combined with semi-supervised learning driven by sequence confidence, the framework extracts high-quality samples with complete spatial relationships from unlabeled data. These two components work synergistically to form a closed loop of "cognitive prior guidance-data integrity constraint," enabling small language models to approximate the performance of LLMs in spatial semantic reasoning tasks. Experimental results demonstrate that the proposed framework also significantly improves the performance of small language models in low-resource scenarios, offering a novel paradigm for semantic understanding in resource-constrained settings.