Relational Integrated Cross-modal Scene Fusion with CLIP for Fine-grained Indoor Scene Recognition in Intelligent Manufacturing Systems
B. Hu, Yiping Gao, Xinyu Li, Zerui Xi · Chinese Journal of Mechanical Engineering · 2026
Scene recognition stands as a fundamental pillar of manufacturing systems and intelligence, with critical applications spanning augmented reality, robotic navigation, and smart environments. While significant progress has been made in general scene understanding, indoor scene recognition presents unique challenges due to complex spatial layouts, severe intra-class diversity, and subtle inter-class distinctions. Many existing models are unable to accurately distinguish fine-grained scenes due to the lack of modeling of object-to-scene relationships and alignment between low-level visual patterns and high-level scene semantics. To overcome these challenges, this paper proposes relational integrated cross-modal scene fusion (RICSF), a novel framework that achieves comprehensive scene understanding by synergizing geometric object relationships with vision-language interaction through three key innovations: a geometric relational transformer (GRT), an adaptive feature gating mechanism (AFG), and a bidirectional cross-attention mechanism (BCA). Extensive experiments have been conducted to evaluate RICSF on four widely used scene recognition datasets, achieving 92.09% Top-1 accuracy on MITIndoor67, 80.17% on SUN397, 94.29% on Places7, and 89.28% on Places14, outperforming most of the state-of-the-art approaches. The experimental results demonstrate the superiority of RICSF in classification accuracy and its improvement of robustness in complex scene recognition. The unified framework explicitly models object-level geometric relationships and dynamically fuses them with cross-modal vision-language features to achieve robust indoor scene recognition.