RDeFormer: Transformer framework for visual loss dominated indoor mobile robot object detection under multi-source perceptual attention fusion

Jiangxun Liu, Hui Liu · Applied Soft Computing · 2026

In noisy indoor environments, accurate detection of the objects around the mobile robot is crucial for effective navigation planning. Achieving accurate and efficient multi-object detection in complex situations such as occlusion or unequal object scales remains challenging. Many mobile robot object detection studies rely on a single visual image to provide information, ignoring the clues brought by other sensors. In this paper, we propose a novel end-to-end transformer framework, called RDeFormer, for processing multi-source sensor data from the camera and inertial measurement unit. The proposed dual-level multidimensional attention mechanism focuses on the region of interest in the feature and integrates multimodal structure information to adjust the weighted feature vector. The visually dominant multimodal loss function is our loss computation module, which enhances the feature aggregation center constraint of the visual modality and reduces the feature distance within the numerical modality. To evaluate the model performance, we construct a mobile robot multimodal dataset, IAIR RGB-D, recorded during the robot's indoor operation. The proposed method achieves state-of-the-art performance compared with other high-performance methods on the self-built dataset and two public datasets. The [email protected]/accuracy values are 0.849/0.880, 0.930/0.942, and 0.477/0.424 on the three datasets, respectively. In addition, comprehensive ablation experiments are conducted to evaluate the impact of RDeFormer's primary design choices. The experimental results show that the proposed framework is practically applicable to indoor mobile robots to accurately perform workspace object localization and classification tasks.

Read the paper · More papers on PaperTik