A Tri-modal Fusion Network for Object Detection Using Small Amounts of Low-Quality Data

Yusuke Watanabe, Yuma Yoshimoto, Hakaru Tamukoh · 2026

By employing multiple sensors, a robot can improve its ability to detect objects. In this paper, we propose a tri-modal fusion network, which accepts common layout tri-modal data of RGB, depth, and thermal data as input data and is composed of three sub-networks fused into one network. We also propose a new non-random initialization method of the weights and biases of the tri-modal fusion network. Results of object detection experiments on a dataset which is composed of small amounts of low-quality data showed that the proposed tri-modal network outperformed uni-modal networks and bi-modal fusion networks. We conclude that the proposed tri-modal network can be trained on small amounts of low-quality data without diverging and achieving high performance in environments where data of multiple sensors are required for accurate object detection.

Read the paper · More papers on PaperTik