LTN-VINS: A Robust VSLAM in Blurring Environment Based on Lightweight Transformer Network
Yunyang Wang, Pin Lyu, Jizhou Lai, Cheng Yuan · IEEE Sensors Journal · 2024
The visual simultaneous location and mapping (VSLAM) method has been widely proposed to estimate the mobile equipment’s position. However, in dim indoor application scenarios, there exists general image blur, which highly affects the precision and reliability of the VSLAM method. In this article, we proposed a novel pipeline named LTN-VINS based on a lightweight transformer network (LTN) and a self-adaptive point spread function (PSF) estimator, to improve VSLAM localization performance when motion blur occurs in images. First, based on modulated deformable convolution network (MDCN) and inertial measurement unit (IMU) pre-integration between image frames, a self-adaptive image PSF estimator (SAPE) is proposed to preprocess the input image. Second, considering that the feature point methods commonly used in most VSLAM systems are usually related to edge learning in images, we designed a differential high-pass filtering (DHF) module to further enhance the network’s ability to extract image edges. Lastly, to lighten the computational burden and improve the network’s ability to extract semantics between long-distance pixel dependencies, we propose a multihead linear self-attention (MLSA) mechanism. Extensive experiments in the public Euroc datasets and real-world underground parking environments demonstrate that the proposed method can effectively alleviate the image blur influence in VSLAM and achieve remarkable localization optimization performance compared to other SOTA deblurring network methods in terms of accuracy and robustness.