Environment Recognition in Spherical Video Images Using Multi-Attention DeepLab

Yuta Nishida, Guangxu Li, Huimin Lu, Tohru Kamiya · 2025

Due to the aging population in Japan, the demand for electric wheelchairs is increasing accordingly. To reduce the cost of an autonomous recognition of environment, the spherical camera is widely used to replace the expensive LiDAR, which can also provide a 360-degree view with a single device. Automatic environment recognition has become an auxiliary method to reduce traffic accidents and improve the driving experience of electric wheelchairs. However, environment recognition remains a major challenge due to the severe distortion of images from spherical cameras. In this paper, we propose a convolutional neural network-based method for semantic segmentation in the field of image recognition. Multiattention pyramid upsampling is introduced to improve accuracy by dealing with image distortion. Our method was applied to the SYNTHIA dataset and to a dataset created by cropping images from videos recorded at our campus. The mean intersection over union (mIoU) was 72.2%, and the processing speed was 28.2 frames per second in the combined dataset.

Read the paper · More papers on PaperTik