An ESP-Based Lightweight Model for Joint Object Detection and Affordance Segmentation

Chi‐Yi Tsai, Han-Po Lin, Yu‐Chen Chiu · 2021

Affordance segmentation is one of the most challenging and popular research topics in the field of robotic vision in recent years. Although the existing affordance segmentation methods can achieve accurate segmentation results, the model parameters and network architecture are quite large, and most of these methods cannot achieve real-time performance. In order to reduce the number of model parameters and simplify network architecture, this paper presents a lightweight affordance segmentation model based on ESPNetv2, which can effectively increase the processing speed and reduce the computational requirements required at runtime. The proposed method adopts a one-stage anchor-based object detection model as the backbone network for the integration with a semantic segmentation branch. Because of the advantage of one-stage network architecture, the proposed network model can be implemented by a relatively simpler architecture. In addition, we also use a lightweight ESP module as the basic module in the object detection and affordance segmentation model to reduce the number of model parameters and computational complexity. Experimental results show that in the IIT-AFF data set, the proposed method reaches a high segmentation accuracy of 61 % mIOU and the object detection accuracy of 90% mAP. In the test of processing speed, when the network takes 512×512 RGB image as the input, the proposed method achieves a real-time processing speed of 35 frames per second on a platform with NVIDIA GeForce 1080Ti, which is five times faster than the existing AffordanceNet.

Read the paper · More papers on PaperTik