ConvLSTM-based Neural Network for Video Semantic Segmentation
Lan Zhou, Hui Yuan, Chuan Ge · 2021 International Conference on Visual Communications and Image Processing (VCIP) · 2021
We propose a convolutional long short-term memory(ConvLSTM)-based neural network for video semantic segmentation. The network can capture the timing information between frames through the ConvLSTM module to improve the prediction accuracy. The back-bone network uses dense connection, atrous convolution, and pooling pyramid structure to expand the receptive field. During the training, to avoid over fitting, data augmentation and learning rate attenuation strategies were used. The proposed method is end-to-end trainable and evaluated on the street scene benchmark Cityscapes dataset. Experimental results show that, benefit from the ConvLSTM module, the proposed network can extract temporal information between frames effectively, and thus improve the accuracy of video semantic segmentation, especially for dynamic objects and small obiects. such as truck. pedestrian and pole.