Video Text Detection with Fully Convolutional Network and Tracking

Yang Wang, Lan Wang, Feng Su, Jiahao Shi · 2019

Scene text in videos carries rich semantic information that is of great value in various content-based video applications. In this paper, we propose an effective fully convolutional network model for detecting text in videos based on a novel refine block structure. The model hierarchically exploits low-level features from earlier convolutions to refine high-level semantic features, thereby fusing multi-resolution features extracted from the frame to generate high-resolution semantic feature maps for better capturing widely varied appearances of video text. We further complement the individual-frame detection with an efficient correlation filter based text tracking mechanism, and enhance the overall detection performance by matching and combining detection and tracking results. Experiments on public scene text video datasets demonstrate the state-of-the-art performance of the proposed method.

Read the paper · More papers on PaperTik