Temporally-aware Convolutional Block Attention Module for Video Text Detection

Masato Fujitake, Hongpeng Ge · 2021 IEEE International Conference on Systems, Man, and Cybernetics (SMC) · 2021

Scene text in video carries rich semantic information that is of great value in various content-based video applications. Existing methods have been proposed to improve accuracy, such as combining tracking; however, many have complex structures and are difficult to execute in real-time. Therefore, to run in real-time, this paper proposes a simple and practical feature refinement module, named Temporally-aware Convolutional Block Attention Module (TCBAM), based on a novel self-attention recurrent neural network. The model exploits still-image-based feature maps to refine temporal constant feature maps for better capturing widely varied appearances of video text. For better generalization, we also provide the flow-based data augmentation method with artificial data. Explements on the scene text video datasets including ICDAR2013 Video, Minetto, and RoadText-1K demonstrate that the proposed methods perform the competitive accuracy to the state-of-the-art models within real-time running. Our method with ResNext-50 can run at 17 FPS with 73.11 F-score on ICDAR 2013 Video without complex tracking methods.

Read the paper · More papers on PaperTik