Attention After Attention: Reading Text in the Wild with Cross Attention
Yunlong Huang, Canjie Luo, Lianwen Jin, Qingxiang Lin, Weiying Zhou · 2019
Recent methods mostly regarded scene text recognition as a sequence-to-sequence problem. These methods roughly transform the image into a feature sequence and use the algorithms for sequence-to-sequence problem like CTC or attention to decode the characters. However, text in images is distributed in a two-dimensional (2D) space and roughly converting the features of text into a feature sequence may introduce extra noise, especially if the text is irregular. In this paper, we propose a novel framework named cross attention network, which learns to attend to local features of a 2D feature map corresponding to individual characters. The network contains two 1D attention networks, which operates harmoniously in two directions. Thus, one of the attention modules vertically attends to the features corresponding to the whole text of 2D features and the other horizontal module selects the local features to decode individual characters. Extensive experiments are performed on various regular benchmarks, including SVT, ICDAR2003, ICDAR2013, and IIIT5K-Words, which demonstrate that the proposed model either outperforms or is comparable to all previous methods. Moreover, the model is evaluated on irregular benchmarks including SVT-Perspective, CUTE80 and ICDAR 2015. The performance on irregular benchmarks shows the robustness of our model.