Collaborative Learning Network for Scene Text Detection
Xiaoye Zhang, Yuanhao Yue, Yingyi Yang, Xining Zhang, Wei Wang, Qin Zou · 2020
Text detection in the wild has been a popular research topic, and previous approaches to scene text detection have achieved promising performances across various benchmarks. In general, these methods use a large number of labels with precise character positions as training data. However, accurate annotation of large amounts of training data is a challenge. On the other hand, image-level annotation can be quickly obtained from rich media in a tag retrieval manner. Therefore it is essential to investigate how to improve the performance of strongly supervised text detection models through image-level annotation data. In this work, we propose a training framework for collaborative learning of a weakly supervised text classification network and a strongly supervised text detection network. The collaborative learning of the two sub-task networks is achieved by constraining the consistency of the two networks at the perceptual level. Experiments on standard data set ICDAR2015 show that the framework can significantly improve the performance of the strongly supervised text detection network by training using image-level annotation data.