A Novel Method for Text Detection in Arbitrary Scenes Based on Multi-scale Segmentation Networks
Kai Dai, Jun‐Guo Lu, Shaohui Ruan · 2019
In recent years, most of the scene text detection algorithms are benefited from the development of deep learning methods which depend on bounding box regression. These methods usually perform two kinds of predictions: text/non-text classification and location regression. Regression may be an important part when predicting bounding boxes, but it is not necessary because the text/non-text prediction can be used as a semantic segmentation which contains location information. However, text instances usually distribute closely in scene images, which is difficult to separate. Therefore, instance segmentation is proposed to fix the problem. In this paper, we propose a pixel-based method based on an efficient segmentation network with multi-scale features. Firstly, the high-level and low-level features are extracted from the modified Xception network. The positive links between the pixels are then calculated to complete the instance segmentation. After we obtain the two predictions of both features, the standard Non-Maximum Suppression (NMS) is used to remove the redundant frames and keep the best output of the bounding boxes. The results show that our method performs well on several public benchmarks, and this method requires less training time and eliminates the trouble of setting anchors.