PolarText: Single-stage Scene Text Detection with Polar Representation

Qiran Kong, Yirui Wu, Shaohua Wan · 2021

Although deep learning has achieved great success in object detection recently, scene text detection is still a challenging task, due to inherent difficulties of locating texts in complex scenes. Many approaches adopt inspirations from segmentation to detect arbitrary shaped scene text. However, most segmentation based methods have high computation cost and generally needs a lot of refinements to get accurate results. To ease this problem, we propose a novel single-stage method, i.e., PolarText network, which detects text regions by generating contour points in polar coordinates. PolarText not only relieves the burden of high computation cost by directly regressing contour points instead of pixels, but also fits with intrinsic characteristics of text instances by centers and contours, thus suppressing mislabeling boundary pixels caused by pixel-level labeling. To cope with polar representation, PolarText utilizes Polar IoU loss and polar centerness to generalize effective paradigms from box representation for polar representation. In addition, we add a dedicated bounding box branch to work with text detection since most text instances are approximately rectangular in shape. Compared with the existing methods, the proposed method achieves superior results in both accuracy and efficiency by testing on CTW 1500 and ICDAR 2015 datasets.

Read the paper · More papers on PaperTik