A Robust and Binarization-Free Approach for Text Line Detection in Historical Documents
Tobias Gruuening, Gundram Leifert, Tobias Strauß, Roger Labahn · 2017
Text line extraction from complex handwritten documents, especially for historical collections, is still an unsolved problem. There is a strong demand for reliable and robust approaches since text line extraction is a crucial pre-processing step for modern text recognition and keyword spotting systems. We propose a binarization-free system which employs a newly developed clustering approach based on so-called 'superpixels'. Although multiple ways of generating superpixels were developed in the past, we demonstrate that even a standard method yields impressive results. Our clustering approach is applicable to various scenarios by making use of general characteristics of text lines (e.g., curvilinearity, interline spacings, local homogeneity), and without adapting its parametrization. State-of-the-art results are achieved by the same parametrization for 8 different well-established benchmarking datasets. These datasets cover historical and modern texts as well as images with diverse resolutions and fonts. The system is developed for detecting text lines in complex scenarios. It is not tuned to assign foreground pixels to detected text lines. Thus, superior performance is achieved for the historical datasets for which no pixel hit accuracy of 95% is required. Remarkably, for the dataset of the ICDAR 2015 Competition on Text Line Detection in Historical Documents, the average cost per text line was reduced from 9.77 (winning team) to 8.19.