An algorithm for extracting cursive text lines

Elisabetta Bruzzone, M.C. Coffetti · 1999

In this paper a new algorithm for extracting text lines from a cursive image field is described. The proposed algorithm is a fast and satisfactorily accurate procedure for isolating text lines without loss of information. The algorithm is based on the analysis of horizontal run projections and connected component grouping and splitting on a partition of the input image into vertical strips, in order to deal with undulating or skewed text. The goal of the algorithm is to prevent the ascending and descending characters from being corrupted by arbitrary cuts. The algorithm has been designed for cursive text and can also be applied to handwritten text. It maintains punctuation to allow a better performance word extraction in a subsequent phase of handwritten line processing.

Read the paper · More papers on PaperTik