Text string extraction within mixed-mode documents
Frank Hönes, J�rgen Lichter · 2002
Digitized images of printed documents typically consist of a mixture of text, graphics, and image elements. For proper processing and efficient representation, these elements have to be separated. For most applications it is sufficient to separate between text and non-text, because text captures the most information. The authors describe the implementation and performance of a robust algorithm for text string extraction which is completely independent from text orientation and may deal with text in various font styles and sizes. Text objects may be nested in non-text areas and inverse printing can also be analyzed. It should be mentioned that no recognition of individual characters is performed. The classification is only based on rough image features.>