Mixed text/graphics images: automated text separation and graphics representation
Rangachar Kasturi, James Gattiker, J. Shah · Annual Meeting Optical Society of America · 1987
A method for the separation of graphics and text from digitized documents for automatic conversion and compression of paper-based information for data base storage is described. First, the text portion of the digitized document image is separated from the graphics by a robust algorithm which classifies text based on the properties of the connected components of the document. General characteristics of text are employed to categorize and remove it from the image through the utilization of image-dependent size filters, Hough domain grouping, and the application of heuristic knowledge of text attributes. Once the text is removed, the graphics portion of the image is converted to a higher-level representation in the form of a list of segments (straight lines, curves) and their corresponding thickness.