Extracting line features from images of business forms and tables

Adrian Pizano · 2003

Business forms and tables are special document classes typically used to collect or distribute data; they are characterized by the presence of horizontal and vertical lines that delimit the usable space. The paper describes an algorithm that identifies these lines in binary digital images. This algorithm can be used to separate text from graphics before applying optical character recognition, or as a feature extractor in a form classification system. The approach presented differs from exiting vectorization, line extraction, and text-graphics separation methods, in that it focuses exclusively on the recognition of horizontal and vertical lines.>

Read the paper · More papers on PaperTik