An Algorithm for identifying, extracting and converting a table structure from a document inage into LaTeX format

San Sethasopon, Chidchanok Lursinsap · 2002

Table analysis is one of the attractive and challenging problems in document image analysis that encompasses table identification and table recognition. Table identification is based on the techniques of page segmentation and classification, whereby the results so extracted are analyzed and stored in some prearranged structures. This study proposes an algorithm for table analysis that starts from separating a document image into individual blocks. A non-tabled block is determined by the arrangement of data inside the block and the position of lines. Then, the recognized table blocks are converted into LaTeX formatted tables suitable for subsequent modification, storage, retrieval and transmission. The algorithm was tested with image blocks extracted from actual document images and synthesis samples. Various styles of tabled block-lines and data arrangement were correctly identified and analyzed. The algorithm gave good results for input samples having less skewed angle and noise.

Read the paper · More papers on PaperTik