Geological Document Layout Analysis via Synthetic Dataset Creation

C.H. Lun, S. Hou · 2022

Summary A document page may contain not only text, but also figures, tables, titles and captions. Locating these components which make up a page, a task called document layout analysis, allows further processing of the document, for example, table extraction and image classification for the located tables and figures respectively, and is therefore vital towards converting unstructured documents into knowledge. Computer vision can be used to perform document layout analysis, but publicly available datasets are sourced from domains outside of oil and gas and thus are not directly applicable. Manually labelling datasets is time consuming; therefore, we present a method of creating a synthetic dataset to address the issue of limited labelled data. Finally, as a downstream task, we discuss the problem of table cell type classification, which is a first step towards table understanding and extraction of knowledge from tables.

Read the paper · More papers on PaperTik