Towards Structured Information Extraction Using OCR Free Transformer Model for Energy Documents
N. El Kamel, P. Malige · 2024
Summary Technical energy documents often contain specific terms, intricate layouts, and extensive page counts. Daily drilling reports (DDRs) as an example detail field operations across several tables; they come in diverse layouts depending on client specifications. This paper presents a novel approach for extracting structured information from images of activity tables in daily drilling reports (DDRs). It is crucial to detect both the table structure and textual information for DDRs understanding and information processing such as visualization or classification. Yet, manual extraction is time-intensive and error-prone and relying on Optical Character Recognition (OCR) for textual extraction is time and resource consuming. In this work, we will fine-tune Document Understanding Transformer (Donut) as an OCR free end-to-end model for table information extraction (table IE) from an activity table image enabling the simultaneous extraction of both text and positional information. We will also introduce a new task, table visual question answering (table VQA) task and fine-tune Donut on it to extract detailed information from the activity tables. The synthetic data generator we established and the freshly created table VQA dataset enhance the model performance. Our model exhibits strong generalization capabilities: effectively adapting to various document layouts and performing well on previously unseen documents.