A Hybrid Approach for Table Information Extraction: Combining Google Document AI with Custom Object Detection Models
Muthukumaraswamy Balakrishnan, R. Rajagopal · 2024
In today's era of document digitization, the extraction of table data from both structured and unstructured tables documents holds significant importance. Tables found within documents such as scientific articles, invoices, and others often exhibit diverse formats and designs, varying from those with borders to those without, or with minimal borders. Addressing this variability necessitates the development of a custom mechanism for extracting data from these tables. This paper proposes a tailored solution to address such extraction challenges. The proposed approach involves the collection of custom data specifically tailored for extraction purposes. Subsequently, three object detection models (Yolov5) are trained to identify regions of interest (ROIs), annotate rows and columns within tables, and train with the collected custom data. Pretrained models are used to ease the job of annotation. Through inference and post-processing of the model results, cell bounding boxes are generated. Leveraging these bounding boxes, we can align the output of Google Document AI with the specific document, facilitating accurate extraction of desired results. This custom solution offers a robust method for efficiently extracting table data from documents with diverse formats and designs. The results and performance evaluation proves that the method proposed is competent in real time.