YOLOv7-DocInstSeg: Efficient Instance Segmentation Framework for Document Analysis and Recognition
Vamshi Krishna Kancharla, Neelam Sinha · 2023
In this study, we present an “Instance Segmentation” framework for document layout analysis, in order to automate document understanding. We propose a YOLOv7 based network and exploit its object detection abilities to identify document constituents such as tables, text, title, list and figures. This is in contrast with the existing state of the methods that require a computationally expensive two-stage network and large datasets (100GB) for training. The performance of the proposed framework shows improved accuracy for instance segmentation over publicly available PubLayNet dataset, of which 95,916 (about 28% of the available) document images have been utilized for training. We report average precision of 98.5 on detection and 97.8 on segmentation on test dataset (document images containing 124899 instances of document-constituents listed above), moreover 98.7 on detection and 98.5 on segmentation on validation dataset (document images containing 120761 instances of document-constituents listed above) with an improvement of 6.7% in detection and 9.1% in segmentation on validation dataset as compared to SOTA.