ADVANCED SYSTEM FOR AUTOMATED PDF PARSING AND CONTENT EXTRACTION
International Journal of Progressive Research in Engineering Management and Science · 2024
Extracting structured information, particularly tables and graphs, from Portable Document Format (PDF) documents remains a challenging task due to layout variations and the complexity of these elements.Traditional methods often lack flexibility or accuracy.This paper proposes a novel PDF parser that leverages the power of Machine Learning for efficient and accurate information extraction.The proposed approach utilizes YOLOv8, a state-of-the-art object detection model, to identify tables and graphs within PDFs.YOLOv8 is fine-tuned using a high-quality dataset to enhance its ability to detect these specific elements.Once identified, the coordinates of the tables and graphs are utilized by Camelot-py, a Python library specifically designed for table data extraction from PDFs.Camelot-py extracts the data from the identified tables and converts it into a structured format, such as a Data Frame.This work evaluates the performance of the proposed parser on a benchmark dataset and demonstrates its effectiveness in achieving accurate and efficient information extraction from various PDF documents.