Layout Aware Research Paper Parsing and Draft Research Paper Layout-Error Detection Using NLP and Rule-based Techniques
Naveen Hedalla Arachchi, Ranul Navojith Dayarathne, Nipuna Dilshan Aluthdeniya, Wanuja Ranasinghe, Gamage Upeksha Ganegoda · 2024
The growing demand for precise formatting and content organization in academic writing has created a need for tools to help undergraduates prepare high-quality research papers for better conference acceptance. This paper introduces a layout-aware content extraction and error detection system, employing a fine-tuned layout parser classification model alongside rule-based algorithms to enhance section classification accuracy and formatting compliance in research papers. The proposed system achieves an average accuracy of 0.8737 and demonstrates better performance across various sections with an average cosine similarity exceeding 0.8900, reflecting its effectiveness in preserving the content. Additionally, the error detection component identifies layout issues, such as section availability, reference duplication, and column misalignment, with high accuracy and F1 scores, showcasing its potential to streamline the drafting process. Unlike existing methods focused solely on text or entity extraction, this system integrates layout-aware content extraction with error detection to provide a comprehensive solution. It addresses diverse academic publisher standards and layout complexities, assisting students in refining their drafts to ensure higher quality and compliance. Future enhancements will incorporate transformer-based models and expanded datasets to improve adaptability, semantic understanding, and the detection of additional formatting errors, further advancing research paper quality.