LaTeX Rainbow: Universal LaTeX to PDF Document Semantic & Layout Annotation Framework

Changxu Duan, Zhiyin Tan, Sabine Bartsch · 2023

Machine Learning models in the field of Information Extraction for Scientific Publications require high-quality labeled data.The large amount of easily accessible L A T E X source code is a treasure trove of high-quality labeled data.However, existing datasets comprised of document collections and PDF extraction tools have limitations: (1) The hierarchical structure of papers is lost because labeling is done in terms of pages rather than documents; (2) The reading order is not extracted, which potentially muddles the extracted contextual structure; (3) Papers included in the datasets are not likely to be up-to-date.To address these challenges, we propose L A T E X Rainbow, a framework that bridges L A T E X to PDF that can automatically annotate and extract semantic and layout information from L A T E X source code.This framework extends existing annotation methods by taking into account the properties of different existing approaches.It can produce token-level semantic structure annotations, preserve the paper's reading order, and extract the table of contents, i.e., the article's section structure.L A T E X Rainbow enables anyone to extend their datasets with the latest documents.The project is opensourced on GitHub 1 for community contributions and use.

Read the paper · More papers on PaperTik