Fine-Grained, Accurate Data Generation and Multimodal Layout Analysis for Academic Papers

Dehao Ying, Fengchang Yu, Haihua Chen, Wei Ping Lu · 2024

Layout analysis of academic papers aims to identify various components within unstructured papers, benefiting researchers in quickly locating and extracting critical information. The effectiveness of this process depends heavily on the datasets and models used for training. However, existing datasets often have issues with annotation accuracy, granularity, scale, and acquisition cost. Current models treat each document image in isolation, ignoring the position information of a page within the entire paper. To address these challenges, we propose DLAgen, a method for rapidly, accurately, and cost-effectively generating fine-grained annotated paper datasets. DLAgen uses context-free grammar to generate textual content in LaTeX format, and incorporates visual content, such as images, tables, and formulas, from real papers, thus creating synthetic papers with accurate annotations. Concurrently, to leverage the high correlation between page numbers and components in academic papers and to make better use of textual information, we introduce MDT, a multimodal academic paper layout analysis model that utilizes page position information and correctly ordered text. Experiments show that MDT trained with data generated by DLAgen achieves higher accuracy in fine-grained layout analysis of real academic papers compared to existing state-of-the-art models. The mAP is improved from 85.13 to 88.61, which is a 4.09% enhancement, validating the effectiveness of our approach. Both the model and dataset will be released to the public.

Read the paper · More papers on PaperTik