End-to-end Information Extraction from Archival Records with Multimodal Large Language Models

Mahsa Vafaie, Sven Hertling, Inger Banse-Strobel, Kevin Dubout, Harald Sack · 2025

Semi-structured Document Understanding presents a challenging research task due to the significant variations in layout, style, font, and content of documents.This complexity is further amplified when dealing with born-analogue historical documents, such as digitised archival records, which contain degraded print, handwritten annotations, stamps, marginalia and inconsistent formatting resulting from historical production and digitisation processes.Traditional approaches for extracting information from semi-structured documents rely on manual labour, making them costly and inefficient.This is partly due to the fact that within document collections, there are various layout types, each requiring customised optimisation to account for structural differences, which substantially increases the effort needed to achieve consistent quality.The emergence of Multimodal Large Language Models (MLLMs) has significantly advanced Document Understanding by enabling flexible, prompt-based understanding of document images, needless of OCR outputs or layout encodings.Moreover, the encoder-decoder architectures have overcome the limitations of encoder-only models, such as reliance on annotated datasets and fixed input lengths.However, there still remains a gap in effectively applying these models in real-world scenarios.To address this gap, we first introduce BZKOpen, a new annotated dataset designed for key information extraction from historical German index cards.Furthermore, we systematically assess the capabilities of several state-of-the-art MLLMs-including the open-source InternVL2.0and InternVL2.5 series, and the commercial GPT-4o-mini-on the task of extracting

Read the paper · More papers on PaperTik