Text Extraction for Complex Historical Documents: A Modular Approach to Layout Detection and OCR
David Fleischhacker, Wolfgang Thomas Göderle, Roman Kern · 2024
We present a modular approach for high-precision extraction of data from retro-digitized historical texts with complex layouts. Our two-stage process combines AI-driven layout recognition using YOLOv9 with a fine-tuned Kraken OCR engine. By leveraging synthetic training data and custom fonts, we achieve low single-digit Character Error Rates (CER) for 19th-century documents like the Schematismus. Our approach is particularly effective for processing large-scale historical collections with intricate layouts and nested structures, demonstrating significant improvements over existing solutions in both accuracy and processing efficiency. The system's modular design allows for easy adaptation to different historical document types while maintaining high performance levels.