Converting German Historical Legal Documents to TEI XML including challenges with Table Extraction

Thomas Reiser, Petra Steiner · Annals of Computer Science and Information Systems · 2024

The occupations archive at the Federal Institute for Vocational Education and Training contains thousands of historical German Vocational Education and Training (VET) and Continuing Vocational Education and Training (CVET) regulations from the last 100 years.However, these are hardly accessible because they are currently only available in their original paper form.We present a workflow that transcribes images of these regulations into the TEI XML format which preserves the logical document structure and stores metadata.This paper addresses issues caused by poor page segmentation of the applied optical character recognition (OCR) methods and presents rules that can reconstruct a large part of the documents' hierarchy.A straightforward table recognition method for tables with borders is presented, as well as a metadata extraction procedure for the selected data set.While our approach is generic and functional, further research is necessary to develop a fully automated and more robust workflow.

Read the paper · More papers on PaperTik