Automated assignment of topics to OCRed historical texts
Florian Fink, Christoph Ringlstetter, Klaus U. Schulz · 2014
During the last decade, a huge amount of OCRed historical texts has been made available on the Internet. For most of these documents meta data are missing that assign topic categories from library classification systems to texts. Data of this form would offer a much better access to these collections. We report on an experiment where we used a completely automated system for topic assignment, originally designed for modern texts, and apply it to OCRed texts from an 18th century German lexicon (Zedler). Lexicon pages/images used in the experiment lead to poor OCR quality and are full of historical spelling variants. In order to measure the influence of OCR errors and historical orthography on topic detection, we created ground truth versions and in addition ground truth versions with modernized orthography for all texts. We found that automated topic assignment leads to useful results for both the OCR output and the two ground truth versions. Difficulties arise from a "changing world" in a (e.g., technical) field as well as language changes beyond simple orthography.