Analysis of Part-Of-Speech Tagging of Historical German Texts

Markus Paluch, Gabriela Rotari, David Steding, Maximilian Weß, Maria Moritz, Marco Büchler · 2017

The amount of data in contemporary digital corpora is too large to be processed manually, which increases the necessity for computer linguistic tools in humanities. However, the processing of natural languages is a challenge for automatic tools, because languages are used heterogeneously. To process a text, often taggers are used that are trained on a standardized language variety (e.g. recent newspaper articles). Unfortunately, these training data often differ from the target texts (i.e. the text on which a trained model later is applied) in terms of language variety and register, which is especially the case for historical texts. Therefore, additional, manual analyses are usually inevitable. Training tools on the target language variety, however, can improve the results of these tools so that the manual prost-processing could be avoided. Thus, the need to process large datasets of diachronic texts and to obtain accurate results in a short time-span requires an adaptable approach.

Read the paper · More papers on PaperTik