Lemma Hunting: Automatic Spelling Normalization for CMC Corpora
Eckhard Bick · University of Southern Denmark Research Portal (University of Southern Denmark) · 2022
This paper presents and evaluates a method for automatic orthographic normalization and the treatment of out-of-vocabulary words (OOV) in German social media data. The system uses a cascade of spellchecking operations including casing-, sound- and keyboard-based letter permutations, as well as letter context likelihoods, and combines partial and root spellchecking with compound analysis and heuristic inflection analysis in novel ways. The system also handles contractions, elisions and some tokenization errors. In addition, pattern-based recognition of foreign words and abbreviations is attempted, supported by jargon-informed lexicon expansion. Contextual Constraint Grammar (CG) disambiguation is used to resolve possible ambiguity. For Twitter data, F-scores of 87.3 and 77.1 were achieved for the identification and correct lemmatization, respectively, of German spelling errors and non-standard abbreviations. 77.6% of foreign words were recognized with 86.5% precision and 1/3 POS errors.