Letter Sequence Labeling for Compound Splitting
Jianqiang Ma, Verena Henrich, Erhard Hinrichs · 2016
For languages such as German where compounds occur frequently and are written as single tokens, a wide variety of NLP applications benefits from recognizing and splitting compounds.As the traditional word frequency-based approach to compound splitting has several drawbacks, this paper introduces a letter sequence labeling approach, which can utilize rich word form features to build discriminative learning models that are optimized for splitting.Experiments show that the proposed method significantly outperforms state-ofthe-art compound splitters.