Classification of noun-noun compound semantics in Dutch and Afrikaans

Ben Verhoeven, Walter M. P. Daelemans, Gerhard B. Van Huyssteen · 2012

Abstract—This article presents initial results on a supervised machine learning approach to determine the semantics of noun compounds in Dutch and Afrikaans. After a discussion of previous research on the topic, we present our annotation methods used to provide a training set of compounds with the appropriate semantic class. The support vector machine method used for this classification experiment utilizes a distributional lexical semantics representation of the compound’s constituents to make its classification decision. The collection of words that occur in the near context of the constituent are considered an implicit representation of the semantics of this constituent. F-scores were reached of 47.8 % for Dutch and 51.1 % for Afrikaans. Keywords—compound semantics; Afrikaans; Dutch; machine

Read the paper · More papers on PaperTik