Affix Discovery based on Entropy and Economy Measurements

Alfonso Medina · 2008

This paper briefly describes an entropy and economy-based word segmenta-tion method. Results of its application to discover items belonging to affix subsystems of two unrelated American languages are presented; namely, a variant of Ralámuli or Tarahumara (Uto-Aztecan) and one of Chuj (Mayan). More importantly, an attempt is made to compare these experiments in order to evaluate this approach by means of precision and recall measurements. The following sections (7.1, 7.1.1, 7.1.2) present the method. Data ob-tained from the experiments is shown and discussed in section 7.2. The eval-uation is presented in section 7.3. 7.1 Method There are several prominent approaches to word segmentation. The earli-est one is due to Zellig Harris, who first examined corpus evidence for the automatic discovery of morpheme boundaries for various languages, Harris (1955). His approachwas based on counting phonemes preceding and follow-ing a possible morphological boundary: the more variety of phonemes, the more likely a true morphological border occurs within a word. Later, Nikolaj Andreev designed in the sixties the first automatic method based on character string frequencies which applied to various languages. His work was oriented towards the discovery of whole inflectional paradigms and applied to Russian and several other languages, Cromm (1996); and that of Kock and Bossaert (1974, 1978) in the seventies for French and Spanish. More recent promi-nent approaches deal with bigram statistics, see for instance Kageura (1999);

Read the paper · More papers on PaperTik