Guessing morphology from terms and corpora

Christian Jacquemin · 1997

This study proposes an algorithm for automatically acquiring morphological Iinks between words.This algorithm relies on the concurrent use of a corpus and a list of multi-word terms, and does not require any prior linguistic knowledge.The four steps of the algorithm are (1) single-word truncation, (2) conflation of multi-word terms, (3) classification and filtering, and (4) clustering of contiation clasea.At each step a precise evaluation is performed in order to chose the optimal parameters.The final results indicate a clustering of 45% of the classes with a prectilon of 87Y0.The derivational knowledge acquired through this method can be used for conceiving a domain-oriented stemmer for scientific and technical corpora.

Read the paper · More papers on PaperTik