A Logical Approach to the Lemmatisation of Computational Lexica

Jon Mills · Kent Academic Repository (University of Kent) · 1999

Lemmatisation is a crucial part of the compilation of a computational lexicon; it is the process which determines the selection and presentation of lemmata. Lemmatisation is a non-trivial task; it consists of a good deal more than deinflection to identify a baseform that can serve as a headword. The lemma has three functions: to uniquely identify the lexical unit, to locate it in the system, and to describe its form. The computational lexicographer is confronted by a number of problems. Homographs need to be distinguished. Several variants of the baseform may exist from which a preferred form will have to be chosen. Compounds may be written as solid, hyphenated or as two words. A way has to be found to treat multi-word lexemes. It may be necessary to give some very common affixes main-entry status. A decision has to be made whether to treat derivatives as main entries with cross reference to the baseform or regroup them under the baseform. A solution is suggested in which the preferred baseform together with a number of other distinguishers may be satisfactorily employed to fulfil all the functions of the lemma. Next it is shown how these elements may be placed within a logical framework to implement computational lexica. The model is then extended to deal with problems of asymmetry in interlingual lemmatisation.

Read the paper · More papers on PaperTik