GéDériF: Automatic Generation and Analysis of Morphologically Constructed Lexical Resources

Fiammetta Namer, Georgette Dal · 2000

One of the major frequent problems in text retrieval comes from large number of words encountered which are not listed in general language dictionaries.However, it is very often the case that these words are morphologically complex, and as such have a meaning which is predictable on the basis of their structure.Furthermore, such words typically belong to specialized language uses (e.g.scientific, philosophical or media technolects).Consequently, tools for listing and analysing such words can help enrich a terminological database.The purpose of this paper is to present a system that automatically generates morphologically complex lexical French items which are not listed in dictionaries, and that furthermore provides a structural and semantic analysis of these items.The output of this system is a morphological database (currently in progress) which forms a powerful lexical resource.It will be very useful in Natural Language Processing (NLP) and in IR (Information Retrieval) applications.Indeed the system generates a potentially infinite set of complex (derived) lexical units (henceforth CLUs) automatically associated with a rich array of morpho-semantic features, and is thus capable of dealing morphologically complex structures which are unlisted in dictionaries. Progress Status, or : Processing of CLUs not Listed in Dictionaries? Analysers Based on DictionariesMost of the (rare) automatic systems that give information, no matter how minimal, about CLUs start with closed lexicons 5 .Such is the case, for example, of the French system developed by (Grabar N. & Zweigenbaum P. 1999).Its goal is to constitute a morphological database using the SNODEM medical termonology.3 This project, which brings together Ch.Jacquemin, N. Hathout as well as the two authors of the present work, is funded by the Ministère de l'Education Nationale, de la Recherche et de la Technologie français [French National Ministry of Education, Research and Technology], as part of the program Actions Concertées Incitatives 1999 [Concerted Incitement Actions]. 4 This number corresponds to an estimation of the number of derivatives produced by the affixes -(a)tion, -(at)eur, -able, -age, -aire, -al, dé-, -et(te), -eux, -ifi(er), -is(er), -ité and -oir(e).5 In French, NLP attaches little attention to constructional information (Bouillon P. 1998: 48), which is considered less adapted to the field than inflectional information (Sproat R.W. 1992; Fradin B. 1994).

Read the paper · More papers on PaperTik