Extending an existing specialized semantic lexicon

Benoît Habert, Adeline Nazarenko, Pierre Zweigenbaum, Jacques Bouaud · 1998

There is a constant need to extend and tune specialized vocabu-laries to account for new words and new word usages. This paper addresses the issue of characterizing the semantic class of such words. We test the hypothesis that the analysis of word distribu-tion in a representative corpus, as obtained by robust NLP tools, can help identify words with similar meanings, and to decide on the most likely category for a given word based on the categories of its neighbors. We report on an experiment with a moderate-size corpus of patient discharge summaries collected during the MENELAS project, taking as categories the high-level axes of the SNOMED nomenclature, and processing the corpus with the ZEL-LIG suite of tools. We attempt to quantify the extent to which this process succeeds in proposing a correct category for a given word of the corpus while we vary several parameters of the method. The percentage of correctly categorized words (precision) ranges between 50 and 75 %, while the best percentage of categorized words (recall) is 37 % for the whole categorization process. Cate-gorization results are significantly above chance, but not sufficient for a fully-automated process. We discuss possible uses of such a categorization help and identify further directions for improve-ment. 1

Read the paper · More papers on PaperTik