ARC A3: A Method for Evaluating Term Extracting Tools and/or Semantic Relations between Terms from Corpora
Christophe Jouis, ARC A3 · 2000
This paper describes an ongoing project evaluating Natural Language Processing (NLP) systems 1 .The aim of this project is to test software capabilities in automatic or semi-automatic extraction of terminology from French corpora in order to build tools used in NLP applications.We are putting forward a strategy based on qualitative evaluation.The idea is to submit the results to specialists (i.e.field specialists, terminologists and/or knowledge engineers).Building terminology (terms or concept names and the logic-semantic relations they hold) from extensive textual data is not a simple task when the designer has to examine a new field of knowledge.The designer may not be acquainted with the representation of the field, its structures and the articulations between its objects.To make the designer's task easier, natural language processing systems can be of help particularly those dedicated to the identification of terms or concepts names related to a specific field of knowledge (construction of a reference terminology) and the logic-semantic relations they contain.These systems can be applied to the modeling and designing of the following types of systems : (1) The modeling of an object-oriented database design (static aspects: i.e. describing the structure), (2) Knowledge-based systems : modeling the hierarchies between classes and the relations between the objets concerned by a set of rules, (3) Modeling the conceptual design of a relational database (domains, relations, coherence maintenance), (4) Thesaurus construction (documentary databases, Information Retrieval, ...), (5) Terminological database construction, and so on.The Natural Language Processing systems we are evaluating use various modules in order to identify terms or concept names and the logic-semantic relations they hold.The approaches involved in corpus analysis are either based on morphosyntactic analysis, statistical analysis, semantic analysis, recent connectionist models or any combination of two or more of these approaches.Most of these systems need, in addition, a general language dictionary, a glossary of technical terms covering the relevant field, etc.The identification of terms is in fact an extraction of noun phrases corresponding to the concepts representing the field of knowledge.In their current state, these systems are mostly semi-automatic processing tools.In this paper, we will examine the evaluation problem. 1 The research we are conducting is sponsored by the "Association des Universites Francophones" (AUF) an international Organisation whose mission is to promote the dissemination of French as a scientific medium.Software submitted to this evaluation are conceived by French, Canadian and US research institutions (National Scientific Research Centre and Universities) and/or companies : CNRS (France), XEROX, and LOGOS Corporation among others.