Automatic Retrieval of Definitions in Texts, in Accordance with a General Linguistic Ontology.

Charles Teissèdre, Brahim Djioua, Jean-Pierre Desclés · The Florida AI Research Society · 2008

A semantics of definition category and its sub-categories used in texts, organized in a semantic map, requires a more complex structuring of the domain underlying the meaning representations than is commonly assumed. This paper proposes a three-layer ontology in which the notion of definition takes part and indicates how it can be used in Information Retrieval. The first part describes an automatic process to annotate definitions based on linguistic knowledge, in accordance with a general linguistic ontology, and the second part shows a practical use of its semantic and discourse organizations in retrieving information through the Web. General introduction A large variety of information processing applications deals with natural language texts. Many of these applications require extracting and processing the meanings of texts, in addition to processing their morphosyntactic forms. In order to extract meanings from texts and manipulate them, a natural language processing system must have a significant amount of knowledge about the organization of semantic and discourse notions. We can also observe that the focus of modern information systems is moving from ”data processing” and ”concept processing” towards ”relation between concept processing”, which means that the basic unit of processing is less and less an atomic piece of data and tends to be some more general semantic and discourse organization of texts. We already made a process for automatic building of domain ontology with semantic and discourse relations related to a linguistic general ontology and terminology for a specific domain [Le Priol et ali., 07]. The debate about general and domain ontology is about which approach to domain categorization is ’best’ ? [Poesio, 05]: Copyright © 2008, Association for the Advancement of Artificial Intelligence (www.aaai.org). All rights reserved. • Designing a clean, elegant ontology with a clear semantic based on sound philosophical principles and scientific evidence, • Relying on evidence from psychology and corpora, and on machine learning techniques, to acquire automatically, as far as possible a domain structure that in most cases will be rather messy. This paper describes the semantic and discourse organization of the linguistic notion of ”definition” in accordance with a general linguistic ontology and its use in Information Retrieval. After a presentation of some general ontologies, we describe the notion of definition in texts, before showing a practical use in information retrieval through the Web. General ontologies versus Domain ontologies According to Wikipedia a general ontology is defined as following: In information science, an upper ontology (toplevel ontology, or foundation ontology) is an attempt to create an ontology which describes very general concepts that are the same across all domains... The goal of this is to construct broad accessible ontologies resulting from these Upper-Ontologies. An UpperOntology is often presented in the form of a hierarchy of entities and of their associated rules (theorems and constraints) which try not to hold account of a particular issue in specific domain. It appears increasingly that more than the domain ontologies, general ontologies are economically relevant. 1 http://en.wikipedia.org/wiki/Upper_ontology_\%28computer_science\% 29) A Proceedings of the Twenty-First International FLAIRS Conference (2008)

Read the paper · More papers on PaperTik