Construction d'une cartographie de domaine à partir de ressources sémantiques hétérogènes
Sarra Ben Abbès · HAL (Le Centre pour la Communication Scientifique Directe) · 2013
In recent years, a large effort has been devoted to the sharing and reuse of knowledge in the semantic web. A number of semantic resources have been made available significantly but this richness and diversity complicate the reuse of existing resources. Before trying to reuse resources, the knowledge engineer must begin by drawing up an inventory of the domain's interest, identify available resources and put it against each other. He must locate in a "teeming" web where the semantic heterogeneity is growing. This is the challenge that this thesis has sought to address. We propose to map the semantic web on a particular domain and we have developed an automatic method to create such "domain cartographies" for engineers wishing to gain insught in the available resources on a particular domain. In order to have these cartographies suited to their interests, engineers give as an input a text which is representative of the domain and the target application and the text is a pivot to the whole the cartography process. We have defined an overall process of building cartographies that is composed of three steps, assuming resolved the prior selection step of semantic resources. The text acquisition is a starting point, the first step is to link the selected resources to the text : it is a step of semantic resources annotation which projects ontological entities on the text in order to identify those which are mentioned. The second step is to align different resources and the alignment process itself is also based on the text : the entities of different resources are aligned based on their distributional proximity in the text. The third step is to detect and correct the mapping anomalies obtained after the alignment step ; this step can also detect and display the remarkable mappings that can help the engineer to take knowledge of the existing resources. The originality of our ontology alignment method is to use the text. The alignment consists typically on identifying relations or mappings between entities from different ontologies. It is used in several applications, such as the ontology enrichment, the ontology merging or here, to help the knowledge engineer to understand the available resources of a particular domain. Using a text allows closer ontological entities that do not have the same label - as in the lexical alignment methods - or the same position in the original resources - as in structural approaches. We propose to distinguish between two types of mappings with the nature of the relationship between related terms : if the terms or labels tend to appear in the same sentences, we conclude an associative relationship between them ; on the contrary, if they are substituted to each other, we can conclude an equivalence relation. The building cartography process is based on the results of the alignment process. This step consists on identifying anomalies and to present the remarkable mappings to the knowledge engineer. We make the hypothesis that these two types of configurations are important to identify : anomalies show significant differences in the choice of conceptualization represented on the two aligned resources ; remarkable configurations are instead highlighting the most central parts and most consistent of the proposed alignment. Identify and analyze these configurations must allow the engineer to understand how resources are aligned to each other. This overall approach was implemented and tested on different use cases in the biology, geography and food domains.