Cluster-based similarity aggregation for ontology matching

Quang-Vinh Tran, Ryutaro Ichise, Bao-Quoc Ho · 2011

Abstract. Cluster-based similarity aggregation (CSA) is an automatic similarity aggregating system for ontology matching. The system have two main part. The first is calculation and combination of different similarity measures. The second is extracting alignment. The system first calculates five different basic measures to create five similarity matrixes, i.e, string-based similarity measure, WordNet-based similarity measure... Furthermore, it exploits the advantage of each mea-sure through a weight estimation process. These similarity matrixes are combined into a final similarity matrix. After that, the pre-alignment is extracted from this matrix. Finally, to increase the accuracy of the system, the pruning process is applied. 1 Presentation of the system In the Internet, ontologies are widely used to provide semantic to data. Since they are created by different users for different purposes, we need to develop a method to match multiple ontologies for integrating data from different resources [2]. 1.1 State, purpose, general statement CSA (Cluster-based Similarity Aggregation) is the automatic weight aggregating sys-tem for ontology alignment. The system is designed to search for semantic correspon-dence between heterogeneous data sources from different ontologies. The current im-plementation only support one-to-one alignment between concepts and properties (in-cluding object properties and data properties). The core of CSA is utilizing the advan-tage of each basic strategy for the alignment process. For example, the string-based similarity measure works well when the two entities are similar linguistically while the structure-based similarity measure is effective when the two entities are similar in their local structure. The system automatically combines many similarity measurements based on the analysis of their similarity matrix. Details of the system are described in the following parts.

Read the paper · More papers on PaperTik