Instance-Based Ontology Matching and the Evaluation of Matching Systems

Katrin Zaiß · Univ. Duesseldorf: Duesseldorfer Dokumenten- und Publikationsserver · 2010

The matching of heterogeneous information sources is a crucial task in many different domains. In order to find relations between the different pieces of information, which are annotated using different structures and formats, matching systems have been developed. In the past two decades, ontologies became more and more important as a way to represent the semantics of information in a machine read- and processable way. Hence, many ontology matching systems have been developed as well, which make use of the different parts of ontologies to resolve the heterogeneities. Most systems focus on the exploit of schema or structure information, but ontologies also provide instances, which express the semantics of a concept independent of its meta information. Current instance-based matching methods give room for improvements in several aspects. Matching Systems also need to be evaluated using appropriate test data. Existing benchmarks are not sufficient for testing instance-based methods. In this thesis, we focus on the development of instance-based matching methods, their combination with schema- and structure-based methods and their evaluation. We introduce two novel instance-based matching methods. The first method makes use of regular expressions or sample values to characterize the concepts of an ontology by their instance sets. The second approach uses the instance sets to calculate many different features like average length or the set of frequent values. Both approaches finally compare the characterizations, i.e. the regular expressions or the features, to obtain similarities between the entity sets of two (or more) ontologies. An alignment between the ontologies is then obtained by examining the similarity set. In order to test single matching methods or complex matching systems well-defined test benchmarks have to be available, preferably including the correct alignments to facilitate the evaluation. Current benchmarks do not enable extensive studies on instance-based methods, because the number of instances is significantly too low. We present an additional benchmark, ONTOBI, which can be used to test instance-based methods, but also all other kinds of matching algorithms or systems. Finally, we present MICU, a complex matching system which unifies the advantages of instance-, schema- and structure-based matching methods combined with an efficient user feedback interaction. In order to speed up the process alignments of previous matching cycles are reused.

Read the paper · More papers on PaperTik