Automatic annotation for mapping discovery in data integration systems.

Sonia Bergamaschi, Laura Po, Serena Sorrentino · SEBD · 2008

We propose a CWSD (Combined Word Sense Disambiguation) algorithm for the automatic annotation of structured and semi-structured data sources. Rather than being targeted to textual data sources like most of the traditional WSD algorithms found in the literature, our algorithm can exploit information coming from the structure of the sources together with the lexical knowledge associated with the terms (elements of the schemata). We integrated CWSD in the MOMIS system (Mediator EnvirOment for Multiple Information Sources) [1], which is an I3 framework designed for the integration of data sources, where the lexical annotation of terms was performed manually by the user. CWSD combines a structural disambiguation algorithm that starts the disambiguation of the terms using the semantic relationships extracted from the schemata structural relationships with a WordNet Domains based disambiguation algorithm to re ne terms disambiguation by using domains information. Structural relationships are stored in a Common Thesaurus (CT) generated by the MOMIS system. The CT is a set of relationships describing interand intra-schema knowledge among the source schemas. From a source schema we extract the following relationships: SYN (Synonym-of), de ned between two terms (term is the name of a class/attribute of a schema) that are considered synonyms/equivalent; BT (Broader Terms), de ned between two terms such as the rst one is more general than the second one (the opposite of BT is NT, Narrower Terms); RT (Related Terms) de ned between two terms that are generally used together in the same context. The extracted ODLI3 relationships can be used in the disambiguation process according to a lexical database (in our approach we used WordNet). The algorithm tries to nd a lexical relationship when a CT relationship exists among two terms; in this case we choose the meanings connected by this relationship as the correct ones to disambiguate the terms. The same holds if we nd a chain of lexical relationships that connect terms meanings. The WordNet Domains disambiguation algorithm exploits the information from WordNet Domains. WordNet Domains [2] can be considered an extended version of

Read the paper · More papers on PaperTik