Semi-Automatically Generated Hybrid Ontologies for Information Integration.
Lisa Ehrlinger, Johannes Kepler, Wolfram Wöß · 2015
Large and medium-sized enterprises and organizations are in many cases characterized by a heterogeneous and distributed information system infrastructure. For data processing activities as well as data analytics and mining, it is essential to establish a correct, complete and efficient consolidation of information. Information integration and aggregation are therefore fundamental steps in many analytical workflows. Furthermore, in order to evaluate and classify the result of an integrated data query and thus, the quality of resulting data analytics, it is previously necessary to determine the data quality of each processed data source. This paper aims mainly at the first aspect of the mentioned twofold challenge. Both, data dictionary as well as information source content are analyzed to derive the conceptual schema, which is then provided as a machine-readable description of the information source semantics. Several descriptions of the semantics can be integrated to a global view by eliminating possible redundancies and by applying ontology similarity measures. Attributes for data quality metrics are included in the descriptions but not yet determined. The implementation of the presented approach is evaluated by extracting the semantics of a specific MySQL database, represented as RDF triples.