Contributions à la conception et l’exploitation de systèmes d’intégration de données

Ladjel Bellatreche · HAL (Le Centre pour la Communication Scientifique Directe) · 2009

With the development of the Internet and Intranets, it has become crucial to exchange and share an enormous quantity of information from various data sources scattered across the Web or within different organizations. To meet these needs, integration solutions have been proposed along three dimensions: data integration, application integration, and platform integration. The work presented in this thesis aims to propose innovative solutions for building a data integration system. A comprehensive approach to the development of an integration system is presented. It is structured around three main phases: building an integration system, operating it, and customizing it.For the construction phase, we have proposed an automatic semantic integration approach, while leaving each of the sources likely to be integrated with significant autonomy in terms of both its structure and its evolution. It assumes that each source contains both its own ontology and the semantic relations that link it a priori with one or more shared ontologies. Such a source is called an ontology-based data source (OBDS). To implement our integration approach, we first proposed a model and architecture for managing ontology-based data sources. This architecture is made up of four parts: the first two correspond to the usual database structure: data based on a logical data schema, and a meta-base describing the entire table structure. The other two, original parts, respectively represent ontologies and the ontology meta-model within a reflexive meta-model. Abstraction and naming mechanisms enable each piece of data to be associated with the ontological concept that defines its meaning, and data to be accessed from concepts, without having to worry about data representation.For the data exploitation phase, we presented solutions to provide administrators with query optimization structures. Since we have been working on this phase since 1996, we have proposed optimization solutions that can be applied to integration systems following an architecture materialized in the form of a traditional database or a relational data warehouse. We have identified two types of optimization structure selection: isolated selection and multiple selection. For isolated selection, we presented algorithms for horizontal fragmentation and join indexes. For multiple selection, we studied the problem of selecting binary join indexes and derived fragmentation by exploiting the similarities between them. Other problems are also presented, such as parallel processing and resource allocation between redundant structures (materialized views and indexes). To facilitate administration tasks, we have developed a tool to assist administrators in their tasks, which can be used before or after the creation of an integration system. Personalization is a recent phase in our work. We first studied its effect on the selection of optimization structures. Recently, we have proposed solutions for the representation of user profiles within a database.Personalization is a recent phase in our work. We first studied its effect on the selection of optimization structures. More recently, we have proposed solutions for representing user profiles within a BDBO to facilitate their sharing and exchange.

Read the paper · More papers on PaperTik