Optimisation de la performance des entrepôts de données XML par fragmentation et répartition
Hadj Mahboubi · HAL (Le Centre pour la Communication Scientifique Directe) · 2008
XML data warehouses form an interesting basis for decision-support applications that exploit heterogeneous data from multiple sources. However, XML-native database systems currently suffer from limited performances, both in terms of manageable data volume and response time for complex analytical queries. It is therefore necessary to design methods to optimize performances. In this thesis, we propose to address both these issues by fragmenting and distributing XML data warehouses on grids. To the best of our knowledge, we propose the first fragmentation methods for XML data warehouses. These methods exploit an XQuery workload and output a derived horizontal fragmentation schema. We first adapted the most efficient fragmentation methods from the relational context to XML, and then proposed an original k-means-based fragmentation method that allows mastering the number of fragments. We finally propose an approach aimed at distributing XML data warehouses on grid architectures. Our proposals exploit a unified XML warehouse reference model that we propose to synthesize and enhance related work from the literature. Finally, we experimentally validate our proposal and compare our fragmentation and distribution methods. For this purpose, we designed and developed an XML data warehouse benchmark: XWeB. Our results show that our methods help overcome the data volume and query execution time limitations. They also show that our k-means-based fragmentation method outperforms classical derived horizontal fragmentation methods, both in terms of performance gain and overhead.