A Data Mining-based Blocks Placement Optimization for Distributed Data Warehouses
Billel Arres, Nadia Kabachi, Omar Boussaïd, Fadila Bentayeb · 2014
The amount of data that is captured and generated by modern computing devices has augmented exponentially over the last years. The Hadoop framework - an open source project based on the MapReduce paradigm - is a popular choice for processing these large volumes of data or big data. However, the performance gained from Hadoop's features is currently limited by its default block placement policy, which does not take any data characteristics into account. This is particularly true for relational data bases and data warehouses. Indeed, the efficiency of many operations can be improved by a careful data placement, including indexing, grouping, aggregation and joins. In this paper we propose a data warehouse distribution strategy to improve query gain performances on multi-nodes clusters, especially Hadoop clusters. Based on k-means clustering method that allows to master the number of clusters through its k parameter, we investigate the performance gain for OLAP cube construction with and without data organization. And this, by varying the number of clusters and data warehouse size. Our experiments suggest that a good data placement on a cluster during the implementation of the data warehouse increase significantly the OLAP cube construction and querying performances.