Query-load Agnostic Pre-processing for Multi-dimensional Data Sets
Aditya Manthramurthy · 2010
Implementing a tera-scale, native multi-dimensional database for data cube based analytics applications requires new ways of loading data, materialising views and partitioning it on multiple machines to achieve efficient query processing. In this work, we give a heuristic based technique to choose views for partial materialisation and partitioning data on commodity clusters to minimise query processing time, while being query-load agnostic. The technique is based on using association rule mining to find correlations between attributes in the database. Since the pre-processing is query load agnostic, we have the advantage of not having to redo pre-processing of data for different query patterns, and remaining flexible enough to be able to apply other techniques for distributed databases such as horizontal partitioning on top of this form of partitioning.