Computing data cubes over GPU clusters.

Lucas Henrique Moreira Silva · 2018

The data cube is a fundamental relational operator for decision support systems, thus very important for analytics. Unfortunately, a full data cube with all of its tuples has exponential complexity in terms of runtime and memory consumption as the dimensions increase linearly, so algorithms to reduce query response times continue under development. The problem stated in this work is: how can we reduce complex multidimensional queries response times from high dimensional data cubes? The problem is aggravated if recurrent updates occur and if there is a huge volume of high dimensional data to be managed. The hypothesis of this work is that clusters of CPU-GPU devices can speedup queries from high dimensional holistic data cubes that are updated constantly. The alternative solution presented in this work, named JCL-GPU-Cubing, partitions the base relation into multiple independent sub- cubes. These multiple sub-cubes represent a partial data cube to reduce the exponentiality and they are used to perform queries in CPU or in CPU-GPU computer architectures efficiently. The experimental evaluations using complex multidimensional queries demonstrated that the CPU cluster version scaled well when the base relation increased and the CPU-GPU version outperformed the CPU only version in certain scenarios.

Read the paper · More papers on PaperTik