Green- and heterogeneity-aware partitioning for data analytics

Aniket Chakrabarti, Srinivasan Parthasarathy, Christopher Stewart · 2016

Distributed algorithms for analytics partition their input data across many machines for parallel execution. At scale, it is likely that some machines will perform worse than others because they are slower, power constrained or dependent on undesirable, dirty energy sources. It is even more challenging to balance analytics workloads as the algorithms are sensitive to statistical skew across data partitions and not just partition size. Here, we propose a lightweight framework that controls the statistical distribution of each partition and sizes partitions according to the heterogeneity of the environment. We model heterogeneity as a multi-objective optimization, with the objectives being functions for execution time and dirty energy consumption. We use stratification to control data skew and model heterogeneity-aware partitioning. We then discover Pareto-optimal partitioning strategies. We built our partitioning framework atop Redis and measured its performance on data mining workloads with realistic data sets. Our framework simultaneously achieved 34% reduction in time and 21% reduction in dirty energy usage for a popular webgraph compression algorithm using 8 partitions.

Read the paper · More papers on PaperTik