The Power of Nested Parallelism in Big Data Processing Hitting Three Flies with One Slap

Gábor Gévay, Jorge-Arnulfo Quiané-Ruiz, Volker Markl · 2021

Many common data analysis tasks, such as performing hyperparameter optimization, processing a partitioned graph, and treating a matrix as a vector of vectors, offer natural opportunities for nested-parallel operations, i.e., launching parallel operations from inside other parallel operations. However, state-of-the-art dataflow engines, such as Spark and Flink, do not support nested parallelism. Users must implement workarounds, causing orders of magnitude slowdowns for their tasks, let alone the implementation effort.

Read the paper · More papers on PaperTik