The Power of Nested Parallelism in Big Data Processing Hitting Three Flies with One Slap
Gábor Gévay, Jorge-Arnulfo Quiané-Ruiz, Volker Markl · 2021
Many common data analysis tasks, such as performing hyperparameter optimization, processing a partitioned graph, and treating a matrix as a vector of vectors, offer natural opportunities for nested-parallel operations, i.e., launching parallel operations from inside other parallel operations. However, state-of-the-art dataflow engines, such as Spark and Flink, do not support nested parallelism. Users must implement workarounds, causing orders of magnitude slowdowns for their tasks, let alone the implementation effort.