Selecting Efficient Cluster Resources for Data Analytics: When and How to Allocate for In-Memory Processing?
Jonathan Will, Lauritz Thamsen, Dominik Scheinert, Odej Kao · 2023
Distributed dataflow systems such as Apache Spark or Apache Flink enable parallel, in-memory data processing on large clusters of commodity hardware. Consequently, the appropriate amount of memory to allocate to the cluster is a crucial consideration.