An Approach for Optimizing the Performance for Apache Spark Applications
Preeti Gupta, Arun Sharma, Rajni Jindal · 2018
Apache Spark has become the de-facto processing framework for big data analytics. The main challenges for big data analytics is to manage diverse varieties, store huge volumes and process data at high speed. Apache Spark provides a number of advantages over MapReduce due to its in-memory processing. The default parameters for running the Spark applications are based on commodity hardware and may not provide one-size-fits-all solution for every configuration and environment. Tuning resource allocation for Spark applications is of paramount importance to achieve optimal performance. This article discusses various parameters/options such as caching, broadcast variables, repartitioning and number of executors that may be tuned according to the environment resulting in better performance of Spark applications.