A Comparative Study of Spark Schedulers' Performance
Ashmita Raju, Ramya Ramanathan, Hemavathy R. V. · 2019
Big data applications have become an integral part of many intelligent systems, enabling better business decision making, by extracting useful information from historical data. This involves a large amount of computational tasks and algorithms. In such scenarios, the time taken by the processing become an important concern, and needs to be optimized. Apache Spark is a high speed big data analytics engine, best known for its fast computation speeds, due to the use of in-memory data structures and its efficient schedulers. Further optimization to enhance the speed of Spark jobs will be valuable. In this paper, the efficiencies of the four standard schedulers of Apache Spark i.e., Standalone Scheduler, Mesos, YARN(Yet Another Resource Negotiator) and the recently introduced Kubernetes Scheduler are compared. The speeds of these schedulers are compared for jobs of different sizes, and the best alternative is identified.