dSpark: Deadline-Based Resource Allocation for Big Data Applications in Apache Spark

Muhammed Tawfiqul Islam, Shanika A. Karunasekera, Rajkumar Buyya · 2017

Large-scale data processing framework like Apache Spark is becoming more popular to process large amounts of data either in a local or a cloud deployed cluster. When an application is deployed in a Spark cluster, all the resources are allocated to it unless users manually set a limit on the available resources. In addition, it is not possible to impose any user-specific constraints and minimize the cost of running applications. In this paper, we present dSpark, a lightweight, pluggable resource allocation framework for Apache Spark. In dSpark, we have modelled the application completion time with respect to the number of executors and application input/iteration. This model is further used in our proposed resource allocation model where a deadlinebased, cost-efficient resource allocation scheme can be selected for any application. As opposed to the existing frameworks that focus more on modelling the number of VMs to use for an application, we have modelled both the application cost and completion time with respect to executors, hence providing a finegrained resource allocation scheme. In addition, users do not need to specify any application types in dSpark. We have evaluated our proposed framework through extensive experimentation, which shows significant performance benefits. The application completion time prediction model has a mean relative error (RE) less than 7% for different types of applications. Furthermore, we have shown that our proposed resource allocation model minimizes the cost of running applications and selects effective resource allocation schemes under varying user-specific deadlines.

Read the paper · More papers on PaperTik