Wide-Area Spark Streaming: Automated Routing and Batch Sizing
Wenxin Li, Di Tao Niu, Yinan Liu, Shuhao Liu, Baochun Li · IEEE Transactions on Parallel and Distributed Systems · 2018
Modern stream processing frameworks, such as Spark Streaming, are designed to support a wide variety of stream processing applications, such as real-time data analytics in social networks. As the volume of data to be processed increases rapidly, there is a pressing need for processing them across multiple geo-distributed datacenters. However, these frameworks are not designed to take limited and varying inter-datacenter bandwidth into account, leading to longer query latencies. In this paper, we present the design and implementation of an extended Spark Streaming framework to automatically and optimally schedule tasks, select data flow routes and determine micro-batch sizes across geo-distributed datacenters in wide-area networks. To make these decisions, we propose a sparsity-regularized ADMM algorithm to efficiently solve a nonconvex optimization problem, based on readily measurable operating traces. Toward incremental real-world deployment, we take a non-intrusive approach to support flexible routing of micro-batches by adding a new DStream transformation we have developed to the existing Spark Streaming framework. As a result, our implementation can enforce scheduling decisions by modifying application workflows only. We have deployed our implementation on Amazon EC2 with emulated bandwidth constraints, and our experimental results on various types of queries have demonstrated the effectiveness of our proposed framework, as compared to the existing Spark Streaming scheduler and other data-locality-based heuristics.