An adaptive SLA-based data flow mechanism for stream processing engines

Muhammad Iftikhar Hanif, Hyungduk Yoon, Sunglim Jang, Choonhwa Lee · 2017

With the upsurge in the volume of data and profuse cloud applications, big data analytics becomes very popular in the research communities in industry as well as in academia. This led to the emergence of real-time distributed stream processing systems such as Flink, Storm, Spark, and Samza. These systems sanction complex queries on streaming data to be distributed across multiple worker nodes in a cluster. Some of the stream processing systems provide basic supports for controlling the latency and throughput of the system as well as correctness of the results. We present an intelligent and efficient adaptive watermarking and dynamic buffering timeout mechanism for modern distributed stream processing engines. It is designed to increase the overall throughput by making the watermarks of the system adaptive according to incoming workload streams, and dynamically scale the buffering timeout for every task tracker on the fly while maintaining the SLA-based end-to-end latency of the system. Apache Flink is used as testing distributed processing engine in the paper. However, the proposed mechanism can be applied to other streaming frameworks. Our preliminary results indicate that the proposed system outperforms the existing stream processing framework.

Read the paper · More papers on PaperTik