Optimizing performance of Real-Time Big Data stateful streaming applications on Cloud

Amit Gupta, Sushant Jain · 2022

Exponential growth in the volume of data generated over the last decade has triggered massive research and adoption of distributed big data analytics platforms. In real-time streaming analytics, data is received as a continuous event stream. These events are iteratively correlated and analyzed in micro-batches till the processing reaches a logical conclusive stage. For every iteration, data received as part of the new batch is correlated with the results of the previous iteration and cached as interim results. This real-time stateful analysis of streaming data has been made possible by distributed data-intensive computing frameworks like Apache Spark. Though the spark state storage is highly optimized, its access is limited within the context of the Spark job. Since, these interim results may need to be consumed by applications running external to Spark context, saving these results in external memory involves a significant overhead due to the high volume of I/O. Similar challenges are observed when the spark job needs to consume from an external data store. These challenges get compounded in cloud-based deployment. This paper presents the optimized data storage and retrieval pattern from the spark job to/from the independent storage using the standard off the shelf tools. Our focus on standard off the shelf storage is to ensure that the design patterns presented as part of this paper could be adopted for production deployment.

Read the paper · More papers on PaperTik