Benchmarking Distributed Stream Processing Frameworks for Real Time Classical Machine Learning Applications

Merlin Sundar, Sriram Kailasam, Timothy A. Gonsalves · 2020

As the volume of data generated is growing at an unprecedented rate, it becomes important to analyze this data in real-time. To handle the huge volume of data streaming at a high velocity, we not only require powerful machines but also means to distribute the computation involved on the multiple machines. There are several open-source distributed stream processing frameworks such as Apache {Storm, Flink, Spark} and Confluent Kafka for building real-time machine learning applications. Prior works benchmarked some of these platforms using low-level operations like filters, joins, windowed computations etc. Our work includes benchmarking these popular frameworks for their applicability to classical machine learning models: Online K-Means, Online Linear Regression and Online Logistic Regression. We study the following quantitative metrics of evaluation: throughput, latency, CPU, and memory usage. The experiments were conducted in both standalone and clusters setups to determine the scalability of the models. This study will help system designers choose the right model and the right framework, given a specific configuration on streaming data.

Read the paper · More papers on PaperTik