Managing Big Data Stream Pipelines Using Graphical Service Mesh Tools

Muhammad Faizan, Christian Prehofer · 2021

Current big data frameworks like Apache Flink and Spark enable efficient processing of large-scale streaming data in a distributed setup. For the management of such data pipelines and the computing resources, we propose a combination of a graphical tool for pipeline management, Apache StreamPipes, and container management tools like Kubernetes. For evaluation, we implemented a use case with data preprocessing, vehicle power consumption, and driving behavior services in StreamPipes. We discuss the capabilities of StreamPipes in managing and executing complex stream processing pipelines and also evaluate the possible integration of container and service mesh tools (i.e., Istio) with StreamPipes. Furthermore, we implemented and evaluated a service management layer in our system design to provide extended features. In particular, we evaluated the delay when such a complex pipeline is restarted, e.g. for updates or reconfiguration.

Read the paper · More papers on PaperTik