Practical Approach to Monitor Runtime Engine Statistics for Apache Spark and Docker Swarm using Graphite, Grafana, and CollectD
Franz Frederik Walter Viktor Walter Tscharf · 2021
Monitoring runtime engine statistics to identifying and manage bottlenecks and resource availability is a current and heavily debated topic in academics as well as in the industry. Due to the application variety and the complexity of distributed computer clusters, this problem is still important, and many proofs of concept emerge to challenge this issue. This paper aims to summarize the system’s specification, component design, implementation, and evaluation of chosen infrastructure to develop a monitoring engine for exposing JVM and network-related statistics of the mentioned system using Docker Swarm, Apache Spark, Graphite, and Grafana.