TOPOSCH: Latency-Aware Scheduling Based on Critical Path Analysis on Shared YARN Clusters
Chunming Hu, Jianyong Zhu, Renyu Yang, Hao Ming Peng, Tianyu Wo, Shiqing Xue, Xiaoqiang Yu, Jie Xu, Rajiv Kumar Ranjan · 2020
Balancing resource utilization and application QoS is a long-standing research topic in cluster resource management. Big data YARN clusters need to co-schedule diverse workloads on shared resources including batch processing jobs, streaming jobs, and other long-running applications such as web services, database services, etc. Current resource managers are only responsible for resource allocation among applications/jobs but completely unaware of runtime QoS requirements of interactive and latency-sensitive applications. Prior works to maximize the QoS of monolithic applications ignore inherent dependencies and temporal-spatio performance variability of components, characteristics of distributed applications primarily driven by microservices. In this paper, we present Toposch, a new resource management system to adaptively co-locate batch tasks and microservices by harvesting runtime latency. In particular, Toposch tracks full footprints of every request across microservices over time. A latency graph is periodically generated for identifying victim microservices through an end-to-end latency critical path analysis. We then exploit per-microservice and per-node risk assessment to gauge the visible resources to the capacity scheduler in YARN. Execution of batch tasks are adaptively throttled or delayed, thereby avoiding latency increase due to node over-saturation. TOPOSCH is integrated with YARN and experiments show that the latency of DLRAs can be reduced by up to 39.8% against the default capacity scheduling in YARN.