A Greedy Consensus-Based Approach to Distributed Job Selection: Toward Fully-Decentralized Workload Management System
Komal Thareja, Krishnan Raghavan, Anirban Mandal, Pawel Zuk, Imtiaz Mahmud, Mariam Kiran, Ewa Deelman · 2025
Current approaches to resilience for highly distributed, heterogeneous, large-scale scientific workflows are limited. Most existing workflow and resource management systems have a single point of failure and resilience strategies are often static, depend on a centralized control, and require considerable design effort from experts. The increasing scale and complexity of workflows coupled with limited resilience capabilities in centralized systems necessitates a fully decentralized, adaptive resource management approach. This paper addresses a very important slice of the overall problem by leveraging the advances in multi-agent systems (MAS). In particular, we explore the suitability of a MAS consisting of globally distributed agents to perform distributed job selection from a dynamic job pool in a truly decentralized, performant, and resilient manner. We present a novel consensus formulation of the distributed job selection problem. By introducing a cost function encapsulating the requirements and constraints of the job and resource loads, we design a novel, greedy consensus algorithm leveraging the Practical Byzantine Fault Tolerance (PBFT)-based consensus method, allowing agents to collectively select jobs in a resilient manner. We compared our algorithms with other state of the art approaches by deploying them in a network testbed infrastructure to emulate distributed job selection. Our evaluation results demonstrated that our greedy consensus algorithm employing the cost-function and PBFT-based consensus method outperforms the ones using the vanilla PBFT-based consensus method - improving scheduling latency by as much as 63.5 % and reducing resource idle time by as much as 63.8 %, with benefits increasing with higher numbers of agents emulated.