Leader Set Selection for Low-Latency Geo-Replicated State Machine
Shengyun Liu, Marko Vukolić · IEEE Transactions on Parallel and Distributed Systems · 2016
Modern planetary scale distributed systems largely rely on a State Machine Replication protocol to keep their service reliable, yet it comes with a specific challenge: latency, bounded by the speed of light. In particular, clients of asingle-leaderprotocol, such as Paxos, must communicate with the leader which must in turn communicate with other replicas: inappropriate selection of a leader may result in unnecessary round-trips across the globe. To cope with this limitation, severalall-leaderandleaderlessalternatives have been proposed recently. Unfortunately, none of them fits all circumstances. In this article we argue that the “right” choice of the number of leaders depends on a given replica configuration and the workload. Then we present${\mathsf {Droopy}}$and${\mathsf {Dripple}}$, two sister approaches built upon state machine replication protocols.${\mathsf {Droopy}}$dynamically reconfigures the set of leaders. Whereas,${\mathsf {Dripple}}$coordinates state partitions wisely, so that each partition can be reconfigured (by${\mathsf {Droopy}}$) separately. Our experimental evaluation on Amazon EC2 shows that,${\mathsf {Droopy}}$and${\mathsf {Dripple}}$reduce latency under imbalanced or localized workloads, compared to their native protocol. When most requests are non-commutative, our approaches do not affect the performance of their native protocol and both outperform a state-of-the-art leaderless protocol.