Performance analysis on Raft-based fault-tolerant and scalable applications
Yu Zhang · 2024
Fault tolerance is a cornerstone of distributed systems, ensuring their uninterrupted operation and resilience in the face of failures. This thesis investigates fault tolerance within distributed key-value storage systems through the lens of the Raft consensus algorithm. The Raft consensus algorithm, recognized for its design simplicity and effectiveness in fault tolerance, stands out as a crucial component in these systems. Yet, its practical application in dynamic environments, marked by variable network and disk latencies, presents intricate challenges that are not fully understood. This thesis critically examines the impact of network and disk latencies on the operational performance and scalability of Raft-based key-value storage systems. It delves into how these latencies affect the key performance metrics of the consensus process, including leader election speed, client request handling, and data shard redistribution during scaling processes. This exploration is pivotal for the design of resilient distributed storage solutions that can maintain high performance under the influence of inherent latency challenges. Furthermore, this study assesses the potential of Raft as a comprehensive solution for achieving fault tolerance and scalability in distributed storage services. Through empirical analysis, the research aims to clarify Raft’s ability to support a system’s expansion and manage data efficiently amidst growth, while concurrently ensuring system reliability against node failures and network partitions. The findings aim to contribute significant empirical insights and practical guidance to the field, helping understand how disk and network configuration can affect the performance of fault-tolerant applications.