Using Comprehensive Analysis for Performance Debugging in Distributed Storage Systems

Albert Wingnang Leung, E. Lalonde, J. Telleen, James Davis, Carlos Maltzahn · 2007

Achieving perfonnance, reliability, and scalability presents a unique set of challenges for large distributed storage. To identify problem areas, there must be a way for developers to have a comprehensive view of the entire storage system. That is, users must be able to understand The current standard for visualizing system perfor both node specific behavior and complex relationships be mance is to log and graph relevant performance counters. tween nodes. We present a distributed file system profiling This is appropriate when seeking knowledge of individual method that supports such analysis. Our approach is based metrics, but as system size grows, the usefulness of these on combining node-specific metrics into a single cohesive techniques diminishes. For example, users can easily view system image. This affords users two views of the storage the throughput of any single node as a graph and be satis system: a micro, per-node view, as well as, a macro, multified, but the log and graph approach fails when the goal is node view, allowing both node-specific and complex inter to convey more complex concepts such as how an individnodal problems to be debugged. We visualize the storage ual node failure impacts overall resource availability. system by displaying nodes and intuitively animating their metrics and behavior allowing easy analysis of complex problems. 1.

Read the paper · More papers on PaperTik