Performance Analysis of Large Scale Distributed Systems by Ranking Dominant Features

Debessay Fesehaye, Lenin Singaravelu, Amitabha Banerjee, Ruijin Zhou, Xiaobo Huang, Chien‐Chia Chen, Rajesh Somasundaran · 2017

Large scale distributed systems generate so many metrics/features to monitor and analyze their performances. Most of the analysis to detect performance regressions and root causes of the regressions is done by manually looking at graphs of the metric values. This approach is error prone and doesn't scale. There are some recent works to automate root cause analysis. However these schemes either rely on specific network metrics or aggregate values such as median. Such approach, which doesn't use various physical and virtual device specific metrics, is ineffective in detecting anomalies and their root causes for large scale distributed systems/clusters such as VMware Virtual Storage Area Network (vSAN) and vSphere.

Read the paper · More papers on PaperTik