Performance Analysis of Large Scale Distributed Systems by Ranking Dominant Features
Debessay Fesehaye, Lenin Singaravelu, Amitabha Banerjee, Ruijin Zhou, Xiaobo Huang, Chien‐Chia Chen, Rajesh Somasundaran · 2017
Large scale distributed systems generate so many metrics/features to monitor and analyze their performances. Most of the analysis to detect performance regressions and root causes of the regressions is done by manually looking at graphs of the metric values. This approach is error prone and doesn't scale. There are some recent works to automate root cause analysis. However these schemes either rely on specific network metrics or aggregate values such as median. Such approach, which doesn't use various physical and virtual device specific metrics, is ineffective in detecting anomalies and their root causes for large scale distributed systems/clusters such as VMware Virtual Storage Area Network (vSAN) and vSphere.