Early Work on Characterizing Performance Anomalies in Hadoop

Puja Makhanlal Gupta, Christopher Stewart · 2016

Hadoop processes a huge variety of data in diverse operating environments, ranging from large, heterogeneous wharehouse-scale computers to enterprise server rooms. There is not a unique execution setting for these systems which can guarantee best performance for each of the query. In this paper, we establish how significant a particular execution parameters are Hadoop clusters. Specifically, we characterize the parameter space using using decision trees. We studied parameter features with finite discrete domains. Our tree learns the importance of each parameter by splitting source sets, each set is comprises performance data of many executions, into subsets based on attribute values. The attribute to split dataset on is selected based on maximum information gain or lowest entropy. We have found that the impact an execution parameter can have on target attribute can be related to distance of that feature node from the root of the constructed decision tree. From initial results the percent change in values of target attribute for various value of feature node which is closer to root is 6 times larger than when that same feature node is one more level away from root node.

Read the paper · More papers on PaperTik