A workload aware model of computational resource selection for big data applications
Amit Kumar Gupta, Weijia Xu, Natalia Ruiz Juri, Kenneth A. Perrine · 2016
Workload characterization of Big Data applications has always been a challenging research problem. Big data applications often have high demands on multiple computing components in concert, such as storage, memory, network and processors and have evolving performance characteristics along with the scale of the workload. To further complicate the problem, the increasing diversity of hardware technologies available makes side-by-side comparisons hard. Choosing right resources among a wide array of available systems is a decision that is likely to plague both end users and resources providers. In this paper, we propose a workload aware model for the computational infrastructure selection problem for a given application. Our model considers both features of the workload and features of the computational infrastructure and predicts expected performance for a given workload, based on historical performance results using Support Vector Machines (SVM). We tested our model with a practical application from the domain of Transportation research on two distinct computing resources. The application has significant requirements on both memory availability and processing power. Therefore the optimal performance of the application is a dedicated trade-off between different types of resources and it is workload specific. The two testing systems represent two main trends in high performance computing resources. One infrastructure is a traditional high end computing cluster consisting of moderate number of CPUs and memories running at high frequency and high bandwidth. The other system, based on the latest Intel Knights Landing processor, is a good representation of the trending Many-Core technology in which high number of processing cores running at lower frequencies are available. The memory allocation models are also often different between the two systems. Our results show that our proposed model can achieve over 90% accuracy in performance prediction with small training data sets for our test application. The results also indicate that our model is a viable approach to be extended to other classes of applications and to be potentially adopted by high performance computing resource providers.