Minimizing Data Access Latencies via Virtual Machine Placement Method in Datacenter

Yong Zhang, Xinyan Zhang · 2017

The processing framework of large-scale data is becoming a major concern due to an explosive growth of data intensive applications in the cloud environment, such as MapReduce/Hadoop architecture. Many virtual machines (VMs) are used for processing large-scale data of cloud applications. Therefore, the total completion time of a task is an important index to evaluate the cloud performance. The access latency between nodes is one of the key factors affecting the task completion time for computing-intensive applications. Additionally, minimizing total access time can reduce the overall bandwidth cost of running the job. This paper proposes an optimization model focused on optimizing VMs placement so as to minimize the total data access latency where the data sets have been located. According to the proposed model, our VMs optimization problem is linear programming. Therefore, we obtain the optimum solution of our model by the branch-andbound algorithm that its time complexity is O(2NM) (where N is the number of the data nodes and M is the number of VMs). Simultaneously, we also present a greedy algorithm, which has O(NM) of time complexity, to solve our model. Finally, the simulation results show that all of the solutions of our model are superior to existing models and close to the optimal value.

Read the paper · More papers on PaperTik