Enhanced Cost Estimation Model for Replica Selection

Kama Hazira Abdul Kadir, Gian Chand Sodhy · 2014

A data grid allows large amount of data to be stored and shared among users at geographically different locations. However, when required, data has to be copied to the location where it is needed. This means a lot of network bandwidth will be used for the data transfer. One strategy employed in data grids is to make multiple copies of the data, called replicas, and place them in strategic locations so that the data is not too far from where it might be needed. When data is required, a user now has to choose among the different replicas. Hence, the need for best replica selection to reduce time, effort and resources needed to bring that data to where it is required. Best replica selection tries to estimate the cost involved in copying the data from the alternative replica sites. Most replica cost estimation models use only a few parameters to make the calculations, hence limiting their usage for specific purposes. In this work, we include multiple parameters, taking into account the characteristics of the replica sites as well the network links between the user site and replica sites. We combine the parameters into a formula, and test various scenarios in a simulator. Our results show that this enhanced model can consistently select best replica, and is comparable to other models. Keywords—replica selection, cost estimation model, data replication, data grids

Read the paper · More papers on PaperTik