Automated Data Management In Cloud Computing
Arkaitz Ruiz Alvarez · 2012
Scientists are increasingly relying on computational resources, both compute and storage, to expand scientific knowledge.For example, the data deluge is quickly overcoming the capacity of storage systems and the increasing use of simulation requires large compute capabilities.Thus, scientists need to expand their local resources with highly available and scalable systems.We consider cloud computing to be the solution that provides scientific applications with the computational resources needed.However, the services offered by the cloud providers do not address several important issues: how to meet the data requirements with the storage systems available, and how to optimize cost and other performance metrics.The variety of storage and compute choices with different characteristics and prices, the growth of the data stored in terms of size and number and the data management requirements make these tasks overwhelmingly complex for individual users.To address these challenges, we focus on four key elements of data management: the analysis of current storage services, the expression of data requirements and storage capabilities, data management algorithms and data-aware scheduling algorithms.We combine the information from our analysis of the storage services with their capabilities in a machine-readable format that can be processed by our implementation of the user's data requirements.Thus, we can obtain within a few milliseconds a list of storage services per application dataset that meet the user's requirements, and provide cost and performance estimates.Our unique approach to data management generates an integer linear programming problem with this list.The solution to this problem is an optimal assignment of the application's data to cloud services.Our implementation can provide optimal Recently, several companies like Amazon, Google, and Microsoft have started providing remote computational resources on a pay-as-you-go basis, which has been introduced as cloud computing.Computational resources are allocated in their huge datacenters and provide the users with great flexibility: resource usage can scale up or down, resource capacity appears to be unlimited to individual users, there are no upfront costs, and the pricing is competitive with local resources.One of the main motivations for the introduction of cloud computing is the low average utilization of datacenters.Datacenters are sized to handle peak loads (e.g.customer orders during the holiday season for Amazon), which leads to resources being underutilized most of the time since the average load is much smaller than peak loads.If there are enough idle resources in Amazon's datacenters then it makes economic sense to rent some of the excess capacity to external users.Also, cloud providers can take advantage of the economies of scale to build datacenters that support the computational needs