Resource allocation in parallel shared-nothing database systems

Manish Mehta · 1995

Parallel database management systems are increasingly being used for high-performance applications that require efficient access to large amounts of data, e.g. large-scale transaction processing, decision-support systems, database mining, and multimedia systems. Typically, these systems are designed using a shared-nothing architecture. The modular design of shared-nothing architecture enables incremental growth and scalability to hundreds of nodes. However, the performance potential of such large configurations can be fully exploited only through efficient resource management. This thesis explores three key resource allocation issues in parallel shared-nothing database systems: Memory Management, Data Placement and Resource Allocation. The algorithms developed for solving each of these issues can be combined to develop a comprehensive resource allocation algorithm for parallel shared-nothing database systems. The emphasis of the thesis is on developing techniques that can efficiently execute complex workloads as well as adapt to dynamic changes in the workload. The first part of the thesis studies memory management for workloads consisting of multiple classes, each with varying resource requirements and response time constraints. Since memory management for such complex workloads has not yet been adequately addressed for centralized database systems, the thesis studies memory management in a centralized database system. The thesis presents two feedback-oriented algorithms for managing workloads that consist of multiple query classes. The algorithms can dynamically adapt to different workloads by automatically determining the multi-programming levels and memory allocation for each query class. The second part of the thesis investigates data placement in parallel shared-nothing database systems. The thesis investigates methods to determine the number of sites on which to partition each relation and selecting the particular sites to place each partition. Full declustering, which places each relation on all sites, is shown to provide the best overall performance because it improves resource utilization and provides the best load balancing. The third part of the thesis investigates algorithms for processor allocation. It is shown that algorithms that dynamically adapt to different workloads and configurations can improve performance significantly.

Read the paper · More papers on PaperTik