Load Balancing Using Adaptive Overlapped Data Chained Declustering For Distributed File Systems
Vidya Shitole, N. P. Karlekar · 2016
Today in the world of cloud and grid computing integration of data from heterogeneous databases is inevitable. Virtual Database Technology (VDB) is one of the effective solutions for integration of data from heterogeneous sources. This will become complex when size of the database is very large. MapReduce is a new framework specifically designed for processing huge datasets on distributed sources. Apaches Hadoop is an implementation of MapReduce. Distributed file systems (DFS) are key building blocks for cloud computing applications based on the MapReduce programming paradigm. In such file systems, nodes simultaneously serve computing and storage functions; a file is partitioned into a number of chunks allocated in distinct nodes so that MapReduce tasks can be performed in parallel over the nodes. However, in a cloud computing environment, failure is the norm, and nodes may be upgraded, replaced, and added in the system. Files can also be dynamically created, deleted, and appended. This results in load imbalance; that is, the file chunks are not distributed as uniformly as possible in the nodes. The performance of the proposal implemented in the Hadoop distributed file system is further investigated in a cluster environment.