Analysis of Data Placement Strategy based on Computing Power of Nodes on Heterogeneous Hadoop Clusters

Sanket Chintapalli · 2014

Hadoop and the term ’Big Data ’ go hand in hand. The information explosion caused due to cloud and distributed computing lead to the curiosity to process and analyze massive amount of data. The process and analysis helps to add value to an organization or derive valuable information. The current Hadoop implementation assumes that computing nodes in a cluster are homogeneous in nature. Hadoop relies on its capability to take computation to the nodes rather than migrating the data around the nodes which might cause a significant network overhead. This strategy has its potential benefits on homogeneous environment but it might not be suitable on an heterogeneous environment. The time taken to process the data on a slower node on a heterogeneous environment might be significantly higher than the sum of network overhead and processing time on a faster node. Hence, it is necessary to study the data placement policy where we can distribute the data based on the processing power of a node. The project explores this data placement policy and notes the ramifications of this strategy based on running few benchmark applications. ii Acknowledgments I would like to express my deepest gratitude to my adviser Dr. Xiao Qin for shaping my career. I would like to thank Dr. James Cross and Dr. Jeffrey Overbey for serving on my advisory committee. Auburn University has my regard for making me a better engineer and I would like to thank all my professors and staff for shaping my intellect and personality. I would like to thank my friends and family for supporting me throughout my tenure at Auburn. I had great fun exploring various technologies, meeting interesting people, taking up challenges and expanding my perspective on software development and engineering. iii

Read the paper · More papers on PaperTik