Analysis of HDFS RPC and Hadoop with RDMA by evaluating write performance

Somya Singh, Gaurav Raj, Gurneet Kaur · 2016

In the era of data explosion, one of the crowd pleasing words is Big Data. For large scale data handling and processing, Hadoop is in the mainstream. The Hadoop Distributed File Storage deals with the issue of data availability by replicating the data over multiple servers. When the data has to be written over multiple remote locations, the write performance is a major consideration. The objective of this paper is to evaluate the HDFS write performance using both the replication schemes, i.e., the default Pipelined Replication and Parallel Replication. The TestDFSIO Benchmark is used for benchmarking the Hadoop cluster with both replication schemes. It is observed that the write throughput of the Hadoop cluster using the parallel replication is 8% more than that of pipeline replication. It is also observed that the performance of RDMA based HDFS is far better than HDFS RPC.

Read the paper · More papers on PaperTik