Efficient scheme for compressing and transferring data in hadoop clusters

Seungyeon Lee, Jusuk Lee, Yongmin Kim, Ki-Cheol Park, Jiman Hong, Junyoung Heo · 2020

The size of data collected by public institutions and industries is rapidly exploding. As the data that needs to be processed grows larger, there is a limit to processing big data simply by using scale-up servers. To address this limitation, distributed cluster computing systems that use scale-out servers have emerged. However, if the network bandwidth is not used efficiently, the distributed cluster computing systems can not maximize the performance of the scale-out servers. In this paper, we propose an efficient scheme for compressing and transferring data in Hadoop clusters. The proposed method selects an appropriate compression algorithm by calculating the data transfer cost model based on the information entropy of data and network bandwidth. Experimental results show that the data transfer time and the amount of data transfer between the data nodes of the proposed scheme are significantly reduced.

Read the paper · More papers on PaperTik