Programming model based on MapReduce for importing big table into HDFS

Chen Ji-ron · Journal of Computer Applications · 2013

To solve the problems of instability and inefficiency when data from a relation database system are transferred into Hadoop Distributed File System( HDFS) using Sqoop, the authors proposed and implemented a new programming model based on MapReduce framework. The algorithm splitting a big table in this model was as follows: firstly a step was calculated by dividing the total lines by the mapper number, then a SQL statement corresponding to each split could be constructed with a start line index and a span range equal to the above step, so this approach could guarantee that each mapper task would issue identical SQL workload. In map phrase, a mapper would only call map function once, with the single key-value pair below:the key was the above SQL statement corresponding to a split, and the value was null. The comparison experiments show that,for two different big tables with the same number of records, the respective importing time was approximately identical regardless of the records distribution, while using two different splitting fields in one big table, the importing time was also the same. At the same time, when applying two different approaches to one big table, the importing efficiency using the model was largely promoted than that using Sqoop.

Read the paper · More papers on PaperTik