Research on private cloud platform of seed tracing based on Hadoop parallel computing

Dongming Li, Yan Li, Yuan Chao, Haochuan Chen, Zhang Li-juan · 2015

It is important for the management of seeds tracing which includes the collection and calculation of seeds sale. In the paper, firstly, we designed and implemented the seeds Hadoop-based trace data processing model, through the processing of the crawl data of seed, such as ETL process, parallel computing, distributed storage. Secondly, we studied the Consistent Hash Algorithm to optimize the database cluster configuration and the great data parallel computing of MapReduce. The experiment results show that compared with the traditional single node approach, using private cloud parallel calculation can greatly improve parallel computing, the efficiency of storage and load capacity of the platform. The processing efficiency of cloud platform is increased by 33.3%.

Read the paper · More papers on PaperTik