Research on private cloud platform of seed tracing based on Hadoop parallel computing
Dongming Li, Yan Li, Yuan Chao, Haochuan Chen, Zhang Li-juan · 2015
It is important for the management of seeds tracing which includes the collection and calculation of seeds sale. In the paper, firstly, we designed and implemented the seeds Hadoop-based trace data processing model, through the processing of the crawl data of seed, such as ETL process, parallel computing, distributed storage. Secondly, we studied the Consistent Hash Algorithm to optimize the database cluster configuration and the great data parallel computing of MapReduce. The experiment results show that compared with the traditional single node approach, using private cloud parallel calculation can greatly improve parallel computing, the efficiency of storage and load capacity of the platform. The processing efficiency of cloud platform is increased by 33.3%.