A MapReduce‐based parallelK‐meansclustering for large‐scale CIM data verification

Chuang Deng, Yang Liu, Lixiong Xu, Jie Yang, Junyong Liu, Siguang Li, Maozhen Li · Concurrency and Computation Practice and Experience · 2015

Summary The Common Information Model (CIM) has been heavily used in electric power grids for data exchange among a number of auxiliary systems such as communication systems, monitoring systems, and marketing systems. With a rapid deployment of digitalized devices in electric power networks, the volume of data continuously grows, which makes verification of CIM data a challenging issue. This paper presents a parallelK‐meansclustering algorithm for large‐scale CIM data verification. The parallelK‐meansbuilds on the MapReduce computing model which has been widely taken up by the community in dealing with data‐intensive applications. A genetic algorithm‐based load‐balancing scheme is designed to balance the workloads among the heterogeneous computing nodes for a further improvement in computation efficiency. The performance of the parallelK‐meansis initially evaluated in a small‐scale in‐house MapReduce cluster and subsequently evaluated in a commercial cloud computing platform. Finally, the parallelK‐meansis evaluated in large‐scale simulated MapReduce environments. Both the experimental and simulation results show that the parallelK‐meansreduces the CIM data‐verification time significantly compared with the sequentialK‐meansclustering, while generating a high level of precision in data verification. Copyright © 2015 John Wiley & Sons, Ltd.

Read the paper · More papers on PaperTik