Accelerating EM Missing Data Filling Algorithm Based on the K-Means
SUN Hua-Yan, Yeli Li, Yunfei Zi, Xu Han · 2018
In the whole process of data mining, the EM algorithm is widely applied to dealing with incomplete data for its numerical stability, simplicity of implementation, reliable global convergence. the main disadvantage of the EM is slow convergence speed, the algorithm is highly dependent on the initial value of the option, In this paper, the clustering results use K-means algorithm as the initial scope of EM algorithm, according to the different choice of different characteristics of mining purposes, then use incremental EM algorithm (IEM) step by step EM iterative refinement repeatedly, it obtains the optimal value of filling missing data quickly and efficiently. it is concluded that the optimal value of filling missing data experimental results show that the algorithm of this paper to speed up the convergence rate, strengthened the stability of clustering, data filling effect is remarkable.