Distrim: Parallel GMM learning on multicore cluster

Renyong Yang, Tengke Xiong, Tao Chen, Zhexue Huang, Shengzhong Feng · 2012

Learning GMM model on extreme large data is challenging. We provide theoretical support for the feasibility of parallel EM-based GMM learning via distributed computing, and also design and implement a distributed memory sharing GMM learning system on multicore clusters, which is named as Distrim. Distrim aims to maximize the usage of computational power and minimize the communication overheads as much as possible. The experimental results show that Distrim is much more efficient than Hadoop, and also has a good scalability with respect to the number of computing nodes.

Read the paper · More papers on PaperTik