Improving matrix factorization recommendations for problems in big data

Xiaohan Tu, Siping Liu, Renfa Li · 2017

In the big data environment, recommender systems are facing many problems, such as poor extendibility, data sparseness and low efficiency. In this paper, a new collaborative filtering parallel algorithm named NALS-WR is designed in the Linux clusters by using Spark to solve these problems, especially aiming at the bottleneck of processing speed and resource allocation of traditional matrix factorization algorithm in massive data information. Our experiments were on real MovieLens datasets. Compared with other such recommendation systems based on ALS-WR or sigular value decomposition (SVD), our accuracy rate was improved. The running efficiency was much higher than ALS-WR in Hadoop and SVD. The bigger the scale of data, the more efficient it is. Owing to the distributed computing framework and cloud computing environment, our model is more sufficiently scalable. NALS-WR can improve the execution efficiency of collaborative filtering recommendation algorithm at large data scale. The experiments also indicated that NALS-WR algorithm was better, whether in extendibility, sparseness resistance or efficiency in the implementation.

Read the paper · More papers on PaperTik