Distributed Recommendation Algorithm Based on Fuzzy Clustering
Jiali Zhang, Haohua Qing · Journal of Physics Conference Series · 2021
In recent years, the Internet is sneaking into every corner of users' lives at a speed visible to the naked eye. Information has shown a geometric growth with the development of the Internet. This article mainly studies the application of distributed recommendation algorithm based on fuzzy clustering. In this paper, the data set used is divided into training set, test set and validation set in a random manner. The ratio of the three number sets is 7:1:2. This paper randomly selects 70% of consumers' historical data as the training data set, and the last 10% as the test set. The USCensus 1990 raw data set is divided into three groups, numbered 1, 2, and 3, and the number of samples is 400000, 600000 and 800000 respectively. Under the Hadoop platform, the MrDPCND-FCM algorithm is used to cluster the three sets of data to be tested, the number of nodes is increased from 2 to 5, and the time taken to cluster different data sets to be tested under different numbers of nodes is recorded. Therefore, offline testing uses items that the user has already evaluated. If the user has not scored the item, it cannot be evaluated. When the number of clusters is 60, the data sparsity drops from 92% to 51%. When the number of clusters reaches 5, the sparsity drops to around 3%. The algorithm in this paper solves the problem of large-scale data to a certain extent, and improves the efficiency and accuracy.