EFFECTIVE AND EFFICIENT COLLABORATIVE FILTERING
Yi Ding · 2011
Collaborative filtering is regarded as one of the most promising approaches in recommender systems. To date, it is best known for its use on e-commerce web sites. It has also been widely used in other areas, for example, filtering Usenet News, recommending TV shows and web personalization. However, a survey of existing algorithms shows there remain some fundamental research questions in overcoming some challenges for collaborative filtering system. These questions influence the prevalence of the recommendation systems to a great extent. • Scalability: A large online retailer might have huge amounts of data, tens of millions of customers and millions of distinct catalogue items. These “long user rows” slow down the performance of the recommender system, further reducing scalability. • Accuracy: Accuracy of collaborative filtering considers the problem of how the system would successfully measure the similarity between users (in user-based approaches) or items (in item-based approaches) in order to discover the intrinsic properties that exist amongst users and/or amongst item. Users need recommendations of high quality and they can trust to help them. • Robustness: Robustness of collaborative filtering systems is defined as the ability to provide accurate predictions given some degree of noise in the data. • Sparsity: It refers to the fact that most users do not rate many items and hence the user-item rating matrix is very sparse and insufficient to identify similarities in consumer interests. In many commercial recommender systems, even active users may have purchased less than 1% of the items, (1% of 2 million books is 20,000 books). As a result, the accuracy of recommendations may be poor. • Cold Start: It has been used to describe the situation when almost nothing is known about the new users or new items. This dissertation focuses on effectiveness and efficiency of collaborative filtering technologies. We proposed a novel collaborative filtering framework which is capable of dealing with an immense and dynamic dataset effectively and efficiently. Specifically, we improved the existing algorithms from the following aspects: 1. Generic Model: Integrate the content-based filtering and collaborative filtering by unifying the external attributes of users and items with rating information in a generic model. 2. Effectiveness: 1)Propose a new method of computing similarity in collaborative filtering to better reflect the reality. 2)Consider the interest drift in collaborative filtering. We introduce the time recency to tackle it. 3. Efficiency: Propose a number of solutions towards traditional item-based and user-based collaborative filtering which can handle a large scale of data in the dynamical environment. The experiments have shown that our proposed framework can substantially improve the performance of traditional collaborative filtering algorithms. The main contributions of this dissertation include: demonstrating a generic model, improving the accuracy of collaborative filtering through the new similarity computation and addressing interest drift in the traditional collaborative filtering and providing a highly efficient incremental framework which can be easily used for the online applications and an effective indexing which can reduce the time complexity of the traditional algorithms dramatically.