Recommender systems over distributed architectures
Παναγιώτης Γιαννικόπουλος · 2012
The widespread adoption of social networks, where the logs to apply a frequentpattern mining algorithm are scattered in many servers has driven the need for distributed techniques for FPM.Frequent-pattern mining algorithms belong to two major categories: either collaborative filtering, where the posts are clustered according to the interests of each user, in order to present content preferred by the group he/she belongs to, or content-based filtering, where the clickstreams of the surfers are monitored so as to discover frequently visited sets of posts.Given a frequently visited set of pages F containing k pages, a page PF can be proposed to a user, provided that the rest k-1 pages belonging to F have been requested.Nevertheless, since the current implementations are usually designed to work in systems consisting of a single server which reads the involved posts and implements one of the aforementioned kinds of FPM algorithms, they exhibit limited scalability when the number of users or the volume of data substantially increases, requiring us to move to a distributed architecture.Furthermore, taking into account that there is an increasing number of applications (e.g. in data stream mining), where new information arrives constantly, there is a need to process it quickly, without the need to reprocess or reaccess the transaction database prior to the arrival of the update portion.Moreover, in order to augment the quality of recommendations offered, the taxonomy information inherent in data considered by FPM applications (e.g.news sites, digital libraries) could be exploited by adding information about the categories the web pages belong to.Taking the aforementioned unified framework into account, this thesis presents algorithms in the fields of centralized, distributed, as well as incremental frequentpattern mining, focusing on our contributions in the final two fields (FGP and FGP+, our implementation in centralized, taxonomy-aware FPM, have already been presented in our previous work; nevertheless, they will be also outlined in brief).The incorporation of the taxonomy information is given special attention, given its wide applicability in social media.We simply note that, despite the fact that we have conducted extensive experiments with our algorithms in web log files, they can be applied in a number of contemporary scenarios, such as digital libraries, social media, server farms, as well as content distribution networks.