Scalable Real-Time Product Recommendation based on Users Activity in a Social Network
Michael Haspra · Repository for Publications and Research Data (ETH Zurich) · 2011
The web has become a real-time communication medium, used by a large amount of people, in ever-increasing parts of their daily life.This new usage pattern gives advertisers and marketeers great opportunities to learn about their customers' temporal interests, thoughts and current context.Despite it's a well known fact that this information is extremely valuable for advertising and product recommendation, online advertising is adapting only slowly.This is mainly due to the fact that it's not clear what information is valuable to use, the huge amount of produced data and the lack of efficient models to process this data.This thesis describes an approach to implement Scalable Real-Time Product Recommendation based on Users Activity in a Social Network.The products are taken from Amazon.com and the used social network is the microblogging platform Twitter.It presents an implementation of this approach on top of the key-value database Cassandra, using a system called Triggy.Triggy extends Cassandra with incremental Map-Reduce tasks for push-style data processing.Use cases that require high-performance analysis of large amounts of data, are the showpiece of every stream processing engine.These engines are built to process massive amounts of data in very short time.Therefore, this thesis contains a comparison between four state-of-the-art distributed stream processing engines and the implementation with Triggy.It's showed that the analyzed use case has various properties that make it's implementation in a stream processing engine impossible.Finally, a demo application is presented, to show the described approach. List of Figures Motivating Example:We want to find overlappings between user interests and interests that are related to a product . . . . . . . .Conceptual implementation of the matching algorithm . . . . . .Differences between the two Map-Reduce programming models .System architecture overview . . . . . . . . . . . . . . . . . . . .Implementation of the matching algorithm with Map-Reduce tasks Description of the score-calculation for a product-user matching .Column families before any tweet was processed . . . . . . . . . .Column families after the first tweet has been processed . . . . .Column families after the second tweet has been processed . . . .Column families after the third tweet has been processed . . . . .Column families after the recommendation has been made and the model has been resetted . . . . . . . . . . . . . . . . . . . . .Screenshot of the input form for the simulation . . . . . . . . . .Screenshot of the running demo application . . . . . . . . . . . .Screenshot of a recommendation together with it's justification .Querying vs. Pushing test results . . . . . . . . . . . . . . . . . .Comparison of different numbers of Map-Reduce tasks test results Scalability test results . . . . . . . . . . . . . . . . . . . . . . . .