Scalable High-Performance Architecture for Evolving Recommender System

Ravi Kumar Singh, Mayank Mishra, Rekha Singhal · 2023

Recommender systems are expected to scale to the requirement of the large number of recommendations made to the customers and to keep the latency of recommendations within a stringent limit. Such requirements make architecting a recommender system a challenge. This challenge is exacerbated when different ML/DL models are employed simultaneously. This paper presents how we accelerated a recommender system that contained a state-of-the-art Graph neural network (GNN) based DL model and a dot product-based ML model. The ML model was used offline, where its recommendations were cached, and the GNN-based model provided recommendations in real time. The merging of offline results with the results provided by the real-time session-based recommendation model again posed a challenge for latency. We could reduce the model's recommendation latency from 1.5 seconds to under 65 milliseconds with careful re-architecting. We also improved the throughput from 1 recommendation per second to 1500 recommendations per second on a VM with 16-core CPU and 64 GB RAM.

Read the paper · More papers on PaperTik