Improving Out-of-Vocabulary Hashing in Recommendation Systems
William Shiao, Mingxuan Ju, Zhichun Guo, Xin Chen, Evangelos E. Papalexakis, Tong Zhao, Neil Shah, Yozen Liu · 2025
Recommendation systems (RS) are an increasingly relevant area for both academic and industry researchers, given their widespread impact on the daily online experiences of billions of users. In real applications, one common challenge is recommending new users and items unseen (out-of-vocabulary, or OOV) at training time, i.e. the inductive setting. Additionally, modern RS also faces challenges in ID embedding table size. To handle large cardinality user/item embeddings, memory intensive embedding tables are required. The size of OOV user/item IDs are often large and varies. As a result of both issues, existing solutions applied in practice are often naïve, such as assigning OOV or hashed users/items to a fix set of random buckets. In this work, we tackle the cold-start OOV and ID embedding memory problem and propose approaches that better leverage available user/item features and memory-efficient hashing at the embedding table level. We discuss plug-and-play approaches that are easily applicable to RS models and improve inductive performance without negatively impacting transductive performance. Through our extensive evaluation, we find that proposed methods that exploit feature similarity using LSH consistently outperform alternatives on a majority of model-dataset combinations, with the best one showing a mean improvement of 3.74% over the industry standard baseline in recommendation performance. We release our code and hope our work helps practitioners make more informed decisions for efficiently hashing OOV in their RS and further inspires academic research into improving OOV support in RS.