Hetero-Rec: Optimal Deployment of Embeddings for High-Speed Recommendations
Chinmay Mahajan, Ashwin M. Krishnan, Manoj Nambiar, Rekha Singhal · 2022
We see two trends emerging due to exponential increase in AI research- rise in adoption of AI based models in enterprise applications and development of different types of hardware accelerators with varying memory and computing architectures for accelerating AI workloads. Accelerators may have different types of memories, varying on access latency and storage capacity. A recommendation model’s inference latency is highly influenced by the time to fetch embeddings from the embedding tables.