Towards Fully Disaggregated Recommendation Model Serving
Yibo Huang, Yiming Qiu, Zhenning Yang, Yi Dai, Dingming Wu, Fan Lai, Jiarong Xing, Ang Chen · 2025
Serving embedding-based recommendation (EMR) models requires a mix of GPUs, CPUs, and DRAM. Current systems typically provision these resources on monolithic servers with a fixed ratio across resource types, leading to inefficient resource utilization and inflated operational costs. To solve this problem, we propose FlexEMR, a system architecture that fully disaggregates these resources, and interconnects them via an optimized RDMA network. This design enables independent scaling of resources, enhances failure isolation, improves overall resource efficiency, and reduces operational costs. We achieve this by introducing two classes of techniques to address the networking challenges introduced by disaggregation: (1) optimizing embedding lookup communication by leveraging workload locality, and (2) improving network transport through a high-performance, multithreaded RDMA engine. We detail our design considerations and share early performance insights, highlighting the potential of FlexEMR for enabling fully disaggregated EMR model serving.