Rabbitail: A Tail Latency-Aware Scheduler for Deep Learning Recommendation Systems with Hierarchical Embedding Storage

Hu Wan, Yun Huang, Shuhan Bai, Xuan Sun, Tei‐Wei Kuo, Chun Jason Xue · 2025

Deep learning-based recommendation systems are critical for many online platforms, but face challenges in managing large embedding tables while meeting strict latency requirements. This paper presents Rabbitail, a novel inference scheduler designed for recommendation systems utilizing hierarchical DRAM-SSD embedding storage. Rabbitail employs a cache-aware approach, classifying inferences into hit and miss categories based on embedding cache lookup results. This allows hit inferences to proceed immediately to top MLPs without waiting for slower SSD retrievals. For miss inferences, Rabbitail implements an on-demand embedding lookup strategy and a reordering mechanism to optimize SSD retrieval. Additionally, it uses dedicated resource allocation for prompt processing of miss queue MLP tasks and employs batch splitting to manage maximum execution times. Evaluations using real-world datasets demonstrate that Rabbitail significantly reduces end-to-end model inference tail latency, achieving a 53.7% lower p99 tail latency compared to the baseline while maintaining throughput.

Read the paper · More papers on PaperTik