Alleviating All-to-All Communication for Deep Learning Recommendation Model Inference
Songjun Huang, Yihong Li, L P Chen, Xiaoxi Zhang, Shuo Liu, Jingpu Duan, Wenfei Wu, Xu Chen · 2024
Massive DLRMs require large-scale multi-node systems for distributed training and inference, thus suffering from the all-to-all communication bottleneck. We propose an architecture, EmbedSwitch, that offloads the cache function of the embedding table vectors to a programmable switch, to overcome this bottleneck and provide switch-level response latency for embedding table vector requests.