Designing the Most Energy-Efficient Recommendation Inference Chip
Youn-Long Lin, Joe Kao, Kinny Chen · 2023
NEUCHIPS' purpose-built recommender system accelerator, RecAccel™ N3000, includes optimized hardware and software stack and delivers unparalleled performance, energy efficiency – and the lowest TCO. Recommendation systems predict the “rating” or “preference” a user would give to an item. Most internet service providers employ recommendation systems to deliver their products or services. Recommendations usually account for most of AI inference workload in data centers. A typical neural personalized recommendation model consists of multiple embedding table lookup, feature interaction, and multiple multilayer perceptron (MLP) [1]. Different models employ different architecture for various applications and service level agreements (SLAs) target [2]. Some models are compute-intensive while others are memory-intensive. Therefore, in addition to efficient high performance GEMM engines, an accelerating solution must provide high-capacity, high-bandwidth, and low-latency storage as well as efficient allocation and access strategies.