LCM: LLM-focused Hybrid SPM-cache Architecture with Cache Management for Multi-Core AI Accelerators

Chengtao Lai, Zhongchun Zhou, Akash Poptani, Wei Zhang · 2024

The proliferation of large language models (LLMs) with substantial computational requirements and memory footprints has necessitated the design of more capable AI accelerators. Given the long compilation time of scratchpad memory-based (SPM-based) AI accelerators and the challenges brought by LLMs, we have explored the other side of the tradeoff - a multi-core AI accelerator system that incorporates a shared cache and application-specific management strategies - to provide significantly shorter compilation time at the cost of sometimes slightly lower performance than SPM-based systems. Besides, state-of-the-art mixed precision quantization methods also bring dynamic and irregular memory access patterns that do not fit SPMs well.

Read the paper · More papers on PaperTik