Towards Real-Time LLM Inference on Heterogeneous Edge Platforms
Rakshith Jayanth, Neelesh Gupta, Souvik Kundu, Deepak A. Mathaikutty, Viktor K. Prasanna · 2024
The deployment of computationally intensive workloads, such as Large Language Models (LLMs), at the edge presents significant challenges due to resource constraints and limited computing power. State-of-the-art edge platforms address this with integrated high-performance accelerators optimized for machine learning workloads. While numerous high-performance edge platforms are now available in the market, there is limited exploration of methods to maximize performance and optimize resource utilization at the edge. Our work aims to address this gap by proposing methodologies for maximizing performance through the distribution of model inference on a heterogeneous edge platform. This poster presents preliminary findings from our ongoing research project funded by NSF IUCRC IDEAS Center. The Center is focused on Edge Intelligence and applications that can be potentially mapped to Edge platforms.