Knowledge Graph-Enhanced Semantic Cache for Low-Latency and Cost-Effective Inference in Large Language Models

Nicholas Dominic, Bens Pardamean · 2024

In organizational knowledge management, Large Language Model (LLM) caches act as a semantic repository gathered from previous LLM responses. Due to intensive calls from multiple users, LLM may suffer from high inference latency. While there are many prior available approaches to solve this problem, most of them are inherently complex. This paper introduced a Knowledge Graph-enhanced Semantic Cache mechanism as an alternative, lightweight technique to boost retrieval for similar prompts. The latest state-of-the-art open-source LLM, named Google's Gemma-2B-it, was used to generate sample prompts and responses as a draft, while a knowledge graph (KG) was built from Wikipedia sentences. To create embeddings of prompts and KG, all-MiniLM-L6-v2 from SentenceTransformer was used. This new cache system resulted in up to 28% improvement over a standard model. In particular, reinforcement with KG cache embeddings yielded more than 85% semantic cache accuracy. To map the next trajectory of this pilot study, an overview of the extended framework for LLM knowledge management was also presented in this paper. The framework includes the new KG- enhanced cache system equipped with scalable security and fallback mechanisms that can promote green technology through substantial improvements in latency, throughput, and overall LLM costs.

Read the paper · More papers on PaperTik