Context-based Semantic Caching for LLM Applications

Ramaswami Mohandoss · 2024

Large Language Models (LLM), aided by the popularity of ChatGPT, have provided a paradigm shift to engineering AI Chatbots. LLMs offer many conveniences to the AI engineer, and the most important benefit is their robustness in handling semantic matches of user queries. In other words, an engineer building an AI conversational assistant does not need to train the agent for semantic user query matches explicitly. This benefit comes with a cost, which is felt in two ways. Firstly, LLMs need an expensive infrastructure like GPUs, large RAMs, etc. Secondly, even with all the cutting-edge infrastructure, their response time will not be sub-second anytime soon. So, AI applications that use LLMs are likely to be expensive and suffer from high latency. One way to reduce cost and response time is by introducing a caching solution between the LLMs and the UX layer. Such a caching layer can help minimize calls to LLMs when the queries are similar and repetitive. However, traditional caching methods cannot be used as they are for LLM-based applications. The reason is that queries handled by AI-based conversational assistants are unstructured, i.e., free-flowing user-generated text, and will not always be context-free. In other words, end users tend to query from their point of view. Existing solutions like GPTCache work well for context-free questions, but caching context-sensitive user queries needs an evolved design. In this paper, we shall explore a novel design that, by exploiting the power of context, shall provide effective caching solutions to user-generated queries (both context-free and context-sensitive) that offer improved performance without compromising on the quality of response.

Read the paper · More papers on PaperTik