A New Performance Analysis Method for Semantic Caching for Large Language Models

Ahmet Zahit Ak, Osman Büyük, Mustafa Erden · 2025

Popularity of large language models (LLMs) has impacted a diverse range of fields in our lives, leading to increased demand for using LLMs. However, the computational complexity of LLMs has also increased, resulting in higher costs and longer response times. One solution to this problem is semantic caching, a novel approach designed optimize the performance of systems using LLMs. In this study, we propose an evaluation method for semantic caching systems based on continuous scoring of question-answer pairs, as opposed to binary labels. In addition, we analyze the effects of similarity thresholds and different embedding models on the performance of semantic caching systems in terms of accuracy and response time. Using a dataset derived from GPTCache, our work reveal a trade-off between accuracy and response time, with proprietary embedding models showing higher latency due to external server requests. These findings provide valuable insights for optimizing semantic caching systems in real-time LLM applications, aiming to balance accuracy, cost, and efficiency.

Read the paper · More papers on PaperTik