Exploring Efficient Optimization Techniques in Online Retrieval-Augmented Generation Application

Yining Zhang, Yinan Peng, C.W. Tu, Zherui Zhang, Hongfei Yan, Chong Chen, Hao Ma, Jia Yang, Yan Zhang, Rikun Liao · 2024

Recent advances in large language models (LLM) have brought an explosive growth to chat-bot applications. Among them, retrieval-augmented generation, which provides extra context to make LLM capable of answering out-of-domain questions is becoming a popular method. However, naive implementation of RAG usually cannot reach ideal answer quality in complicated real-world scenarios. Researchers have proposed a number of methods to improve RAG, but many of them involves extra LLM calls which is too time-consuming for online application. In this paper, we explored practical techniques and designs in RAG that improve answers to user-satisfying quality while keeping the response latency at a moderate level in the scenario of a research data QA application in university. Our main findings include introducing a relevance judge with small-scale LLM for retrieved documents can effectively filter out less relevant ones, which can otherwise disrupt the generated answer greatly, and decomposing the generation task into multiple independent sub-tasks can reduce the chance of hallucination and also accelerates the generation. As for model performance, prompt engineering and fine-tuning (through learning from strong LLM) are effective yet simple ways to enhance answer quality. Our results and experience provide insights for building future real-world LLM applications.

Read the paper · More papers on PaperTik