TeaRAG:A Token-Efficient Agentic Retrieval-Augmented Generation Framework
CHAO; id_orcid 0009-0007-2579-8783 ZHANG, YUHAO; id_orcid 0000-0002-6051-8659 WANG, DERONG; id_orcid 0000-0002-3971-9907 XU, HAOXIN ZHANG, Yuanjie Lyu, CHEN, YUHAO, SHUOCHEN LIU, Tong Xu, XIANGYU; id_orcid 0000-0003-2926-4416 ZHAO, Yan Gao, YAO HU, Enhong Chen · CityU Scholars · 2026
Retrieval-Augmented Generation (RAG) utilizes external knowledge to augment Large Language Models’ (LLMs) reliability. For flexibility, agentic RAG employs autonomous, multi-round retrieval and reasoning to resolve queries. Although recent agentic RAG has improved via reinforcement learning, they often incur substantial token overhead from search and reasoning. This tradeoff prioritizes accuracy over efficiency. To address this issue, this work proposes TeaRAG, a Token-efficient agentic RAG framework capable of compressing both retrieval content and reasoning steps. (1) First, the retrieved content is compressed by augmenting chunk-based semantic retrieval with a graph retrieval using concise triplets. A knowledge association graph is then built from semantic similarity and co-occurrence. Finally, Personalized PageRank is leveraged to highlight key knowledge within this graph, reducing the number of tokens per retrieval. (2) Besides, to reduce reasoning steps, Iterative Process-aware Direct Preference Optimization (IP-DPO) is proposed. Specifically, our reward function evaluates the knowledge sufficiency by a knowledge matching mechanism, while penalizing excessive reasoning steps. This design can produce high-quality preference-pair datasets, supporting iterative DPO to improve reasoning conciseness. Across six datasets, TeaRAG improves the average Exact Match by 4% and 2% while reducing output tokens by 61% and 59% on Llama3-8B-Instruct and Qwen2.5-14B-Instruct, respectively. Code is available at https://github.com/Applied-Machine-Learning-Lab/TeaRAG. © 2026 Copyright held by the owner/author(s). Publication rights licensed to ACM.