GRAT: Guiding Retrieval-Augmented Reasoning through Process Rewards Tree Search
Xianshu Peng, Wei Wei · 2025
Enhancing large models for complex multihop question-answering has become a research focus in the Retrieval-augmented generation (RAG) area.Many existing approaches aim to mimic human thought processes by enabling large models to perform retrieval-augmented generation step by step.However, these methods can only perform single chain reasoning, which lacks the ability for multi-path exploration, strategic look-ahead, stepwise evaluation, and global selection.In addition, to effectively decompose complex problems, these methods can only rely on labor-intensive intermediate annotations for supervised fine-tuning.To address these issues, we propose GRAT, an algorithm guided by Monte Carlo Tree Search (MCTS) and process rewards.GRAT not only enables self-evaluation and self-correction but also assigns fine-grained rewards to each intermediate step in the search path.These finegrained annotations can be used for model selftraining, which enables GRAT to continuously self-update its problem analysis and reasoning capabilities.We conducted experiments on four multihop QA datasets: HotPotQA, 2WikiMul-tiHopQA, MuSiQue, and Bamboogle, demonstrating that GRAT outperforms various RAGbased methods.Additionally, incorporating self-training significantly enhances GRAT's reasoning performance.