Search Augmented Instruction Learning

Hongyin Luo, Tianhua Zhang, Yung-Sung Chuang, Yuan Gong, Yoon Kim, Xixin Wu, Helen M. L. Meng, James Glass · 2023

It is widely believed that connecting large language models with search engines can improve their transparency, truthfulness, and accessing to up-to-date information.However, we show that search grounding introduces new challenges to language models because of distracting, misleading, and untrustworthy information.To deal with these difficulties, we propose search-augmented instruction learning (SAIL), which allows a fine-tuned language model to source, denoise, and reason based on a mixed set of helpful and distracting search results.With an instruction tuning corpus, we collect search results for each training case from different search APIs and domains, and construct a new search-grounded training set containing (instruction, grounding information, response) triplets.We then fine-tune the LLaMA-7B model on the constructed training set.Since the collected search results contain distracting and disputing languages, the model needs to learn to ground on trustworthy search results, filter out distracting passages, and generate the target response.The search result-denoising process entails explicit trustworthy information selection and multi-hop reasoning, since the retrieved passages might be informative but not contain the instruction-following answer.Experiments show that the fine-tuned SAIL-7B model has a strong instruction-following ability, and it performs significantly better on transparencysensitive tasks, including question answering and fact checking.

Read the paper · More papers on PaperTik