Don’t Just Say “I don’t know”! Self-aligning Large Language Models for Responding to Unknown Questions with Explanations

Yang Deng, Yong Hui Zhao, Moxin Li, See-Kiong Ng, Tat‐Seng Chua · 2024

Despite the remarkable abilities of Large Language Models (LLMs) to answer questions, they often display a considerable level of overconfidence even when the question does not have a definitive answer.To avoid providing hallucinated answers to these unknown questions, existing studies typically investigate approaches to refusing to answer these questions.In this work, we propose a novel and scalable self-alignment method to utilize the LLM itself to enhance its response-ability to different types of unknown questions, being capable of not just refusing to answer but further proactively providing explanations to the unanswerability of unknown questions.Specifically, the Self-Align method first employ a two-stage classaware self-augmentation approach to generate a large amount of unknown question-response data.Then we conduct disparity-driven selfcuration to select qualified data for fine-tuning the LLM itself for aligning the responses to unknown questions as desired.Experimental results on two datasets across four types of unknown questions validate the superiority of the Self-Aligned method over existing baselines in terms of three types of task formulation. 1 * Equal contribution.Q: What animal can be found at the top of the men's Wimbledon trophy? Direct AnswerA: The animal that can be found at the top of the men's Wimbledon trophy is a falcon.

Read the paper · More papers on PaperTik