FaCTQA:Detecting and Localizing Factual Errors in Generated Summaries Through Question and Answering from Heterogeneous Models

Trina Dutta, Xiuwen Liu · 2024

With the advancements of pre-trained large language models, it has become easier to generate fluent abstractive summaries automatically. However, these generated summaries often suffer from factual inconsistencies, or incorrect information, known as hallucinations. Even though there are several methods on identifying hallucinations and hallucinated quantities in machine-generated summaries, detecting hallucination accurately and precisely is still challenging. In this paper, we propose a new pipeline, Factual Check Through Question-Answering (FaCTQA) to detect hallucinations and localize hallucinated quantities for automatically evaluating machine-generated summaries. State of the art (sota) language models fine-tuned with benchmark datasets have better performance on the application problem. Hence, we use question-answering techniques with heterogeneous fine-tuned models to interrogate a summary and its corresponding text document to verify its factual consistencies and extract the probable causes of hallucinations. We have also implemented a variant of the pipeline, FaCTQA* where all the components are replaced by GPT 4. Our experimental results show that FaCTQA outperforms the previous sota models for detecting and localizing hallucination. FaCTQA achieves 90.3% accuracy on the benchmark data XSum and 82.62% accuracy on CNN/DM. FaCTQA has 12.5% better accuracy than the sota models and 4% better performance that FaCTQA*.

Read the paper · More papers on PaperTik