Can Large Language Models Fix Data Annotation Errors? An Empirical Study Using Debatepedia for Query-Focused Text Summarization

Md Tahmid Rahman Laskar, Mizanur Rahman, Israt Jahan, Enamul Hoque, Jimmy Xiangji Huang · 2023

Debatepedia is a publicly available dataset consisting of arguments and counter-arguments on controversial topics that has been widely used for the single-document query-focused abstractive summarization task in recent years.However, it has been recently found that this dataset is limited by noise and even most queries in this dataset do not have any relevance to the respective document.To this end, this paper aims to study whether large language models (LLMs) can be utilized to clean the Debatepedia dataset to make it suitable for query-focused abstractive summarization.More specifically, we harness the language generation capabilities of two LLMs, namely, ChatGPT 1 , and PaLM 2 to regenerate its queries.Based on our experiments, we find that only fixing the queries in Debatepedia via LLMs may not be useful.However, leveraging a rule-based approach via filtering out noisy instances followed by query regeneration using LLMs for the sampled instances may ensure a higher quality version of this dataset suitable for the development of more generalized query-focused text summarization models.

Read the paper · More papers on PaperTik