RAGing Against the Literature: LLM-Powered Dataset Mention Extraction
Priyangshu Datta, Suchana Datta, Dwaipayan Roy · 2024
Dataset Mention Extraction (DME) is a critical task in the field of scientific information extraction, aiming to identify references to datasets within research papers. In this paper, we explore two advanced methods for DME from research papers, utilizing the capabilities of Large Language Models (LLMs). The first method employs a language model with a prompt-based framework to extract dataset names from text chunks, utilizing patterns of dataset mentions as guidance. The second method integrates the Retrieval-Augmented Generation (RAG) framework, which enhances dataset extraction through a combination of keyword-based filtering, semantic retrieval, and iterative refinement. We observe that both of the proposed methods achieve more than a 25% improvement in recall compared to the baselines. Further, the RAG-based model achieves an extensive 26% improvement over the baselines. We also propose exData, a web-based tool for extracting dataset name mentions from a given article.