Instruct-to-SPARQL: A text-to-SPARQL dataset for training SPARQL Agents

Mehdi Ben Amor, Alexis Strappazzon, Michael Granitzer, Elöd Egyed‐Zsigmond, Jelena Mitrović · 2025

The rapid adoption of Large Language Models (LLMs) for search engines and fact-checking platforms necessitates enhancing their output accuracy.Retrieval Augmented Generation (RAG) mitigates hallucinations but requires semantically rich repositories like Wikidata.However, there is a lack of high-quality data to fine-tune LLMs for querying such knowledge bases.To address this gap, we propose a curated dataset with 2,771 unique queries for fine-tuning LLMs to generate accurate and syntactically valid SPARQL queries from natural language instructions.This dataset, customized for interaction with Wikidata, also serves as a robust benchmark for text-to-SPARQL task evaluation.Key findings show that models generally perform better on queries with lower complexity.

Read the paper · More papers on PaperTik