Investigating Large Language Models for Text-to-SPARQL Generation

Jacopo D’Abramo, Andrea Zugarini, Paolo Torroni · 2025

Large Language Models (LLMs) have demonstrated strong capabilities in code generation, such as translating natural language questions into SQL queries.However, state-of-the-art solutions often involve a costly fine-tuning step.In this study, we extensively evaluate In-Context Learning (ICL) solutions for textto-SPARQL generation with different architectures and configurations, based on methods for retrieving relevant demonstrations for few-shot prompting and working with multiple generated hypotheses.In this way, we demonstrate that LLMs can formulate SPARQL queries achieving state-of-the-art results on several Knowledge Graph Question Answering (KGQA) benchmark datasets without finetuning.* Work done while being at expert.ai. the following key aspects: (1) the influence of various In-Context Learning strategies on the quality of the generated queries; (2) the impact of different state-of-the-art model backbones, varying in architecture, size, and training data; (3) the potential of beam search to generate multiple query candidates, thereby enhancing the results;(4) a comparison between ICL and specialized models finetuned for the task.The code is publicly available at https://github.com/jacopodabramo/DFSL.In the interest of reproducibility, as backbones, we use three state-of-the-art open-weight LLMs: Mixtral 8x7B, Llama-3 70B, and CodeLlama 70B.We run experiments on two widely-used Knowledge Bases, DBpedia and Wikidata, using four publicly available datasets: QALD-9, based on DBpedia, and QALD-9 plus, QALD-10 and LC-QuAD 2.0, based on Wikidata.Our experimental results demonstrate that LLMs In-Context Learning solutions achieve state-of-theart results, without the need of any fine-tuning.The injection of demonstrations similar to the input question into the prompt combined with the generation of multiple query candidates directly from beam search hypotheses, yield the best results, exceeding in most of the benchmarks state-of-the-art models fine-tuned for the task.Finally, we also run ablation studies to gauge the effectiveness of the approach without gold information from the EL and RL modules. Related workWe first provide an overview of most related In-Context-Learning approaches.Then, we discuss text-to-SPARQL methods, including KGQA systems that typically make use of text-to-SPARQL techniques to tackle the problem. 66 In-Context PromptThe task involves translating questions from English into SPARQL queries for the Wikidata knowledge graph.The queries must follow specific guidelines to ensure accuracy and correct execution: 1. Enclose

Read the paper · More papers on PaperTik