Dataset for legal question answering system in the Indian judiciary context
K. Veningston, Apratim Mishra · Data in Brief · 2025
Legal documents, such as court judgments and statutes, are vital for understanding judicial decisions, legal principles, and procedural details. However, these documents are often dense, complex, and abundant, making it challenging for lawyers, researchers, and citizens to quickly and easily locate and retrieve relevant information. The Legal Question Answering (LQA) [ 1 ] task involves developing systems that can automatically answer legal questions based on relevant legal documents centred on the constitution and law , preferably from delivered judgments that are considered public property of the nation. The need for specialized datasets in LQA is particularly pressing in countries like India, where legal texts follow distinct judicial structures, specialized terminologies, and procedural intricacies [ 2 ]. Due to the lack of a relevant dataset for an LQA system [ 3 ], this paper presents a comprehensive dataset designed for LQA in the Indian judiciary context that facilitates efficient legal information retrieval. The dataset comprises 10,000 question-answer pairs derived from 1,256 Indian Supreme Court judgments across various legal domains, including 538 criminal and 718 civil cases available on Mendeley Data [ 4 ]. Each QA pair is derived from detailed legal judgments from the Apex court (i.e. Supreme Court of India), with the questions framed to capture essential legal issues, principles, or facts, and answers extracted directly from the text. The dataset covers a balanced mix of legal topics in criminal and civil law, such as constitutional matters, property disputes, criminal offences, procedural matters, family disputes, employment matters, financial and taxation issues, and public welfare concerns. Additionally, it includes metadata such as case name and judgement date. This dataset supports the development of AI-driven LQA systems to enhance access to precise legal information and aid legal professionals/common citizens about India's complex legal system. To evaluate its effectiveness for legal question-answering tasks, the IndicLegalQA Dataset is fine-tuned on the "meta-llama/Llama-2-7b-hf" model [ 5 ] using Parameter-Efficient Fine-Tuning (PEFT), specifically the Low-Rank Adaptation (LoRA) technique [ 6 , 7 , 8 ]. The fine-tuned model is evaluated using Sentence-BERT (SBERT) [ 9 ], with the "paraphrase-MiniLM-L6-v2" model embedding. Cosine similarity measures how well the model captures the nuances of legal language between actual and generated answers. This ensures the dataset is well-suited for real-world legal applications, making it a valuable resource for improving AI-driven legal information retrieval systems.