Related Experiment Video
Updated: Sep 19, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Dataset for legal question answering system in the Indian judiciary context
Veningston K1, Apratim Mishra1
1Department of Computer Science and Engineering, National Institute of Technology Srinagar, Jammu and Kashmir 190006, India.
Abstract:
Legal documents, such as court judgments and statutes, are vital for understanding judicial decisions, legal principles, and procedural details. However, these documents are often dense, complex, and abundant, making it challenging for lawyers, researchers, and citizens to quickly and easily locate and retrieve relevant information. The Legal Question Answering (LQA) [1] task involves developing systems that can automatically answer legal questions based on relevant legal documents centred on the constitution and law, preferably from delivered judgments that are considered public property of the nation. The need for specialized datasets in LQA is particularly pressing in countries like India, where legal texts follow distinct judicial structures, specialized terminologies, and procedural intricacies [2]. Due to the lack of a relevant dataset for an LQA system [3], this paper presents a comprehensive dataset designed for LQA in the Indian judiciary context that facilitates efficient legal information retrieval. The dataset comprises 10,000 question-answer pairs derived from 1256 Indian Supreme Court judgments across various legal domains, including 538 criminal and 718 civil cases available on Mendeley Data [4]. Each QA pair is derived from detailed legal judgments from the Apex court (i.e. Supreme Court of India), with the questions framed to capture essential legal issues, principles, or facts, and answers extracted directly from the text. The dataset covers a balanced mix of legal topics in criminal and civil law, such as constitutional matters, property disputes, criminal offences, procedural matters, family disputes, employment matters, financial and taxation issues, and public welfare concerns. Additionally, it includes metadata such as case name and judgement date. This dataset supports the development of AI-driven LQA systems to enhance access to precise legal information and aid legal professionals/common citizens about India's complex legal system. To evaluate its effectiveness for legal question-answering tasks, the IndicLegalQA Dataset is fine-tuned on the "meta-llama/Llama-2-7b-hf" model [5] using Parameter-Efficient Fine-Tuning (PEFT), specifically the Low-Rank Adaptation (LoRA) technique [[6], [7], [8]]. The fine-tuned model is evaluated using Sentence-BERT (SBERT) [9], with the "paraphrase-MiniLM-L6-v2" model embedding. Cosine similarity measures how well the model captures the nuances of legal language between actual and generated answers. This ensures the dataset is well-suited for real-world legal applications, making it a valuable resource for improving AI-driven legal information retrieval systems.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
07:50A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
Related Concept Videos
Sources of Law
Constitutional law is foundational, deriving from federal and state constitutions, and...
Legal Guidelines for Documentation
Torts I
Intentional...
Torts II
Statically Indeterminate Problem Solving
Torts III
Quasi-intentional torts in healthcare involve acts where intent is not directed to harm an individual but results in harm due to careless or reckless speech.