Related Experiment Video
Updated: Sep 19, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Dataset Readiness Assessment With Large Language Model (DRAFT-LLM): A Multi-Axis Audit Guided by LLM
Guillaume Guerard1, Sonia Djebali1
1De Vinci Higher Education, De Vinci Research Center, Paris, France.
Abstract:
This article details the Dataset Readiness Assessment for Training (DRAFT), a systematic method for determining whether a high-dimensional biological dataset is suitable for developing reliable, equitable (i.e., the extent to which model performance, error patterns, and potential benefits or harms are evaluated and found to be acceptably distributed across relevant demographic, biological, clinical, and contextual subgroups), and scientifically meaningful machine-learning models, and DRAFT Large Language Model (DRAFT-LLM), its optional human-in-the-loop extension for calibrating study-specific audits through structured, critically reviewed LLM guidance. Standard model validation often fails to detect when apparent performance is driven by spurious correlations, technical artifacts, or hidden stratification, leading to irreproducible and inequitable findings. DRAFT-LLM addresses this gap by shifting the focus from model tuning to structured dataset auditing, organized around Support Protocols 1 to 4 that capture the scientific intent, data structure, and governance constraints of a given study. These Support Protocols: (1) elicit and formalize investigator input into a study intake and dataset card; (2) compute standardized dataset statistics and structural summaries suitable for downstream analysis and LLM context; (3) configure the language model using form-based responses, safety guardrails, and governance rules; and (4) generate personalized instructions, prompts, and code templates for running DRAFT audits. Basic Protocols 1 to 3 are instantiated from this support layer for generalization, equity, and stability: they are reusable execution patterns whose concrete behavior is determined by the cards, statistics, and configurations defined in the Support Protocols. DRAFT-LLM and DRAFT are demonstrated in this article through an end-to-end case study on The Cancer Genome Atlas (TCGA). © 2026 Wiley Periodicals LLC. Support Protocol 1: Study intake and dataset card construction Support Protocol 2: Dataset structure and advanced summary statistics for LLM context Support Protocol 3: LLM configuration using structured form responses Support Protocol 4: Generation of personalized instructions for DRAFT audits Basic Protocol 1: Generalization audit Basic Protocol 2: Equity audit Basic Protocol 3: Stability audit.