Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same journal

NaP-TRAP: A Versatile and Accessible Workflow to Dissect Principles of Translational Regulation and mRNA Stability.

Current protocols·2026
Same journal

High-Throughput Focal Adhesion Analysis Software for Automated Detection, Tracking, and Classification of Cell-Matrix Adhesions.

Current protocols·2026
Same journal

Beyond Fear Learning in Isolation: Integration of Social Dimensions into Aversive Learning.

Current protocols·2026
Same journal

An Optimized and Improved Protocol for Efficient Isolation of Extracellular RNA from Human Serum.

Current protocols·2026
Same journal

A Biomedical Researcher's Guide for Analyzing Short Tandem Repeat (STR) Genotypes of Human Cell Lines and in Vitro Tissue Samples Using Three Standard Authentication Algorithms.

Current protocols·2026
Same journal

Application of Microbial and Parasitology Techniques for Diagnosis in Laboratory Rodents.

Current protocols·2026

Related Experiment Video

Updated: Jan 8, 2026

Databases to Efficiently Manage Medium Sized, Low Velocity, Multidimensional Data in Tissue Engineering
09:43

Databases to Efficiently Manage Medium Sized, Low Velocity, Multidimensional Data in Tissue Engineering

Published on: November 22, 2019

6.7K

Dataset Readiness Assessment for Training (DRAFT): A Protocol for Auditing High-Dimensional Biological Data.

Guillaume Guerard1, Sonia Djebali1

  • 1De Vinci Higher Education, De Vinci Research Center, Paris, France.

Current Protocols
|December 16, 2025
PubMed
Summary

The Dataset Readiness Assessment for Training (DRAFT) protocol ensures biological datasets are suitable for reliable machine learning. It identifies potential issues like bias and spurious correlations before modeling.

Keywords:
algorithmic fairnesscomputational reproducibilitydata auditingfeature stabilitymachine learning

More Related Videos

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

7.9K
Author Spotlight: Advancing Alzheimer's Research – Exploring Early Detection and Multi-Omics Approaches
09:47

Author Spotlight: Advancing Alzheimer's Research – Exploring Early Detection and Multi-Omics Approaches

Published on: December 15, 2023

1.7K

Related Experiment Videos

Last Updated: Jan 8, 2026

Databases to Efficiently Manage Medium Sized, Low Velocity, Multidimensional Data in Tissue Engineering
09:43

Databases to Efficiently Manage Medium Sized, Low Velocity, Multidimensional Data in Tissue Engineering

Published on: November 22, 2019

6.7K
Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

7.9K
Author Spotlight: Advancing Alzheimer's Research – Exploring Early Detection and Multi-Omics Approaches
09:47

Author Spotlight: Advancing Alzheimer's Research – Exploring Early Detection and Multi-Omics Approaches

Published on: December 15, 2023

1.7K

Area of Science:

  • Bioinformatics
  • Computational Biology
  • Machine Learning in Biology

Background:

  • Standard machine learning model validation often overlooks spurious correlations in biological datasets, leading to irreproducible results.
  • High-dimensional biological data requires specialized methods to ensure model reliability and equity.
  • Existing validation techniques may not adequately detect dataset limitations before extensive computational modeling.

Purpose of the Study:

  • To introduce the Dataset Readiness Assessment for Training (DRAFT) protocol for evaluating high-dimensional biological datasets.
  • To provide a systematic methodology for assessing dataset integrity prior to machine learning model development.
  • To ensure the development of reliable, equitable, and scientifically useful computational models.

Main Methods:

  • DRAFT employs a multiaxis assessment framework.
  • It includes protocols for generalization (cross-validation), equity (subgroup performance analysis), and scientific utility (feature stability).
  • A case study using The Cancer Genome Atlas Lung Adenocarcinoma (TCGA-LUAD) dataset demonstrates the protocol's application.

Main Results:

  • DRAFT identified potential issues in the TCGA-LUAD dataset, highlighting its capacity to generate fragile, inequitable, and uninformative models.
  • The protocol revealed limitations even when preliminary models showed high performance metrics.
  • The assessment framework effectively probes dataset characteristics crucial for robust model development.

Conclusions:

  • The DRAFT protocol is essential for vetting biological datasets before machine learning.
  • Implementing DRAFT ensures that computational models are robust, equitable, and yield genuine scientific insights.
  • This framework sets a standard for dataset integrity in biological research, promoting reproducible and reliable AI applications.