Related Experiment Video
Updated: May 15, 2025

Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER
Published on: June 23, 2012
Data splitting to avoid information leakage with DataSAIL
Roman Joeres1,2,3,4, David B Blumenthal5, Olga V Kalinina6,7,8
1Helmholtz Institute for Pharmaceutical Research Saarland (HIPS), Helmholtz Centre for Infection Research (HZI), Saarbrücken, Germany. roman.joeres@helmholtz-hips.de.
Abstract:
Information leakage is an increasingly important topic in machine learning research for biomedical applications. When information leakage happens during a model's training, it risks memorizing the training data instead of learning generalizable properties. This can lead to inflated performance metrics that do not reflect the actual performance at inference time. We present DataSAIL, a versatile Python package to facilitate leakage-reduced data splitting to enable realistic evaluation of machine learning models for biological data that are intended to be applied in out-of-distribution scenarios. DataSAIL is based on formulating the problem to find leakage-reduced data splits as a combinatorial optimization problem. We prove that this problem is NP-hard and provide a scalable heuristic based on clustering and integer linear programming. Finally, we empirically demonstrate DataSAIL's impact on evaluating biomedical machine learning models.
Related Concept Videos
Censoring Survival Data
Data Reporting and Recording
Leaky Scanning
Shear Diagram
First, a free-body diagram of the beam is drawn, representing all the external forces and internal reactions acting on the beam. One can calculate the reaction forces at each support by employing the equilibrium equations of force and moment. The vertical component...
Data: Types and Distribution
Distributions in...
Standard Deviation of Calculated Results
A broad Gaussian distribution curve has a wider standard deviation, representing a data set with...

