Related Experiment Videos
Deep learning perturbation models can outperform baselines on calibrated metrics
Henry E Miller1, Gabriel M Mejia2, Francis J A Leblanc2
1Shift Bioscience Ltd., Cambridge, UK. henry@shiftbioscience.com.
Nature Biotechnology
|October 1, 2026
Abstract:
Recent benchmarks report that deep-learning-based genetic perturbation models fail to outperform uninformative baselines. Here we introduce a positive control baseline and metric calibration framework, showing-across 14 datasets and 18 metrics-that common benchmarking metrics (for example, mean squared error and Pearson Δ) are frequently miscalibrated, with reduced sensitivity to measure positive model performance. Under well-calibrated metrics, we find that deep-learning-based genetic perturbation models can outperform uninformative baselines.