Related Experiment Video
Updated: Sep 6, 2026

Biobank for Translational Medicine: Standard Operating Procedures for Optimal Sample Management
Published on: November 30, 2022
Benchmarking biomedical foundation models
Julio Saez-Rodriguez1,2, Philipp Sven Lars Schäfer3, Nikolas Kalavros4,5
1European Molecular Biology Laboratory, European Bioinformatics Institute, Hinxton, UK. saezlab@ebi.ac.uk.
Abstract:
A transparent evaluation and proof of reproducibility, generalization and replicability of algorithms are the bedrock of method development in computational biology. Many benchmarking efforts have been developed for problems ranging from structural biology to translational biomedicine. Rigor is relatively controllable for tasks such as the prediction of patient outcomes or the outcomes of biological assays, but the problem is exacerbated when the aim is to benchmark foundation models. The parameters constituting them are supposed to capture the patterns underlying the data; therefore, the models are parameterized embodiments of the phenomena that gave rise to the data. How can we test the limitations of these models? Here, we discuss the epistemological value of foundation models; whether they can be refuted, verified or evaluated primarily on the basis of utility; what principles should guide their benchmarking; and what role the scientific community should play in that benchmarking process.

