Related Experiment Videos
Architectural good practices for reproducible benchmarking in protein machine learning
Julián García-Vinuesa1, Diego Fernández-Villegas1, Michelle Soto-García1
1Departamento de Ingeniería En Computación, Universidad de Magallanes, Punta Arenas, Chile.
Abstract:
Reliable benchmarking in protein machine learning requires biological tasks, datasets, representations, execution conditions, and evaluation regimes to be explicitly defined and traceable. Protein-specific dependencies, including label semantics, negative-class construction, evolutionary relationships, similarity control, inferred structures, protein language model configurations, and potential pretraining contamination, can alter benchmark interpretation. This Perspective provides an operational, protein-specific synthesis of established reproducibility practices through an artefact-linked architecture. We define five core requirements for verifiable benchmarking, including deterministic curation, persistent and versioned artefacts, explicit interface contracts, reproducible execution, and transparent evaluation. Prediction provenance, uncertainty, calibration, and explainability are treated as an additional decision-readiness layer. These requirements are translated into minimum records, provenance links, and verification checks. An antimicrobial peptide classification case study demonstrates their application to alternative task formulations, negative-class definitions, dataset compositions, and partitioning strategies within a controlled workflow.