GOLDMARK: Governed Outcome-Linked Diagnostic Model Assessment Reference Kit
Chad Vanderbilt1, Gabriele Campanella2, Siddharth Singi1
1Department of Pathology and Laboratory Medicine, Memorial Sloan Kettering Cancer Center, New York, NY, USA.
Background:
Computational biomarkers (CBs) are histopathology-derived patterns extracted from hematoxylin-eosin (H&E) whole-slide images (WSIs) using artificial intelligence (AI) to predict therapeutic response or prognosis. Recently, slide-level multiple-instance learning (MIL) with pathology foundation models (PFMs) has become the standard baseline for CB development. While these methods, with architectural and optimization advances, have improved predictive performance, computational pathology lacks standardized intermediate data formats, provenance tracking, checkpointing conventions, and reproducible evaluation metrics required for clinical-grade deployment. Consequently, discipline-level standardization, including data representation, model versioning, evaluation protocols, and auditability, is essential to enable reliable, scalable, and regulatory-ready clinical translation of CBs.
Methods:
We introduce GOLDMARK: Governed Outcome-Linked Diagnostic Model Assessment Reference Kit (www.artificialintelligencepathology.org), a standardized benchmarking framework built on a curated TCGA cohort with clinically anchored OncoKB level 1-3 biomarker labels. GOLDMARK distributes structured intermediate outputs, including tile coordinates, per-slide feature embeddings from canonical PFMs, embedding-level quality-control metadata, trained slide-level weights, and reference code. Multiple publicly available PFMs are benchmarked under a unified attention-based MIL head using predefined patient-level splits. Models are trained on TCGA and evaluated on an independent MSKCC cohort with reciprocal testing.
Results:
We evaluated 33 tumor-biomarker tasks; aggregate summaries over the 33 tasks with complete reciprocal metric coverage yielded mean AUROC of 0.689 (TCGA) and 0.630 (MSKCC). Restricting analysis to the eight highest-performing tasks yielded mean AUROCs of 0.831 and 0.801, respectively. These tasks correspond to established morphologic-genomic associations (e.g., LGG IDH1, COAD MSI/BRAF, THCA BRAF/NRAS, BLCA FGFR3, UCEC PTEN) and showed the most stable cross-site performance. Differences between canonical encoders were modest relative to task-specific variability.
Conclusions:
Computational pathology is entering a translational phase in which reproducibility, transparency, and cross-institutional robustness are prerequisites for clinical trust. GOLDMARK establishes a reference framework that separates dataset curation from model evaluation and introduces structured intermediate artifacts, quality-control metadata, and symmetric cross-dataset testing as core components of benchmarking. Such infrastructure is essential for transforming computational biomarkers from research demonstrations into reproducible, clinically trusted workflows.
Related Concept Videos
Receiver Operating Characteristic Plot
Automated Microbial Diagnostics
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
Data Validation
Nursing assessment guides are generally based on holistic models rather than medical...
Myasthenia Gravis: Diagnostic Tests
The edrophonium test is a diagnostic tool for myasthenia gravis. It involves...


