Related Experiment Video
Updated: Aug 27, 2026

Using Optical Coherence Tomography and Optokinetic Response As Structural and Functional Visual System Readouts in Mice and Rats
Published on: January 10, 2019
A benchmark study of vision and pathology foundation models for computational pathology
Rohan Bareja1, Francisco Carrillo-Perez1,2, Yuanning Zheng1
1Division of Computational Medicine, Department of Medicine, Stanford University, Stanford, CA, USA.
Abstract:
To advance precision medicine in pathology, artificial intelligence (AI)-driven foundation models must generalize across diverse datasets, tissues, and clinical tasks. However, their comparative performance and generalizability in computational pathology remain incompletely characterized. Here, we benchmark 32 AI foundation models across four categories, including general vision models (VM), general vision-language models (VLM), pathology-specific vision models (Path-VM), and pathology-specific vision-language models (Path-VLM), using slide- and patch-level tasks from The Cancer Genome Atlas (TCGA), Clinical Proteomic Tumor Analysis Consortium (CPTAC), external benchmarking datasets, and out-of-domain datasets. Across TCGA tasks, Path-VMs consistently rank among the strongest performers. Evaluation across CPTAC and out-of-domain datasets reveals more nuanced generalization behavior, with model rankings showing modest but consistent shifts across datasets and task categories. Pairwise statistical comparisons indicate that differences among top-performing models are often small and task dependent. Path-VMs outperform Path-VLMs and remain competitive with VMs. Model size and pretraining dataset scale do not consistently predict downstream performance. Finally, late decision-level ensembling improves aggregate performance across external datasets and tissue types, highlighting complementary strengths across foundation models. PathBench: https://pathbench.stanford.edu/.

