Established machine learning matches tabular foundation models in clinical predictions.
Lawrence A Shaktah1, Marco Gustav1, Tim Lenz1
1Else Kroener Fresenius Center for Digital Health, Faculty of Medicine, TUD Dresden University of Technology, Dresden, Germany.
BMC Medical Informatics and Decision Making
|July 3, 2026
Summary
Foundation models (FMs) show limited clinical value for tabular data prediction. TabPFN, a leading FM, did not consistently outperform established machine learning methods and incurred higher computational costs.
Area of Science:
- Clinical informatics
- Machine learning
- Biostatistics
Background:
- Foundation models (FMs) offer potential for standardized predictive modeling.
- The clinical utility of FMs for tabular data, particularly in healthcare, remains largely unproven.
- Evaluating FMs against established methods is crucial for understanding their practical application.
Purpose of the Study:
- To benchmark the performance of TabPFN, a prominent FM for tabular data, against traditional machine learning (ML) methods.
- To assess the clinical value of TabPFN across diverse binary prediction tasks using real-world patient data.
- To evaluate the efficiency and computational costs associated with using TabPFN in clinical settings.
Main Methods:
- A large-scale, reproducible benchmark comparing TabPFN to twelve ML methods.
- Evaluation across twelve distinct binary clinical prediction tasks with patient cohorts ranging from 788 to 139,528.
- Standardized preprocessing, bootstrapping, and multiple performance metrics including AUROC were employed.
Main Results:
- TabPFN demonstrated competitive performance but did not consistently surpass strong ML baselines.
- TabPFN outperformed the best ML model in only 16.7% of tasks, with minimal AUROC differences (±0.02).
- TabPFN exhibited significantly higher computational costs (5.5x longer median runtimes) and required GPU acceleration.
Conclusions:
- For routine binary clinical prediction, TabPFN offers marginal performance improvements over optimized ML methods.
- The adoption of TabPFN in clinical practice may be limited by its substantial efficiency trade-offs.
- Further research is needed to identify specific clinical scenarios where FMs provide a definitive advantage over traditional ML approaches.


