Related Experiment Video
Updated: Feb 28, 2026

07:31
Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
8.2K
Extending BEHRT to UK Biobank: assessing transformer model performance in clinical prediction
Yusuf Yildiz1, Goran Nenadic2, Meghna Jani3,4
1Faculty of Biology, Medicine and Health, School of Health Sciences, Division of Informatics, Imaging and Data Sciences, University of Manchester, Manchester, United Kingdom.
Frontiers in Digital Health
|February 26, 2026
Summary
Transformer models show promise for clinical prediction using electronic health records. Model size and medical terminology significantly impact long-term diagnostic prediction accuracy.
Area of Science:
- Artificial Intelligence in Healthcare
- Clinical Informatics
- Biomedical Data Science
Background:
- Transformer models demonstrate significant potential for clinical prediction tasks utilizing electronic health record (EHR) data.
- Model performance is known to be influenced by various factors including modeling decisions and data characteristics.
Purpose of the Study:
- To evaluate the performance of a BERT-based model for Healthcare (BEHRT) on clinical prediction tasks using UK Biobank data.
- To assess the impact of model size, medical terminology (CALIBER vs. ICD-10), and data splitting strategies on prediction accuracy.
Main Methods:
- Trained a BEHRT model on hospital-based UK Biobank data.
- Evaluated model performance on four clinical prediction tasks, including next-visit and up to five-year diagnosis prediction.
- Systematically analyzed the effects of model size, vocabulary choice, and data partitioning.
Main Results:
- Larger transformer models achieved superior performance in long-term prediction tasks (5-year AUROC = 0.874 vs. 0.858 for smaller models).
- Performance differences were minimal for short-term (6-month) prediction tasks.
- Utilizing the CALIBER medical terminology resulted in higher average precision scores compared to ICD-10 (0.773 vs. 0.678).
Conclusions:
- Transformer models can achieve high predictive performance across various clinical prediction scenarios.
- Model outcomes are highly dependent on specific modeling choices, especially for long-term predictions.
- Careful consideration of model size and medical terminology is crucial for optimizing clinical prediction performance.