Related Experiment Videos
High-Recall Biomedical Language Models for Radiation Oncology Evidence Synthesis: An Integrated Systematic Review of
Guang-Zhi Lin1, Yang-Wei Hsieh1,2, Chih-Yu Cheng1
1Medical Physics and Informatics Laboratory of Electronic Engineering, National Kaohsiung University of Science and Technology, Kaohsiung 80778, Taiwan.
Background:
Radiotherapy-related complications can impair long-term outcomes in nasopharyngeal carcinoma (NPC), while high-volume literature screening remains a bottleneck in evidence synthesis.
Methods:
Four databases were searched through January 2026. After deduplication, 4916 records underwent two-stage manual screening, yielding 38 studies. Three Bidirectional Encoder Representations from Transformers (BERT) models, BERT-base, BioBERT, and PubMedBERT, were fine-tuned using stratified five-fold cross-validation, and screening performance was assessed using the area under the receiver operating characteristic curve (ROC-AUC), the area under the precision-recall curve (PR-AUC), and work saved over sampling (WSS) at the 100% and 85% recall thresholds (WSS@100% and WSS@85%).The clinical evidence was evaluated using the Prediction Model Risk of Bias Assessment Tool (PROBAST) and synthesized narratively according to endpoint, discrimination metric, time horizon, unit of analysis, validation level, and uncertainty reporting.
Results:
At complete recall, BioBERT and PubMedBERT achieved mean WSS values of 97.8% and 97.5%, respectively, compared with 89.7% for BERT-base. Separately, the clinical evidence review identified 39 discrimination estimates: 33 conventional binary ROC-AUCs, three Harrell's C-indices, and three time-dependent AUCs. Only two studies contributed external-validation estimates and one contributed a temporal-validation estimate. Sixteen estimates lacked a usable measure of uncertainty, and PROBAST rated all studies as having a high overall risk of bias, driven by the analysis domain.
Conclusions:
Within the retrieved and manually annotated corpus, biomedical pretrained language models showed the potential to reduce the title-and-abstract screening burden while maintaining complete recall. Independently, the clinical evidence map identified promising discrimination but insufficiently robust and transportable evidence for routine clinical implementation.