Related Experiment Video
Updated: May 16, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Using large language models for automated assessment of reporting quality and completeness of prediction model
I Spiero1, A M Leeuwenberg1, J Kamperman1
1Julius Center for Health Sciences and Primary Care, University Medical Center Utrecht, Utrecht University, Utrecht, The Netherlands.
Background And Objectives:
Transparent and complete reporting in scientific papers is important for interpretation of study results and for downstream evidence generation, such as systematic reviews and clinical guidelines. Many reporting checklists and tools have been developed for various types of biomedical research studies, but adherence assessment to these checklists is costly and laborious. Recent developments in large language models (LLMs) can accelerate assessment by automatically scoring the items of a checklist. We aimed to evaluate whether LLMs can accurately and efficiently assess the quality and completeness of reporting in prediction model studies based on the Transparent Reporting of a multivariable prediction model for Individual Prognosis or Diagnosis (TRIPOD) reporting guideline.
Methods:
We selected and evaluated 5 LLMs (Gemini-2.5-pro, Gemma3, GPT-5, Granite3.3, and Llama3.2) in their ability to automatically assess the 93 items of the TRIPOD checklist of 70 manually scored papers. The LLMs were asked to score each item with "reported" or "not reported" and to provide a supporting text quote. We evaluated the LLMs in terms of score correctness (sensitivity, precision, and F1-score), quote correctness, and resources required, using the manual double-reviewer scored papers as reference.
Results:
We found that Gemini-2.5-pro and GPT-5 performed best in scoring items with F1-scores of 0.74 and 0.73, respectively. With regard to quote correctness, the Gemini-2.5-pro model returned correct quotes in 80% to 97% of the cases based on manual evaluation. The LLMs took 20 to 30 minutes to conduct one complete checklist assessment.
Conclusion:
We conclude that Gemini-2.5-pro and GPT-5 are suitable to implement in a tool in a semiautomated manner rather than in full automation to assist researchers, reviewers, and editors to check the quality and completeness of reporting of prediction model studies according to the TRIPOD reporting checklist. By assisting in the reporting quality and adherence scoring, the time spent on adherence assessments can be drastically decreased, and thereby reporting in the literature can be improved.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy