Related Experiment Video
Updated: Apr 15, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Zero-shot interpretable biomedical literature appraisal with generative large language models
Fangwen Zhou1, Muhammad Afzal2, Ashirbani Saha3
1Health Information Research Unit, Department of Health Research Methods, Evidence, and Impact, Faculty of Health Sciences, McMaster University, Hamilton, ON L8S 4K1, Canada.
Generative Pre-trained Transformer (GPT) models like GPT-4o can automate randomized controlled trial (RCT) appraisal, showing performance comparable to specialized models when using full text. This AI-driven approach enhances transparency in critical appraisal.
Area of Science:
- Artificial Intelligence in Medicine
- Biomedical Informatics
- Clinical Trial Methodology
Background:
- Automating the critical appraisal of randomized controlled trials (RCTs) is crucial for efficient knowledge synthesis.
- Large language models (LLMs) offer potential for automating complex scientific text analysis.
Purpose of the Study:
- To evaluate the performance of two decoder-based Generative Pre-trained Transformer (GPT) models (GPT-4o and GPT-3.5-mini) in automating RCT methodological appraisal.
- To compare GPT models against a fine-tuned encoder-only BioLinkBERT model using various prompting strategies.
Main Methods:
- A stratified random sample of 800 RCT articles was appraised.
- Two prompting schemes were used: classifier (independent assessment) and verifier (validation of BioLinkBERT).
- Assessments considered either title/abstract (TIAB) or full text, with performance measured against human assessments using Matthews correlation coefficient (MCC).
Main Results:
- GPT-4o as a classifier using full text achieved an MCC of 0.429, comparable to BioLinkBERT (MCC 0.466).
- GPT-4o as a verifier using full text showed similar performance (MCC 0.391).
- GPT models provided transparent, criterion-specific justifications, but performance significantly decreased when using only TIAB (MCC ≤0.100).
Conclusions:
- GPT-4o effectively automates RCT critical appraisal when full text is available, offering comparable performance to specialized models.
- GPT models enhance interpretability and transparency through explicit justifications.
- Fine-tuned models may complement GPTs when full texts are unavailable, and prompt optimization is key for clinical adoption.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
Genomics
Genetic Lingo
Leaky Scanning
Genetic Screens
Forward genetic screens
Forward or “classical” genetic screens involve creating random mutations in an organism’s DNA using radiation, mutagens, or insertion of additional bases, which...

