Related Experiment Video
Updated: May 31, 2026

08:57
Identifying, Diagnosing, and Grading Malignant Peripheral Nerve Sheath Tumors in Genetically Engineered Mouse Models
Published on: May 17, 2024
LLM-Generated Transplant Algorithms: Comparing GPT and Gemini in Producing Step-Wise Diagnostic and Management
1Cam Sakura Training and Research Hospital, Department of General Surgery, University of Health Sciences, Istanbul, Türkiye.
The American Surgeon
|May 29, 2026
Summary
Gemini 2.5 Pro demonstrated superior performance over GPT-4 in generating transplant algorithms. This study highlights the potential of large language models (LLMs) for clinical protocol drafting, emphasizing expert validation.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Medicine
- Transplant Surgery
Background:
- Solid organ transplantation requires dynamic, evidence-based clinical algorithms.
- Large language models (LLMs) offer potential for rapid protocol drafting.
- Evaluating LLMs for educational reliability and feasibility in transplant medicine is crucial.
Purpose of the Study:
- To compare the performance of GPT-4 and Gemini 2.5 Pro in generating diagnostic and management algorithms for post-transplant scenarios.
- To assess LLMs as tools for synthesizing existing guidance, not for creating new clinical practice guidelines.
Main Methods:
- Ten high-stakes post-transplant scenarios were used for evaluation.
- Both LLMs received identical, structured prompts.
- Three transplant specialists independently scored algorithms for Clinical Concordance, Logical Flow, and Completeness.
Main Results:
- Gemini 2.5 Pro achieved higher median scores for Clinical Concordance, Logical Flow, and Completeness compared to GPT-4.
- Performance differences were notable in complex scenarios differentiating infectious and alloimmune causes of graft dysfunction.
- High inter-rater agreement (κ = 0.81) was observed among specialists.
Conclusions:
- Gemini 2.5 Pro outperformed GPT-4 in generating clinically concordant, structured, and complete transplant algorithms.
- LLM outputs require expert validation and should be used within a governance framework.
- LLMs are valuable tools for synthesizing guidance but not as autonomous clinical decision-makers.

