Related Experiment Video
Updated: Aug 8, 2026

07:15
Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
Published on: August 16, 2020
6.6K
Evaluating the reference accuracy of large language models in radiology: a comparative study across subspecialties
Yasin Celal Güneş1, Turay Cesur2, Eren Çamur3
1Kırıkkale Yüksek İhtisas Hospital, Clinic of Radiology, Kırıkkale, Türkiye.
Summary
Claude 3.5 Sonnet excels at generating accurate radiology references, significantly outperforming other large language models like ChatGPT and Google Gemini. This advancement offers dependable citations for radiology research and education, mitigating risks of misinformation.
Area of Science:
- Artificial Intelligence in Medical Imaging
- Natural Language Processing for Scientific Literature
- Radiology Informatics
Background:
- Large language models (LLMs) are increasingly used in academic and clinical settings.
- Accurate generation of scientific references is crucial for research integrity and avoiding misinformation.
- The performance of LLMs in generating domain-specific, verifiable citations remains largely unevaluated.
Purpose of the Study:
- To compare the accuracy, fabrication rates, and bibliographic completeness of six leading LLMs in generating radiology references.
- To assess the performance of models including ChatGPT variants, Google Gemini 1.5 Pro, Claude 3.5 Sonnet, and Claude 3 Opus.
Main Methods:
- A cross-sectional observational study administered 120 open-ended questions across eight radiology subspecialties.
- LLMs were prompted to generate four references with full bibliographic details and in-text citations for each question.
- References were rigorously verified using multiple databases and search engines, with accuracy scored on a 5-point Likert scale.
Main Results:
- Claude 3.5 Sonnet achieved the highest reference accuracy (80.8% fully accurate) and lowest fabrication rate (3.1%), significantly outperforming all other models.
- Claude 3 Opus showed moderate performance (59.6% accurate, 18.3% fabrication).
- ChatGPT models and Google Gemini 1.5 Pro demonstrated substantially lower accuracy and higher fabrication rates, with Gemini 1.5 Pro performing the worst.
Conclusions:
- Claude 3.5 Sonnet is the leading LLM for generating accurate radiology references, enhancing literature reviews and educational materials.
- The high fabrication rates of other models, including ChatGPT and Google Gemini, pose risks of misinformation in clinical and academic contexts.
- Further refinement of LLMs is necessary to ensure their safe and effective application in academic citation generation.

