Related Experiment Video
Updated: Aug 11, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Reporting Quality of Large Language Model Studies: A Cross-Sectional Audit of High-Ranking Radiology and Medical
Ismail Mese1, Saime Turgut Gunes2, Ozge Coskun2
1Department of Radiology, Üsküdar State Hospital, Istanbul, Türkiye.
Objective:
To evaluate adherence to the Minimum Reporting Items for Clear Evaluation of Accuracy Reports of Large Language Models in Healthcare (MI-CLEAR-LLM) in radiology and medical imaging studies involving large language models (LLMs).
Materials And Methods:
We conducted a cross-sectional audit of original LLM research studies published between January 1 and December 26, 2025, in Q1 journals within the Web of Science "Radiology, Nuclear Medicine, and Medical Imaging" category. PubMed and Scopus were searched to identify eligible studies. A quota-based subsampling strategy, based on journal publication volume, was used to select approximately 100 studies. All four eligible articles from the Korean Journal of Radiology (KJR) were additionally included as a benchmark. Adherence to the 2025 update of MI-CLEAR-LLM was scored through a two-round, consensus-based process: an initial assessment by one reviewer followed by a critical re-evaluation by secondary reviewers, with consensus adjudication by an additional reviewer when needed. Between-journal differences were analyzed with the Kruskal-Wallis test, followed by Dunn post hoc pairwise comparisons with Holm-adjusted P-values.
Results:
Of 201 eligible studies identified, 102 were finally analyzed after applying the subsampling strategy. Overall adherence to MI-CLEAR-LLM was moderate (mean, 51.2% ± 14.7%; range, 22.2%-84.2%). Adherence was highest for input data type (100%), test-data independence (80.2%), and adaptation strategy (78.1%), and lowest for prompt execution setup (29.4%) and stochasticity management (33.1%). The least frequently reported items were training-data cutoff date (9.8%) and rationale for prompt wording (15.6%). Adherence varied significantly across journals (P = 0.011), with KJR showing the highest mean adherence (72.8% ± 2.7%).
Conclusion:
Reporting transparency in radiology and medical imaging LLM studies published in 2025 was inconsistent across reporting items and journals, with substantial deficiencies in some reproducibility-critical elements. Broader adoption of reporting standards is essential to improve the reproducibility and interpretability of future accuracy evaluations.
Related Concept Videos
Radiological Investigation I: X-ray and CT
Positron Emission Tomography
One of the main requirements of a PET scan is a positron-emitting radioisotope, which is produced in a cyclotron and then attached to a substance used by the part of the body being...
Radiological Investigation II: MRI and Ventilation Perfusion Scan
Magnetic Resonance Imaging (MRI) and Ventilation Perfusion Scans are two radiological investigations that offer detailed diagnostic images of the body, particularly lung structures.
MRI
MRI uses magnetic fields and radiofrequency signals to distinguish between normal and abnormal tissues. This technology provides a more detailed diagnostic image than CT scans, enabling it to characterize pulmonary nodules, stage bronchogenic carcinoma, and evaluate inflammatory activity in...