Related Experiment Video
Updated: Apr 16, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Fine-Tuned Large Language Model for Automated Radiology Impression Generation: A Multicenter Evaluation
Mingyang Li1, Yaning Wang1, Zheng Miao1
1Department of Radiology, The First Hospital of Jilin University, No. 79 Xinmin St, Changchun 130021, China.
None:
Purpose To develop a fine-tuned large language model (Medical Imaging Report Assistant [MIRA]) and evaluate its performance in generating radiology impressions from multicenter data with respect to accuracy, reporting efficiency, and clinical applicability. Materials and Methods A retrospective multicenter dataset comprising 1.87 million radiology reports (including CT, MRI, and digital radiography data) from 42 hospitals across 22 provinces in China (January 2019-August 2024) was compiled. The dataset was used to fine-tune a large language model via a prompt-based strategy. The evaluation framework incorporated both automated and human evaluation metrics. Radiologists evaluated internal and external datasets and three open-source datasets to compare impressions generated by the fine-tuned large language model and GPT-4o (Open AI). Twenty-four radiologists from six centers performed blinded comparisons of MIRA-generated and reference impressions to assess interrater consistency and drafting efficiency. Data were analyzed using appropriate parametric and nonparametric tests and χ2 tests, with Holm-Bonferroni correction for multiple comparisons. Results The internal test set included data for 78 544 reports (median age, 52 years; IQR, 35-65 years; 39 351 male), and the external test set included data for 27 471 reports (median age, 53 years; IQR, 37-66 years; 13 955 male). Site- and modality-aware prompting improved similarity (internal BERTScore F1 and sentence similarity, 0.92 and 0.92, respectively; external BERTScore F1 and sentence similarity, 0.82 and 0.80, respectively, under optimal settings; P < .001). Human evaluation (n = 2327) showed MIRA beat GPT-4o on both similarity and F1 score (P < .001). MIRA-generated impressions were rated as at least as good as the reference impressions in 69.0% (1657 of 2400) of blinded comparisons, reduced draft time by 0.46 minutes per report, and increased interradiologist agreement (P < .001). Conclusion MIRA, a fine-tuned large language model using a prompt-based strategy, generated clinically aligned radiology impressions in multicenter settings, improving accuracy, efficiency, and reporting consistency. Keywords: Computer-aided Diagnosis, CAD, Supervised Learning, Transfer Learning, Conventional Radiography, MRI, Computer Applications-General Informatics, Statistics Supplemental material is available for this article. © The Author(s) 2026. Published by the Radiological Society of North America under a CC BY 4.0 license.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy