Related Experiment Video
Updated: Mar 16, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Objective quality assessment of neuroradiology reports using large language models
P López-Úbeda1, T Martín-Noguerol2, A Luna2
1NLP Department HT Médica Carmelo Torres n°2, 23007, Jaén, Spain.
Aim:
Radiology reports are essential for clinical decision-making and must meet standards of clarity, completeness, and diagnostic accuracy. However, quality assessment is often subjective, time-consuming, and dependent on expert reviewers. Large language models (LLMs) offer a promising alternative for automating this process. This study evaluates whether LLMs can objectively and consistently assess the formal and diagnostic quality of neuroradiology reports.
Materials And Methods:
We analysed 277 neuroradiology reports, originally authored by 10 radiologists and subsequently evaluated by a different radiologist to assess formal and diagnostic quality. Reports were annotated using six quality criteria: diagnostic discrepancy, report structure, clarity, completeness, grading/staging system use, and recommendation of additional tests. Three locally deployed LLMs (LLaMA 3.2, DeepSeek R1:7B, and Gemma 3:4B) were evaluated for their ability to classify reports according to these criteria.
Results:
LLMs showed strong performance in recognising positive categories, with DeepSeek achieving the highest accuracy in report structure (70.40%) and test recommendations (68.59%). However, performance declined significantly in detecting issues such as report completeness or diagnostic discrepancies, reflected in low F1-scores for those categories.
Conclusion:
While current models have limitations in identifying subtle errors or complex clinical nuances, LLMs may assist in automating aspects of radiology report quality control. With further refinement, they could help improve consistency and support clinical decision-making.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy