Related Experiment Video
Updated: Aug 21, 2026

Lesion Explorer: A Video-guided, Standardized Protocol for Accurate and Reliable MRI-derived Volumetrics in Alzheimer's Disease and Normal Elderly
Published on: April 14, 2014
Automatic extraction of structured information from brain MRI reports using an open-weight large language model
Kaouther Mouheb1, Amos Pomp2, Antoine Manenti3,4
1Department of Radiology & Nuclear Medicine, Erasmus MC, Rotterdam, The Netherlands. k.mouheb@erasmusmc.nl.
Objectives:
Automatic data extraction from free-text radiology reports enables large-scale research, but few studies assessed the performance of large language models (LLMs) on Dutch neuroradiology reports.
Materials And Methods:
We analyzed 947 brain MRI reports from a tertiary memory clinic (2016-2021), authored by consultant neuroradiologists. Trained medical students annotated thirty variables; 100 reports were double annotated to assess inter-rater reliability. We evaluated the performance of the open-weight LLM LLaMA 3.1 using different languages (Dutch vs English translation) and few-shot prompting with different example selection strategies. Performance was evaluated using balanced accuracy for categorical variables, accuracy and mean absolute error for counts, and text similarity for free text. Metrics were computed across 10 random splits of the 947 reports.
Results:
LLaMA 3.1 demonstrated high zero-shot performance for visual rating scores (mean [95% CI]): medial temporal atrophy: 0.90 [0.77-1.00] on the left and 0.96 [0.94-0.99] on the right, Global Cortical Atrophy: 0.87 [0.83-0.91], and Fazekas: 0.94 [0.93-0.96]. Microbleed mentions were detected with 0.93 [0.92-0.95] accuracy, and infarct mentions with 0.82 [0.80-0.84]. Text similarity for lesion location reached 0.95 [0.95-0.96]. Performance was lower for numerical variables: 0.80 [0.78-0.82] for the number of microbleeds and 0.66 [0.63-0.68] for infarcts. English translation yielded comparable results. Few-shot prompting improved performance for numerical variables, achieving 0.92 [0.90-0.93] for microbleeds and 0.81 [0.77-0.85] for infarcts using structural similarity-based selection.
Conclusion:
LLaMA 3.1 shows strong potential for extracting data from Dutch neuroradiology reports. Few-shot prompting enhances performance for numerical variables, whereas challenges remain for location-specific variables.
Key Points:
Question The free-text format of radiology reports limits the accessibility of structured data; can LLM LLaMA 3.1 automatically structure free-text Dutch neuroradiology reports? Findings LLaMA 3.1 accurately extracted 30 variables from Dutch neuroradiology reports, with few-shot prompting significantly improving performance compared to zero-shot prompting, particularly with similarity-based example selection. Clinical relevance By enabling accurate, automated post-hoc structuring of free-text radiology reports, open-weight LLMs facilitate large-scale data reuse, improve consistency in reporting, and support clinical research and decision-making in dementia care without disrupting existing clinical workflows.
