Related Experiment Video
Updated: Jul 17, 2026

14:27
Identification of Disease-related Spatial Covariance Patterns using Neuroimaging Data
Published on: June 26, 2013
Using a Large Language Model-Generated Prompt to Extract Features from Synthetic MRI Brain Scan Reports: A
John J Hanna1,2,3, Christopher S Evans2,4, Christopher R Dennis2
1Department of Internal Medicine, ECU Brody School of Medicine, Greenville, North Carolina, United States.
Methods of Information in Medicine
|February 19, 2026
Summary
Large language models (LLMs) show promise for extracting features from MRI brain reports. Newer models like GPT-4 perform well, and LLMs can even generate effective prompts for this task.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Medicine
Background:
- Automated feature extraction from medical reports is crucial for clinical, operational, and research purposes.
- Large Language Models (LLMs) offer potential for automating feature extraction and category assignment from clinical text.
Purpose of the Study:
- To compare the accuracy of feature extraction from MRI brain scan reports using clinician-engineered versus LLM-generated prompts.
- To evaluate the performance of five OpenAI LLMs in extracting predefined features from synthetic MRI reports.
Main Methods:
- Five OpenAI LLMs were tested on their ability to extract nine binary features from synthetic MRI brain reports.
- Two prompt types were used: clinician-engineered and LLM-generated. Performance was assessed using recall, precision, accuracy, and F1 score.
Main Results:
- High overall performance was observed across all models and prompts, with average recall of 0.956, precision of 0.9347, accuracy of 0.982, and F1 score of 0.9431.
- GPT-3.5-turbo showed better performance with an LLM-generated prompt, while GPT-4 models consistently outperformed others regardless of prompt type.
Conclusions:
- LLMs demonstrate significant potential for accurate feature extraction from MRI brain reports, with newer models like GPT-4 showing robust performance.
- The choice of LLM and prompt engineering strategy significantly impacts the efficacy of automated feature extraction from medical imaging reports.

