Related Experiment Video
Updated: Jun 27, 2026

A Multicenter MRI Protocol for the Evaluation and Quantification of Deep Vein Thrombosis
Published on: June 2, 2015
Large language models for structuring knee ultrasound reports-A comparative evaluation of DeepSeek R1, gemini 2.5
Min Tang1, Le Xu1, Libiao Zhu2
1Department of Ultrasound, The First Affiliated Hospital of USTC, Division of Life Sciences and Medicine, University of Science and Technology of China, Hefei, Anhui, 230001, China.
GPT-4o demonstrated superior performance in generating structured, clinically interpretable outputs from knee ultrasound reports compared to DeepSeek R1 and Gemini 2.5 Flash. This indicates its potential to enhance the consistency and clinical utility of ultrasound reporting.
Area of Science:
- Medical imaging analysis
- Artificial intelligence in healthcare
- Natural Language Processing (NLP)
Background:
- Free-text clinical reports pose challenges for data extraction and standardization.
- Large Language Models (LLMs) offer potential for automating the interpretation of unstructured medical data.
- Evaluating LLM performance in specialized medical domains like musculoskeletal ultrasound is crucial.
Purpose of the Study:
- To compare the efficacy of three leading LLMs (DeepSeek R1, Gemini 2.5 Flash, GPT-4o) in processing knee ultrasound reports.
- To assess the ability of LLMs to generate structured, clinically relevant outputs from free-text reports.
- To evaluate LLM-generated diagnostic impressions and clinical recommendations.
Main Methods:
- Retrospective analysis of 359 knee ultrasound reports.
- Utilized four prompt strategies for information extraction, structured reporting, impression generation, and recommendation generation.
- Employed physician-based subjective ratings and objective metrics for performance evaluation, including schema compliance, repeatability, and diagnostic accuracy.
Main Results:
- GPT-4o exhibited the highest overall performance across all evaluated tasks, including named entity recognition (NER) and diagnostic summarization.
- GPT-4o achieved superior diagnostic performance (F1-score=0.79) compared to DeepSeek R1 (0.54) and Gemini 2.5 Flash (0.45).
- DeepSeek R1 showed competitive structured reporting capabilities, while Gemini 2.5 Flash underperformed in key areas like NER accuracy.
Conclusions:
- GPT-4o consistently outperformed DeepSeek R1 and Gemini 2.5 Flash in generating clinically interpretable outputs from knee ultrasound reports.
- The findings support the potential of GPT-4o to improve the consistency and clinical utility of musculoskeletal ultrasound reporting.
- Further research can explore integrating advanced LLMs into clinical workflows for enhanced medical report analysis.
Related Concept Videos
Ultrasonography
During an ultrasonography procedure, a handheld device called a...
Ultrasound II: Endoscopic Ultrasound and FibroScan
Endoscopic Ultrasound (EUS):
Imaging Studies II: Ultrasonography

