Related Experiment Video
Updated: Aug 23, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Error Detection and Correction in Chinese Radiology Reports Using Large Language Models: Real-World Clinical
Jiafeng Zhou1, Yuxin Wei2, Qian Cai3
1Radiology Department, The First Affiliated Hospital of Wenzhou Medical University, Nanbaixiang, Ouhai District, Wenzhou, Zhejiang, China.
Background:
Large language models (LLMs) show promise in automatically detecting errors in radiology reports, but their performance remains insufficiently validated in large-scale, real-world clinical datasets.
Objective:
This study aimed to systematically evaluate the performance of LLMs in detecting and correcting errors in Chinese radiology reports derived from authentic clinical data.
Methods:
A large-scale dataset of 4480 Chinese radiology reports with modification records containing real clinical practice-generated errors was retrospectively collected between January 2023 and June 2024 at a single institution. After exclusions, 1363 reports containing 1551 errors were included. The dataset covers various anatomical parts of the body from different imaging modalities and was randomly divided into a test set (n=1263) and an internal validation set (n=100). Additionally, 100 error-free reports were added to the internal validation set. An additional 200 English-language reports from the Medical Information Mart for Intensive Care (MIMIC-III) were used for external validation. Eight human readers and 8 widely adopted LLMs, enhanced by prompt engineering, were tasked with error detection. Overall and subgroup detection performance and reading time were evaluated. Correction suggestions from the 2 best-performing LLMs were reviewed by a senior radiologist.
Results:
On the test set, DeepSeek-R1 achieved the highest overall detection rate at 89% (95% CI 87%-90%), significantly better than the other 7 models (P=.001-.007). On the internal validation set, DeepSeek-R1 and Claude-3.5-Sonnet achieved detection rates of 83% (100/120; 95% CI 76%-89%) and 80% (96/120; 95% CI 72%-86%), respectively. DeepSeek-R1 showed performance comparable to radiologists (83%, 95% CI 76%-89% vs 80%, 95% CI 72%-86% for junior radiologists and 78%, 95% CI 70%-85% for senior radiologists; P=.39 and P=.19, respectively) and significantly better performance than that of nonradiologists and nonphysicians (83%, 95% CI 76%-89% vs 66%, 95% CI 57%-74% and 38%, 95% CI 30%-47%; P<.001, respectively). DeepSeek-R1 showed a false-positive rate comparable to radiologists (DeepSeek-R1 vs senior radiologists and junior radiologists, 3% vs 0% and 1%; P=.25 and P=.61, respectively) and a significantly lower rate than nonradiologists and nonphysicians (3% vs 13% and 17%; P=.02 and P=.002, respectively). On the external validation set, DeepSeek-R1 and Claude-3.5-Sonnet achieved detection rates of 94% (95% CI 89%-97%) and 93% (95% CI 88%-97%), respectively. The correction accuracy of DeepSeek-R1 and Claude-3.5-Sonnet was 95% and 91%, respectively.
Conclusions:
Enhanced LLMs, particularly DeepSeek-R1, demonstrated robust performance in error detection and correction within real-world Chinese radiology reports, supporting their clinical use for automated quality assurance and integration into workflows to improve reporting accuracy and efficiency.