Related Experiment Video
Updated: Jan 14, 2026

Introduction of an Integrated Pathology Image Management, Artificial Intelligence, and Reporting System
Published on: July 11, 2025
Concordance between artificial intelligence and radiologists in BIRADS classification of breast ultrasound: A study
Nigar Erkoc1, Elif Hazal Karlı1, Emre Tunay1
1Department of Radiology, Bagcilar Training and Research Hospital, Istanbul, Turkey.
Purpose:
To evaluate the diagnostic concordance and consistency of ChatGPT-4o, a multimodal large language model, in assigning BI-RADS categories on breast ultrasound images, and to compare its performance with that of experienced radiologists.
Materials And Methods:
In this retrospective, single-center study, 405 breast ultrasound images from 350 patients were analyzed. Two board-certified radiologists (8-10 years of experience) independently reviewed the images and assigned BI-RADS categories. ChatGPT-4o evaluated the same images in isolated sessions, using a standardized prompt, without access to clinical data or dynamic scanning features. Cohen's kappa was used to assess interobserver agreement between radiologists; Fleiss' kappa was used to measure agreement among the radiologists and ChatGPT-4o.
Results:
Interobserver agreement between the two radiologists was almost perfect (Cohen's κ = 0.832; p < 0.001). ChatGPT-4o showed moderate agreement with Radiologist 1 (κ = 0.593) and substantial agreement with Radiologist 2 (κ = 0.621). The highest concordance was observed in BI-RADS 1 (κ = 0.848) and BI-RADS 5 (κ = 0.894) categories, while agreement was lower in BI-RADS 3 (κ = 0.487). Overall agreement among all three readers was substantial (Fleiss' κ = 0.682; 95 % CI: 0.639-0.725). ChatGPT-4o occasionally upstaged borderline BI-RADS 3 cases to BI-RADS 4 and tended to misclassify anatomical structures, such as ribs or fibroglandular tissue, as lesions.
Conclusion:
ChatGPT-4o demonstrated promising diagnostic performance in breast ultrasound interpretation, particularly for clearly benign and malignant lesions. However, its limitations in intermediate-risk classification and artifact interpretation indicate that it should be used as an adjunct rather than a replacement for expert radiologist evaluation.

