Related Experiment Video
Updated: Jun 27, 2026

Guidelines and Experience Using Imaging Biomarker Explorer IBEX for Radiomics
Published on: January 8, 2018
Comparative Evaluation of BI-RADS Classification, Clinical Management, and Diagnostic Performance in Breast
Duygu Erkal1, Mehmet Tonkaz2, Tümay Bekci2
1Department of Radiology, Giresun Education and Research Hospital, Giresun, Turkey.
None:
This study aimed to compare the Breast Imaging Reporting and Data System (BI-RADS) classifications, clinical management recommendations, and diagnostic performance of two artificial intelligence (AI) systems-ChatGPT-4o and DeepSeek-V3-with those of an experienced radiologist in interpreting breast ultrasound reports. The clinical inconsistency and referral safety of the AI systems were assessed. A total of 595 breast ultrasound reports were independently and blindly assessed by a radiologist and two AI systems-ChatGPT-4o and DeepSeek-V3. Observers provided BI-RADS categories and management recommendations based on the findings section. Diagnostic performance was assessed using histopathological diagnosis or ≥ 24-month follow-up as the standard. Clinical inconsistency means discordance between the BI-RADS categories and corresponding recommendations. Referral safety was quantified using a standardized referral safety score. Interobserver agreement was substantial for both BI-RADS classification (κ = 0.65) and management recommendations (κ = 0.79). The radiologist recorded the highest malignancy detection sensitivity (0.989), followed by DeepSeek-V3 (0.968) and ChatGPT (0.927). Specificity (0.735-0.753) and negative predictive values (0.981-0.997) were comparable across the models, while positive predictive values were limited (0.403-0.436). BI-RADS and management recommendations for ChatGPT were consistent, whereas DeepSeek-V3 had a 0.4% rate (p >.05). Referral safety scores were similar (ChatGPT: 0.44; DeepSeek-V3: 0.46). ChatGPT-4o and DeepSeek-V3 demonstrated substantial consistency with expert interpretations and acceptable diagnostic performance in US breast evaluations. Although radiologists demonstrated the most accurate and balanced performance, these findings suggest that ChatGPT-4o and DeepSeek-V3 may aid radiology decision support, pending further validation.

