Related Experiment Video
Updated: Sep 30, 2026

Single-Port Robotic-assisted Transaxillary Breast-conserving Surgery: A Prospective, Single-arm, Non-randomized Phase IIa Clinical Trial
Published on: August 19, 2025
Reducing Uncertainty in Posttreatment Breast Cosmesis Evaluation With Vision-Language Intelligence
Sangjoon Park1, Hwa Kyung Byun2, Yeona Cho3
1Department of Radiation Oncology, Yonsei Cancer Center, Heavy Ion Therapy Research Institute, Yonsei University College of Medicine, Seoul, South Korea; Yonsei Institute for Digital Health, Yonsei University, Seoul, South Korea.
Purpose:
Aesthetic outcomes following breast-conserving therapy are critical determinants of quality of life, yet current evaluation methods are limited by the subjectivity of existing scales and substantial interobserver variability. We aimed to evaluate the concordance of large vision-language model assessments with expert consensus, their run-to-run repeatability, and their performance across independent external cohorts as a potential adjunct for breast cosmesis assessment.
Methods:
We conducted a cross-sectional analysis of 89 frontal-view breast images from the Radiation Therapy Oncology Group (RTOG) 1014 trial as the development cohort and evaluated the locked configuration in 2 independent external cohorts: the National Surgical Adjuvant Breast and Bowel Project B-39/RTOG 0413 (n = 1850) and RTOG 1005 (n = 1728). Expert-consensus labels were established by 6 radiation oncologists using the Global Cosmetic Score. A subset of 34 randomly selected images was evaluated by 7 clinicians to assess exposure-dependent calibration patterns. Multiple large vision-language models, including the GPT-5 family, Gemini 2.0 Flash, and Claude-4.5 Haiku, were evaluated after systematic optimization of prompt complexity, scoring scales, and few-shot strategies using the development cohort. We assessed concordance with expert consensus using Kendall τ and Spearman ρ, run-to-run repeatability using the intraclass correlation coefficient, metric-level performance, and performance across the 2 external cohorts.
Results:
In the development cohort, the optimized GPT-5-based framework showed moderate concordance with expert consensus (Kendall τ = 0.567; Spearman ρ = 0.697). Concordance was lower in the 2 external cohorts: the National Surgical Adjuvant Breast and Bowel Project B-39/RTOG 0413 (τ = 0.385; ρ = 0.485) and RTOG 1005 (τ = 0.371; ρ = 0.469). In the shared 34-case subset, human experts showed higher concordance (τ = 0.732; ρ = 0.808) than the large vision-language model framework (τ = 0.572; ρ = 0.725). Run-to-run repeatability was high (intraclass correlation coefficient [2,1] = 0.982).
Conclusions:
The evaluated large vision-language model framework demonstrated high run-to-run repeatability and moderate concordance with expert consensus. These findings support further evaluation of the framework as an adjunct to expert assessment in breast cosmesis research.
