Related Experiment Video
Updated: Sep 30, 2026

Single-Port Robotic-assisted Transaxillary Breast-conserving Surgery: A Prospective, Single-arm, Non-randomized Phase IIa Clinical Trial
Published on: August 19, 2025
Reducing Uncertainty in Post-Treatment Breast Cosmesis Evaluation with Vision-Language Intelligence
Sangjoon Park1, Hwa Kyung Byun2, Yeona Cho3
1Department of Radiation Oncology, Heavy Ion Therapy Research Institute, Yonsei Cancer Center, Yonsei University College of Medicine, Seoul, South Korea; Yonsei Institute for Digital Health, Yonsei University, Seoul, South Korea.
Purpose:
Aesthetic outcomes following breast-conserving therapy are critical determinants of quality of life, yet current evaluation methods are limited by the subjectivity of existing scales and substantial inter-observer variability. We aimed to evaluate the concordance of large vision-language model assessments with expert consensus, their run-to-run repeatability, and their performance across independent external cohorts as a potential adjunct for breast cosmesis assessment.
Methods:
We conducted a cross-sectional analysis of 89 frontal-view breast images from the RTOG 1014 trial as the development cohort and evaluated the locked configuration in two independent external cohorts: NSABP B-39/RTOG 0413 (n = 1,850) and RTOG 1005 (n = 1,728). Expert-consensus labels were established by six radiation oncologists using the Global Cosmetic Score. A subset of 34 randomly selected images was evaluated by seven clinicians to assess exposure-dependent calibration patterns. Multiple large vision-language models, including the GPT-5 family, Gemini 2.0 Flash, and Claude-4.5 Haiku, were evaluated after systematic optimization of prompt complexity, scoring scales, and few-shot strategies on the development cohort. We assessed concordance with expert consensus using Kendall's τ and Spearman's ρ, run-to-run repeatability using the intraclass correlation coefficient, metric-level performance, and performance across the two external cohorts.
Results:
In the development cohort, the optimized GPT-5-based framework showed moderate concordance with expert consensus (Kendall's τ = 0.567; Spearman's ρ = 0.697). Concordance was lower in the two external cohorts: NSABP B-39/RTOG 0413 (τ = 0.385; ρ = 0.485) and RTOG 1005 (τ = 0.371; ρ = 0.469). On the shared 34-case subset, human experts showed higher concordance (τ = 0.732; ρ = 0.808) than the LVLM framework (τ = 0.572; ρ = 0.725). Run-to-run repeatability was high (ICC[2,1] = 0.982).
Conclusions:
The evaluated LVLM framework demonstrated high run-to-run repeatability and moderate concordance with expert consensus. These findings support further evaluation of the framework as an adjunct to expert assessment in breast cosmesis research.
