Related Experiment Video
Updated: Jul 7, 2026

Ultrasound-guided Botulinum Toxin-A Injections: A Method of Treating Sialorrhea
Published on: November 9, 2016
Evaluation of Aesthetic Outcomes Following Botulinum Toxin Treatment Using Multimodal Large Language Models: A Paired
Background:
Multimodal large language models (MLLMs) are increasingly applied to visual assessment tasks, yet their ability to evaluate aesthetic treatment outcomes remains unclear.
Objectives:
The aim of this study was to assess whether contemporary MLLMs can identify treatment state and detect region-specific aesthetic improvement following botulinum neurotoxin treatment.
Methods:
In this observational study, 23 paired facial image cases (46 images; 460 model evaluations) were analyzed. Four MLLMs (GPT-5.4 Pro [OpenAI, San Francisco, CA], Grok 4.1 [xAI, San Francisco, CA], Gemini 3.1 Pro [Google DeepMind, Mountain View, CA], and Claude Opus 4.6 [Anthropic, San Francisco, CA]) performed 5 independent inference runs per case. Models identified the posttreatment image and assessed regional improvement (forehead, glabella, and periorbital). Accuracy, sensitivity, specificity, balanced accuracy, Matthews correlation coefficient, and Fleiss' κ were calculated descriptively. Performance was compared with majority-class baselines. Exploratory outputs included aesthetic scores and apparent age estimates.
Results:
All models identified the posttreatment image (100% accuracy). Region-specific improvement detection frequently failed to exceed majority-class baselines (65.2%-91.3%). Gemini 3.1 Pro showed the highest performance for the forehead (74.8%) and glabella (63.5%), whereas no model reached the periorbital baseline. Inter-run reliability varied widely (κ -0.113 to 0.719). High reliability did not imply correctness. All models systematically overestimated improvement (62.6%-94.8% of predictions exceeded ground truth). False positives exceeded false negatives in all 12 model-task combinations. Exploratory outputs indicated perceived rejuvenation.
Conclusions:
MLLMs recognize the format of aesthetic change but not its clinical nuance. Bridging this gap requires more than improved accuracy: a coordinated agenda of domain-specific fine-tuning, expert-rater benchmarking, and structured outcome frameworks, supported by governance of training-data provenance and clear safeguards before any clinical or research deployment.
Level Of Evidence 5 Therapeutic:
For image description, please refer to the figure legend and surrounding text.