Related Experiment Video
Updated: Jul 7, 2026

07:05
Ultrasound-guided Botulinum Toxin-A Injections: A Method of Treating Sialorrhea
Published on: November 9, 2016
Evaluation of Aesthetic Outcomes Following Botulinum Toxin Treatment Using Multimodal Large Language Models: A Paired
Aesthetic Surgery Journal. Open Forum
|July 6, 2026
Summary
Multimodal large language models (MLLMs) can identify post-treatment images but struggle with nuanced aesthetic improvement detection. Further development is needed for reliable clinical application in aesthetic medicine.
Area of Science:
- Artificial Intelligence in Medicine
- Medical Aesthetics
- Computer Vision
Background:
- Multimodal large language models (MLLMs) are increasingly used for visual assessment.
- Their efficacy in evaluating aesthetic treatment outcomes is not well-established.
Purpose of the Study:
- To evaluate if current MLLMs can identify treatment status and regional aesthetic improvements after botulinum neurotoxin injections.
- To assess MLLM performance against established benchmarks.
Main Methods:
- Observational study analyzing 23 paired facial image cases (46 images).
- Four leading MLLMs (GPT-4.5 Pro, Grok 4.1, Gemini 3.1 Pro, Claude Opus 4.6) performed multiple inference runs.
- Models assessed post-treatment status and regional improvement (forehead, glabella, periorbital).
- Performance metrics included accuracy, sensitivity, specificity, and Fleiss' kappa.
Main Results:
- All MLLMs accurately identified post-treatment images (100% accuracy).
- Region-specific improvement detection often failed to surpass majority baselines (65.2%-91.3%).
- Gemini 3.1 Pro demonstrated superior forehead and glabella performance; no model met the periorbital baseline.
- Models systematically overestimated improvement, with false positives outweighing false negatives.
Conclusions:
- MLLMs can recognize aesthetic changes but lack clinical nuance.
- Domain-specific fine-tuning, expert benchmarking, and structured outcome frameworks are crucial for clinical deployment.
- Governance of training data and safeguards are essential before MLLM use in clinical or research settings.