Related Experiment Video
Updated: Sep 30, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Automating Clavien-Dindo classification with large language models in percutaneous nephrolithotomy: an exploratory
Gustavo Perrone1, Alexandre Danilovic2, Carlos Batagello3
1Faculdade de Medicina FMUSP, Universidade de Sao Paulo, Sao Paulo, SP, Brazil.
Purpose:
To explore the use of large language models for Clavien-Dindo classification of complications following percutaneous nephrolithotomy by benchmarking their performance against physician grading.
Methods:
We evaluated ChatGPT, Copilot, DeepSeek, Gemini, Grok and Llama for defining Clavien-Dindo grades, classifying complications according to the Clavien-Dindo classification, and classifying case summaries with and without in-context prompting. Vignettes were obtained from a published consensus categorization of percutaneous nephrolithotomy complications. We measured agreement with Gwet's AC2 coefficient and compared it with attending urologists and urology residents.
Results:
All models accurately defined the Clavien-Dindo classification and graded with almost perfect agreement complication descriptions and case summaries (0.97 to 0.98 overall agreement). Six attendings and 15 residents classified the same vignettes, demonstrating lower agreement than the chatbots (0.95, p < 0.001 for residents; 0.96, p = 0.03 for attendings; 0.95, p < 0.001 for all participants). Models' latency ranged from 0.77 to 3.90 min, while the fastest physician required 17 min to complete the same tasks. Models' intrarater reliability was almost perfect: from 0.97 to 0.99. Interrater reliability between chatbots (0.97, 95% confidence interval: 0.97-0.98) was significantly higher than among physicians (0.93, 95% confidence interval: 0.92-0.94).
Conclusion:
Large language models show promise in being faster, more accurate, and precise than physicians in Clavien-Dindo classification in an experimental setting. We support the need for a prospective agreement study using real-world clinical records. We urge discussion of the legal and ethical implications.