Related Experiment Video
Updated: Sep 30, 2026

Use of Ultra-high Field MRI in Small Rodent Models of Polycystic Kidney Disease for In Vivo Phenotyping and Drug Monitoring
Published on: June 23, 2015
Large Language Models as Diagnostic Tools in Nephrology: Helpful, With Limitations
Background:
Large language models (LLMs) are increasingly being used to support diagnostic decision-making. It is unclear, however, whether their performance on knowledge and examination tasks translates to interpreting clinical case vignettes, especially when a complete diagnosis requires specifying several related elements, such as triggers, organ manifestations, or complications.
Methods:
We carried out a vignette-based benchmark study using 101 case reports from the German-language journal Die Nephrologie to assess six current language models from the Anthropic, Google, and Open AI companies. The strongest model of each company as of March 2026 was tested, as well as one smaller model from each. The models were required to provide a differential diagnosis with five disease entities and to name the most likely diagnosis. The answers were evaluated using a three-level scale. We drew a distinction between diagnoses with a single component and diagnoses with multiple components.
Results:
The most likely diagnosis was most often correctly named by Claude Opus 4.6 (86.1%, 95% confidence interval [78.1; 91.6]), with GPT-5.4 close behind (85.1% [76.9; 90.8]). When partly correct answers were also counted, the three best-performing models all scored in the narrow range of 91.1-92.1%. The answers were highly stable. For diagnoses with multiple components, incorrect answers were more often incomplete than entirely wrong.
Conclusion:
LLMs have high diagnostic accuracy when challenged with published nephrological case scenarios. The interpretation of these findings is limited, however, because previous exposure of the models to individual vignettes, or parts thereof, cannot be ruled out. For clinical use, it must be borne in mind that plausible answers can be incomplete. Global accuracy measures may overestimate the diagnostic reliability of AI language models. The completeness and causal plausibility of the answers and their relevance to clinical action should be systematically investigated in benchmark studies.
Related Concept Videos
Drug Dosing in Renal Diseases: Measurement of Glomerular Filtration Rate
Drug Dosing in Renal Diseases: Estimation of Glomerular Filtration Rate Based on Serum Creatinine Concentration
Acute Kidney Injury IV: Diagnostic Studies and Prevention
