Related Experiment Video
Updated: Aug 5, 2026

07:30
Evaluation of the Cognitive Performance of Hypertensive Patients with Silent Cerebrovascular Lesions
Published on: April 23, 2021
Diagnostic Performance and Workup Efficiency of Large Language Models in Secondary Hypertension: A Blinded
Asena Gökçay Canpolat1, Özge Baş Aksu1, Rıfat Emral1
1Department of Endocrinology and Metabolism, Ankara University School of Medicine, Ankara 06230, Turkey.
Diagnostics (Basel, Switzerland)
|July 28, 2026
Summary
Large language models (LLMs) show varied performance in diagnosing and managing secondary hypertension (SH). Claude Sonnet 4.6 outperformed GPT-5.2 and Gemini 3 Pro in complex clinical reasoning tasks for SH.
Area of Science:
- Artificial Intelligence in Medicine
- Clinical Decision Support Systems
- Natural Language Processing
Background:
- Secondary hypertension (SH) presents diagnostic and management complexities for AI systems.
- Evaluating AI performance in SH clinical reasoning is crucial for developing effective decision support.
Purpose of the Study:
- To comparatively assess three large language models (LLMs) in diagnostic reasoning, management, and communication for secondary hypertension.
- To establish a performance hierarchy among LLMs for SH-related clinical tasks.
Main Methods:
- A blinded, cross-sectional study evaluated GPT-5.2, Claude Sonnet 4.6, and Gemini 3.0 Pro.
- Ten expert-developed SH clinical vignettes were used, with outputs assessed by senior clinicians on a 7-point Likert scale across five domains.
- Kruskal-Wallis tests and intraclass correlation coefficients were used for statistical analysis.
Main Results:
- Claude Sonnet 4.6 demonstrated superior performance across most domains, followed by GPT-5.2, then Gemini 3 Pro.
- Performance differences were most significant in complex clinical reasoning tasks.
- Efficiency of diagnostic workup showed comparable scores among the models.
Conclusions:
- LLM performance in secondary hypertension clinical tasks is heterogeneous and model-dependent.
- A clear clinician-rated performance hierarchy exists, with significant variations in complex reasoning capabilities.
- Further validation in larger studies is needed before routine clinical implementation.