Related Experiment Video
Updated: Jul 16, 2026

Methods for Acute and Subacute Murine Hindlimb Ischemia
Published on: June 21, 2016
The Illusion of Precision: Artificial Intelligence Cannot yet Reliably Predict Individual Outcomes in Infrapopliteal
Oliver Antonio Gómez Gutiérrez1, Franco Imbrognio2, Oscar A De la Torre1
1Vascular Medicine Centre, Tecnologico de Monterrey, School of Medicine and Health Sciences, Monterrey, Mexico.
Background:
The integration of large language models (LLMs) into vascular surgery promises scalable clinical decision support, yet their prognostic reliability remains unproven. This study evaluates the accuracy, calibration, and stability of generative artificial intelligence (AI) in predicting outcomes for chronic limb-threatening ischemia (CLTI).
Methods:
In a retrospective single-center study, we evaluated 49 patients undergoing revascularization for CLTI. Two LLM architectures (GPT-5.5 and DeepSeek-V4 [DS]) were tested using "naive" (unstructured) and "structured" (tabular) zero-shot prompting strategies to predict major adverse events (MAEs) at 6-month, 1-year, and 5-year horizons. Performance was assessed using Brier scores for calibration, sensitivity for discrimination, and Wasserstein distance for stability.
Results:
A critical dissociation between calibration and discrimination was observed. DS naive achieved a superior Brier score of 0.068 at 6 months, suggesting high probabilistic accuracy. However, this mathematical precision masked a failure in clinical utility: all models exhibited a sensitivity of 0.00 at 6-month and 1-year horizons. The models adopted an "actuarial hedging" strategy, clustering predictions around the population mean to minimize mathematical error rather than identifying specific high-risk patients. While internal reproducibility (intraclass correlation coefficient [ICCs] >0.500) and distributional stability across prompts were surprisingly high, this consistency proved deceptive, based entirely on nondiscriminatory safety biases.
Conclusion:
While LLMs demonstrate semantic competence in medical discourse, they currently lack the pragmatic utility required for high-stakes prognostication. The observed "hallucinated calibration" and deceptive stability creates a dangerous illusion of reliability. Until issues of discriminatory failure are resolved, LLMs should remain restricted to administrative rather than predictive roles in patient decision-making.
More Related Videos
08:16High-Resolution Three-Dimensional Imaging of the Footpad Vasculature in a Murine Hindlimb Gangrene Model
Published on: March 16, 2022
07:25Predicting Amputation using Local Circulating Mononuclear Progenitor Cells in Angioplasty-treated Patients with Critical Limb Ischemia
Published on: September 22, 2020
Related Concept Videos
Peripheral Arterial Disease II: Clinical Manifestations and Diagnostic Evaluation
Peripheral Artery Disease III: Interprofessional Care
Peripheral Artery Disease I: Introduction