Related Experiment Video
Updated: Feb 19, 2026

The WATCHMAN Left Atrial Appendage Closure Device for Atrial Fibrillation
Published on: February 28, 2012
Performance evaluation of 9 large language models in anticoagulation decision-making for nonvalvular atrial
Yujie Wen1, Zhenzhen Deng1, Xiaoyan Wang2
1Department of Pharmacy, The Third Xiangya Hospital, Central South University, Changsha, Hunan, China.
Background:
Large language models (LLMs) are increasingly applied in clinical decision support, but their reliability in guiding anticoagulation management for nonvalvular atrial fibrillation (NVAF) remains unclear.
Objectives:
To systematically evaluate the performance of 9 widely used LLMs in guiding anticoagulation therapy decisions for NVAF.
Methods:
A multidisciplinary team constructed 100 virtual NVAF patient cases with comprehensive demographic, clinical, laboratory, and medication information. Nine LLMs (DeepSeek-V3, DeepSeek-R1, Qwen3, Qwen3 Deep Think, ChatGPT-4o, ChatGPT-4o Deep Research, Gemini 2.5, Gemini 2.5 Deep Research, and Claude Sonnet 4) were prompted to provide individualized Congestive heart failure, Hypertension, Age ≥75 years, Diabetes mellitus, Stroke, Vascular disease, Age 65-74 years, Sex category (CHA2DS2-VASc) and Hypertension, Abnormal Renal/Liver Function, Stroke, Bleeding History or Predisposition, Labile INR, Elderly, Drugs/Alcohol Concomitantly (HAS-BLED) scores, anticoagulation contraindications, anticoagulation decisions, regimens, and monitoring recommendations. Outputs were independently evaluated against the multidisciplinary team consensus using accuracy rates and Likert-scale scores, with statistical comparisons performed across models.
Results:
Across 900 responses, performance varied substantially across 6 clinical tasks. ChatGPT-4o Deep Research achieved the highest accuracy in CHA2DS2-VASc scoring (100%) and medication monitoring (mean score, 4.98), ranking first overall. Claude Sonnet 4 demonstrated the best performance in anticoagulation regimen selection (89% accuracy), ranking second, while DeepSeek-V3 showed the lowest overall performance. Across all models, HAS-BLED scoring accuracy was consistently low (8%-51%). Notably, prompt optimization markedly improved HAS-BLED scoring accuracy, reaching 95% with ChatGPT-4o Deep Research and 89% with DeepSeek-R1 (P < .001).
Conclusion:
There was significant variability among LLMs in their decision support for NVAF anticoagulation. Prompt optimization has the potential to markedly enhance accuracy. At present, LLMs may serve as valuable adjuncts to clinical practice but remain unsuitable for use as stand-alone decision-makers.
Related Concept Videos
Anticoagulant Drugs: Vitamin K Antagonists and Direct Oral Anticoagulants
Warfarin, a prominent vitamin K antagonist family member, exerts its effect by inhibiting the enzyme VKORC1 (vitamin K epoxide reductase complex 1). By hindering this enzyme, warfarin...
Anticoagulant Drugs: Low-Molecular-Weight Heparins
Impact of Pharmacokinetic–Pharmacodynamic Models: Regulatory Decisions
Venous Thrombosis III: Interprofessional Care

