General-purpose large language models outperform specialized clinical AI tools on medical benchmarks

Krithik Vishwanath1,2,3, Anton Alyakin4,5,6, Mrigayu Ghosh7,8

  • 1Department of Neurological Surgery, NYU Langone Health, New York, NY, USA. krithik.vish@utexas.edu.

Nature Medicine
|June 12, 2026
PubMed
Summary

Clinical artificial intelligence (AI) tools require rigorous evaluation before widespread adoption. Frontier large language models (LLMs) significantly outperform specialized clinical AI in medical knowledge and real-world query tasks.