Related Experiment Video
Updated: Jun 14, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
General-purpose large language models outperform specialized clinical AI tools on medical benchmarks
Krithik Vishwanath1,2,3, Anton Alyakin4,5,6, Mrigayu Ghosh7,8
1Department of Neurological Surgery, NYU Langone Health, New York, NY, USA. krithik.vish@utexas.edu.
Nature Medicine
|June 12, 2026
Summary
Clinical artificial intelligence (AI) tools require rigorous evaluation before widespread adoption. Frontier large language models (LLMs) significantly outperform specialized clinical AI in medical knowledge and real-world query tasks.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Healthcare
- Clinical Decision Support
Background:
- Specialized clinical artificial intelligence (AI) tools are increasingly integrated into medical practice.
- There is a notable lack of independent, quantitative evaluations for these AI tools.
- Ensuring the reliability and efficacy of AI in clinical settings is paramount.
Purpose of the Study:
- To quantitatively evaluate two clinical AI tools (OpenEvidence, UpToDate Expert AI) against leading frontier large language models (LLMs).
- To assess AI performance across medical knowledge, clinician alignment, and real-world clinical query scenarios.
- To provide evidence-based insights into the comparative effectiveness of clinical AI versus general-purpose LLMs.
Main Methods:
- A three-stage evaluation framework was employed, including medical knowledge testing (MedQA), clinician alignment assessment (HealthBench), and a real clinical queries (RCQ) benchmark.
- The RCQ benchmark utilized 100 de-identified physician queries, with outputs from six AI models reviewed by 12 US clinicians.
- A total of 1,800 model-question annotations were collected for the RCQ benchmark analysis.
Main Results:
- Frontier LLMs (GPT-5.2, Gemini 3.1 Pro, Claude Opus 4.6) demonstrated superior performance compared to specialized clinical AI tools across all three evaluation stages.
- Clinical AI tools exhibited performance comparable to Google Search AI Overview on the real clinical queries benchmark.
- The study identified significant performance gaps between advanced general-purpose LLMs and current clinical AI applications.
Conclusions:
- Independent, real-world evaluations are crucial for validating clinical AI tools prior to their deployment in healthcare.
- Frontier LLMs show potential for surpassing existing specialized clinical AI in handling complex medical queries.
- The findings underscore the need for enhanced validation methodologies for AI in medicine.
