Related Experiment Video
Updated: Sep 9, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Rapidly Benchmarking Large Language Models for Diagnosing Comorbid Patients: Comparative Study Leveraging the
Peter Sarvari1, Zaid Al-Fagih1
1Rhazes AI, First Floor, 85 Great Portland Street, London, W1W 7LT, United Kingdom.
Gemini 2.5 demonstrated superior diagnostic accuracy among 21 large language models (LLMs) evaluated on real patient data. This study highlights LLMs' potential in improving medical diagnostics, though further research is needed.
Area of Science:
- Artificial Intelligence in Medicine
- Clinical Decision Support Systems
- Natural Language Processing for Healthcare
Background:
- Diagnostic errors contribute to significant patient mortality, representing the third leading cause of death in the United States.
- Large Language Models (LLMs) show promise for assisting clinicians with diagnoses, but comparative performance data on real-world patient cohorts is lacking.
Purpose of the Study:
- To compare the diagnostic capabilities of 18 popular Large Language Models (LLMs) using a large, real-world patient dataset.
- To evaluate the impact of different prompts and temperature settings on LLM diagnostic performance.
- To assess the effectiveness of retrieval-augmented generation (RAG) in enhancing LLM diagnostic accuracy.
Main Methods:
- Evaluated 21 LLMs on 1000 randomly selected Medical Information Mart for Intensive Care-IV (MIMIC-IV) hospital admissions.
- Employed an LLM-as-a-judge approach for automated evaluation, comparing LLM-generated diagnoses against final diagnostic codes from patient records.
- Calculated diagnostic hit rates and assessed statistical significance using pooled z-tests for proportions.
Main Results:
- Gemini 2.5 achieved the highest diagnostic hit rate (97.4%) when assessed by GPT-4.1, outperforming other leading models like GPT-4.1 and Claude-4 Opus.
- GPT-4.1 showed the highest performance in separate evaluations by GPT-4 Turbo, indicating variability in assessment depending on the judge LLM.
- Retrieval-augmented generation (RAG) significantly improved GPT-4o 05-13's hit rate by 0.8% (P<.006), and performance varied across different prompts.
Conclusions:
- LLMs demonstrate significant potential for improving diagnostic accuracy in clinical settings.
- Further research with diverse datasets and clinical validation is essential to fully understand and implement LLM diagnostic tools.
- Close collaboration between AI developers and physicians is crucial for the responsible integration of LLMs into healthcare.
More Related Videos
06:48Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
Published on: June 25, 2019
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Related Concept Videos
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
On...
Mechanistic Models: Compartment Models in Individual and Population Analysis