Related Experiment Video
Updated: Jul 4, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Physicians and artificial intelligence diverge in evaluating large language models on real clinical cases
Peilun Shi1, Jian Li2, Ziqi Yang3
1Department of Biomedical Engineering, The Chinese University of Hong Kong, Hong Kong SAR, China.
NPJ Digital Medicine
|July 2, 2026
Summary
Evaluating large language models (LLMs) in healthcare requires real clinical cases. Physician assessments showed variability, indicating AI agents can assist but not replace human clinical judgment for medical applications.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Medicine
- Clinical Evaluation Methodologies
Background:
- Multimodal large language models (LLMs) show promise in healthcare but lack robust clinical utility appraisal.
- Existing evaluations often use limited human expertise, narrow scopes, and synthetic data, potentially overestimating performance.
Purpose of the Study:
- To assess the clinical utility of LLM-generated responses using real patient cases and diverse physician evaluators.
- To compare human physician assessments with AI agent evaluations of LLM outputs.
Main Methods:
- A multicenter, multidisciplinary study involving over 400 physicians across seven specialties.
- Physicians evaluated LLM-generated free-text responses to de-identified clinical cases.
- A matched-control design used AI agents to mimic physician assessment characteristics.
Main Results:
- Physician assessments displayed significant heterogeneity based on clinical seniority and practice environment.
- AI agents provided efficient, aligned assessments but missed nuances of human clinical judgment.
- AI agents could not substitute for physician-centered evaluation.
Conclusions:
- Human physician assessment of LLM outputs in healthcare is complex and varies significantly.
- AI agents show potential as assistive tools for triaging or pre-screening LLM outputs to reduce physician workload.
- Further research is needed to refine AI tools for reliable clinical decision support.