Related Experiment Video
Updated: Apr 17, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Multidisciplinary blinded randomized expert evaluation of large language models for clinical diagnosis and management
Peikai Chen1,2,3,4, Jifu Cai5,6, Jiaying Zhou5,7
1Department Orthopedics, The University of Hong Kong - Shenzhen Hospital, Shenzhen, China. pkchen@hku-szh.org.
Background:
Direct clinical uses of large language models (LLMs) remain controversial, partly because of the lack of methodological rigor in assessing their risks and benefits in medicine.
Methods:
We developed Medieval, a multidisciplinary, randomized, and blinded expert evaluation framework. A ten-point Dreyfus-based scoring scale linked to career stages of human physicians was designed to reflect response qualities. Seven advanced LLMs or their distilled versions that were released within a short time-frame ( ≤ 45 days) in early 2025 were tested. Incidence of fabricated medical facts were documented. Linear mixed-effects models and variance-stabilizing Bayesian generalized linear mixed models were employed to perform statistical analyses.
Results:
We first develop a high-quality question bank comprising 685 real and simulated clinical cases across 13 specialties. An expert panel of 27 clinicians (average years of services: 25.9) evaluated the 4795 model responses. We show that these LLM ratings (n = 9856) have excellent reliability (intraclass correlation coefficients 0.9). Among the seven LLMs tested, Gemini 2.0 Flash achieved the highest raw scores. However, after adjusting for confounders, DeepSeek-R1 was the top-performing model with a mean score of 6.36 (95% confidence interval 6.03 - 6.69), a performance level equivalent to an early-career physician. Despite these strengths, 3-19% LLM responses were rated as incompetent and 40 instances of LLM hallucination were also identified.
Conclusions:
Our study shows that in spite of LLMs' substantial potentials in medicine, their unguarded clinical application could present serious risks, which must be continuously monitored by human expert panels. The evaluation framework developed and validated in this study will facilitate such efforts.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
06:48Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
Published on: June 25, 2019