Related Experiment Video
Updated: Apr 29, 2026

07:14
Virtual Agent for Real-Time Motivational Interviewing by Integrating Adaptive Nonverbal Behavior and Language Models
Published on: December 23, 2025
1.1K
AgentClinic: a multimodal benchmark for tool-using clinical AI agents
Samuel Schmidgall1, Rojin Ziaei2, Carl Harris3
1Department of Electrical and Computer Engineering, Johns Hopkins University, Baltimore, MD, USA. sschmi46@jhu.edu.
NPJ Digital Medicine
|April 27, 2026
Summary
AgentClinic, a new benchmark for evaluating large language models (LLMs) in clinical settings, reveals significant challenges in sequential decision-making. Claude-3.5 agents generally outperform others, but tool utilization varies greatly among LLMs.
Area of Science:
- Artificial Intelligence in Medicine
- Clinical Decision Support Systems
- Natural Language Processing
Background:
- Current benchmarks for large language models (LLMs) in healthcare often use static question-answering formats.
- These static formats fail to capture the dynamic, sequential nature of clinical decision-making.
- There is a need for more realistic evaluations of LLM clinical utility.
Purpose of the Study:
- To introduce AgentClinic, a novel multimodal agent benchmark for assessing LLMs in simulated clinical environments.
- To evaluate LLM performance in complex clinical scenarios involving patient interaction and tool usage.
- To compare the capabilities of different LLM backbones in a clinical context.
Main Methods:
- Development of AgentClinic, a benchmark featuring simulated patient interactions, multimodal data, and tool integration.
- Evaluation of LLMs across nine medical specialties and seven languages.
- Assessment of diagnostic accuracy in sequential decision-making tasks.
- Analysis of LLM tool utilization, including note-taking and retrieval.
Main Results:
- Solving clinical problems in AgentClinic's sequential format significantly reduces diagnostic accuracy compared to static benchmarks.
- Claude-3.5 agents demonstrate superior performance across most evaluated settings.
- LLMs exhibit substantial variability in their ability to effectively utilize tools like experiential learning and reflection cycles.
- Llama-3 showed notable improvement (up to 92%) with a persistent notebook tool.
Conclusions:
- AgentClinic provides a more challenging and realistic evaluation of LLMs for clinical applications.
- LLM performance in clinical settings is highly dependent on the benchmark design and the model's ability to integrate tools.
- Further research is needed to optimize LLMs for complex clinical workflows and patient-centric outcomes.
