Related Experiment Video
Updated: Feb 20, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Benchmarking large language model-based agent systems for clinical decision tasks
Yunsong Liu1,2, Zunamys I Carrero2, Xiaofeng Jiang2,3
1Department of Radiation Oncology, National Cancer Center/National Clinical Research Center for Cancer/Cancer Hospital, Chinese Academy of Medical Sciences and Peking Union Medical College, Beijing, China.
Agentic artificial intelligence (AI) systems show limited performance gains in healthcare despite advanced tools. Current systems offer modest benefits with high computational costs, highlighting the need for improved AI solutions.
Area of Science:
- Artificial Intelligence
- Medical Informatics
- Computational Medicine
Background:
- Agentic AI systems, capable of autonomous reasoning and tool use, show potential in healthcare applications.
- Systematic real-world performance evaluation of these advanced AI systems in medicine is currently limited.
- Existing benchmarks do not fully capture the complexities of clinical decision-making and tool integration.
Purpose of the Study:
- To systematically benchmark the real-world performance of two agentic AI systems in healthcare settings.
- To evaluate the efficacy of agentic AI across diverse medical tasks, including diagnostics, QA, and complex examinations.
- To assess the trade-offs between performance gains, resource utilization, and hallucination rates in medical AI agents.
Main Methods:
- Evaluated OpenManus (Llama-4 based) and Manus (proprietary multistep architecture) on AgentClinic, MedAgentsBench, and Humanity's Last Exam (HLE) benchmarks.
- Assessed performance on text-based and multimodal medical question-answering and diagnostic simulations.
- Quantified accuracy, token usage, latency, and hallucination rates, with in-agent safeguards.
Main Results:
- Agentic AI systems provided modest accuracy improvements over baseline LLMs, with significant increases in token usage and latency.
- Accuracy on AgentClinic MedQA reached 60.3%, MedAgentsBench 30.3%, and HLE text 8.6%.
- Multimodal accuracy was low (15.5% on HLE, 29.2% on AgentClinic NEJM), and hallucinations persisted despite safeguards.
Conclusions:
- Current agentic AI designs offer limited performance benefits in healthcare relative to their substantial computational and workflow costs.
- There is a critical need for the development of more accurate, efficient, and clinically viable agent systems for medical applications.
- Further research is required to optimize agentic AI architectures for practical healthcare deployment.
Related Concept Videos
Methods of Documentation VI: Case Management Model
For example, a patient with a chronic...
SBAR II: Application of SBAR
SBAR Report from a Nurse to a Health Care Provider
S: "Hello, Dr. Smith. This is Jane, RN, from the Med Surg unit. I am calling to tell you about Ms. White in Room 210, who is experiencing increased pain and redness at her incision site. Her recent...
Decision Making: Traditional Method
First, a specific claim about the population parameter is decided based on the research question and is stated in a simple form. Further, an opposing statement to this claim is also stated. These statements can act as null and alternative hypotheses, out of which a null hypothesis would be a...
Pharmacokinetic Models: Comparison and Selection Criterion
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.
