Related Experiment Video
Updated: Feb 21, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Agentic memory-augmented retrieval and evidence grounding for medical question-answering tasks
Shuyue Jia1, Subhrangshu Bit2, Varuna H Jasodanand3
1Department of Electrical and Computer Engineering, Boston University, Boston, MA 02215, USA.
A new tool-using agent system built on large language models (LLMs) shows improved performance on medical question-answering tasks. This AI system enhances medical reasoning by integrating tools and evidence, outperforming standalone LLMs on key exams.
Area of Science:
- Artificial Intelligence
- Medical Informatics
- Natural Language Processing
Background:
- Large language models (LLMs) show promise in medical question-answering.
- Standalone LLMs often struggle with complex, multi-step medical reasoning.
- The integration of external tools and evidence grounding is crucial for enhancing AI reliability in medicine.
Purpose of the Study:
- To evaluate if a tool-using agent-based system built on LLMs outperforms standalone LLMs in medical question-answering.
- To develop and assess an open-source LLM-based agentic system for dynamic, multi-step medical reasoning.
Main Methods:
- Developed a unified, open-source LLM-based agentic system with document retrieval, reranking, evidence grounding, and diagnosis generation.
- Implemented a retrieval-augmented generation pipeline and a cache-and-prune memory bank for efficient long-context inference.
- The system autonomously invoked specialized tools, bypassing manual prompt engineering.
Main Results:
- The agentic system achieved high accuracies on medical question-answering benchmarks: 82.98% on USMLE Step 1 and 86.24% on Step 2.
- Performance surpassed GPT-4 on USMLE Steps 1 and 2, while closely matching on Step 3.
- The system demonstrated superior or comparable performance to state-of-the-art models in both multiple-choice and open-ended formats.
Conclusions:
- Combining tool-augmented and evidence-grounded reasoning strategies is valuable for developing reliable and scalable medical AI systems.
- Agent-based LLM systems offer a promising approach to enhance AI capabilities in complex medical reasoning tasks.
- The developed system provides an open-source framework for advancing AI in medical question answering.
Related Concept Videos
The Availability Heuristic
Retrieval
Recall involves accessing information without cues, such as during an essay test, where individuals must retrieve facts and concepts from memory unaided. Another example is remembering the name of a colleague...
Higher Mental Functions of Brain: Learning and Memory
Eyewitness Memory
One such error is memory distortion, which occurs because human memory does not function...
Explicit Memories
Episodic memory contains information about personally experienced events and is reported as a story. An example of episodic memory is recalling a birthday celebration. This type of memory includes the what, where, and when of an event, as...
Role of Amygdala in Memory
One of the...

