Related Experiment Video
Updated: Oct 10, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Ensemble multi-retrieval methodologies with reasoning model decision in multi-step biomedical QA
Jongmyung Jung1, Hyeongsoon Hwang1,2, Yein Park1,2
1Department of Computer Science and Engineering, Korea University, 145 Anam-ro, Seongbuk-gu, Seoul, 02841, Republic of Korea.
Abstract:
Robust and trustworthy biomedical question answering (QA) remains a critical challenge for large language models (LLMs), especially in complex domains such as rare diseases, where information is fragmented across multiple sources. MedHopQA 2025 benchmark introduces 10 000 multi-step reasoning questions curated from Wikipedia, requiring systems to extract, connect, and synthesize biomedical knowledge across interlinked documents. In this work, we present a retrieval-augmented generation (RAG) and decision-making framework that integrates diverse retrieval strategies, including Query2Doc-based, Rationale-based, and Web-augmented retrieval, and employs a dedicated Decision Maker model to select or directly generate the most accurate and well-reasoned answers. Our system not only leverages evidence from both Wikipedia and the web but also explicitly evaluates and compares candidate answers to ensure answer reliability and reasoning transparency. Experiments on the test set demonstrate that our ensemble approach achieves state-of-the-art performance, yielding an Exact Match score of 0.84, and highlighting the importance of hybrid retrieval and robust decision-making in advancing biomedical multi-step QA. Furthermore, comprehensive analyses across retrieval stages reveal that mitigating model self-bias and ensuring a balanced integration of diverse sources are crucial for maximizing overall system performance, which ultimately reaches a score of 0.87.