Related Experiment Video
Updated: Jul 3, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
MedHopQA: A Disease-Centered Multi-Hop Reasoning Benchmark and Evaluation Framework for LLM-Based Biomedical Question
Rezarta Islamaj1, Robert Leaman1, Joey Chan2
1National Library of Medicine, Division of Intramural Research, Bethesda, MD, US.
MedHopQA is a new benchmark for evaluating large language models (LLMs) in biomedicine. It tests multi-hop reasoning, crucial for clinical tasks, using expert-curated, open-ended questions to resist pattern matching.
Area of Science:
- Biomedical Informatics
- Artificial Intelligence
- Natural Language Processing
Background:
- Evaluating large language models (LLMs) in biomedicine is challenging.
- Existing benchmarks struggle to differentiate reasoning from pattern matching and are prone to data contamination.
- Multi-hop reasoning is vital for clinical applications but underrepresented in current benchmarks.
Purpose of the Study:
- Introduce MedHopQA, a novel benchmark for assessing LLM multi-hop reasoning in the biomedical domain.
- Provide a robust evaluation framework resistant to common benchmark limitations like answer elimination and data contamination.
- Facilitate the development of LLMs with superior reasoning capabilities for clinical support and discovery.
Main Methods:
- Developed MedHopQA, a disease-centered benchmark with 1,000 expert-curated question-answer pairs.
- Questions require synthesizing information from two distinct Wikipedia articles.
- Answers are open-ended free-text, augmented with ontology-grounded synonyms (MONDO, NCBI Gene, NCBI Taxonomy).
- Constructed via human annotation, triage, verification, and LLM-as-a-judge validation.
- Embedded scored questions within a larger set to mitigate leaderboard gaming and contamination.
Main Results:
- MedHopQA comprises 1,000 expert-curated, disease-centered questions.
- The benchmark necessitates multi-hop reasoning, integrating information from multiple sources.
- Open-ended answers and ontology-grounded annotations support comprehensive evaluation.
- A structured validation process ensured quality and reduced contamination risk.
Conclusions:
- MedHopQA offers a robust benchmark for evaluating biomedical LLMs, focusing on multi-hop reasoning.
- The benchmark design addresses limitations of existing QA datasets, promoting saturation and contamination resistance.
- MedHopQA serves as a foundation for future biomedical QA dataset development, prioritizing compositional reasoning.
Related Concept Videos
Issues And Trends In Healthcare Delivery System
Cost Containment
Payment for healthcare services has historically promoted adoption of costly and often unnecessary or inefficient...
Investigation of Disease Outbreaks
Health Information Technology and Healthcare Information System
Health Information Technology, commonly called HIT, integrates advanced information systems and technology in healthcare settings. Its primary functions include:
