Related Experiment Video
Updated: Sep 16, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
MedHopQA: a disease-centred multi-hop reasoning benchmark and evaluation framework for LLM-based biomedical question
Rezarta Islamaj1, Robert Leaman1, Joey Chan2
1Division of Intramural Research, National Library of Medicine, 8600 Rockville Pike, Bethesda, MD 20894, United States.
Abstract:
Evaluating large language models (LLMs) in the biomedical domain requires benchmarks that can distinguish genuine reasoning from pattern matching and that remain discriminative as model capabilities improve. Existing biomedical question answering (QA) benchmarks are limited in this respect: multiple-choice formats allow models to succeed by answer elimination rather than inference, and widely circulated exam-style datasets are subject to performance saturation and training data contamination. Multi-hop reasoning, i.e. the ability to integrate information across multiple sources to derive an answer, is central to clinically meaningful tasks such as diagnosis support, literature-based discovery, and hypothesis generation, yet remains underrepresented in current biomedical QA benchmarks. We present MedHopQA, a disease-centred multi-hop reasoning benchmark of 1000 expert-curated question-answer pairs introduced as a shared task at BioCreative IX. Each question requires synthesis of information across two distinct Wikipedia articles, and answers are provided in open-ended free-text format rather than as multiple-choice selections. Gold annotations are augmented with ontology-grounded synonym sets (MONDO, NCBI Gene, and NCBI Taxonomy) to support both lexical and concept-level evaluation. The dataset was constructed through a multi-stage human-AI pipeline combining structured human annotation, triage, iterative verification, and LLM-as-a-judge validation. To reduce leaderboard gaming and contamination risk, the 1000 scored questions are embedded within a publicly downloadable set of 10 000 questions, with answers withheld, on a CodaBench leaderboard. Evaluation of four frontier LLMs under a zero-shot setting (GPT-5.1, Gemini 2.5 Pro, Claude Sonnet 4.5, and GPT-4o) reveals performance variation across answer types, with overall accuracy ranging from 66.3% to 83.4%. Performance is strongest on chemical and anatomical questions and most variable on disease and gene/protein categories, where fine-grained semantic discrimination is required. MedHopQA provides both a benchmark and a reusable framework for constructing future biomedical QA datasets that prioritize compositional reasoning, saturation resistance, and contamination mitigation as design constraints.
Related Concept Videos
Investigation of Disease Outbreaks
Issues And Trends In Healthcare Delivery System
Cost Containment
Payment for healthcare services has historically promoted adoption of costly and often unnecessary or inefficient...

