MedHopQA: A Disease-Centered Multi-Hop Reasoning Benchmark and Evaluation Framework for LLM-Based Biomedical Question

Rezarta Islamaj1, Robert Leaman1, Joey Chan2

  • 1National Library of Medicine, Division of Intramural Research, Bethesda, MD, US.

Arxiv
|July 2, 2026
PubMed
Summary

MedHopQA is a new benchmark for evaluating large language models (LLMs) in biomedicine. It tests multi-hop reasoning, crucial for clinical tasks, using expert-curated, open-ended questions to resist pattern matching.