Related Experiment Video
Updated: Jan 9, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Language Models are Surprisingly Fragile to Drug Names in Biomedical Benchmarks
Jack Gallifant1, Shan Chen2,3,4, Pedro Moreira1,5
1MIT.
Abstract:
Medical knowledge is context-dependent and requires consistent reasoning across various natural language expressions of semantically equivalent phrases. This is particularly crucial for drug names, where patients often use brand names like Advil or Tylenol instead of their generic equivalents. To study this, we create a new robustness dataset, RABBITS, to evaluate performance differences on medical benchmarks after swapping brand and generic drug names using physician expert annotations. We assess both open-source and API-based LLMs on MedQA and MedMCQA, revealing a consistent performance drop ranging from 1-10%. Furthermore, we identify a potential source of this fragility as the contamination of test data in widely used pre-training datasets.
Related Concept Videos
Drug Nomenclature
Pharmacokinetic Models: Overview
There are three primary types of models: empirical, compartment, and physiological. Empirical models, with minimal...
Drug Discovery: Overview
Drug Biotransformation: Overview
Drug Regulation
Prescription, Nonprescription and Orphan Drugs
The misuse and addiction to prescription drugs is a growing problem that can affect people of all age groups, specifically teenagers. This can happen when prescription medications are used in ways not intended by the prescriber, such as taking someone else's prescription or using medication for...

