Related Experiment Video
Updated: Jan 13, 2026

05:47
Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
1.3K
Evaluating Web Retrieval-Assisted Large Language Models With and Without Whitelisting for Evidence-Based Neurology:
Lars Masanneck1, Paula Zoe Epping1, Sven G Meuth1
1Department of Neurology, Medical Faculty and University Hospital Düsseldorf, Heinrich Heine University Düsseldorf, Dusseldorf, Germany.
Journal of Medical Internet Research
|October 29, 2025
Summary
Restricting large language models (LLMs) to authoritative sources significantly improved accuracy in retrieving medical evidence. This targeted approach enhances reliability for evidence-based care, making LLMs more trustworthy decision-support tools.
Area of Science:
- Artificial Intelligence in Medicine
- Natural Language Processing
- Biomedical Informatics
Background:
- Large language models (LLMs) with web retrieval are becoming primary information sources for medical evidence.
- Open-web retrieval risks factual errors and hallucinations from nonprofessional sources, potentially harming evidence-based care.
- Ensuring the accuracy of LLM-generated medical information is crucial for clinical decision-making.
Purpose of the Study:
- To evaluate the impact of restricting retrieval to guideline domains (whitelisting) on the answer quality of publicly available LLMs.
- To compare the performance of whitelisted LLMs against a specialized biomedical literature retrieval system (OpenEvidence).
Main Methods:
- A 130-item question set from American Academy of Neurology (AAN) guidelines was used.
- Three Perplexity LLMs were queried with both open-web and whitelisted retrieval (AAN/neurology domains).
- Responses were evaluated by blinded neurologists for accuracy, with disagreements resolved by a third expert.
Main Results:
- Whitelisting improved correct-answer rates by 8-18 percentage points across Perplexity models.
- LLMs using only AAN/neurology sources had double the odds of higher accuracy compared to those including nonprofessional sources.
- Factual questions were answered more accurately than case-based questions on Perplexity models.
Conclusions:
- Restricting LLM retrieval to authoritative domains significantly enhances answer correctness and reduces variability.
- This source control transforms consumer search assistants into reliable decision-support tools comparable to specialized systems.
- Lightweight source control is a vital safety mechanism for web-based LLMs in evidence-based neurology.
Keywords:
artificial intelligenceevidence-based medicineinformation retrievallarge language modelsmedical guidelinesneurology
