Related Experiment Video
Updated: Aug 5, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Evidence Use and Identifier-Conditioned Prior Knowledge in Large Language Model Classification of Oncology Trials
Paul Windisch1,2, Carole Koechli2, Fabio Dennstädt2
1Department of Radiation Oncology, Kantonsspital Winterthur, Brauerstrasse 15, Winterthur, Zurich, 8401, Switzerland, 41 052 266 26 53.
Large language models (LLMs) accurately classify oncology trial outcomes based on abstract evidence. However, identifiers like DOIs can introduce bias, requiring careful evaluation to ensure predictions are text-grounded.
Area of Science:
- Biomedical Informatics
- Artificial Intelligence in Medicine
- Clinical Trial Analysis
Background:
- Large language models (LLMs) demonstrate proficiency in classifying biomedical documents.
- Benchmark performance does not guarantee predictions are text-grounded, as identifiers may trigger prior knowledge.
- Biomedical literature tasks can be influenced by pre-trained knowledge from titles, abstracts, DOIs, and trial identifiers.
Purpose of the Study:
- To determine if oncology randomized trial success classification relies on abstract evidence or identifier-conditioned prior knowledge.
- To assess LLM adherence to counterfactual outcome evidence when it conflicts with original trial identifiers.
Main Methods:
- Evaluated 250 oncology randomized controlled trials (RCTs) with adjudicated outcomes.
- Tested GPT-5.2, Gemini 3 Flash, and Claude Opus 4.5 using title+abstract, title only, and DOI only inputs.
- Introduced counterfactual title+abstract inputs with flipped outcomes and paired them with original DOIs to create identifier-text conflicts.
Main Results:
- All models achieved high accuracy (0.96-0.97) with title+abstract inputs.
- Performance decreased with evidence removal: title-only (0.79-0.88) and DOI-only (0.63-0.67).
- Models followed counterfactual evidence accurately (0.96-0.99), with minor performance drops when original DOIs were reintroduced.
Conclusions:
- LLMs reliably follow explicit endpoint statements, even when contradicting trial identifiers.
- Identifiers (DOIs, titles) provide a predictive signal that can compete with textual evidence.
- A reproducible audit using content removal and counterfactual conflicts can assess LLM grounding in biomedical evaluations.
Related Concept Videos
Cancer Survival Analysis
Combination Therapies and Personalized Medicine
The combination of the drug acetazolamide and sulforaphane is a good example of combination therapy to treat cancer. The cells in the interior of a large tumor often die due to the hypoxic and...
Clinical Trials: Overview
Targeted Cancer Therapies
There are several types of targeted therapies against specific...
Clinical Trials
There are four phases in a clinical trial. A phase one...
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe and...