Related Experiment Video
Updated: Aug 6, 2026

A Study on an Intelligent Diagnosis and Treatment Assistant System for Acupuncture in Diminished Ovarian Reserve Based on a Knowledge Graph
Published on: May 29, 2026
Guideline-anchored retrieval-augmented generation outperforms baseline and literature-only configurations in
Danielle Dukes1, Connor Yost2, Runzhi Wang1
1Creighton University School of Medicine Phoenix- Department of OBGYN, 500 West Thomas Road, Suite 660, Phoenix, AZ 85013, United States of America.
Introduction:
Large language models (LLMs) are being studied as oncology decision-support tools but can produce inaccurate outputs. We compared LLM performance in gynecologic oncology across three knowledge-integration configurations differing in retrieval strategy and underlying model, using the modified Generative Performance Score (mGPS) as the primary outcome.
Methods:
Fifty de-identified gynecologic oncology cases were submitted (October-November 2025) to three LLMs: baseline GPT-5, an NCCN-anchored GPT-5 retrieval-augmented generation (RAG) configuration, and OpenEvidence (a literature-anchored clinical AI without NCCN access at that time). Three gynecologic oncologists independently scored outputs using the mGPS (range - 1 to +1; Guideline Concordance plus Hallucination Penalty). Wilcoxon signed-rank tests and mixed-effects ordered logistic regression were used.
Results:
GPT-RAG produced the highest mGPS (0.83, SD 0.26), followed by OpenEvidence (0.70, SD 0.27) and baseline GPT-5 (0.65, SD 0.31). GPT-RAG exceeded baseline (W = 189.5, Z = -3.42, P < .001, r = 0.49) and OpenEvidence (W = 254.0, Z = -2.64, P = .008); OpenEvidence and baseline did not differ (P = .22). Mixed-effects modeling confirmed higher mGPS for GPT-RAG (OR 3.74; 95% CI, 1.57-8.90). Inter-rater agreement (ICC) was 0.49 for mGPS, 0.30 for Hallucination Penalty, and 0.70 for Readability and Rationality.
Conclusion:
NCCN-anchored RAG outperformed both baseline GPT-5 and a literature-anchored clinical AI without direct guideline access. OpenEvidence's subsequent NCCN integration (April 27, 2026) provides external validation of guideline anchoring's operational importance. Findings reflect benchmark performance, not clinical safety or improved patient outcomes.