Related Experiment Video
Updated: Jun 13, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Reasoning or reciting? A temporal contamination audit of large language models in clinical medicine
Alexander P Sheppert1,2,3, Brian Adams4,5, Andrew D Sheppert1,2
1Internal Medicine, Legacy Health System, Portland, OR 98686, United States.
Objective:
Evaluate whether large language models reason or simply regurgitate training data in clinical diagnosis.
Materials And Methods:
We audited 2000 clinical case reports from PubMed Central: 1000 from 2021 to 2022 (within training data) and 1000 from 2025 (after training cutoffs). Five frontier LLMs generated diagnoses evaluated by an independent AI judge validated against physician consensus (n = 10 000 evaluations).
Results:
Diagnostic accuracy was virtually identical across temporal cohorts (66.8% contaminated vs 66.9% clean), directly contradicting the memorization hypothesis. Lexical similarity was uniformly low (mean ROUGE-L 0.057), and semantic similarity measured by BERTScore showed no memorization signal (F1 0.8182 contaminated vs 0.8195 clean, Δ = +0.0013), confirming that models generate novel reasoning rather than regurgitating training data.
Discussion:
This large-scale audit, using both lexical and semantic similarity metrics, provides compelling evidence that LLMs engage in genuine clinical reasoning rather than regurgitating memorized training data.
Conclusion:
Models demonstrated equivalent accuracy on cases they could not have seen during training, suggesting they have internalized generalizable medical knowledge rather than memorizing specific cases.
Related Concept Videos
Introduction to Language of Pathophysiology ll
Introduction to Language of Pathophysiology l
Language and Cognition
