Related Experiment Video
Updated: Aug 22, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large Language Models as Peer Reviewers: Prompt Sensitivity and Model-Dependent Reproducibility
Sukru Mehmet Erturk1, Mustafa Durmaz1
1Department of Radiology, Istanbul University, Istanbul Faculty of Medicine, Istanbul, Turkiye (S.M.E., M.D.).
Rationale And Objectives:
To evaluate the reproducibility of editorial re--ations by Large Language Models, agreement across different models, prompt-sensitivity, and fidelity of critique statements to source manuscripts in a simulated peer-review setting.
Materials And Methods:
Fifteen open-access radiology manuscripts were anonymized and reviewed by eight large language models (LLMs) across four developer families (ChatGPT, DeepSeek, Gemini, Grok) with two different prompts. Each manuscript-model-prompt condition was repeated across three independent runs, yielding 720 reviews. Intra-model stability was defined as identical decisions across runs. Inter-model agreement was assessed with Fleiss' kappa. A stratified random sample of 128 reviews underwent manual verification against the source manuscripts and was categorized as grounded, distorted, or hallucinated.
Results:
Across all reviews, decisions were Minor Revision in 51.3% (369 of 720), Major Revision in 43.9% (316 of 720), Accept in 4.9% (35 of 720), and Reject in 0% (0 of 720). Prompt strictness shifted decision severity (p < 0.001): Prompt 1 yielded 9.7% Accept, 64.7% Minor Revision, and 25.6% Major Revision, whereas Prompt 2 eliminated Accept and increased Major Revision to 62.2%. Inter-model agreement was fair (Fleiss' κ = 0.25). In the audit, 94.0% of statements were grounded and 6.0% were distorted, with no hallucinated statements observed.
Conclusion:
Large language model editorial re--ations were prompt sensitive and showed fair agreement across models despite critique statements that were largely grounded in manuscript text, supporting assistive use with human oversight.
Related Concept Videos
Stereotype Content Model
Improving Translational Accuracy
