Related Experiment Videos
Guideline-augmented prompting improves comparative preference and response consistency of large language model
Anita Széll1,2, Yinan Yu3, Jacob F Oeding4
1Department of Anesthesiology and Intensive Care Medicine, Institute of Clinical Sciences Sahlgrenska Academy Gothenburg Sweden.
Purpose:
Large language models (LLMs) are increasingly used in clinical contexts; however, performance in complex perioperative decision-making remains uncertain. Orthopaedic anaesthesia presents a demanding test case due to comorbidity burden and guideline-dependent management. Whether successive LLM generations and guideline-augmented prompting improve clinical alignment, comparative performance and response consistency was evaluated in this study.
Methods:
In this controlled prompting study, 34 orthopaedic anaesthesia questions spanning seven clinical subdomains were independently answered by three LLMs (GPT-3.5-turbo, GPT-4o and GPT-5.2) and three expert anesthesiologists. GPT-4o and GPT-5.2 each generated responses with and without access to relevant clinical practice guidelines. All responses were anonymized and evaluated in blinded pairwise comparisons by an independent guideline-informed LLM-as-a-judge (GPT-5.2) using the same guideline material as the reference standard. The judge recorded preference, response consistency and confidence. Intra-model consistency was assessed from repeated independent generations of each question.
Results:
All LLMs were preferred over human experts in pairwise comparisons without guideline augmentation, with win rates of 67.6% (GPT-3.5-turbo), 79.4% (GPT-4o) and 95.1% (GPT-5.2). Performance was improved by guideline augmentation, most notably for GPT-4o (91.2%, +11.8 percentage points), while GPT-5.2 approached ceiling performance (97.1%). Response consistency varied between models. GPT-4o showed the highest baseline consistency (80.4%), whereas GPT-5.2 demonstrated no fully contradictory outputs but showed greater partial variability. Guideline augmentation numerically improved GPT-5.2 consistency (66.7% to 80.4%) and reduced inter-model differences. Directed qualitative analysis suggested that reviewer preference for later GPT generations was associated with greater completeness, explicit clinical reasoning and guideline-oriented responses.
Conclusion:
Successive LLM generations demonstrated progressively improved performance in orthopaedic anaesthesia reasoning. Guideline-augmented prompting enhanced response quality, particularly for intermediate models. Guideline-informed LLM-as-a-judge evaluation appears promising for comparative assessment but requires further validation.
Level Of Evidence:
NA.