Related Experiment Video
Updated: Sep 30, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating reasoning-tuned large language models for clinical decision-making in spine surgery
Chuck H Lam1, Conor T Boylan2, Adit Ravishankar3
1Cedars-Sinai Spine, Cedars-Sinai Medical Center, 8700 Beverly Blvd, Los Angeles, CA, 90048, USA. chuck.lam@cshs.org.
Purpose:
Most clinical evaluations of large language models assess factual recall rather than the multi-step reasoning behind operative plans. Reasoning-tuned models, post-trained to generate explicit intermediate reasoning before answering, may better approximate surgical decision-making. We compared two such models from different development ecosystems, OpenAI o1 and DeepSeek R1, on complex spine cases.
Methods:
Ten synthetic vignettes spanning deformity, degenerative, urgent, and cervical pathology were analysed; an eleventh was excluded owing to a survey error. Both models received identical prompts. Outputs were anonymised and randomised. Eight fellowship-trained spine surgeons (five consultants, three fellows) scored diagnostic accuracy, reasoning and thoroughness, surgical plan appropriateness, and clarity on 5-point Likert scales, giving 78 paired evaluations. Analysis used linear mixed-effects models with crossed random intercepts for rater and vignette (Bonferroni-adjusted alpha 0.0125); reliability was assessed by intraclass correlation.
Results:
o1 scored higher in all four domains, significantly for diagnostic accuracy (4.6 vs. 4.3; mean difference 0.24, 95% CI 0.11-0.38) and clarity (4.4 vs. 4.2; 0.26, 0.07-0.44); reasoning and thoroughness (0.21, 0.03-0.38), and surgical plan appropriateness (0.17, - 0.03 to 0.36) did not meet the adjusted threshold. Inter-rater reliability was poor (ICC 0.03-0.07). Clarity correlated strongly with the clinical domains (rho 0.74-0.79); adjusting for clarity attenuated the o1 advantage by 55% to over 100%, leaving no detectable clinical-domain difference.
Conclusion:
Both models produced plans surgeons usually rated good or excellent. o1 was rated higher, significantly only for diagnostic accuracy and clarity, most of it reflecting presentation rather than clinical content. Surgeon oversight and format-normalised evaluation designs remain essential.