Related Experiment Video
Updated: Aug 13, 2026

Introduction of an Integrated Pathology Image Management, Artificial Intelligence, and Reporting System
Published on: July 11, 2025
Improved accuracy and efficiency of guideline-based orthopedic disability assessment with a retrieval-augmented AI
Ismail Duran1, Zilan Karadağ1, Mehmet Selçuk Şenol1
1Department of Orthopedics and Traumatology, Istanbul Lutfi Kirdar City Hospital, Istanbul, Turkey.
Background:
/aim: Traditional orthopedic disability assessment relies on complex guidelines and is prone to evaluator error and inefficiency. This study evaluated whether a retrieval-augmented generation (RAG) AI assistant (NotebookLM) improves accuracy and efficiency in guideline-based impairment rating compared with manual consultation.
Materials And Methods:
In this prospective comparative study, 50 anonymized case-based clinical scenarios representing common musculoskeletal impairments were developed. Four orthopedic specialists evaluated all scenarios using two methods: manual guideline consultation and NotebookLM-assisted assessment restricted to the same regulation. The starting method was randomized, with each evaluator assessing 25 scenarios per method. The primary outcome was accuracy against an expert reference standard; the secondary outcome was time per scenario.
Results:
AI-assisted evaluation achieved higher accuracy than manual consultation (91% vs 68%; absolute difference, 23 percentage points, 95% CI 12.3 to 33.7; χ2(1) = 14.849, p < .001). Generalized estimating equations confirmed an independent association between the AI-assisted method and correct ratings (adjusted OR 4.76, 95% CI 3.24-6.98, p < .001). Mean evaluation time decreased from 131 to 37 seconds per scenario (mean difference, 94.36 seconds, 95% CI 74.38 to 114.34; p < .001).
Conclusion:
A RAG-based AI assistant improved both accuracy and efficiency in orthopedic impairment rating under controlled conditions. These findings suggest meaningful clinical utility in structured, guideline-bound workflows, although further external validation in routine clinical practice is necessary before broader adoption can be recommended.