Related Experiment Videos
Are artificial intelligence-generated recommendations clinically safe? Evaluation of thromboembolic prophylaxis
Ömer Polat1, Alper Dünki1, Mehmet Talha Aydin1
1University of Health Sciences, Umraniye Education and Research Hospital, Department of Orthopaedics and Traumatology - Istanbul, Türkiye.
Objective:
Venous thromboembolism prophylaxis remains a complex and controversial issue in orthopedic practice, with persistent variability in recommendations even among experts. With the growing use of online platforms and artificial intelligence, the aim of this study was to evaluate the quality and clinical adequacy of artificial intelligence-generated recommendations regarding venous thromboembolism.
Methods:
Thirty-three guideline-based questions derived from International Consensus Meeting, National Institute for Health and Care Excellence, American College of Chest Physicians, and American Academy of Orthopaedic Surgeons recommendations were independently submitted to ChatGPT (Generative Pre-trained Transformer [GPT]-4.0) by two orthopedic surgeons, resulting in 66 generated responses for analysis. Responses were assessed using the DISCERN instrument, Journal of the American Medical Association benchmark criteria, and a novel clinically oriented scale, the Thromboembolism Digital Knowledge Level. Inter-rater reliability was evaluated using intraclass correlation coefficients.
Results:
According to DISCERN, 62.1% of responses were classified as good quality, whereas Thromboembolism Digital Knowledge Level evaluation demonstrated that only 45.45% of responses were clinically adequate, with no responses classified as excellent. The same responses were systematically categorized into higher quality levels by DISCERN compared to Thromboembolism Digital Knowledge Level. All scoring systems demonstrated excellent inter-rater reliability (Intraclass Correlation Coefficient range: 0.897-0.961). Although differences between scoring systems were statistically significant, the effect size was moderate (r=0.39). Strong correlations were observed between evaluators across all scales (p<0.001).
Conclusion:
ChatGPT-generated responses demonstrated good structural quality but limited clinical adequacy for thromboembolic prophylaxis after arthroplasty. These findings suggest that artificial intelligence systems should be used cautiously and may not be sufficient as standalone clinical decision-support tools.
Related Concept Videos
Venous Thrombosis III: Interprofessional Care
Peripheral Artery Disease III: Interprofessional Care
Pulmonary Embolism II: Diagnostic Studies and Interprofessional Care
Peripheral Artery Disease V: Postoperative Nursing Management