Related Experiment Video
Updated: Jun 7, 2025

05:57
A Teleoperated Robotic System-Assisted Percutaneous Transiliac-Transsacral Screw Fixation Technique
Published on: January 6, 2023
2.2K
How soon will surgeons become mere technicians? Chatbot performance in managing clinical scenarios
Darren S Bryan1, Joseph J Platz2, Keith S Naunheim2
1Department of Surgery, University of Chicago, Chicago, Ill.
The Journal of Thoracic and Cardiovascular Surgery
|November 13, 2024
Summary
Four popular chatbots performed significantly worse than board-certified surgeons on a thoracic surgery exam. AI implementation in clinical decision-making requires caution due to accuracy concerns.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Medicine
- Surgical Education
Background:
- Chatbots are increasingly used in medicine for clinical decision support.
- Concerns exist regarding the accuracy of information provided by artificial intelligence (AI) platforms.
- This study evaluates the performance of popular chatbots against expert surgeons.
Purpose of the Study:
- To assess the accuracy and performance of leading AI chatbots in a medical context.
- To compare the diagnostic, evaluative, and treatment-related responses of chatbots with those of board-certified thoracic surgeons.
- To identify potential risks associated with AI implementation in clinical decision-making.
Main Methods:
- Developed clinical scenarios based on the American Board of Thoracic Surgery (ABTS) Qualifying Exam.
- Utilized the Key Feature methodology for scenario construction, focusing on diagnosis, evaluation, and treatment.
- Administered 10 scenarios to four chatbots (ChatGPT-4, Gemini, Perplexity, Claude 2) and 21 board-certified surgeons, comparing scores using the Mann-Whitney U test.
Main Results:
- Chatbots achieved a median score of 1.06 per scenario, significantly lower than surgeons' median score of 1.88 (P=.019).
- Surgeons outperformed chatbots in all but two scenarios.
- Chatbot responses were more likely to result in critical failures (median 0.50) compared to surgeon responses (median 0.19; P=.016).
Conclusions:
- Popular AI chatbots demonstrate significantly lower performance levels compared to board-certified thoracic surgeons.
- The findings underscore the need for caution when implementing AI in clinical decision-making processes.
- Further research is warranted to improve AI accuracy and safety in medical applications.
Related Concept Videos
Issues And Trends In Healthcare Delivery System
5.6K
The issues and trends in healthcare delivery are constantly changing. The COVID-19 pandemic is one recent issue that wreaked havoc on healthcare systems, causing a shortage of healthcare workers, high demand for medicines and supplies, and increased medical expenditure due to a lack of insurance. Other issues include rising healthcare costs and care fragmentation.
Cost Containment
Payment for healthcare services has historically promoted adoption of costly and often unnecessary or inefficient...
Cost Containment
Payment for healthcare services has historically promoted adoption of costly and often unnecessary or inefficient...
5.6K
Current Trends in Nursing II
1.2K
Trends in nursing are multifactorial and associated with changes in society, within the nursing profession, and in other professions. Notably, telehealth and remote nursing contribute to successful healthcare delivery for numerous patients and help reduce stress for nurses due to nursing shortages. Nurses can reach patients, monitor their conditions, and interact with them using computers, audio, visual accessories, and telephones—for example, remote patient monitoring systems. Likewise,...
1.2K

