Related Experiment Video
Updated: Feb 3, 2026

Introduction of an Integrated Pathology Image Management, Artificial Intelligence, and Reporting System
Published on: July 11, 2025
Performance of artificial intelligence in addressing questions regarding management of clavicle fractures
John D Milner1, Matthew Quinn1, Ashley Knebel1
1Department of Orthopaedic Surgery, Brown University, Warren Alpert Medical School, Providence, RI, USA.
Objectives:
Artificial intelligence (AI) has revolutionized public access to extensive information with large language model (LLM)-based chatbots allowing users to receive comprehensive, individualized responses. In this study, we aimed to evaluate the quality of LLM responses to questions about common orthopedic conditions. We hypothesized that both ChatGPT and Gemini would demonstrate high quality, evidence-based responses across evaluation criteria.
Methods:
Responses from ChatGPT and Gemini to prompts based on the 14 AAOS Clinical Practice Guidelines for clavicle fracture management were evaluated on six criteria by seven fellowship-trained shoulder and trauma orthopedic surgeons. Statistical analyses including mean scoring, standard deviation and two-sided t-tests were calculated to compare performance between ChatGPT and Gemini. Scores were then evaluated for inter-rater reliability (IRR).
Results:
ChatGPT and Gemini demonstrated overall mean scores greater than 3.5 for both platforms. Mean overall score for ChatGPT was highest in evidence-based (4.52 ± 0.16) and lowest in clarity (4.22 ± 0.19). Mean overall score for Gemini was highest in clarity (4.31 ± 0.17) and lowest in evidence-based (3.81 ± 0.22). ChatGPT had significantly better performance in the overall completeness category (4.50 ± 0.17 vs 4.11 ± 0.19, p < 0.005) than Gemini but scores were otherwise not significantly different. Over 70 % of respondents rated the responses of ChatGPT as higher quality than Gemini.
Conclusions:
ChatGPT and Gemini produced responses that were generally in line with the 2022 AAOS guidelines on the treatment of clavicle fractures. Scores were comparable in every overall category except completeness, with ChatGPT outperforming Gemini. These results suggest that both LLMs are capable of providing clinically relevant responses to questions related to clavicle fracture management.
Related Concept Videos
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Intelligence
Measures of Intelligence
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
Multiple Intelligences Theory
Fractures: Bone Repair
Minor fractures with no bone displacement are treated by immobilizing the fractured bone using a cast or splint. However, in the case of fractures with displaced bones, the broken bones are repositioned before immobilization to ensure successful healing without deformation and loss of function. The realignment of fractured bone ends is performed through a process called reduction. If the...
Cattell's Theory of Intelligence
Fluid intelligence involves the capacity to solve new problems and adapt to unfamiliar situations. It's the type of intelligence individuals use when they encounter a novel problem or puzzle that requires innovative thinking. For instance, figuring out how to operate a new gadget relies heavily on...

