Related Experiment Video
Updated: Oct 10, 2026

Subject-specific Musculoskeletal Model for Studying Bone Strain During Dynamic Motion
Published on: April 11, 2018
Large language models show distinct multidimensional performance profiles and variable reproducibility in
Silvia Gandini1,2, Jakob Lindskog3,4, Yinan Yu5
1Milan Lab Research Department AC Milan S.p.A Milan Italy.
Purpose:
This study evaluated five publicly available large language model-based systems using a multidimensional, expert-based assessment of consensus-based clinical questions on bone stress injuries.
Methods:
Ten standardised questions derived from a 2025 international Delphi consensus on bone stress injuries were submitted to ChatGPT-5, Perplexity AI, Claude 4, Google Gemini 2.5 and Microsoft Copilot. Responses were independently rated by clinical experts using a multidimensional evaluation framework covering accuracy, comprehensiveness, understanding, reasoning, clarity, potential harm and trust, each rated on a 5-point Likert scale. Median scores were calculated for each question and model, and comparisons were performed using Kruskal-Wallis tests followed by Dunn's post hoc tests with Bonferroni correction. Interrater agreement was assessed using the weighted Gwet's agreement coefficient. Reproducibility was assessed by repeatedly submitting nine questions representing no, partial, and full expert consensus to each model, with response consistency scored from 1 to 3.
Results:
Significant differences between models were observed for comprehensiveness, reasoning, clarity and trust (all p < 0.001), with no differences in accuracy or harm. ChatGPT-5 and Claude 4 achieved significantly higher comprehensiveness scores, while ChatGPT-5 and Google Gemini 2.5 had significantly higher reasoning scores than Microsoft Copilot. Claude 4 demonstrated significantly lower clarity than all other models. ChatGPT-5 achieved significantly higher trust scores and showed the highest mean reproducibility score, followed by Perplexity AI and Gemini 2.5; Microsoft Copilot showed the lowest reproducibility. Interrater agreement ranged from fair to almost perfect, with substantial overall agreement.
Conclusions:
Publicly available large language model-based systems demonstrated moderate-to-high performance in addressing expert-derived, consensus-based questions on bone stress injuries, suggesting potential to support the communication of consensus-based knowledge. However, performance varied across evaluation dimensions and reproducibility, indicating that large language models may support, but should not replace, primary consensus documents or expert clinical judgement.
Level Of Evidence:
N/A.