Related Experiment Video
Updated: May 4, 2026

Treatment of Ankle Osteoarthritis with Total Ankle Replacement Through a Lateral Transfibular Approach
Published on: January 24, 2018
Large language models are comparable with commonly used statistical software: A validation of GPT 5.1 for frequentist
Mikhail Salzmann1,2, Nikolai Ramadanov1,2, Robert Prill1,2
1Centre of Orthopaedics, Traumatology and Plastic Surgery, Brandenburg Medical School, University Hospital Brandenburg an der Havel, Brandenburg an der Havel, Germany.
Purpose:
The purpose of this study was to evaluate whether Chat Generative Pre-trained Transformer (ChatGPT; Version 5.1) can reproduce frequentist meta-analytic calculations with an accuracy comparable to established statistical software in orthopaedic research.
Methods:
In this methodological comparison study, data from two previously published orthopaedic meta-analyses with identical statistical architectures as reference standards were used. Between-study variance (τ2) was estimated using the Sidik-Jonkman method and uncertainty was quantified using the Hartung-Knapp adjustment for the random-effects models, while common-effect models assume τ2 = 0. Original data extraction tables were provided to ChatGPT-5.1, which was instructed to perform the same analyses. ChatGPT-generated pooled mean differences, confidence intervals and heterogeneity statistics (I2, τ2, p values) were compared with verified reference results obtained using the meta and metafor packages in R.
Results:
Across seven evaluated outcomes, ChatGPT-5.1 reproduced the direction of effects in all cases. Deviations compared with reference meta-analyses were classified as minor in three outcomes (43%), moderate in one outcome (14%) and major in three outcomes (43%). Agreement was highest in low-heterogeneity settings, whereas substantial deviations occurred in outcomes with pronounced between-study heterogeneity, particularly under random-effects models.
Conclusion:
ChatGPT-5.1 demonstrates emerging capability to approximate frequentist meta-analytic calculations, particularly in low-heterogeneity settings. However, its tendency to underestimate between-study variability and to deviate in complex random-effects scenarios limits its reliability as a standalone tool. At present, large language models may support exploratory analyses but cannot fully replace dedicated statistical software for meta-analyses in orthopaedic research.
Level Of Evidence:
Level III.
Related Concept Videos
Statistical Methods to Analyze Parametric Data: Student t-Test and Goodness-of-Fit Test
The Student's t-test is a statistical test that examines if there is a statistically significant difference between the means of two groups. This test is instrumental when dealing with...
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
Statistical Software for Data Analysis and Clinical Trials
Introduction to R
