Related Experiment Video
Updated: Sep 25, 2026

An R-Based Landscape Validation of a Competing Risk Model
Published on: September 16, 2022
Variability among large language models in assessing CONSORT compliance of published randomized clinical trials
Daniel Y Tsybulnik1, Justin J Gillette1, Thomas F Heston1,2
1Department of Medical Education and Clinical Sciences, Elson S. Floyd College of Medicine, Washington State University, Spokane, Washington, United States of America.
Background:
Peer review processes may inadequately assess compliance with established reporting guidelines such as the Consolidated Standards of Reporting Trials (CONSORT) criteria. Large language models (LLMs) demonstrate potential for systematic manuscript evaluation; however, how consistently different models assess adherence to CONSORT guidelines in published clinical trials remains unexplored.
Methods:
Twenty randomized controlled trials published in immunology journals between 2015 and 2016 were identified through PubMed. Three LLMs (ChatGPT-4o, Gemini 2.5 Flash, and Claude Sonnet 4.6) independently assessed compliance across 37 CONSORT 2010 subpoints. The primary endpoint was the difference between models in mean CONSORT compliance score. Secondary endpoints included inter-model agreement and the proportion of articles meeting a 90% compliance threshold. Statistical analysis employed repeated measures analysis of variance (ANOVA) with post-hoc pairwise comparisons (α = 0.05).
Results:
Mean CONSORT compliance rates were: ChatGPT-4o 80.8% (95% CI 75.8-85.8%), Gemini 2.5 Flash 64.4% (95% CI 58.0-70.8%), and Claude Sonnet 4.6 54.7% (95% CI 48.0-61.3%). Using a 90% compliance threshold as a quality benchmark, ChatGPT-4o identified 25% of papers (5/20) as meeting this standard, while Gemini 2.5 Flash and Claude Sonnet 4.6 each identified none (0/20) as meeting this standard. Repeated-measures ANOVA demonstrated significant differences between models (F(2,38) = 43.01, p < 0.001, partial η² = 0.694). All pairwise comparisons were statistically significant (ChatGPT-4o versus Gemini 2.5 Flash and ChatGPT-4o versus Claude Sonnet 4.6, both p < 0.001; Gemini 2.5 Flash versus Claude Sonnet 4.6, p = 0.014).
Conclusions:
LLMs varied substantially in their assessment of CONSORT compliance in published randomized trials, with a consistent ordering: ChatGPT-4o scored compliance highest, Gemini 2.5 Flash intermediate, and Claude Sonnet 4.6 lowest. Item-level agreement between models was only moderate, with no pairwise weighted kappa reaching the 0.61 substantial-agreement threshold. This inter-model variability indicates the need for standardized evaluation protocols before LLM-assisted manuscript screening is adopted.
Related Concept Videos
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast, controlled...
Clinical Trials
There are four phases in a clinical trial. A phase one...
Bioequivalence studies: Biowaivers
Bioequivalence Experimental Study Designs: Completely Randomized and Randomized Block Designs
Bioequivalence Data: Statistical Interpretation
Confounding in Epidemiological Studies
