Related Experiment Video
Updated: May 3, 2026

05:37
An R-Based Landscape Validation of a Competing Risk Model
Published on: September 16, 2022
2.2K
Large Language Model Analysis of Reporting Quality of Randomized Clinical Trial Articles: A Systematic Review
Apoorva Srinivasan1, Jacob Berkowitz1, Nadine A Friedrich1
1Department of Computational Biomedicine, Cedars Sinai Medical Center, Los Angeles, California.
JAMA Network Open
|August 28, 2025
Summary
A new large-language-model (LLM) pipeline effectively assesses randomized clinical trial (RCT) reporting quality, identifying persistent gaps in crucial elements like allocation concealment and external validity discussions.
Area of Science:
- Clinical Trial Reporting
- Artificial Intelligence in Medicine
- Biomedical Research Transparency
Background:
- Incomplete reporting in randomized clinical trials (RCTs) hinders bias assessment and reproducibility.
- Manual audits of RCTs for adherence to Consolidated Standards of Reporting Trials (CONSORT) guidelines are not scalable.
- There is a need for automated methods to assess RCT reporting quality.
Purpose of the Study:
- To develop and validate a zero-shot large-language-model (LLM) pipeline for automated CONSORT guideline assessment.
- To analyze trends in reporting quality over time, across biomedical disciplines, and in relation to trial characteristics.
Main Methods:
- Systematic review of RCTs indexed on PubMed (1966-2024), converted from PDF to XML.
- LLM (Chat GPT-4o-mini) tested on a CONSORT-Text Classification Model (CONSORT-TM) benchmark and validated against expert reviews.
- LLM applied to assess CONSORT adherence across a large sample of RCTs, analyzing reporting of 21 CONSORT items.
Main Results:
- The LLM demonstrated high performance, matching expert reviews with 91.7% accuracy and achieving a macro F1 score of 0.86 on the CONSORT-TM benchmark.
- Mean CONSORT compliance increased from 27.3% (1966-1990) to 57.0% (2010-2024), but critical elements like allocation concealment (16.1%) and external validity discussion (1.6%) remain poorly reported.
- Reporting compliance varied significantly across disciplines (35.2% in pharmacology to 63.4% in urology) with minimal association with trial characteristics.
Conclusions:
- A zero-shot LLM can effectively audit CONSORT adherence at scale, providing insights into reporting quality.
- Persistent gaps in reporting critical trial elements necessitate targeted interventions.
- The findings highlight the need for enhanced editorial actions to improve transparency and reproducibility in biomedical research.

