Related Experiment Video
Updated: Jun 30, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
SPIRIT-CONSORT-ELM: Element-Level Assessment of Randomized Controlled Trial Reporting Using Large Language Models
Lan Jiang1, Xiangji Ying2, Andrew W Brown3,4
1School of Information Sciences, University of Illinois Urbana-Champaign, IL, US.
Medrxiv : the Preprint Server for Health Sciences
|June 29, 2026
Summary
This study introduces SPIRIT-CONSORT-ELM, a new dataset for assessing randomized controlled trial (RCT) reporting completeness at the element level. An automated pipeline using machine learning accurately evaluates RCT transparency beyond checklist items.
Area of Science:
- Medical research methodology
- Clinical trial reporting standards
- Natural Language Processing in healthcare
Background:
- Incomplete reporting in randomized controlled trials (RCTs) hinders verification and utility.
- SPIRIT and CONSORT guidelines aim to improve protocol and results reporting, but completeness remains a challenge.
- Automated checking of manuscripts could enhance reporting quality before publication.
Purpose of the Study:
- To extend the SPIRIT-CONSORT-TM corpus with element-level annotations (SPIRIT-CONSORT-ELM) for assessing reporting completeness.
- To develop and evaluate an automated machine reading comprehension pipeline for element-level assessment of RCT reports.
- To establish a benchmark for evaluating reporting guideline completeness at a granular level.
Main Methods:
- Extended the SPIRIT-CONSORT-TM corpus with element-level annotations, formulating assessment as a machine reading comprehension task with 119 questions.
- Developed an automated pipeline using PubMedBERT to identify relevant sentences and a generative large language model (GPT-5) with chain-of-thought reasoning to answer element-level questions.
- Annotated 50 articles (25 pairs) by two independent annotators, with remaining 150 articles (75 pairs) assessed by one annotator; calculated inter-annotator agreement (Gwet's AC1: 0.782).
Main Results:
- The automated pipeline achieved high accuracy in identifying element-level reporting evidence (F1: 0.822, Gwet's AC1: 0.796).
- Ablation studies confirmed that chain-of-thought reasoning and in-context examples modestly improved the large language model's performance.
- SPIRIT-CONSORT-ELM provides a publicly available benchmark for detailed reporting completeness assessment.
Conclusions:
- SPIRIT-CONSORT-ELM enables a more nuanced assessment of RCT transparency than item-level checks alone.
- The automated pipeline offers a robust baseline for evaluating RCT reporting completeness.
- The developed system has the potential to serve as a practical tool for authors, reviewers, and editors to improve RCT report transparency.