Related Experiment Video
Updated: Jan 15, 2026

Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
Published on: June 25, 2019
Large Language Models in Randomized Controlled Trials Design: Observational Study
Liyuan Jin1, Jasmine Chiat Ling Ong1,2, Kabilan Elangovan3
1Duke-NUS Medical School, 8 College Road, Singapore, 169857, Singapore, 65 66016503.
Large language models (LLMs) show promise in improving randomized controlled trial (RCT) design, enhancing recruitment and generalizability. While effective in intervention planning, LLMs require expert oversight for eligibility criteria and outcome measures to ensure safety and ethical standards.
Area of Science:
- Clinical Trial Design
- Artificial Intelligence in Healthcare
- Medical Research Methodology
Background:
- Randomized controlled trials (RCTs) face significant challenges including limited generalizability, insufficient participant diversity, and high failure rates.
- These limitations often stem from restrictive eligibility criteria and inefficient patient selection processes.
- Large language models (LLMs) show potential in clinical applications, but their role in optimizing RCT design is largely unexplored.
Purpose of the Study:
- To investigate the capability of LLMs, specifically GPT-4-Turbo-Preview, in assisting the design of RCTs.
- To assess LLM's potential to improve RCT generalizability, recruitment diversity, and reduce failure rates.
- To evaluate LLM-assisted RCT design while upholding clinical safety and ethical standards.
Main Methods:
- An observational study analyzed 20 parallel-arm RCTs (10 completed, 10 registered) published after January 2024.
- LLMs generated RCT designs based on provided criteria, including eligibility, recruitment, interventions, and outcomes.
- Quantitative assessment by clinical experts and NLP metrics (BLEU, ROUGE-L, METEOR) evaluated LLM design accuracy against ClinicalTrials.gov data; qualitative assessments used Likert scales for safety, accuracy, bias, pragmatism, inclusivity, and diversity.
Main Results:
- LLMs achieved 72% overall accuracy in replicating RCT designs, with high accuracy in recruitment (88%) and intervention (93%) design.
- Lower accuracy was observed in designing eligibility criteria (55%) and outcomes measurement (53%).
- Qualitative evaluations indicated strong clinical alignment, with LLM-generated designs ranking similarly to original designs in safety, accuracy, and objectivity, while enhancing diversity and pragmatism.
Conclusions:
- LLMs demonstrate significant potential to enhance RCT design, particularly in recruitment and intervention strategies, improving generalizability and diversity.
- Expert oversight and regulatory frameworks are crucial for ensuring patient safety and ethical compliance in LLM-assisted RCT design.
- Further refinement of LLMs is needed to overcome limitations in eligibility criteria and outcomes measurement for broader clinical trial application.
More Related Videos
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
08:05Measuring Statistical Learning Across Modalities and Domains in School-Aged Children Via an Online Platform and Neuroimaging Techniques
Published on: June 30, 2020
Related Concept Videos
Bioequivalence Experimental Study Designs: Repeated Measures, Cross-Over, Carry-Over, and Latin Square Designs
Group Design
Study Designs in Epidemiology
Observational studies are those where the researcher does not intervene but rather observes natural variations. They include cross-sectional, cohort, and...
Bioequivalence Experimental Study Designs: Completely Randomized and Randomized Block Designs
Randomized Experiments
Simple randomization
Simple...
Observational Studies
There are three types of observational studies – Prospective, retrospective, and cross-sectional.
Prospective Study
Prospective studies, also known as longitudinal or cohort studies, are carried out by collecting future data from groups sharing similar characteristics. One...