Related Experiment Video
Updated: Jan 14, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large Language Models for Clinical Trial Protocol Assessments
Euibeom Shin1, Amruta Gajanan Bhat1, Murali Ramanathan1
1Artificial Intelligence & Clinical Pharmacology Laboratory, Department of Pharmaceutical Sciences, University at Buffalo, The State University of New York, Buffalo, New York, USA.
None:
The purpose was to evaluate the utility of large language models (LLMs) for reviewing the statistical analysis plan (SAP) and pharmacokinetics-pharmacodynamics (PK-PD) components of clinical trial protocols. Clinical trial protocols and SAPs were obtained from clinicaltrials.gov for a testbed of 15 small-molecule drugs, biologics, and global antibiotic and public health interventions. The GPT-4o (ChatGPT) LLM was used to elicit study design attributes, relevant guidelines, and detailed SAP evaluations with prompts engineered to the persona of a regulatory expert. The SAP methodology was assessed against the Food and Drug Administration's (FDA) E9 Statistical Principles for Clinical Trials guidance. The SAP evaluation outputs were assessed in post hoc analyses with ChatGPT and Grok, based on a rubric that evaluated the accuracy of primary outcome identification, the correctness of statistical methodology, compliance with the FDA E9 guidance, and clinical interpretability. PK-PD analysis plans were assessed on the accuracy of PK-PD objectives and measures and PK analysis methods. ChatGPT accurately identified the disease, intervention, and comparator groups for all trials, as well as the study sample size for 14 out of 15 trials. The most frequently cited guidelines were the FDA's E9 guidance for SAP and the FDA Guidance for Industry: Population Pharmacokinetics for PK-PD. ChatGPT outputs of the SAP and PK-PD analysis plans were clear and organized, demonstrating a satisfactory ability to extract and summarize technical details; some limitations in contextual accuracy were observed. LLMs can be effective tools for assessing the SAP, PK-PD, and other aspects of clinical trial protocol reviews.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
12:18A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
Related Concept Videos
Clinical Trials
There are four phases in a clinical trial. A phase one...
Clinical Trials: Overview
Improving Translational Accuracy
Improving Translational Accuracy
Statistical Software for Data Analysis and Clinical Trials