Related Experiment Video
Updated: Jan 14, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Leveraging AI Large Language Models for Writing Clinical Trial Proposals in Dermatology: Instrument Validation Study
Megan Hauptman1, Daniel Copley2, Kelly Young1
1Department of Dermatology, University of Michigan, 1500 E Medical Center Dr, Ann Arbor, MI, 48105, United States, 1 734-936-4054.
Background:
Large language models (LLMs) are becoming increasingly popular in clinical trial design but have been underused in research proposal development.
Objective:
This study compared the performance of commonly used open access LLMs versus human proposal composition and review.
Methods:
A total of 10 LLMs were prompted to write a research proposal. Six physicians and each of the LLMs assessed 11 blinded proposals for capabilities and limitations in accuracy and comprehensiveness.
Results:
ChatGPT-o1 and Llama 3.1 were rated the most and least accurate, respectively, by human scorers. LLM scorers rated ChatGPT-o1 and DeepSeek R1 as the most accurate. ChatGPT-o1 and Llama 3.1 were rated as the most and least comprehensive, respectively, by human and LLM scorers. LLMs performed poorly on scoring proposals and, on average, rated proposals 1.9 points higher than humans for both accuracy and comprehensiveness.
Conclusions:
Paid versions of ChatGPT remain the highest-quality and most versatile option of the available LLMs. These tools cannot replace expert input but serve as powerful assistants, streamlining the development process and enhancing productivity.
More Related Videos
Related Concept Videos
Clinical Trials
There are four phases in a clinical trial. A phase one...
Improving Translational Accuracy
Improving Translational Accuracy
Clinical Trials: Overview
Statistical Software for Data Analysis and Clinical Trials

