Related Experiment Video
Updated: Jan 12, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Learning from experts: A self-improving LLM framework for study population generation in clinical research
Yaoqian Sun1, Zikang Chen1, Hailing Cai1
1College of Biomedical Engineering and Instrument Science, Zhejiang University, Zheda Road, 310027 Hangzhou, Zhejiang Province, China.
Introduction:
The widespread adoption of electronic health records has led to the rapid accumulation of real-world data (RWD), an essential basis for generating real-world evidence (RWE). While large language models (LLMs) have supported multiple stages of RWD-driven research, their application to the design of study populations with both interpretability and credibility remains a challenge, which serves as a bridging role between study objectives and downstream analyses.
Methods:
In this study, we propose CriteriaLLM, a framework that enables LLMs to generate eligible study populations directly from clinical research objectives by incorporating clinician feedback. Inspired by the after-action review method, which facilitates learning from past experiences and feedback, we build an expert knowledge base that records the LLM output study populations and the modifications made by the clinician. A dual-retrieval algorithm, combining disease domain relevance and lexical similarity, then identifies relevant historical cases from the expert knowledge base to guide future generations. To ensure clinical relevance and real-world applicability, we introduce a continuous validation loop where expert feedback is iteratively integrated, refining model performance over time.
Results:
We evaluated our framework on 254 published clinical studies based on the MIMIC-III database using four representative LLMs: GPT-4o, Deepseek-R1, and two LLaMA models. The experimental results indicate that the proposed framework could effectively generate a high-quality study population with the highest Macro F1 score on 0.9180, maintaining generalizability across foundation models with varying parameter sizes and deployment methods.
Conclusion:
Our expert-in-the-loop framework allows LLMs to generate eligible study populations from clinical objectives without additional fine-tuning. By integrating structured expert feedback and retrieval guidance, it enhances the quality and reliability of study criteria. With continuous validation, the framework highlights a scalable approach toward self-improving systems that bridge generative AI with the demands for clinical appropriateness, reliability, and interpretability in clinical research.
More Related Videos
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Statistical Software for Data Analysis and Clinical Trials
Analysis of Population Pharmacokinetic Data

