Related Experiment Video
Updated: Feb 28, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Asking the Right Questions: Benchmarking Large Language Models in the Development of Clinical Consultation Templates
Liam G McCoy1, David Wu2, Sarita Khemani3
1Division of Neurology, Faculty of Medicine and Dentistry, University of Alberta, Edmonton, AB, Canada2Department of Medicine, Beth Israel Deaconess Medical Center, Boston, MA, USA, lmccoy@ualberta.ca.
Large language models (LLMs) can generate clinical consultation templates but struggle with conciseness and prioritization. Further development is needed for effective clinical information exchange in healthcare.
Area of Science:
- Artificial Intelligence in Medicine
- Clinical Informatics
- Natural Language Processing
Background:
- Electronic consultations (eConsults) facilitate efficient physician communication.
- Structured templates are crucial for clear and concise clinical information exchange.
- Large language models (LLMs) show potential for automating template generation.
Purpose of the Study:
- To evaluate the capability of state-of-the-art LLMs in generating structured clinical consultation templates.
- To assess the clinical coherence, conciseness, and prioritization of LLM-generated templates.
- To identify limitations of current LLMs in producing clinically relevant eConsult schemas.
Main Methods:
- Utilized 145 expert-crafted eConsult templates from Stanford's eConsult team.
- Assessed frontier LLMs including o3, GPT-4o, Kimi K2, Claude 4 Sonnet, Llama 3 70B, and Gemini 2.5 Pro.
- Employed a multi-agent pipeline with prompt optimization, semantic autograding, and prioritization analysis.
Main Results:
- Models achieved high comprehensiveness (up to 92.2%) but generated overly long templates.
- LLMs failed to prioritize clinically important questions effectively under length constraints.
- Performance varied by medical specialty, with notable degradation in psychiatry and pain medicine.
Conclusions:
- LLMs show promise for improving structured clinical information exchange.
- Current LLMs require enhanced evaluation methods focusing on clinical salience and prioritization.
- Addressing length constraints and specialty-specific nuances is critical for real-world LLM application in eConsults.
More Related Videos
Related Concept Videos
Methods of Documentation VI: Case Management Model
For example, a patient with a chronic...
Modeling in Therapy
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in...
Improving Translational Accuracy
Improving Translational Accuracy

