Related Experiment Video
Updated: Jun 4, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Clinical outcomes and reporting quality of large language model interventions in practice: a systematic evidence map
Zixuan He1,2, Lan Yang1,2, Zitao Liang1,2
1Institute of Medical Technology, Peking University Health Science Center, Beijing, China.
Abstract:
Large language models (LLMs) are being deployed in clinical settings despite an underdeveloped evidence base regarding their real-world effectiveness. This study employed systematic evidence mapping to characterize outcome measures used in published studies and registered clinical trials (Jan 2022-Jun 2025) evaluating LLM performance. Analysis of 55 included studies revealed a predominance of human-AI collaborative designs (65.5%) for decision support and symptom management. LLM-only interventions focused on functional performance and operational or process impact outcomes (e.g., accuracy and time saving), whereas LLM-assisted interventions showed positive clinical effects, particularly in psychological health endpoints. Critical evidence gaps persist: diagnostic accuracy in randomized trials was notably lower and more variable (range 0.65-0.88) compared to non-randomized studies (typically ≥ 0.80); clinical efficiency impacts were inconsistent, and reporting quality was suboptimal (78.8% mean CONSORT-AI adherence), with critical omissions in handling data quality and performance errors. These findings indicate a heterogeneous and insufficient evidence landscape, necessitating standardized core outcome sets, mandatory use of specialized reporting guidelines, and robust clinical trials to ensure the safe integration of LLMs.
Related Concept Videos
Methods of Documentation VI: Case Management Model
For example, a patient with a chronic illness...
Modeling in Therapy
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in situations...