Related Experiment Video
Updated: Jan 6, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
980
Evaluating clinical AI summaries with large language models as judges
Emma Croxford1, Yanjun Gao2, Elliot First3
1Department of Biostatistics and Medical Informatics, University of Wisconsin, Madison, USA.
NPJ Digital Medicine
|November 5, 2025
Summary
Automating the evaluation of AI-generated clinical summaries using a Large Language Model (LLM) approach significantly improves accuracy assessment. This method offers a scalable and efficient alternative to manual review for Electronic Health Records (EHRs).
Area of Science:
- Artificial Intelligence in Medicine
- Clinical Informatics
- Natural Language Processing
Background:
- Electronic Health Records (EHRs) generate extensive clinical data, posing challenges for healthcare providers to synthesize efficiently.
- Generative AI and Large Language Models (LLMs) offer potential for summarizing EHRs to reduce provider cognitive load.
- Ensuring the accuracy of AI-generated summaries necessitates robust evaluation methods, as human review is time-consuming and expensive.
Purpose of the Study:
- To introduce and validate an automated LLM-based method for assessing the quality of multi-document summaries derived from real-world EHR data.
- To benchmark the LLM-as-a-Judge framework against the established Provider Documentation Summarization Quality Instrument (PDSQI).
Main Methods:
- Developed and validated an LLM-as-a-Judge framework for evaluating EHR multi-document summaries.
- Benchmarked the LLM-based evaluation against human review using the PDSQI.
- Assessed inter-rater reliability and performance metrics including intraclass correlation coefficient (ICC) and evaluation time.
Main Results:
- The LLM-as-a-Judge framework demonstrated strong inter-rater reliability, comparable to human evaluators.
- GPT-o3-mini achieved an ICC of 0.818 and a median score difference of 0 from human assessments, completing evaluations in 22 seconds.
- Reasoning LLMs outperformed other approaches in inter-rater reliability, especially for complex evaluations requiring domain expertise.
Conclusions:
- Automated LLM-based evaluation provides a scalable and efficient method for assessing the accuracy and safety of AI-generated clinical summaries.
- The LLM-as-a-Judge approach can significantly reduce the burden and cost associated with traditional human review.
- This technology facilitates the reliable deployment of AI for synthesizing clinical data in Electronic Health Records.
Related Concept Videos
Improving Translational Accuracy
14.0K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.0K
Improving Translational Accuracy
3.5K
3.5K
Clinical Trials: Overview
4.5K
Clinical development focuses on how the drug will interact with the human body and encompasses four key phases of clinical trials, each serving a specific purpose in assessing the safety and effectiveness of new drugs. These phases overlap and build upon one another. Phase I involves a small group of healthy volunteers (typically 20-80 individuals) or, in cases where significant toxicity is expected, patients with the targeted disease, such as cancer or AIDS. The volunteers are tested for...
4.5K
