Related Experiment Video
Updated: Feb 28, 2026

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
MedFactEval and MedAgentBrief: A Framework and Workflow for Generating and Evaluating Factual Clinical Summaries
Françcois Grolleau1, Emily Alsentzer2, Timothy Keyes3
1Center for Biomedical Informatics Research, Stanford University, Stanford, CA, USA, grolleau@stanford.edu.
Evaluating clinical text from Large Language Models (LLMs) is challenging. A new framework, MedFactEval, uses an LLM Jury for scalable, fact-grounded assessment, achieving high agreement with human experts.
Area of Science:
- Artificial Intelligence
- Medical Informatics
- Clinical Natural Language Processing
Background:
- Evaluating factual accuracy in clinical text generated by Large Language Models (LLMs) is crucial for adoption.
- Manual expert review is not scalable for continuous quality assurance of LLM-generated clinical content.
Purpose of the Study:
- To introduce MedFactEval, a scalable framework for fact-grounded evaluation of LLM-generated clinical text.
- To present MedAgentBrief, a workflow for generating high-quality, factual clinical discharge summaries.
- To validate the MedFactEval framework against a gold standard established by human expert consensus.
Main Methods:
- MedFactEval utilizes clinician-defined key facts and an "LLM Jury" (multi-LLM majority vote) for evaluation.
- MedAgentBrief is a model-agnostic, multi-step workflow for generating discharge summaries.
- A gold standard was created using a seven-physician majority vote on key facts from inpatient cases.
Main Results:
- The MedFactEval LLM Jury demonstrated near-perfect agreement with the human expert panel (Cohen's κ = 81%).
- The LLM Jury's performance was statistically non-inferior to that of a single human expert (κ = 67%, P < 0.001).
Conclusions:
- MedFactEval offers a robust and scalable framework for evaluating factual accuracy in clinical LLM outputs.
- MedAgentBrief provides a high-performing generation workflow for factual clinical summaries.
- The combined approach facilitates the responsible deployment of generative AI in clinical settings.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
09:20Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
Published on: February 23, 2019
Related Concept Videos
Flow Sheet
Here's a closer look at the examples of flowsheets commonly used by nurses:
Graphic Sheet Documentation:
Clinical Trials: Overview
Formats for Nursing Documentation
Nursing Assessment Form:
• A nursing assessment form is a foundational document that captures detailed patient data from physical assessments and nursing histories.
• It includes patient demographics, medical history,...
Discharge Summary Forms
Here's a detailed look at the key components and guidelines for preparing a discharge summary:
Nursing Clinical Information System
A Nursing Clinical Information System (NCIS) is a specialized type of healthcare information system tailored to meet the unique needs of nursing practice. It incorporates the principles of nursing informatics to streamline information management and improve the quality of care delivery.
Critical attributes of NCIS include:
Methods of Documentation I: Source-Oriented Records
In an SOR, each discipline involved in patient care maintains a separate medical record section. This record-keeping method enables easy tracking of patient progress and ensures healthcare staff have access to up-to-date information.
Key Attributes include the following: