Related Experiment Video
Updated: Sep 10, 2026

Fine-Tuning Large Language Models Using Entity Hallucination Index for Text Summarization
Published on: January 9, 2026
Sentence-Level Provenance for AI Medical Record Summarization in a Click-to-Inspect Interface: Formative Usability
Andrew Parambath1, Giordana Pulpo2, Vince Hartman2
1Stanford Medicine, 900 Welch Road, Suite 350, Palo Alto, CA, 94304, United States, 1 2672979144.
Background:
Large language models can generate fluent summaries of longitudinal medical records, but in high-stakes clinical settings, verification burden remains a barrier to trust. Existing provenance mechanisms such as document-level citations and section references often require manual search within long, fragmented notes, limiting their usefulness during time-constrained workflows for clinicians.
Objective:
This study aimed to design and evaluate a sentence-level provenance interface ("click to inspect") that enables rapid verification of AI-generated longitudinal medical record summaries at the level of individual statements.
Methods:
Between November 2023 and January 2024, we conducted a formative usability study using remotely moderated usability sessions via Zoom to evaluate a web-based sentence-level provenance interface for AI-generated longitudinal medical record summaries. A convenience sample of clinicians was recruited through email outreach to academic and professional networks across the United States. Formative usability testing was conducted with 46 clinician interactions using synthetic longitudinal patient charts. Participants included medical students, residents, and attending physicians across multiple specialties, including internal medicine, dermatology, radiology, plastic surgery, anesthesiology, interventional radiology, obstetrics and gynecology, and family medicine. Usability was assessed using the System Usability Scale and net promoter score, alongside qualitative feedback.
Results:
Clinicians reported high usability (mean System Usability Scale score 86.25, SD 7.77; 95% CI 83.96-88.54 from 46 participants) and a positive overall experience (net promoter score of 35; 22/46, 47.8% promoters; 18/46, 39.1% passives; and 6/46, 13% detractors). Participants described rapid access to supporting evidence as critical for trust calibration during first-pass chart review. Qualitative feedback identified friction in traditional citation-based interfaces and supported sentence-level inspectability as a low-friction verification mechanism.
Conclusions:
Sentence-level provenance transforms AI-generated summaries from static narratives into interactive verification tools. An approach that enables rapid, selective inspection of individual claims during longitudinal chart review may reduce verification burden and support calibrated reliance in high-risk clinical contexts.
Related Concept Videos
Purpose of Health Records II
Formulating and Validating Nursing Diagnosis II
Risk nursing diagnoses represent clinical judgments of an individual, family, or community more vulnerable to developing the health problem than others...
Methods of Documentation III: PIE
Formulating and Validating Nursing Diagnosis I
There are thirteen domains for...
Pre-Procedural Guidelines for Assessing Blood Pressure