Related Experiment Video
Updated: Aug 7, 2026

Introduction of an Integrated Pathology Image Management, Artificial Intelligence, and Reporting System
Published on: July 11, 2025
Sentence-Level Provenance for AI Medical Record Summarization: Formative Usability Evaluation of a Click-to-Inspect
Andrew Parambath1, Giordana Pulpo2, Vince Hartman3
1Stanford Medicine, 900 welch roadsuite 350, Palo Alto, US.
Background:
Large language models (LLMs) can generate fluent summaries of longitudinal medical records, but in high-stakes clinical settings, verification burden remains a barrier to trust. Existing provenance mechanisms, such as document-level citations and section references, often require manual search within long, fragmented notes, limiting their usefulness during time-constrained workflows for clinicians.
Objective:
To design and evaluate a sentence-level provenance interface ("click-to-inspect") that enables rapid verification of AI-generated longitudinal medical record summaries at the level of individual statements.
Methods:
Between November 2023 and January 2024, we conducted a formative usability study using remote moderated usability sessions via Zoom to evaluate a web-based sentence-level provenance interface for AI-generated longitudinal medical record summaries. A convenience sample of clinicians was recruited through email outreach to academic and professional networks across the United States. Formative usability testing was conducted with 46 clinician interactions using synthetic longitudinal patient charts. Participants included medical students, residents, and attending physicians across multiple specialties including internal medicine, dermatology, radiology, plastic surgery, anesthesiology, interventional radiology, obstetrics-gynecology, and family medicine. Usability was assessed using the System Usability Scale (SUS) and Net Promoter Score (NPS), alongside qualitative feedback.
Results:
Clinicians reported high usability (mean SUS score 86.25, SD 7.77; 95% CI 83.96-88.54 from 46 participants) and a positive overall experience (NPS 35; 22/46 promoters, 18/46 passives, 6/46 detractors). Participants described rapid access to supporting evidence as critical for trust calibration during first-pass chart review. Qualitative feedback identified friction in traditional citation-based interfaces and supported sentence-level inspectability as a low-friction verification mechanism.
Conclusions:
Sentence-level provenance transforms AI-generated summaries from static narratives into interactive verification tools. An approach that enables rapid, selective inspection of individual claims during longitudinal chart review, may reduce verification burden and support calibrated reliance in high-risk clinical contexts.

